Machine Learning-Based Yield Prediction for First-Row Transition Metal Catalyzed Cross-Coupling Reactions | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Machine Learning-Based Yield Prediction for First-Row Transition Metal Catalyzed Cross-Coupling Reactions Rajalakshmi C, Vivek Vijay, Abhirami Vijayakumar, Parvathi Santhoshkumar, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4011086/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The advent of first-row transition metal-catalyzed cross-coupling reactions has marked a significant milestone in the field of organic chemistry, primarily due to their pivotal role in facilitating the construction of carbon-carbon and carbon-heteroatom bonds. Traditionally, the determination of reaction yields has relied on experimental methods, but in recent times, the integration of efficient machine learning techniques has revolutionized this process. Developing a highly accurate predictive model for reaction yields applicable to diverse categories of cross-coupling reactions, however, remains a formidable challenge. In our study, we curated an extendable dataset encompassing a wide range of yields of cross-coupling reactions catalyzed by first-row transition metals through rigorous literature mining efforts. Using this dataset, we have developed an automated and open-access reaction model, employing both regression and classification methodologies. Our ML model could be used even by non-expert users, who can solely input the reaction components as datasets to predict the yields. We have achieved a correlation of 0.46 using the Random Forest regression approach and an accuracy of 0.54 using the K-Nearest Neighbours (KNN) classification which employs hyperparameter tuning. Considering the vast chemical space of our small dataset encompassing various transition metals catalysts and different categories of reactions, the above results are commendable. By releasing an open-access dataset comprising cross-coupling reactions catalyzed by 3d-transition metal, our study is anticipated to make a substantial contribution to the progression of predictive modeling for sustainable transition metal catalysis, thereby shaping the future landscape of synthetic chemistry. Physical sciences/Chemistry Physical sciences/Chemistry/Catalysis Physical sciences/Chemistry/Organic chemistry Physical sciences/Chemistry/Theoretical chemistry Machine Learning Cross-Coupling First-row transition metals Regression Classification Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Transition metal-catalyzed cross-coupling reactions represent a paradigmatic advancement in organic chemistry, offering unprecedented accessibility for the generation of carbon-carbon and carbon-heteroatom (C-O, C-S) bonds, thereby shaping the trajectory of synthetic methodologies 1 , 2 . These canonical methodologies find widespread utility in the synthesis of essential molecular entities, encompassing biaryls, 1,3-diynes, ethers, thioethers, and analogous structures, serving as indispensable precursors in the intricate assembly of diverse natural products, pharmaceuticals, and agrochemicals 3 , 4 . The development of sustainable and cost-effective reaction protocols is an emerging focus in catalytic research, with a particular interest in novel approaches utilizing catalysts based on first-row earth-abundant transition metals. Generally, the innovation of novel reaction protocols in cross-coupling reactions mandates the intricate optimization of substrates, coupling partners, catalysts, ligands, and reagents. This intricate process exposes practitioners to an expansive chemical space characterized by manifold permutations of reactants, demanding considerable resources and the expertise of synthetic chemists. The conventional empirical approach of validating each reaction through iterative experimentation poses inherent inefficiencies, resulting in notable chemical wastage. Within this context, the judicious application of machine learning (ML) techniques emerges as a salient prospect, leveraging the availability of experimental data. The construction of robust predictive models for cross-coupling reactions through ML holds the promise of precise prediction regarding key reaction outcomes such as yields, enantioselectivities, etc. 5 , 6 Consequently, they hold the potential to provide valuable assistance to chemists, offering benefits such as time and resource efficiency. Although, ML boasts a broad range of applications within the realm of chemistry, encompassing areas such as drug design, molecular design, reaction prediction etc, 7 , 8 its application for the prediction of reaction yields stands out as a less-explored domain and continues to pose a significant challenge. 9 – 12 This is attributed to the fact that the accuracy in predicting the yield for chemical processes largely depends on the data contained in the available datasets. To enable comprehensive reaction yield prediction, it becomes crucial to provide the machine learning system with a diverse training dataset encompassing various types of reactions. Chemists struggling with this challenge frequently encounter the common obstacle of the need for a substantial volume of data to attain accurate predictions. This requirement is often impeded by the scarcity of publicly available data sources. Henceforth, there is an urgent need for publicly available data sets to be used for ML. 13 Moreover, the application of machine learning for yield prediction is more effective when datasets are obtained from literature mining of published research articles as it also encompasses negative results. The most commonly available public databases are the USPTO (United States Patent and Trademark Office), Reaxys, SciFinder, etc 14 , and the data sets obtained from High-throughput experimentation (HTE) data. 15 Yet, the availability of publicly accessible databases is a valuable resource, it is important to note that these are often limited to successful reactions, which may not suffice for the creation of an effective machine-learning model. The HTE has proven to be effective in predicting yields for certain categories of reactions like Buchwald-Hartwig and Suzuki-Miyaura cross-coupling reactions, howbeit, these are not representative of the whole information contained in the published reaction data 16 , 17 . Another category of datasets that recently appeared in ML to overcome the above limitations is the literature-mined datasets. Recently, some studies have reported the literature mined data sets for predicting the yield in cross-coupling reactions. The literature-mined dataset NiCOlit stands as a benchmark for machine learning prediction of chemical yields for C-O coupling reactions catalyzed by Ni catalysts. 18 More recently, an ML-based yield prediction of cobalt and nickel catalyzed borylation reaction based on classification method is reported by Trofymchuk et al . 19 While the aforesaid predictive modeling studies of transition metal-catalyzed cross-coupling reactions showed reasonable performance, they were limited to either a specific transition metal or a particular category of cross-coupling reactions. Therefore, developing a predictive model that encompasses various transition metals for a diverse array of cross-coupling reactions catalyzed is, very essential and could aid in the design and development of economical and sustainable catalytic systems for cross-coupling reactions. In the present study, we have taken the initiative to make available a predictive model for cross-coupling reactions (as depicted in Fig. 1 ) catalyzed by first-row transition metals like Mn, Fe, Co, Cu, and Zn. Our dataset comprises a total of 963 distinct cross-coupling reactions, comprise of Suzuki coupling, Sonogashira coupling, Cadiot-Chodkiewicz coupling, Ullmann-type coupling reactions, etc which are catalyzed by various first-row transition metals. We have chosen the above category of cross-coupling reactions because of their pharmaceutical and agrochemical importance 20 . We meticulously extracted this data from published articles on cross-coupling reactions. Unlike the existing literature-mined data sets, which either focused on a single metal, or a single category of reactions, our dataset is more diverse and spread out. 18 , 21 As a result, predicting reaction yields in our dataset presents a greater challenge. We have performed various ML algorithms and featurisation methods on our dataset. Moreover, considering the lack of publicly available databases for transition metal-catalysed cross-coupling reactions, our primary objective is to create an open-access dataset for different categories of cross-coupling reactions encompassing different transition metal catalysts. By employing supervised learning techniques like classification and regression models, we have designed a predictive model for first-row transition metal-catalyzed cross-coupling reactions 9 . Our model has shown a superior accuracy and correlation with existing predictive models developed for cross-coupling reactions. By doing so, we hope to encourage the creation of more publicly available, machine-readable datasets, which would enhance the effectiveness of various machine learning models. Methodology Dataset preparation We obtained our dataset through literature mining, which encompasses a wide range of information related to metal catalysts and various categories of cross-coupling reactions which include Sonogashira coupling, Suzuki-Miyaura Coupling, Cadiot-Chodkiewicz coupling, Ullmann coupling, and Hiyama coupling reactions 22 – 30 . Among these different categories, to make a homogeneity in our dataset, we have chosen those reactions which give the products biaryl ethers, 1,3-diyynes, biarylacetylenes, biaryls etc. By considering the above conditions, we have a dataset comprised of 963 distinct first-row transition metal-catalyzed cross-coupling reactions. It is formatted for machine-readable access and includes crucial reaction parameters such as temperature and time. To make this data more accessible to other chemists, we have also included details of the research articles such as DOI and specific tables or schemes from which the data was extracted. The dataset employed was curated from academic literature sources and comprised distinct chemical reactions as individual data points. These reactions, which encompassed reactants, products, and related components, were subjected to a standardization procedure involving the use of SMILES (Simplified Molecular Input Line Entry Specification) representations 25 . Subsequently, the dataset was subjected to preprocessing to eliminate duplicate entries and to standardize various physical parameters associated with the reactions, such as temperature and duration, ensuring that the dataset is streamlined for efficient research and analysis. Our reaction data is broadly categorized into two distinct groups: optimization and scope, which are conveniently presented in our dataset (as shown in Fig. 2 ). In the optimization category, specific values are held constant while other parameters, such as reagents, are systematically modified to identify the most yielding reagent. On the other hand, in scope category of reactions, different combinations of substrates and coupling partners are simultaneously evaluated, providing a comprehensive exploration of reaction possibilities. In a cross-coupling reaction, there typically exist two key components: an electrophilic coupling partner, often represented by an organic halide, and a nucleophilic coupling partner, typically an organometallic species. Therefore, within our dataset, we have categorized these reactants into two distinct classes: aryl halides fall under the coupling partner class, while nucleophilic coupling partners are classified as substrate class. Figure 3 illustrates the distribution of counts for various substrate and coupling partner classes in the dataset. Our dataset was meticulously curated to encompass comprehensive information for each reaction, including the equivalence of substrate, coupling partner, ligand, and reagent, as well as details such as temperature, time, mechanism, and reaction conditions. This holistic approach ensures that all pertinent information is considered when attempting to yield predictions for a given reaction. Molecular Feature Engineering To render the pre-processed dataset into machine-readable format, we employed four distinct methods: RDkit fingerprints 31 , Density Functional Theory (DFT), DRFP (Differential Reaction Fingerprint) 32 and RXNFP 33 . These methods generate hash representations of molecules based on their SMILES notations, facilitating the model's ability to distinguish and categorize various aspects of chemical reactions. RDkit fingerprints were generated using the RDkit library. In the case of DRFP, it accepts a reaction SMILES as input and constructs a binary fingerprint derived from the symmetric difference of circular molecular n-grams generated from molecules located to the left and right of the reaction arrow, without differentiation between reactants and reagents. Additionally, one-hot encodings from scikit-learn 34 were employed to convert specific components of the reaction, such as ligands, catalysts, and reagents. physical reaction conditions, such as temperature, time, and equivalent masses, were directly fed to the model as numerical arrays without undergoing further feature engineering. The third approach for featurisation method is the DFT method. In our study, DFT calculations was performed to optimize the geometries of the reactants, products and reagents involved in the cross-coupling reactions, using Gaussian 09 35 employing B3LYP 36 functional. To improve the accuracy of the B3LYP functional, D3 version of Grimme's dispersion correction was added 37 . 6–31 + G(d) basis sets were used to describe C, H, N, Cl, S, F, B and O atoms. The effective core potentials of Hay and Wadt with a double ζ -basis set (LanL2DZ) 38 were used to describe Cu, Co, Zn, Fe, Mn, I, Br, atoms. Vibrational frequency calculations were performed to identify the stationary points as minima (zero imaginary frequency) or saddle points (one imaginary frequency) and to obtain thermal corrections. Then, using the Auto-QChem framework 39 , we have generated DFT-based molecular descriptors tailored for each reaction under consideration. Model Development and Hyperparameter Optimization In our pursuit to establish connections between reaction characteristics and the resulting yields, we explored both classification and regression type models for our dataset. For regression tasks, we utilized the random forest algorithm 28 from scikit-learn and a neural network approach with PyTorch. For classification, we used, the K-Nearest Neighbours (KNN) Algorithm, a supervised learning classifier that relies on proximity to categorize or predict the grouping of individual data points. To find the best hyperparameters for each model, we carried out an exhaustive search using GridSearch CV. During the development of these models, we experimented with different ways of splitting the dataset, including both stratified and non-stratified approaches. Upon partitioning the dataset using stratified approaches, the chosen parameters are the substrate class, coupling partner class, and the category of the reactions which we labelled as “mechanism” in our dataset. The division between the training and test sets was executed using both 7:3 and 8:2 ratios. Importantly, we automated the entire process, covering everything from feature engineering to model evaluation, using Python programming, which makes our model more user friendly. Results and Discussion In the present study, we have used both regression and classification models to predict the reaction yield for the 3d-transition metal-catalyzed cross-coupling reactions. Our dataset comprises information gathered from 22 research articles, all of which centre around coupling reactions utilizing first-row transition metals like Mn, Fe, Co, Cu, and Zn as catalysts. The datasets included reactions with a diverse range of experimental yields, spanning from negligible or zero to those surpassing 80%, and even slightly below the average level. This wide spectrum of yield outcomes indicates the dataset's diverse nature. However, within our dataset, certain reaction components are less represented in terms of frequency. While this expands the chemical space covered, it also has the potential to diminish prediction accuracy. To address this issue, we partitioned our primary dataset into multiple subsets based on criteria related to the frequency of various reaction components. This segmentation aimed to narrow down the chemical space covered by the dataset. A significant aspect of our model is the creation of an automated and openly accessible reaction model incorporating both regression and classification methodologies. Thereby, our model is designed to be user-friendly, allowing non-experts to input reaction components as datasets for yield predictions. The following sections describe the detailed results obtained from our studies based on regression and classification models. Regression Model: In our dataset consisting of 963 reactions, we attained the highest coefficient of correlation R 2 of 0.46, as depicted in Figure 4 . This was realized through the application of a random forest model, with each reaction being featured using the DRFP method. To ensure the inclusion of all types of reactants and products in both the training and testing sets, we classified each reaction based on its substrate, coupling partner, and mechanism or reaction type. This classification played a crucial role in partitioning the data into train-test sets, guaranteeing a balanced distribution of reaction classes across both splits. Remarkably, the most optimal results were obtained when employing a stratified split with the substrate class serving as the stratification criterion. In addition to utilizing the DRFP featurization method, we have investigated various other featurization methods in combination with Random Forest model, in which noteworthy correlations have been observed. Specifically, the utilization of DFT yielded a correlation of 0.45 ( Figure 5d ), demonstrating performance similar to that of the DRFP method. Additionally, employing RDkitFP resulted in a correlation of 0.43, while the RXNFP featurization method produced a correlation of 0.36. Notably, each of these approaches underwent a test-train split with stratification based on mechanism or reaction type, leading to improved outcomes. Furthermore, the tuning of Random Forest hyperparameters was undertaken in our study. we observed that both the DRFP method, utilizing substrate class for stratified sampling, and the RDkitFP method, using coupling partner class for stratified sampling, exhibited a correlation coefficient (R²) of 0.43 ( Figure 5b & 5e ).Top of Form In instances where the stratification was based on substrate class and the DFT method was employed, a correlation of 0.41 was observed. When train-test split was conducted based on the coupling partner class, the Neural Network model exhibited a correlation of 0.28 for the DRFP method. All these regression results are shown in Figure 5 and all other results are provided in Table S1 (see supporting information). Table 1: High correlation results of subsets, where HPT stands for Hyper Parameter Tuning, RMSE (Root Mean Square Error), MAE (mean absolute error). Dataset Model Type Featurisation Test Size Iterations RMSE MAE Correlation Subset Product Random forest HPT DRFP 0.2 10 27.26 21.69 0.30 Random forest DRFP 0.2 5 28.72 23.13 0.30 Subset Ligand Random forest DFT 0.3 10 23.80 18.40 0.45 Random forest RDkitFP 0.2 10 24.93 19.37 0.45 Subset Reagent Random forest HPT DFT 0.2 5 24.25 18.78 0.41 subset catalyst precursor Random forest DRFP 0.2 10 23.94 18.37 0.41 Subset Coupling partner Random forest HPT DFT 0.2 5 25.45 20.50 0.39 Subset Solvent Random forest HPT DFT 0.2 10 25.48 19.65 0.37 Subset Substrate Random forest DRFP 0.2 5 23.98 18.57 0.48 A distinct subset classification is established by directing attention toward specific components, including substrates, solvents, ligands, and similar factors. In doing so, reactions employing a particular element, such as a solvent, only once or twice are systematically excluded. This curation results in a refined dataset comprising only those solvents that are recurrently utilized. This discernment underscores the importance of the frequency of unique values within a dataset, indicating that reactions featuring more frequently occurring values often yield favourable outcomes compared to the broader original dataset, despite potentially containing fewer individual reactions. Remarkably, when we generated a subset by selecting “substrate” as the component with a higher frequency, we observed nearly higher results compared to our entire dataset in regression analyses. Specifically, we found a correlation of 0.48 when implementing no particular train-test split stratification and utilizing a Random Forest model in conjunction with DRFP. Another specific subset, obtained by focusing on the "Ligand" column, exhibits a correlation of 0.45 when evaluated in conjunction with Random Forest and the inclusion of both DRFP and DFT featurisation methods and also shows a better classification result. It's important to emphasize that these subsets were assessed without applying any stratification. Notably, it becomes evident that incorporating specific stratification techniques may significantly enhance the correlation of our regression model. Additional Outcomes of subsets are detailed in Table 1 . To further evaluate the performance of our model, we employed diverse dataset segmentation strategies, within the selected reactions. We created distinct subgroups, each comprising reactions in which the components exhibited varying count thresholds, including counts more than 1, 2, 3, 5, and 7. In the dataset where each component count was greater than 5, we observed a correlation of 0.38 with the Random Forest model utilizing the DRFP featurization method. Other higher correlation results are provided in supplementary information ( Table S3) . Validation of Model Performance on Out-of-Sample Data: To validate the performance of our model, we conducted an assessment of its predictive capabilities on out-of-sample data. We accomplished this by excluding reactions that have used substrate from the class of aryl acetylene, training the model on the remaining dataset, and subsequently evaluating its predictive performance. We achieved a correlation of 0.25 between the predicted and actual yields, demonstrating the robustness of our model when it comes to out-of-sample predictions. This finding underscores the model's capacity for accurate predictions beyond the training data. However, when we made out-of sample prediction by using a transition metal that was not known to training set, we have got poor accuracy. To illustrate this effect, we gathered data specifically for Sonogashira cross-coupling reactions employing Nickel as the catalyst and then train the model using a dataset that included other first-row transition metals and then applied it to predict 37 reactions with Nickel as the catalyst, the outcome was a negative R 2 value (R 2 =-2.94). Henceforth, it is notable that the transition metal choice has a significant impact on the model's performance in out-of-sample prediction. Classification Results In our partitioning of training and testing sets, the implementation of the coupling partner class as a stratification criterion, along with the utilization of Density Functional Theory (DFT) as the featurization approach, yielded a classifier accuracy of 0.48 and a precision of 0.43 for K-nearest neighbors (KNN) classification 29 hyperparameter tuning method as illustrated in (Figure 6c) . However, when KNN Classification with DRFP featurization method, involving no s tratification was considered, we got an accuracy about 0.46 ( Figure 6d) . Furthermore, employing the same classifier and utilizing the DRFP featurization method without the incorporation of a stratification method resulted in an improved accuracy of 0.54 and a precision of 0.53, as depicted in Figure 6a . Thus, the performance of our classification model signifies overall effectiveness for analyzing a literature-mined small dataset encompassing various cross-coupling reactions. A comparable accuracy (0.51) was obtained when KNN Classification using hyperparameter tuning with DFT as featurization method, and a s tratification criterion employing coupling partner class ( Figure 6b) . For detailed classification results of each model with different featurization methods. (Refer to supplementary information, Table S5 ). Table 2 provides additional details on the efficacy of the classification approach on our small data set. It was found that when classification is applied to the subset classes, we have got better results for the subset “ligand”. The classifier demonstrated an accuracy of 0.56 after applying hyperparameter tuning in the case of KNN classification with DRFP for the subset class solvent. We again observed that a subset “Solvent” also demonstrated an accuracy level almost similar to that of the entire dataset (0.51). These indicate the robust performance of our classification model. Notably, this assessment was conducted without the implementation of any specific stratification technique. Our findings underscore the significance of tailoring classification models to the specific attributes of the dataset which indicates that better predictive performance can be achieved by careful feature selection and algorithm customization. In another subset of our dataset, where each component count exceeded one, and eliminating all others where the components appear only once, we observed an accuracy of 0.49 and a precision of 0.51 when employing the K-nearest neighbors (KNN) classification model with the RDKitFP featurization method. More results for subsets characterized by higher correlation are presented in Table S6 (see supporting information). All these results indicate the efficacy of our predictive model in a small-dataset encompassing a wide chemical space, which are sparsely reported in literature. Discussion The results obtained from the present study illustrate the efficacy of our regression and classification models in predicting the yield of transition metal-catalyzed cross-coupling reactions. For our entire dataset comprising of 963 distinct reactions, we got a regression value of 0.46. Our model has exhibited strong performance despite the limited size of the dataset. Notably, when attempting to predict cross-coupling reactions involving Nickel as a catalyst, a metal not present in our dataset, the correlation was Table 2 Higher Classification results of each subset are given, where HPT stands for Hyper Parameter Tuning, . Dataset Model Type Featurisation Test Size Iterations Accuracy Precession Subset Product KNN classification HPT DRFP 0.2 10 0.49 0.47 Subset Ligand KNN classification HPT DRFP 0.2 10 0.56 0.54 Subset Reagent KNN classification HPT DFT 0.2 10 0.49 0.50 Subset Catalyst precursor KNN classification DRFP 0.2 10 0.49 0.52 Subset Coupling partner KNN classification DRFP 0.2 5 0.46 0.47 KNN classification HPT DRFP 0.2 5 0.46 0.47 Subset Solvent KNN classification DRFP 0.2 10 0.51 0.51 KNN classification HPT DRFP 0.2 5 0.51 0.48 Subset Substrate KNN classification DRFP 0.2 10 0.45 0.51 KNN classification HPT DRFP 0.2 10 0.45 0.46 KNN classification RDkitFP 0.2 5 0.45 0.47 KNN classification HPT RDkitFP 0.3 10 0.45 0.44 KNN classification HPT RxnFP 0.2 5 0.45 0.48 Accuracy= (Number of correct predictions)/(Total predictions made) lower (R 2 =-2.94), highlighting the importance of the specific transition metal used in the reactions. These findings suggest that our predictive method is well-suited for this type of analysis, with the caveat that the choice of transition metal in the reactions has a significant impact on prediction accuracy. Thus training of the data set with the different types of transition metal could increase the predictive capability of the model. Our study emphasizes the importance of extracting data from research articles, highlighting the potential for valuable insights using machine-readable data in predicting yields for a range of chemical reactions. The open-access database that we created for this work focussing on transition metal catalyzed cross-coupling reactions will serve as a database for further machine learning applications. The predictive performance of our machine learning-based yield prediction model, trained on a dataset comprising 963 reactions involving diverse first-row transition metals obtained through literature mining, demonstrate notable superiority compared to existing research findings. Upto our knowledge, it’s been the first time investigated on several transition metal-catalyzed reactions while the one reported on the highly heterogeneous USPTO dataset showed a correlation that is less than 0.2. 40 The recently reported literature-mined dataset focusing on borylation reactions with Co and Ni catalysts displayed a correlation of about 0.35 specifically for the Co catalyst and 0.27 correlation for Ni catalyst using regression methodologies. 21 This demonstrates the effectiveness of our machine learning model over the existing models when working with a literature-mined small dataset containing 963 reactions comprising diverse transition metals. Our study also shows the significance of acquiring data from published research to enhance the model's predictive capabilities, especially in the context of yield prediction. Although our regression and classification models have showcased their effectiveness in predicting the yields of cross-coupling reactions, yet there is room for improvement by expanding the dataset. Furthermore, the utilization of advanced models and feature engineering methods, such as trying other descriptors for Density Functional Theory (DFT), holds the promise of enhancing the model's performance, offering exciting prospects for future research endeavours. Conclusion In the present study, we have developed a comprehensive machine-learning model for yield prediction in transition metal-catalyzed cross-coupling reactions involving diverse first-row transition metals. The development of this model involved meticulous literature mining, resulting in a dataset of 936 reactions covering a wide spectrum of yields. The Random Forest regression model exhibited a significant correlation (R 2 = 0.46) within the diverse chemical space of the small dataset. Additionally, the KNN classification model, with hyperparameter tuning, achieved an accuracy of 0.54, indicating effectiveness in predicting favorable yields. Subset analyses further revealed promising results, with a correlation of 0.48 in regression method for reactions excluding those reactions in which substrates occurring once or twice. Similarly, in classification method, a subset focusing on ligands demonstrated an increased accuracy of 0.56. These findings indicate the accuracy of model for small-data sets. This further suggest the potential for performance enhancement with the inclusion of more data within the studied dataset, paving the way for future research. The commendable performance of our predictive model signifies a step toward automating traditional experimental trial-and-error methods, offering substantial resource and time savings. The automated and open accessible framework of our ML model by integrating regression and classification methodologies, made it accessible to non-expert users who can input reaction components as datasets to predict yields. By streamlining research processes, our ML model holds promise for advancing the arena of sustainable organometallic catalysis through more efficient and data-driven practices. Declarations Data availability: Details regarding our dataset are made publicly available and can be found in the repository: https://github.com/VITresearchgroup2024/ML_yield_prediction-/blob/main/DATA/Dataset.csv Code Availability: Additional details regarding our methodology and code implementation can be found in the following repository: https://github.com/VITresearchgroup2024/ML_yield_prediction- Supporting Information Coordinates of optimized geometries of reactants, products, and transition states are given as supporting information. Conflicts of interest The authors declare no competing interests Acknowledgements C. Rajalakshmi thank the University Grants Commission (UGC, India) for Senior Research Fellowship (SRF). Authors also thank the HPC PARAMASTRA, Mahatma Gandhi University Innovation Foundation (MGUIF), Mahatma Gandhi University Campus, Kottayam, for providing the High-Performance Computing facilities (https://hpcparamastra.mguif.com/). Author Contribution C. Rajalakshmi: Conceptualization, Investigation, Formal Analysis, Resources, Data Curation, Validation, writing – Original draft, Methodology, Visualization.Vivek Vijay: Investigation, Writing, Methodology, Validation, Visualization, Data Curation Abhirami Vijayakumar : Investigation, Writing, Methodology, Validation, Visualization, Data Curation Parvathi Santhoshkumar: Validation, Visualization, Methodology, Data CurationG. Krishnaveni: Validation, Visualization, Data Curation. John B Kottooran: Validation, Data Curation, Visualization.Ann Miriam Abraham: Validation, Data Curation, Visualization C S Anjanakutty: Validation, Data Curation, Visualization. Binuja Varghese: Data CurationVibin Ipe Thomas: Conceptualization, Methodology, Resources, Software, Investigation, Supervision, Formal Analysis, Visualization, Project Administration, Validation. References Pérez Sestelo, J. & Sarandeses, L. A. Advances in Cross-Coupling Reactions. Molecules 25 , 4500 (2020). Han, F. S. Transition-metal-catalyzed Suzuki–Miyaura cross-coupling reactions: a remarkable advance from palladium to nickel catalysts. Chem. Soc. Rev. 42 , 5270–5298 (2013). Penn, L. & Gelman, D. Copper-Mediated Cross-Coupling Reactions. in PATAI’S Chemistry of Functional Groups (John Wiley & Sons, Ltd, 2011). doi:10.1002/9780470682531.pat0451. Ayogu, J. I. & Onoabedje, E. A. Recent advances in transition metal-catalysed cross-coupling of (hetero)aryl halides and analogues under ligand-free conditions. Catal. Sci. Technol. 9 , 5233–5255 (2019). Lledós, A. Computational Organometallic Catalysis: Where We Are, Where We Are Going. Eur. J. Inorg. Chem. 2021 , 2547–2555 (2021). Meyer, B., Sawatlon, B., Heinen, S., von Lilienfeld, O. A. & Corminboeuf, C. Machine learning meets volcano plots: computational discovery of cross-coupling catalysts. Chem. Sci. 9 , 7069–7077 (2018). Stocker, S., Csányi, G., Reuter, K. & Margraf, J. T. Machine learning in chemical reaction space. Nat. Commun. 11 , 5505 (2020). Meuwly, M. Machine Learning for Chemical Reactions. Chem. Rev. 121 , 10218–10239 (2021). Stevens, J. M. et al. Advancing Base Metal Catalysis through Data Science: Insight and Predictive Models for Ni-Catalyzed Borylation through Supervised Machine Learning. Organometallics 41 , 1847–1864 (2022). Hueffel, J. A. et al. Accelerated dinuclear palladium catalyst identification through unsupervised machine learning. Science (80-. ). 374 , 1134–1140 (2021). Żurański, A. M., Martinez Alvarado, J. I., Shields, B. J. & Doyle, A. G. Predicting Reaction Yields via Supervised Learning. Acc. Chem. Res. 54 , 1856–1865 (2021). Kovács, D. P., McCorkindale, W. & Lee, A. A. Quantitative interpretation explains machine learning models for chemical reaction prediction and uncovers bias. Nat. Commun. 12 , 1695 (2021). Baldi, P. Call for a Public Open Database of All Chemical Reactions. J. Chem. Inf. Model. 62 , 2011–2014 (2022). Mutton, T. & Ridley, D. D. Understanding Similarities and Differences between Two Prominent Web-Based Chemical Information and Data Retrieval Tools: Comments on Searches for Research Topics, Substances, and Reactions. J. Chem. Educ. 96 , 2167–2179 (2019). Schwaller, P., Vaucher, A. C., Laino, T. & Reymond, J. L. Prediction of chemical reaction yields using deep learning. Mach. Learn. Sci. Technol. 2 , (2021). Ahneman, D. T., Estrada, J. G., Lin, S., Dreher, S. D. & Doyle, A. G. Predicting reaction performance in C–N cross-coupling using machine learning. Science (80-. ). 360 , 186–190 (2018). Perera, D. et al. A platform for automated nanomole-scale reaction screening and micromole-scale synthesis in flow. Science (80-. ). 359 , 429–434 (2018). Schleinitz, J. et al. Machine Learning Yield Prediction from NiCOlit, a Small-Size Literature Data Set of Nickel Catalyzed C-O Couplings. J. Am. Chem. Soc. 144 , 14722–14730 (2022). Pereira, A. & Trofymchuk, O. S. Machine Learning Prediction of High-Yield Cobalt- and Nickel-Catalyzed Borylations. J. Phys. Chem. C 127 , 12983–12994 (2023). Campeau, L. C. & Hazari, N. Cross-Coupling and Related Reactions: Connecting Past Success to the Development of New Reactions for the Future. Organometallics 38 , 3 (2019). Pereira, A. & Trofymchuk, O. S. Machine Learning Prediction of High-Yield Cobalt- and Nickel- Catalyzed Borylations. (2023). Thomas, A. M., Sherin, D. R., Asha, S., Manojkumar, T. K. & Anilkumar, G. Exploration of the mechanism and scope of the CuI/DABCO catalysed C–S coupling reaction. Polyhedron 176 , 114269 (2020). Rohit, K. R., Saranya, S., Harry, N. A. & Anilkumar, G. A Novel Ligand-free Manganese-catalyzed C-O Coupling Protocol for the Synthesis of Biaryl Ethers. ChemistrySelect 4 , 5150–5154 (2019). Asha, S., Thomas, A. M., Ujwaldev, S. M. & Anilkumar, G. A Novel Protocol for the Cu-Catalyzed Sonogashira Coupling Reaction between Aryl Halides and Terminal Alkynes using trans-1,2-Diaminocyclohexane Ligand. ChemistrySelect 1 , 3938–3941 (2016). Asha, S. et al. A convenient route to 1,3-diynes using ligand-free Cadiot-Chodkiewicz coupling reaction at room temperature under aerobic conditions. Synth. Commun. 49 , 256–265 (2019). Krishnan, K. K., Ujwaldev, S. M., Thankachan, A. P., Harry, N. A. & Gopinathan, A. A novel Zinc-catalyzed Cadiot-Chodkiewicz cross-coupling reaction of terminal alkynes with 1-bromoalkynes in ethanol solvent. Mol. Catal. 440 , 140–147 (2017). Krishnan, K. K. et al. Zinc-Catalyzed Etherification Reaction of Aryl Iodides with Phenols. ChemistrySelect 4 , 3984–3988 (2018). Sindhu, K. S. et al. A green approach for arylation of phenols using iron catalysis in water under aerobic conditions. J. Catal. 348 , 146–150 (2017). Thankachan, A. P., Sindhu, K. S., Krishnan, K. K. & Anilkumar, G. A novel and efficient zinc-catalyzed thioetherification of aryl halides. RSC Adv. 5 , 32675–32678 (2015). Thankachan, A. P., Sindhu, K. S., Krishnan, K. K. & Anilkumar, G. An efficient zinc-catalyzed cross-coupling reaction of aryl iodides with terminal aromatic alkynes. Tetrahedron Lett. 56 , 5525–5528 (2015). Lovrić, M., Molero, J. M. & Kern, R. PySpark and RDKit: Moving towards Big Data in Cheminformatics. Mol. Inform. 38 , (2019). Probst, D., Schwaller, P. & Reymond, J. L. Reaction classification and yield prediction using the differential reaction fingerprint DRFP. Digit. Discov. 1 , 91–97 (2022). Schwaller, P. et al. Mapping the space of chemical reactions using attention-based neural networks. Nat. Mach. Intell. 3 , 144–152 (2021). Varoquaux, G. et al. Scikit-learn. GetMobile Mob. Comput. Commun. 19 , 29–33 (2015). M. J. Frisch, G. W. Trucks, H. B. Schlegel, G. E. Scuseria, M. A. Robb, J. R. Cheeseman, G. Scalmani, V. Barone, G. A. Petersson, H. Nakatsuji, X. Li, M. Caricato, A. Marenich, J. Bloino, B. G. Janesko, R. Gomperts, B. Mennucci, H. P. Hratchian, J. V. Ort, and D. J. F. Gaussian 09, Revision D.01. at (2016). Becke, A. B3LYP. J. Chem. Phys. (1993). Grimme, S., Antony, J., Ehrlich, S. & Krieg, H. A consistent and accurate ab initio parametrization of density functional dispersion correction (DFT-D) for the 94 elements H-Pu. J. Chem. Phys. 132 , 154104 (2010). Chiodo, S., Russo, N. & Sicilia, E. LANL2DZ basis sets recontracted in the framework of density functional theory. J. Chem. Phys. 125 , 104107 (2006). Żurański, A. M., Wang, J. Y., Shields, B. J. & Doyle, A. G. Auto-QChem: an automated workflow for the generation and storage of DFT calculations for organic molecules. React. Chem. Eng. 7 , 1276–1284 (2022). Schwaller, P., Vaucher, A. C., Laino, T. & Reymond, J.-L. Prediction of chemical reaction yields using deep learning. Mach. Learn. Sci. Technol. 2 , 015016 (2021). Additional Declarations No competing interests reported. Supplementary Files SupplementaryInformation.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4011086","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":283147839,"identity":"5b5b29fd-e46e-4db4-a40d-b18982efd7bd","order_by":0,"name":"Rajalakshmi C","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Rajalakshmi","middleName":"","lastName":"C","suffix":""},{"id":283147840,"identity":"2e797578-8f33-49f5-b7d8-155072fb2a3d","order_by":1,"name":"Vivek Vijay","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Vivek","middleName":"","lastName":"Vijay","suffix":""},{"id":283147841,"identity":"39c5fb63-bbb9-4c6e-b136-3aa887a7aae0","order_by":2,"name":"Abhirami Vijayakumar","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Abhirami","middleName":"","lastName":"Vijayakumar","suffix":""},{"id":283147842,"identity":"1aec0779-f862-48ab-9837-26d73fd8c42d","order_by":3,"name":"Parvathi Santhoshkumar","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Parvathi","middleName":"","lastName":"Santhoshkumar","suffix":""},{"id":283147843,"identity":"7749ef76-a6e0-4330-9092-6422a36d1c67","order_by":4,"name":"John B Kottooran","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"John","middleName":"B","lastName":"Kottooran","suffix":""},{"id":283147844,"identity":"3e1f7921-49b4-454e-93e1-f95dc9dc05bb","order_by":5,"name":"Ann Miriam Abraham","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Ann","middleName":"Miriam","lastName":"Abraham","suffix":""},{"id":283147845,"identity":"7bd2d275-7ace-4883-ada9-0aecc9b9aedb","order_by":6,"name":"Krishnaveni G","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Krishnaveni","middleName":"","lastName":"G","suffix":""},{"id":283147846,"identity":"ced5f96f-cb86-4b2f-bf71-e048ce576c77","order_by":7,"name":"Anjanakutty C S","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Anjanakutty","middleName":"C","lastName":"S","suffix":""},{"id":283147847,"identity":"215cac08-af92-4912-a2a1-a8309e377ccc","order_by":8,"name":"Binuja Varghese","email":"","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":false,"prefix":"","firstName":"Binuja","middleName":"","lastName":"Varghese","suffix":""},{"id":283147848,"identity":"588f46c2-78c8-40bb-9ee2-587e8655a327","order_by":9,"name":"Vibin Ipe Thomas","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA70lEQVRIiWNgGAWjYDACZh6GAwwVYCbjgQcw0QSCWs5A2AfgKvFqYeABGt+GrgUfMDjOe/Bw4Ty7aPn2wwcOJFTUyjGwH37A8HAHHi2H+RIOz9yWnLvhTFrCgYQzx40ZeNIMGBLP4NYi2cxjcJh324HcDRI8BgcS244lNjDkMDAkthHSMudA7vwZIC3/jtU38L/Br4WfGaSl4UBuww2QloaaBAYJArbwMwP9wnMM5pdjBwzbJJ6BXIhbCxv/2cOfeWrscue3Hz744ENNnTw/f/LDhz/xaEEHhxnYgOQB4jUwMNSRongUjIJRMApGCAAA9ktVwaFehoIAAAAASUVORK5CYII=","orcid":"","institution":"CMS College Kottayam (Autonomous)","correspondingAuthor":true,"prefix":"","firstName":"Vibin","middleName":"Ipe","lastName":"Thomas","suffix":""}],"badges":[],"createdAt":"2024-03-04 08:17:47","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4011086/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4011086/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":53408180,"identity":"65ee8d15-b9f0-48eb-b1d9-09b041dd8ecc","added_by":"auto","created_at":"2024-03-25 15:57:02","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":38708,"visible":true,"origin":"","legend":"\u003cp\u003eCategories of first-row transition metal catalyzed cross-coupling reactions considered in this study.\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/e64d799a6b28f3aedfc27812.jpg"},{"id":53408175,"identity":"c1df6177-b22c-46fe-a970-a582c56393b1","added_by":"auto","created_at":"2024-03-25 15:57:01","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":34217,"visible":true,"origin":"","legend":"\u003cp\u003eYield range of reactions in the dataset which has origin as optimization and scope\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/899368bd8f568ceb6dd34843.jpg"},{"id":53408846,"identity":"aaf9d768-5b6f-4d4d-be37-b7411b4912e2","added_by":"auto","created_at":"2024-03-25 16:05:02","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":82557,"visible":true,"origin":"","legend":"\u003cp\u003eCount of different substrate class and coupling partner class for the C-C and C-X cross coupling reactions reported in literature (X= O, N and S).\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/2002ad01743a15878a69890f.jpg"},{"id":53408181,"identity":"edc2b82c-9648-41cf-8106-1a31fa2ad631","added_by":"auto","created_at":"2024-03-25 15:57:02","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":34660,"visible":true,"origin":"","legend":"\u003cp\u003eReaction yield prediction by Random Forest regression method using DRFP as featurisation technique employing substrate class as stratification criterion.\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/15b124102c467b5e9dcbfad9.jpg"},{"id":53408182,"identity":"2118ed82-fc59-41e8-bee0-0f8af12c6c42","added_by":"auto","created_at":"2024-03-25 15:57:02","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":144420,"visible":true,"origin":"","legend":"\u003cp\u003eRegression results for our dataset explored through different models employing varied featurization methods.\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/89dba67d30f2ee697f5cb615.jpg"},{"id":53408177,"identity":"7d1d8f31-87e4-4e3f-a951-2d401c1c74f8","added_by":"auto","created_at":"2024-03-25 15:57:01","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":104162,"visible":true,"origin":"","legend":"\u003cp\u003eHigh Classification outcomes achieved across diverse models using various featurization techniques.\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/9d426599604c9899683108eb.jpg"},{"id":54956941,"identity":"2aaeca5d-93cd-4fb1-80ab-82cb81b9b337","added_by":"auto","created_at":"2024-04-19 07:31:12","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":734882,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/4da191f4-fe30-4de7-9ede-1d9291256e9b.pdf"},{"id":53408179,"identity":"9871b24a-1f0f-4fe9-8f4a-0f35284b75b4","added_by":"auto","created_at":"2024-03-25 15:57:02","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":714910,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformation.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4011086/v1/1b6f6c08e1202e285f6c0df3.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning-Based Yield Prediction for First-Row Transition Metal Catalyzed Cross-Coupling Reactions","fulltext":[{"header":"Introduction","content":"\u003cp\u003eTransition metal-catalyzed cross-coupling reactions represent a paradigmatic advancement in organic chemistry, offering unprecedented accessibility for the generation of carbon-carbon and carbon-heteroatom (C-O, C-S) bonds, thereby shaping the trajectory of synthetic methodologies\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. These canonical methodologies find widespread utility in the synthesis of essential molecular entities, encompassing biaryls, 1,3-diynes, ethers, thioethers, and analogous structures, serving as indispensable precursors in the intricate assembly of diverse natural products, pharmaceuticals, and agrochemicals\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. The development of sustainable and cost-effective reaction protocols is an emerging focus in catalytic research, with a particular interest in novel approaches utilizing catalysts based on first-row earth-abundant transition metals. Generally, the innovation of novel reaction protocols in cross-coupling reactions mandates the intricate optimization of substrates, coupling partners, catalysts, ligands, and reagents. This intricate process exposes practitioners to an expansive chemical space characterized by manifold permutations of reactants, demanding considerable resources and the expertise of synthetic chemists. The conventional empirical approach of validating each reaction through iterative experimentation poses inherent inefficiencies, resulting in notable chemical wastage. Within this context, the judicious application of machine learning (ML) techniques emerges as a salient prospect, leveraging the availability of experimental data. The construction of robust predictive models for cross-coupling reactions through ML holds the promise of precise prediction regarding key reaction outcomes such as yields, enantioselectivities, etc.\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e Consequently, they hold the potential to provide valuable assistance to chemists, offering benefits such as time and resource efficiency.\u003c/p\u003e \u003cp\u003eAlthough, ML boasts a broad range of applications within the realm of chemistry, encompassing areas such as drug design, molecular design, reaction prediction etc,\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e its application for the prediction of reaction yields stands out as a less-explored domain and continues to pose a significant challenge.\u003csup\u003e\u003cspan additionalcitationids=\"CR10 CR11\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e–\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003eThis is attributed to the fact that the accuracy in predicting the yield for chemical processes largely depends on the data contained in the available datasets. To enable comprehensive reaction yield prediction, it becomes crucial to provide the machine learning system with a diverse training dataset encompassing various types of reactions. Chemists struggling with this challenge frequently encounter the common obstacle of the need for a substantial volume of data to attain accurate predictions. This requirement is often impeded by the scarcity of publicly available data sources. Henceforth, there is an urgent need for publicly available data sets to be used for ML.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e Moreover, the application of machine learning for yield prediction is more effective when datasets are obtained from literature mining of published research articles as it also encompasses negative results. The most commonly available public databases are the USPTO (United States Patent and Trademark Office), Reaxys, SciFinder, etc\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e, and the data sets obtained from High-throughput experimentation (HTE) data.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e Yet, the availability of publicly accessible databases is a valuable resource, it is important to note that these are often limited to successful reactions, which may not suffice for the creation of an effective machine-learning model. The HTE has proven to be effective in predicting yields for certain categories of reactions like Buchwald-Hartwig and Suzuki-Miyaura cross-coupling reactions, howbeit, these are not representative of the whole information contained in the published reaction data\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAnother category of datasets that recently appeared in ML to overcome the above limitations is the literature-mined datasets. Recently, some studies have reported the literature mined data sets for predicting the yield in cross-coupling reactions. The literature-mined dataset NiCOlit stands as a benchmark for machine learning prediction of chemical yields for C-O coupling reactions catalyzed by Ni catalysts. \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e More recently, an ML-based yield prediction of cobalt and nickel catalyzed borylation reaction based on classification method is reported by Trofymchuk \u003cem\u003eet al\u003c/em\u003e.\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e While the aforesaid predictive modeling studies of transition metal-catalyzed cross-coupling reactions showed reasonable performance, they were limited to either a specific transition metal or a particular category of cross-coupling reactions. Therefore, developing a predictive model that encompasses various transition metals for a diverse array of cross-coupling reactions catalyzed is, very essential and could aid in the design and development of economical and sustainable catalytic systems for cross-coupling reactions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the present study, we have taken the initiative to make available a predictive model for cross-coupling reactions (as depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e catalyzed by first-row transition metals like Mn, Fe, Co, Cu, and Zn. Our dataset comprises a total of 963 distinct cross-coupling reactions, comprise of Suzuki coupling, Sonogashira coupling, Cadiot-Chodkiewicz coupling, Ullmann-type coupling reactions, etc which are catalyzed by various first-row transition metals. We have chosen the above category of cross-coupling reactions because of their pharmaceutical and agrochemical importance\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. We meticulously extracted this data from published articles on cross-coupling reactions. Unlike the existing literature-mined data sets, which either focused on a single metal, or a single category of reactions, our dataset is more diverse and spread out. \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e As a result, predicting reaction yields in our dataset presents a greater challenge. We have performed various ML algorithms and featurisation methods on our dataset. Moreover, considering the lack of publicly available databases for transition metal-catalysed cross-coupling reactions, our primary objective is to create an open-access dataset for different categories of cross-coupling reactions encompassing different transition metal catalysts. By employing supervised learning techniques like classification and regression models, we have designed a predictive model for first-row transition metal-catalyzed cross-coupling reactions\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Our model has shown a superior accuracy and correlation with existing predictive models developed for cross-coupling reactions. By doing so, we hope to encourage the creation of more publicly available, machine-readable datasets, which would enhance the effectiveness of various machine learning models.\u003c/p\u003e "},{"header":"Methodology","content":"\u003cp\u003e \u003cb\u003eDataset preparation\u003c/b\u003e \u003c/p\u003e\u003cp\u003eWe obtained our dataset through literature mining, which encompasses a wide range of information related to metal catalysts and various categories of cross-coupling reactions which include Sonogashira coupling, Suzuki-Miyaura Coupling, Cadiot-Chodkiewicz coupling, Ullmann coupling, and Hiyama coupling reactions\u003csup\u003e\u003cspan additionalcitationids=\"CR23 CR24 CR25 CR26 CR27 CR28 CR29\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e–\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. Among these different categories, to make a homogeneity in our dataset, we have chosen those reactions which give the products biaryl ethers, 1,3-diyynes, biarylacetylenes, biaryls etc. By considering the above conditions, we have a dataset comprised of 963 distinct first-row transition metal-catalyzed cross-coupling reactions. It is formatted for machine-readable access and includes crucial reaction parameters such as temperature and time. To make this data more accessible to other chemists, we have also included details of the research articles such as DOI and specific tables or schemes from which the data was extracted.\u003c/p\u003e\u003cp\u003eThe dataset employed was curated from academic literature sources and comprised distinct chemical reactions as individual data points. These reactions, which encompassed reactants, products, and related components, were subjected to a standardization procedure involving the use of SMILES (Simplified Molecular Input Line Entry Specification) representations\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. Subsequently, the dataset was subjected to preprocessing to eliminate duplicate entries and to standardize various physical parameters associated with the reactions, such as temperature and duration, ensuring that the dataset is streamlined for efficient research and analysis. Our reaction data is broadly categorized into two distinct groups: optimization and scope, which are conveniently presented in our dataset (as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003e). In the optimization category, specific values are held constant while other parameters, such as reagents, are systematically modified to identify the most yielding reagent. On the other hand, in scope category of reactions, different combinations of substrates and coupling partners are simultaneously evaluated, providing a comprehensive exploration of reaction possibilities.\u003c/p\u003e\u003cp\u003eIn a cross-coupling reaction, there typically exist two key components: an electrophilic coupling partner, often represented by an organic halide, and a nucleophilic coupling partner, typically an organometallic species. Therefore, within our dataset, we have categorized these reactants into two distinct classes: aryl halides fall under the coupling partner class, while nucleophilic coupling partners are classified as substrate class. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates the distribution of counts for various substrate and coupling partner classes in the dataset. Our dataset was meticulously curated to encompass comprehensive information for each reaction, including the equivalence of substrate, coupling partner, ligand, and reagent, as well as details such as temperature, time, mechanism, and reaction conditions. This holistic approach ensures that all pertinent information is considered when attempting to yield predictions for a given reaction.\u003c/p\u003e\n\u003ch3\u003eMolecular Feature Engineering\u003c/h3\u003e\n\u003cp\u003eTo render the pre-processed dataset into machine-readable format, we employed four distinct methods: RDkit fingerprints\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e, Density Functional Theory (DFT), DRFP (Differential Reaction Fingerprint)\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e and RXNFP\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. These methods generate hash representations of molecules based on their SMILES notations, facilitating the model's ability to distinguish and categorize various aspects of chemical reactions. RDkit fingerprints were generated using the RDkit library. In the case of DRFP, it accepts a reaction SMILES as input and constructs a binary fingerprint derived from the symmetric difference of circular molecular n-grams generated from molecules located to the left and right of the reaction arrow, without differentiation between reactants and reagents. Additionally, one-hot encodings from scikit-learn\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e were employed to convert specific components of the reaction, such as ligands, catalysts, and reagents. physical reaction conditions, such as temperature, time, and equivalent masses, were directly fed to the model as numerical arrays without undergoing further feature engineering.\u003c/p\u003e \u003cp\u003eThe third approach for featurisation method is the DFT method. In our study, DFT calculations was performed to optimize the geometries of the reactants, products and reagents involved in the cross-coupling reactions, using Gaussian 09\u003csup\u003e35\u003c/sup\u003e employing B3LYP\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e functional. To improve the accuracy of the B3LYP functional, D3 version of Grimme's dispersion correction was added\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. 6\u0026ndash;31\u0026thinsp;+\u0026thinsp;G(d) basis sets were used to describe C, H, N, Cl, S, F, B and O atoms. The effective core potentials of Hay and Wadt with a double ζ -basis set (LanL2DZ)\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e were used to describe Cu, Co, Zn, Fe, Mn, I, Br, atoms. Vibrational frequency calculations were performed to identify the stationary points as minima (zero imaginary frequency) or saddle points (one imaginary frequency) and to obtain thermal corrections. Then, using the Auto-QChem framework\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e, we have generated DFT-based molecular descriptors tailored for each reaction under consideration.\u003c/p\u003e\n\u003ch3\u003eModel Development and Hyperparameter Optimization\u003c/h3\u003e\n\u003cp\u003eIn our pursuit to establish connections between reaction characteristics and the resulting yields, we explored both classification and regression type models for our dataset. For regression tasks, we utilized the random forest algorithm\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e from scikit-learn and a neural network approach with PyTorch. For classification, we used, the K-Nearest Neighbours (KNN) Algorithm, a supervised learning classifier that relies on proximity to categorize or predict the grouping of individual data points. To find the best hyperparameters for each model, we carried out an exhaustive search using GridSearch CV.\u003c/p\u003e \u003cp\u003eDuring the development of these models, we experimented with different ways of splitting the dataset, including both stratified and non-stratified approaches. Upon partitioning the dataset using stratified approaches, the chosen parameters are the substrate class, coupling partner class, and the category of the reactions which we labelled as \u0026ldquo;mechanism\u0026rdquo; in our dataset. The division between the training and test sets was executed using both 7:3 and 8:2 ratios. Importantly, we automated the entire process, covering everything from feature engineering to model evaluation, using Python programming, which makes our model more user friendly.\u003c/p\u003e"},{"header":"Results and Discussion","content":"\u003cp\u003eIn the present study, we have used both regression and classification models to predict the reaction yield for the 3d-transition metal-catalyzed cross-coupling reactions. Our dataset comprises information gathered from 22 research articles, all of which centre around coupling reactions utilizing first-row transition metals like Mn, Fe, Co, Cu, and Zn as catalysts. The datasets included reactions with a diverse range of experimental yields, spanning from negligible or zero to those surpassing 80%, and even slightly below the average level. This wide spectrum of yield outcomes indicates the dataset\u0026apos;s diverse nature. However, within our dataset, certain reaction components are less represented in terms of frequency. While this expands the chemical space covered, it also has the potential to diminish prediction accuracy. To address this issue, we partitioned our primary dataset into multiple subsets based on criteria related to the frequency of various reaction components. This segmentation aimed to narrow down the chemical space covered by the dataset. A significant aspect of our model is the creation of an automated and openly accessible reaction model incorporating both regression and classification methodologies. Thereby, our model is designed to be user-friendly, allowing non-experts to input reaction components as datasets for yield predictions. The following sections describe the detailed results obtained from our studies based on regression and classification models.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRegression Model:\u003c/strong\u003e\u003cbr\u003eIn our dataset consisting of 963 reactions, we attained the highest coefficient of correlation R\u003csup\u003e2\u003c/sup\u003e of 0.46, as depicted in \u003cstrong\u003eFigure 4\u003c/strong\u003e. This was realized through the application of a random forest model, with each reaction being featured using the DRFP method. To ensure the inclusion of all types of reactants and products in both the training and testing sets, we classified each reaction based on its substrate, coupling partner, and mechanism or reaction type. This classification played a crucial role in partitioning the data into train-test sets, guaranteeing a balanced distribution of reaction classes across both splits. Remarkably, the most optimal results were obtained when employing a stratified split with the substrate class serving as the stratification criterion.\u003c/p\u003e\n\u003cp\u003eIn addition to utilizing the DRFP featurization method, we have investigated various other featurization methods in combination with Random Forest model, in which noteworthy correlations have been observed. Specifically, the utilization of DFT yielded a correlation of 0.45 (\u003cstrong\u003eFigure 5d\u003c/strong\u003e), demonstrating performance similar to that of the DRFP method. Additionally, employing RDkitFP resulted in a correlation of 0.43, while the RXNFP featurization method produced a correlation of 0.36. Notably, each of these approaches underwent a test-train split with stratification based on mechanism or reaction type, leading to improved outcomes.\u003c/p\u003e\n\u003cp\u003eFurthermore, the tuning of Random Forest hyperparameters was undertaken in our study. we observed that both the DRFP method, utilizing substrate class for stratified sampling, and the RDkitFP method,\u0026nbsp;using coupling partner class for stratified sampling, exhibited a correlation coefficient (R\u0026sup2;) of 0.43 (\u003cstrong\u003eFigure 5b \u0026amp; 5e\u003c/strong\u003e).Top of Form In instances where the stratification was based on substrate class and the DFT method was employed, a correlation of 0.41 was observed. When train-test split was conducted based on the coupling partner class, the Neural Network model exhibited a correlation of 0.28 for the DRFP method. All these regression results are shown in \u003cstrong\u003eFigure 5\u003c/strong\u003e and all other results are provided in \u003cstrong\u003eTable S1\u003c/strong\u003e (see supporting information).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1:\u0026nbsp;\u003c/strong\u003eHigh correlation results of subsets, where HPT stands for Hyper Parameter Tuning,\u0026nbsp;RMSE (Root Mean Square Error), MAE (mean absolute error).\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"601\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003eDataset\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel Type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003e\u003cstrong\u003eFeaturisation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest Size\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003eIterations\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e\u003cstrong\u003eRMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e\u003cstrong\u003eMAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e\u003cstrong\u003eCorrelation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\" rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003eSubset\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eProduct\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest HPT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDRFP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e27.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e21.69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.30\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.48747591522158%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.809248554913296%\"\u003e\n \u003cp\u003eDRFP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"8.477842003853565%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.992292870905588%\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.946050096339114%\"\u003e\n \u003cp\u003e28.72\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.789980732177264%\"\u003e\n \u003cp\u003e23.13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"18.497109826589597%\"\u003e\n \u003cp\u003e0.30\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\" rowspan=\"2\"\u003e\n \u003cp\u003e\u003cstrong\u003eSubset\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eLigand\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDFT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e23.80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e18.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.45\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.48747591522158%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.809248554913296%\"\u003e\n \u003cp\u003eRDkitFP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"8.477842003853565%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.992292870905588%\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.946050096339114%\"\u003e\n \u003cp\u003e24.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.789980732177264%\"\u003e\n \u003cp\u003e19.37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"18.497109826589597%\"\u003e\n \u003cp\u003e0.45\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003eSubset\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eReagent\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest HPT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDFT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e24.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e18.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.41\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003esubset\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003ecatalyst precursor\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDRFP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e23.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e18.37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.41\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003eSubset Coupling partner\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest HPT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDFT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e25.45\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e20.50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.39\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003eSubset\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eSolvent\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest HPT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDFT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e25.48\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e19.65\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.37\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e\u003cstrong\u003eSubset Substrate\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"11.627906976744185%\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eforest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.940199335548172%\"\u003e\n \u003cp\u003eDRFP\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"7.308970099667774%\"\u003e\n \u003cp\u003e0.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.787375415282392%\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"10.299003322259136%\"\u003e\n \u003cp\u003e23.98\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"9.30232558139535%\"\u003e\n \u003cp\u003e18.57\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"15.946843853820598%\"\u003e\n \u003cp\u003e0.48\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eA distinct subset classification is established by directing attention toward specific components, including substrates, solvents, ligands, and similar factors. In doing so, reactions employing a particular element, such as a solvent, only once or twice are systematically excluded. This curation results in a refined dataset comprising only those solvents that are recurrently utilized. This discernment underscores the importance of the frequency of unique values within a dataset, indicating that reactions featuring more frequently occurring values often yield favourable outcomes compared to the broader original dataset, despite potentially containing fewer individual reactions. Remarkably, when we generated a subset by selecting \u0026ldquo;substrate\u0026rdquo; as the component with a higher frequency, we observed nearly higher results compared to our entire dataset in regression analyses. Specifically, we found a correlation of 0.48 when implementing no particular train-test split stratification and utilizing a Random Forest model in conjunction with DRFP. \u0026nbsp;Another specific subset, obtained by focusing on the \u0026quot;Ligand\u0026quot; column, exhibits a correlation of 0.45 when evaluated in conjunction with Random Forest and the inclusion of both DRFP and DFT featurisation methods and also shows a better classification result. It\u0026apos;s important to emphasize that these subsets were assessed without applying any stratification. Notably, it becomes evident that incorporating specific stratification techniques may significantly enhance the correlation of our regression model. Additional Outcomes of subsets are detailed in \u003cstrong\u003eTable 1\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eTo further evaluate the performance of our model, we employed diverse dataset segmentation strategies, within the selected reactions. We created distinct subgroups, each comprising reactions in which the components exhibited varying count thresholds, including counts more than 1, 2, 3, 5, and 7. In the dataset where each component count was greater than 5, we observed a correlation of 0.38 with the Random Forest model utilizing the DRFP featurization method. Other higher correlation results are provided in supplementary information (\u003cstrong\u003eTable S3)\u003c/strong\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eValidation of Model Performance on Out-of-Sample Data:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo validate the performance of our model, we conducted an assessment of its predictive capabilities on out-of-sample data. We accomplished this by excluding reactions that have used substrate from the class of aryl acetylene, training the model on the remaining dataset, and subsequently evaluating its predictive performance. We achieved a correlation of 0.25 between the predicted and actual yields, demonstrating the robustness of our model when it comes to out-of-sample predictions. This finding underscores the model\u0026apos;s capacity for accurate predictions beyond the training data.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHowever, when\u0026nbsp;we made out-of sample prediction by using a transition metal that was not known to training set, we have got poor accuracy. To illustrate this effect, we gathered data specifically for Sonogashira cross-coupling reactions employing Nickel as the catalyst and then train the model using a dataset that included other first-row transition metals and then applied it to predict 37 reactions with Nickel as the catalyst, the outcome was a negative R\u003csup\u003e2\u003c/sup\u003e value (R\u003csup\u003e2\u003c/sup\u003e =-2.94). Henceforth, it is notable that the transition metal choice has a significant impact on the model\u0026apos;s performance in out-of-sample prediction.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClassification Results\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn our partitioning of training and testing sets, the implementation of the coupling partner class as a stratification criterion, along with the utilization of Density Functional Theory (DFT) as the featurization approach, yielded a classifier accuracy of 0.48 and a precision of 0.43 for K-nearest neighbors (KNN) classification\u003csup\u003e29\u003c/sup\u003e hyperparameter tuning method as illustrated in \u003cstrong\u003e(Figure 6c)\u003c/strong\u003e. However, when KNN Classification with DRFP featurization method,\u003cem\u003e\u0026nbsp;\u003c/em\u003einvolving no \u003cem\u003es\u003c/em\u003etratification was considered, we got an accuracy about 0.46 \u003cstrong\u003e(\u003c/strong\u003e\u003cstrong\u003eFigure 6d)\u003c/strong\u003e\u003cstrong\u003e.\u003c/strong\u003e Furthermore, employing the same classifier and utilizing the DRFP featurization method without the incorporation of a stratification method resulted in an improved accuracy of 0.54 and a precision of 0.53, as depicted in \u003cstrong\u003eFigure 6a\u003c/strong\u003e. Thus, the performance of our classification model signifies overall effectiveness for analyzing a literature-mined small dataset encompassing various cross-coupling reactions. \u0026nbsp;A comparable accuracy (0.51) was obtained when KNN Classification using hyperparameter tuning with DFT as featurization method, and a \u003cem\u003es\u003c/em\u003etratification criterion employing\u003cem\u003e\u0026nbsp;\u003c/em\u003ecoupling partner class \u003cstrong\u003e(\u003c/strong\u003e\u003cstrong\u003eFigure 6b)\u003c/strong\u003e. For detailed classification results of each model with different featurization methods. (Refer to supplementary information, \u003cstrong\u003eTable S5\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2\u003c/strong\u003e provides additional details on the efficacy of the classification approach on our small data set. It was found that when classification is applied to the subset classes, we have got better results for the subset \u0026ldquo;ligand\u0026rdquo;. The classifier demonstrated an accuracy of 0.56 after applying hyperparameter tuning in the case of KNN classification with DRFP for the subset class solvent. We again observed that a subset \u0026ldquo;Solvent\u0026rdquo; also demonstrated an accuracy level almost similar to that of the entire dataset (0.51).\u003c/p\u003e\n\u003cp\u003eThese indicate the robust performance of our classification model. Notably, this assessment was conducted without the implementation of any specific stratification technique. Our findings underscore the significance of tailoring classification models to the specific attributes of the dataset which indicates that better predictive performance can be achieved by careful feature selection and algorithm customization. In another subset of our dataset, where each component count exceeded one, and eliminating all others where the components appear only once, we observed an accuracy of 0.49 and a precision of 0.51 when employing the K-nearest neighbors (KNN) classification model with the RDKitFP featurization method. More results for subsets characterized by higher correlation are presented in \u003cstrong\u003eTable S6 (see supporting information).\u003c/strong\u003e All these results indicate the efficacy of our predictive model in a small-dataset encompassing a wide chemical space, which are sparsely reported in literature.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe results obtained from the present study illustrate the efficacy of our regression and classification models in predicting the yield of transition metal-catalyzed cross-coupling reactions. For our entire dataset comprising of 963 distinct reactions, we got a regression value of 0.46. Our model has exhibited strong performance despite the limited size of the dataset. Notably, when attempting to predict cross-coupling reactions involving Nickel as a catalyst, a metal not present in our dataset, the correlation was\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eHigher Classification results of each subset are given, where HPT stands for Hyper Parameter Tuning, .\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eModel Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFeaturisation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTest Size\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eIterations\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ePrecession\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eProduct\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.47\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eLigand\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eReagent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDFT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.50\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eCatalyst\u003c/p\u003e \u003cp\u003eprecursor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.52\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eCoupling\u003c/p\u003e \u003cp\u003epartner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.47\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.47\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eSolvent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e \u003cp\u003eSubset\u003c/p\u003e \u003cp\u003eSubstrate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDRFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRDkitFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.47\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRDkitFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.44\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003cp\u003eclassification HPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRxnFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAccuracy= (Number of correct predictions)/(Total predictions made)\u003c/p\u003e \u003cp\u003elower (R\u003csup\u003e2\u003c/sup\u003e =-2.94), highlighting the importance of the specific transition metal used in the reactions. These findings suggest that our predictive method is well-suited for this type of analysis, with the caveat that the choice of transition metal in the reactions has a significant impact on prediction accuracy. Thus training of the data set with the different types of transition metal could increase the predictive capability of the model.\u003c/p\u003e \u003cp\u003eOur study emphasizes the importance of extracting data from research articles, highlighting the potential for valuable insights using machine-readable data in predicting yields for a range of chemical reactions. The open-access database that we created for this work focussing on transition metal catalyzed cross-coupling reactions will serve as a database for further machine learning applications. The predictive performance of our machine learning-based yield prediction model, trained on a dataset comprising 963 reactions involving diverse first-row transition metals obtained through literature mining, demonstrate notable superiority compared to existing research findings. Upto our knowledge, it\u0026rsquo;s been the first time investigated on several transition metal-catalyzed reactions while the one reported on the highly heterogeneous USPTO dataset showed a correlation that is less than 0.2.\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e The recently reported literature-mined dataset focusing on borylation reactions with Co and Ni catalysts displayed a correlation of about 0.35 specifically for the Co catalyst and 0.27 correlation for Ni catalyst using regression methodologies.\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e This demonstrates the effectiveness of our machine learning model over the existing models when working with a literature-mined small dataset containing 963 reactions comprising diverse transition metals.\u003c/p\u003e \u003cp\u003eOur study also shows the significance of acquiring data from published research to enhance the model's predictive capabilities, especially in the context of yield prediction. Although our regression and classification models have showcased their effectiveness in predicting the yields of cross-coupling reactions, yet there is room for improvement by expanding the dataset. Furthermore, the utilization of advanced models and feature engineering methods, such as trying other descriptors for Density Functional Theory (DFT), holds the promise of enhancing the model's performance, offering exciting prospects for future research endeavours.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn the present study, we have developed a comprehensive machine-learning model for yield prediction in transition metal-catalyzed cross-coupling reactions involving diverse first-row transition metals. The development of this model involved meticulous literature mining, resulting in a dataset of 936 reactions covering a wide spectrum of yields. The Random Forest regression model exhibited a significant correlation (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.46) within the diverse chemical space of the small dataset. Additionally, the KNN classification model, with hyperparameter tuning, achieved an accuracy of 0.54, indicating effectiveness in predicting favorable yields. Subset analyses further revealed promising results, with a correlation of 0.48 in regression method for reactions excluding those reactions in which substrates occurring once or twice. Similarly, in classification method, a subset focusing on ligands demonstrated an increased accuracy of 0.56. These findings indicate the accuracy of model for small-data sets. This further suggest the potential for performance enhancement with the inclusion of more data within the studied dataset, paving the way for future research. The commendable performance of our predictive model signifies a step toward automating traditional experimental trial-and-error methods, offering substantial resource and time savings. The automated and open accessible framework of our ML model by integrating regression and classification methodologies, made it accessible to non-expert users who can input reaction components as datasets to predict yields. By streamlining research processes, our ML model holds promise for advancing the arena of sustainable organometallic catalysis through more efficient and data-driven practices.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDetails regarding our dataset are made publicly available and can be found in the repository:\u003c/p\u003e\n\u003cp\u003ehttps://github.com/VITresearchgroup2024/ML_yield_prediction-/blob/main/DATA/Dataset.csv\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode Availability:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAdditional details regarding our methodology and code implementation can be found in the following repository: \u0026nbsp;https://github.com/VITresearchgroup2024/ML_yield_prediction-\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupporting Information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCoordinates of optimized geometries of reactants, products, and transition states are given as supporting information.\u003c/p\u003e\n\u003cp\u003eConflicts of interest\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests\u003c/p\u003e\n\u003cp\u003eAcknowledgements\u003c/p\u003e\n\u003cp\u003eC. Rajalakshmi thank the University Grants Commission (UGC, India) for Senior Research Fellowship (SRF).\u003c/p\u003e\n\u003cp\u003eAuthors also thank the HPC PARAMASTRA, Mahatma Gandhi University Innovation Foundation (MGUIF), Mahatma Gandhi University Campus, Kottayam, for providing the High-Performance Computing facilities (https://hpcparamastra.mguif.com/).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eC. Rajalakshmi: Conceptualization, Investigation, Formal Analysis, Resources, Data Curation, Validation, writing \u0026ndash; Original draft, Methodology, Visualization.Vivek Vijay: Investigation, Writing, Methodology, Validation, Visualization, Data Curation Abhirami Vijayakumar : Investigation, Writing, Methodology, Validation, Visualization, Data Curation Parvathi Santhoshkumar: Validation, Visualization, Methodology, Data CurationG. Krishnaveni: Validation, Visualization, Data Curation. John B Kottooran: Validation, Data Curation, Visualization.Ann Miriam Abraham: Validation, Data Curation, Visualization C S Anjanakutty: Validation, Data Curation, Visualization. Binuja Varghese: Data CurationVibin Ipe Thomas: Conceptualization, Methodology, Resources, Software, Investigation, Supervision, Formal Analysis, Visualization, Project Administration, Validation.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eP\u0026eacute;rez Sestelo, J. \u0026amp; Sarandeses, L. A. Advances in Cross-Coupling Reactions. \u003cem\u003eMolecules\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 4500 (2020).\u003c/li\u003e\n\u003cli\u003eHan, F. S. Transition-metal-catalyzed Suzuki\u0026ndash;Miyaura cross-coupling reactions: a remarkable advance from palladium to nickel catalysts. \u003cem\u003eChem. Soc. Rev.\u003c/em\u003e \u003cstrong\u003e42\u003c/strong\u003e, 5270\u0026ndash;5298 (2013).\u003c/li\u003e\n\u003cli\u003ePenn, L. \u0026amp; Gelman, D. Copper-Mediated Cross-Coupling Reactions. in \u003cem\u003ePATAI\u0026rsquo;S Chemistry of Functional Groups\u003c/em\u003e (John Wiley \u0026amp; Sons, Ltd, 2011). doi:10.1002/9780470682531.pat0451.\u003c/li\u003e\n\u003cli\u003eAyogu, J. I. \u0026amp; Onoabedje, E. A. Recent advances in transition metal-catalysed cross-coupling of (hetero)aryl halides and analogues under ligand-free conditions. \u003cem\u003eCatal. Sci. Technol.\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 5233\u0026ndash;5255 (2019).\u003c/li\u003e\n\u003cli\u003eLled\u0026oacute;s, A. Computational Organometallic Catalysis: Where We Are, Where We Are Going. \u003cem\u003eEur. J. Inorg. Chem.\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, 2547\u0026ndash;2555 (2021).\u003c/li\u003e\n\u003cli\u003eMeyer, B., Sawatlon, B., Heinen, S., von Lilienfeld, O. A. \u0026amp; Corminboeuf, C. Machine learning meets volcano plots: computational discovery of cross-coupling catalysts. \u003cem\u003eChem. Sci.\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 7069\u0026ndash;7077 (2018).\u003c/li\u003e\n\u003cli\u003eStocker, S., Cs\u0026aacute;nyi, G., Reuter, K. \u0026amp; Margraf, J. T. Machine learning in chemical reaction space. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 5505 (2020).\u003c/li\u003e\n\u003cli\u003eMeuwly, M. Machine Learning for Chemical Reactions. \u003cem\u003eChem. Rev.\u003c/em\u003e \u003cstrong\u003e121\u003c/strong\u003e, 10218\u0026ndash;10239 (2021).\u003c/li\u003e\n\u003cli\u003eStevens, J. M. \u003cem\u003eet al.\u003c/em\u003e Advancing Base Metal Catalysis through Data Science: Insight and Predictive Models for Ni-Catalyzed Borylation through Supervised Machine Learning. \u003cem\u003eOrganometallics\u003c/em\u003e \u003cstrong\u003e41\u003c/strong\u003e, 1847\u0026ndash;1864 (2022).\u003c/li\u003e\n\u003cli\u003eHueffel, J. A. \u003cem\u003eet al.\u003c/em\u003e Accelerated dinuclear palladium catalyst identification through unsupervised machine learning. \u003cem\u003eScience (80-. ).\u003c/em\u003e \u003cstrong\u003e374\u003c/strong\u003e, 1134\u0026ndash;1140 (2021).\u003c/li\u003e\n\u003cli\u003eŻurański, A. M., Martinez Alvarado, J. I., Shields, B. J. \u0026amp; Doyle, A. G. Predicting Reaction Yields via Supervised Learning. \u003cem\u003eAcc. Chem. Res.\u003c/em\u003e \u003cstrong\u003e54\u003c/strong\u003e, 1856\u0026ndash;1865 (2021).\u003c/li\u003e\n\u003cli\u003eKov\u0026aacute;cs, D. P., McCorkindale, W. \u0026amp; Lee, A. A. Quantitative interpretation explains machine learning models for chemical reaction prediction and uncovers bias. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 1695 (2021).\u003c/li\u003e\n\u003cli\u003eBaldi, P. Call for a Public Open Database of All Chemical Reactions. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e \u003cstrong\u003e62\u003c/strong\u003e, 2011\u0026ndash;2014 (2022).\u003c/li\u003e\n\u003cli\u003eMutton, T. \u0026amp; Ridley, D. D. Understanding Similarities and Differences between Two Prominent Web-Based Chemical Information and Data Retrieval Tools: Comments on Searches for Research Topics, Substances, and Reactions. \u003cem\u003eJ. Chem. Educ.\u003c/em\u003e \u003cstrong\u003e96\u003c/strong\u003e, 2167\u0026ndash;2179 (2019).\u003c/li\u003e\n\u003cli\u003eSchwaller, P., Vaucher, A. C., Laino, T. \u0026amp; Reymond, J. L. Prediction of chemical reaction yields using deep learning. \u003cem\u003eMach. Learn. Sci. Technol.\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, (2021).\u003c/li\u003e\n\u003cli\u003eAhneman, D. T., Estrada, J. G., Lin, S., Dreher, S. D. \u0026amp; Doyle, A. G. Predicting reaction performance in C\u0026ndash;N cross-coupling using machine learning. \u003cem\u003eScience (80-. ).\u003c/em\u003e \u003cstrong\u003e360\u003c/strong\u003e, 186\u0026ndash;190 (2018).\u003c/li\u003e\n\u003cli\u003ePerera, D. \u003cem\u003eet al.\u003c/em\u003e A platform for automated nanomole-scale reaction screening and micromole-scale synthesis in flow. \u003cem\u003eScience (80-. ).\u003c/em\u003e \u003cstrong\u003e359\u003c/strong\u003e, 429\u0026ndash;434 (2018).\u003c/li\u003e\n\u003cli\u003eSchleinitz, J. \u003cem\u003eet al.\u003c/em\u003e Machine Learning Yield Prediction from NiCOlit, a Small-Size Literature Data Set of Nickel Catalyzed C-O Couplings. \u003cem\u003eJ. Am. Chem. Soc.\u003c/em\u003e \u003cstrong\u003e144\u003c/strong\u003e, 14722\u0026ndash;14730 (2022).\u003c/li\u003e\n\u003cli\u003ePereira, A. \u0026amp; Trofymchuk, O. S. Machine Learning Prediction of High-Yield Cobalt- and Nickel-Catalyzed Borylations. \u003cem\u003eJ. Phys. Chem. C\u003c/em\u003e \u003cstrong\u003e127\u003c/strong\u003e, 12983\u0026ndash;12994 (2023).\u003c/li\u003e\n\u003cli\u003eCampeau, L. C. \u0026amp; Hazari, N. Cross-Coupling and Related Reactions: Connecting Past Success to the Development of New Reactions for the Future. \u003cem\u003eOrganometallics\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, 3 (2019).\u003c/li\u003e\n\u003cli\u003ePereira, A. \u0026amp; Trofymchuk, O. S. Machine Learning Prediction of High-Yield Cobalt- and Nickel- Catalyzed Borylations. (2023).\u003c/li\u003e\n\u003cli\u003eThomas, A. M., Sherin, D. R., Asha, S., Manojkumar, T. K. \u0026amp; Anilkumar, G. Exploration of the mechanism and scope of the CuI/DABCO catalysed C\u0026ndash;S coupling reaction. \u003cem\u003ePolyhedron\u003c/em\u003e \u003cstrong\u003e176\u003c/strong\u003e, 114269 (2020).\u003c/li\u003e\n\u003cli\u003eRohit, K. R., Saranya, S., Harry, N. A. \u0026amp; Anilkumar, G. A Novel Ligand-free Manganese-catalyzed C-O Coupling Protocol for the Synthesis of Biaryl Ethers. \u003cem\u003eChemistrySelect\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 5150\u0026ndash;5154 (2019).\u003c/li\u003e\n\u003cli\u003eAsha, S., Thomas, A. M., Ujwaldev, S. M. \u0026amp; Anilkumar, G. A Novel Protocol for the Cu-Catalyzed Sonogashira Coupling Reaction between Aryl Halides and Terminal Alkynes using trans-1,2-Diaminocyclohexane Ligand. \u003cem\u003eChemistrySelect\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 3938\u0026ndash;3941 (2016).\u003c/li\u003e\n\u003cli\u003eAsha, S. \u003cem\u003eet al.\u003c/em\u003e A convenient route to 1,3-diynes using ligand-free Cadiot-Chodkiewicz coupling reaction at room temperature under aerobic conditions. \u003cem\u003eSynth. Commun.\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, 256\u0026ndash;265 (2019).\u003c/li\u003e\n\u003cli\u003eKrishnan, K. K., Ujwaldev, S. M., Thankachan, A. P., Harry, N. A. \u0026amp; Gopinathan, A. A novel Zinc-catalyzed Cadiot-Chodkiewicz cross-coupling reaction of terminal alkynes with 1-bromoalkynes in ethanol solvent. \u003cem\u003eMol. Catal.\u003c/em\u003e \u003cstrong\u003e440\u003c/strong\u003e, 140\u0026ndash;147 (2017).\u003c/li\u003e\n\u003cli\u003eKrishnan, K. K. \u003cem\u003eet al.\u003c/em\u003e Zinc-Catalyzed Etherification Reaction of Aryl Iodides with Phenols. \u003cem\u003eChemistrySelect\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 3984\u0026ndash;3988 (2018).\u003c/li\u003e\n\u003cli\u003eSindhu, K. S. \u003cem\u003eet al.\u003c/em\u003e A green approach for arylation of phenols using iron catalysis in water under aerobic conditions. \u003cem\u003eJ. Catal.\u003c/em\u003e \u003cstrong\u003e348\u003c/strong\u003e, 146\u0026ndash;150 (2017).\u003c/li\u003e\n\u003cli\u003eThankachan, A. P., Sindhu, K. S., Krishnan, K. K. \u0026amp; Anilkumar, G. A novel and efficient zinc-catalyzed thioetherification of aryl halides. \u003cem\u003eRSC Adv.\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 32675\u0026ndash;32678 (2015).\u003c/li\u003e\n\u003cli\u003eThankachan, A. P., Sindhu, K. S., Krishnan, K. K. \u0026amp; Anilkumar, G. An efficient zinc-catalyzed cross-coupling reaction of aryl iodides with terminal aromatic alkynes. \u003cem\u003eTetrahedron Lett.\u003c/em\u003e \u003cstrong\u003e56\u003c/strong\u003e, 5525\u0026ndash;5528 (2015).\u003c/li\u003e\n\u003cli\u003eLovrić, M., Molero, J. M. \u0026amp; Kern, R. PySpark and RDKit: Moving towards Big Data in Cheminformatics. \u003cem\u003eMol. Inform.\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, (2019).\u003c/li\u003e\n\u003cli\u003eProbst, D., Schwaller, P. \u0026amp; Reymond, J. L. Reaction classification and yield prediction using the differential reaction fingerprint DRFP. \u003cem\u003eDigit. Discov.\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 91\u0026ndash;97 (2022).\u003c/li\u003e\n\u003cli\u003eSchwaller, P. \u003cem\u003eet al.\u003c/em\u003e Mapping the space of chemical reactions using attention-based neural networks. \u003cem\u003eNat. Mach. Intell.\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 144\u0026ndash;152 (2021).\u003c/li\u003e\n\u003cli\u003eVaroquaux, G. \u003cem\u003eet al.\u003c/em\u003e Scikit-learn. \u003cem\u003eGetMobile Mob. Comput. Commun.\u003c/em\u003e \u003cstrong\u003e19\u003c/strong\u003e, 29\u0026ndash;33 (2015).\u003c/li\u003e\n\u003cli\u003eM. J. Frisch, G. W. Trucks, H. B. Schlegel, G. E. Scuseria, M. A. Robb, J. R. Cheeseman, G. Scalmani, V. Barone, G. A. Petersson, H. Nakatsuji, X. Li, M. Caricato, A. Marenich, J. Bloino, B. G. Janesko, R. Gomperts, B. Mennucci, H. P. Hratchian, J. V. Ort, and D. J. F. Gaussian 09, Revision D.01. at (2016).\u003c/li\u003e\n\u003cli\u003eBecke, A. B3LYP. \u003cem\u003eJ. Chem. Phys.\u003c/em\u003e (1993).\u003c/li\u003e\n\u003cli\u003eGrimme, S., Antony, J., Ehrlich, S. \u0026amp; Krieg, H. A consistent and accurate ab initio parametrization of density functional dispersion correction (DFT-D) for the 94 elements H-Pu. \u003cem\u003eJ. Chem. Phys.\u003c/em\u003e \u003cstrong\u003e132\u003c/strong\u003e, 154104 (2010).\u003c/li\u003e\n\u003cli\u003eChiodo, S., Russo, N. \u0026amp; Sicilia, E. LANL2DZ basis sets recontracted in the framework of density functional theory. \u003cem\u003eJ. Chem. Phys.\u003c/em\u003e \u003cstrong\u003e125\u003c/strong\u003e, 104107 (2006).\u003c/li\u003e\n\u003cli\u003eŻurański, A. M., Wang, J. Y., Shields, B. J. \u0026amp; Doyle, A. G. Auto-QChem: an automated workflow for the generation and storage of DFT calculations for organic molecules. \u003cem\u003eReact. Chem. Eng.\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, 1276\u0026ndash;1284 (2022).\u003c/li\u003e\n\u003cli\u003eSchwaller, P., Vaucher, A. C., Laino, T. \u0026amp; Reymond, J.-L. Prediction of chemical reaction yields using deep learning. \u003cem\u003eMach. Learn. Sci. Technol.\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 015016 (2021).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine Learning, Cross-Coupling, First-row transition metals, Regression, Classification","lastPublishedDoi":"10.21203/rs.3.rs-4011086/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4011086/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe advent of first-row transition metal-catalyzed cross-coupling reactions has marked a significant milestone in the field of organic chemistry, primarily due to their pivotal role in facilitating the construction of carbon-carbon and carbon-heteroatom bonds. Traditionally, the determination of reaction yields has relied on experimental methods, but in recent times, the integration of efficient machine learning techniques has revolutionized this process. Developing a highly accurate predictive model for reaction yields applicable to diverse categories of cross-coupling reactions, however, remains a formidable challenge. In our study, we curated an extendable dataset encompassing a wide range of yields of cross-coupling reactions catalyzed by first-row transition metals through rigorous literature mining efforts. Using this dataset, we have developed an automated and open-access reaction model, employing both regression and classification methodologies. Our ML model could be used even by non-expert users, who can solely input the reaction components as datasets to predict the yields. We have achieved a correlation of 0.46 using the Random Forest regression approach and an accuracy of 0.54 using the K-Nearest Neighbours (KNN) classification which employs hyperparameter tuning. Considering the vast chemical space of our small dataset encompassing various transition metals catalysts and different categories of reactions, the above results are commendable. By releasing an open-access dataset comprising cross-coupling reactions catalyzed by 3d-transition metal, our study is anticipated to make a substantial contribution to the progression of predictive modeling for sustainable transition metal catalysis, thereby shaping the future landscape of synthetic chemistry.\u003c/p\u003e","manuscriptTitle":"Machine Learning-Based Yield Prediction for First-Row Transition Metal Catalyzed Cross-Coupling Reactions","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-03-25 15:56:56","doi":"10.21203/rs.3.rs-4011086/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"cf84e327-140d-45ad-b0fd-c7a9dc6b110e","owner":[],"postedDate":"March 25th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":29810888,"name":"Physical sciences/Chemistry"},{"id":29810889,"name":"Physical sciences/Chemistry/Catalysis"},{"id":29810890,"name":"Physical sciences/Chemistry/Organic chemistry"},{"id":29810891,"name":"Physical sciences/Chemistry/Theoretical chemistry"}],"tags":[],"updatedAt":"2024-04-19T07:23:05+00:00","versionOfRecord":[],"versionCreatedAt":"2024-03-25 15:56:56","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4011086","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4011086","identity":"rs-4011086","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.