Protein structure-based organic chemistry-driven ligand design from ultra-large chemical spaces | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Protein structure-based organic chemistry-driven ligand design from ultra-large chemical spaces Didier Rognan, François Sindt, Anthony SEYLLER, Merveille Eguida This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3687338/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Ultra-large chemical spaces describing several billion compounds are revolutionizing hit identification in early drug discovery. Because of their size, such chemical spaces cannot be fully enumerated and requires ad-hoc computational tools to navigate them and pick potentially interesting hits. We here propose a structure-based approach to ultra-large chemical space screening in which commercial chemical reagents are first docked to the target of interest and then directly connected according to organic chemistry and topological rules, to enumerate drug-like compounds under three-dimensional constraints of the target. When applied to bespoke chemical spaces of different sizes and chemical complexity targeting two receptors of pharmaceutical interest, the computational method was able to quickly enumerate hits that were either known ligands (or very close analogs) of targeted receptors as well as chemically novel candidates that could be experimentally confirmed by in vitro binding assays. The proposed approach is generic, can be applied to any docking algorithm and requires few computational resources to prioritize easily synthesizable hits from billion-sized chemical spaces. Health sciences/Medical research/Drug development Biological sciences/Computational biology and bioinformatics/High-throughput screening Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Identifying the first hit compounds able to target a macromolecule of interest is often achieved by screening experimentally or computationally a library of drug-like compounds, 1 thereby enabling a hit to lead follow-up using classical medicinal chemistry strategies. 2 Until recently, the commercially-available chemical space describing drug-like compounds amenable to screening has been restricted to 10-15 million compounds with a yearly growth of ca. half a million compounds. 3 On-demand compound libraries 4, 5 have completely changed this situation by proposing billions of compounds not yet available but easily synthesizable in a few steps and reproducible parallel synthesis. Early approaches to virtually screen subsets of ultra large chemical spaces led to spectacular successes, 6-9 notably unexpected high hit rates, very high potencies and fine selectivity. 10, 11 Today, ca. 70 billion compounds are accessible on-demand with fast delivery (6-8 weeks) and high-purity grade (> 95%). 12 Due to their huge size, compounds describing these ultra-large chemical spaces cannot be fully enumerated and requires dedicated computational tools for registration, storage and navigation. 13 Usually, large chemical spaces are described in a combinatorial manner from the building blocks and organic chemistry reactions required to synthesize them. 5 If ligand-based approaches are now available to efficiently query these large chemical spaces 14-16 , structure-based approaches including macromolecular target information (e.g. topology of a binding site) still need to be developed to exhaustively mine multibillion chemical spaces. Several computational methods have indeed been described for such a task 17-23 , albeit with moderate to severe restrictions. One the one hand, brute force docking of 1.4 billion compounds 23 has been successfully described with the help of costly dedicated platforms, 23, 24 but will soon reach its limits with next-to-come trillion-sized chemical spaces, 25 since full atomistic docking just scales linearly with the number of compounds to be screened. A workaround consists in the proper selection of seed fragments/scaffolds to screen a representative subset of the entire space. The seed fragment may originate from the early docking of fragment-based representative synthons, 20 X-ray diffraction screening data 21 or medicinal chemistry knowledge. 22 Once a seed fragment has been identified, scaffold-focused two-dimensional (2D) libraries, exploring the corresponding chemical space via a set of organic chemistry reactions, 26 can be enumerated, converted in three-dimensional (3D) atomic coordinates and physically docked to propose novel hits. This approach has been applied with success to a few targets 20-22, 27 but still requires hardware settings enabling docking a significant subset (a few million) of the entire chemical space. Last, fast machine learning approaches may be first trained on a set of representative ligand-annotated docking poses to simply predict docking scores, 17-19, 28, 29 and next be applied to predict docking scores for the remaining space. Even if only a small fraction of the full space (1-5%) has to be docked at the atomic level, this strategy cannot be further applied to trillion- sized chemical spaces since it would require gathering first billion of docking scores on a single target. Moreover, this approach has led to very mitigated results with respect to hit rate and hit potencies, 30 and deserves further experimental validations. Herein, we present a simple and fast computational approach (SpaceDock) avoiding the above-cited drawbacks. It first requires docking commercially available chemical reagents in the target of interest, in order to couple them according to standard organic chemistry reactions to propose multibillion compound libraries in one or two synthetic steps. When applied to two targets of pharmaceutical interest, the method was able to quickly retrieve hits that are chemically identical (or very close) to existing ligands, but also to propose chemically novel and potent ligands. Results Since the SpaceDock method heavily relies on the possibility to accurately dock chemical reagents, we first investigate the best docking prototols for the latter task by setting-up a dedicated benchmarking study. We then describe how chemical reagents are annotated by reactive groups and organic chemistry reactions, to define a chemical space of 53.5 billion synthesizable compounds. Last we present two concrete application of the SpaceDock workflow to two receptors of pharmaceutical interest. Setting up the conditions for accurate docking of chemical reagents To evaluate the feasibility of the SpaceDock approach, we first needed to set-up an archive of reference 3D structures for protein-bound chemical reagents. Since experimental data for such a dataset are missing, we fragmented in 3D space drug-like ligands from known protein-ligand X-ray structures (sc-PDB dataset) 31 using a set of 12 common organic chemistry reactions, then added the 3D atomic coordinates of the missing reactive moieties (e.g. boronic acid, halide; Supplementary Fig. 1 ) and last created on-the-fly "surrogate X-ray poses" for the corresponding reagents expected to yield the parent ligands with the above-described reactions. The final archive of 5,845 reagents was selected after appropriate filtering (Supplementary Table S1 ) and exhibit 13 chemical functions with a prevalence of reactive groups (e.g. amines, aryl halides, boronic acids) reflecting the frequent usage of simple organic chemistry reactions in drug discovery. 32 With a set of reference reagents in hand, we could next verify whether state-of-the-art docking algorithms were able to reproduce the surrogate X-ray poses. Five algorithms relying on different principles (FlexX: 33 incremental construction, GOLD: 34 genetic algorithm, PLANTS: 35 ant colony optimization, RDPSOVina: 36 random drift particle swarm optimization, Surflex: 37 surface-based molecular similarity) were used for that purpose. Since the SpaceDock strategy just need a single pair of complementary reagents to be properly docked to reconstitute a full ligand, the docking performance was measured by computing the root-mean square deviation (rmsd) of the pose found to be the closest (best pose) to that of the surrogate X-ray structure ( Fig. 1 ). All docking tools exhibit an excellent docking performance with 70-80% of chemical reagents being docked within 2 Å rmsd accuracy ( Fig. 1A ). Up to 70% of very high-quality poses (rmsd < 1 Å) could be generated by the apparently best docking/scoring scheme (GOLD docking, PLP scoring; Fig. 1A ). The observed docking accuracy is therefore independent on the chosen docking algorithm, and remains in agreement with docking benchmarks on low molecular weight fragments. 38, 39 Since rmsd is a global measure that does not take into account whether key protein-reagent interactions are verified or not, we additionally computed the similarity of protein-reagent interaction fingerprints (IFPs) 40 between docked and surrogate X-ray poses. Again, an excellent performance could be noticed using this orthogonal quality descriptor, with 75-85 % of chemical reagents for which the IFP similarity to the X-ray pose is deemed acceptable (Tc-IFP > 0.60; 40 Fig. 1B ). To ascertain that all chemical functions are equally suitable for docking, the same analysis was repeated for each of the 13 chemical groups ( Fig. 1C ) present in our library, focusing on the best docking strategy (GOLD docking, PLP scoring). Reassuringly, the docking performance appears to be relatively independent on the chemical function of the reagent ( Fig. 1C ) as well as on the target protein family ( Fig. 1D ). Defining a readily-accessible ultra-large chemical space from simple organic chemistry reactions Starting from the pioneering work of Hartenfeller et al., 26 we selected 36 robust, stereo- and regioselective organic chemistry reactions to define a chemical space of 5.5 billion compounds readily accessible in one or two synthesis steps ( Supplementary Table S2, Supplementary Fig. 2 ). Contrarily to previous similar approaches 26, 41, 42 , chemical reagents were here carefully chosen from specific SMARTS strings in a list of 145,705 commercial chemical reagents contributing to Enamine's REAL space 43 of 36 billion compounds. Moreover, possible side reactions affecting synthesis yields were minored by selecting reagents that are monofunctional for a particular chemical function (e.g. monocarboxylic acid), and lacking additional chemical functions (e.g. nucleophilic groups for an electrophilic reactant) that would decrease the reaction yield ( Supplementary Table S2 ). Altogether, 134,331 commercial reactants could be unambiguously annotated by reaction type, reactant role and reactive atoms yielding a total of 713,155 atomic tags ( Fig. 2 ). Conversion in 3D atomic coordinates provided a total of 176,824 ready-to-dock unique reagents, ionized at pH 7, including stereoisomers for reactants bearing up to two undefined chiral centers. Retrospective chemical space docking of 97 million compounds for human estrogen receptor beta agonists. For a first proof-of-concept, we selected as a target the activated form of the human estrogen receptor beta (ERβ) for the following two reasons: (i) the ligand-binding cavity is nicely druggable with a good hydrophobicity/hydrophilicity balance, (ii) the receptor has been co-crystallized with many high-affinity low molecular-weight agonists, notably compounds sharing a 2-aryl-benzoxazole scaffold 44 whose one-step synthesis from 2-aminophenols and benzaldehydes is one of the 36 reactions that we have encoded. To avoid a possible chemotype bias, we selected an X-ray receptor structure co-crystallized with genistein (PDB 1QKM), a non-benzoxazole high-affinity agonist used from hereon as "reference ligand" ( Fig. 3A ) and asked whether we could recover a "ground truth" benzoxazole agonist (WAY-338, Fig. 3A ) or any close analog, by first docking the necessary reactants (2-aminophenols, benzaldehydes) and then enabling the benzoxazole ring formation within the protein binding site. To this end, 145 commercial 2-aminophenols and 3,874 benzaldehydes were generated in 3D and docked into the 1QKM structure, in order to explore a combinatorial space of 561,730 possible benzoxazoles. Since the later space is small, we additionally considered a much larger space of 97 million sulfonamide decoys synthesizable from 1,275 sulfonyl chlorides and 76,758 amines, thereby strongly minoring the benzoxazole space (0.57%) in the full chemical space to scan. After docking all reagents necessary to mine both chemical spaces according to the previously found best protocol (GOLD docking, PLP scoring), a series of filters of increasing complexity ( Table 1 ) was iteratively passed to a decreasing number of possible solutions, first starting with pairs of potentially reacting reagent poses, then with successfully enumerated ligand poses, and last with quality checked redocking poses. Table 1 | Incremental series of filters applied to prioritize SpaceDock hits Filter Type Criteria Applies to Software used 1 Geometry Distances, angles, clashes Pair of reactant poses this work 2 Interaction Interaction fingerprint similarity to reference Pair of reactant poses IChem 45 3 Energy, geometry Rmsd of refined pose to non-refined pose Fully enumerated ligand Szybki 46 Surflex-Dock 37 4 Interaction, structure Interaction fingerprint similarity (IFP) to reference Number of stereocenters Number of rotatable bonds Drug-likeness Fully enumerated ligand IChem 45 Filter 46 5 Redocking Rmsd to energy-minimized SpaceDock pose IFP similarity to energy-minimized SpaceDock pose Docking poses GOLD 34 Surflex-Dock 37 IChem 45 6 Quality check Number of strained torsions Local and global strain energy Number of unsatisfied H-bond donors and acceptors, number of unsatisfied ionic bonds Docking poses Torsion_analyzer 47 Freeform 46 this work 7 Final selection Duplicates removal HYDEscore Docking poses This work Hydescorer 48 The SpaceDock flowchart is displayed Fig. 3 . In a first step, pure chemical and topological filters ( Supplementary Figs. 3, 4 ) are passed to all docking poses of possible reactant pairs to quickly remove impossible reactions (filter #1). To stay on a safe side, we only considered pairs of bound reactants exhibiting a total interaction fingerprint (IFP) similarity 40 to the genistein X-ray pose above an acceptable threshold 40 (IFP ≥ 0.60 considering all non-bonded interactions, IFP ≥ 0.50 considering polar interactions only; filter #2). The 821,702 remaining pairs of reactants were then converted, in the protein 3D space, into the corresponding benzoxazoles and sulfonamides, respectively and the fully enumerated ligands were quickly minimized in the protein binding site. Only 539,906 poses deviated less than 1.0 Å rmsd from the non-refined poses after energy refinement (filter #3). The remaining minimized poses were filtered again according to IFP similarity of the genistein X-ray pose (IFP ≥ 0.60 considering all non-bonded interactions, IFP ≥ 0.60 considering polar interactions only; filter #4). Compounds with more than 2 stereocenters and 8 rotatable bonds were removed at this stage, leaving 49,569 poses for further processing. To ensure that the selected SpaceDock poses might be recovered by classical docking, all remaining hits were redocked to the ERβ structure, as previously done for the reagents. Only 121,470 poses close to the corresponding energy-minimized SpaceDock poses (rmsd ≤2.0 Å; IFP ≥ 0.60 considering all non-bonded interactions, IFP ≥ 0.60 considering polar interactions only) were retained (filter #5). A quality check of remaining poses (filter #6) was next applied to remove unlikely solutions (≥ 1 strained torsion, local strain energy > 4 kcal/mol, global strain energy > 8 kcal/mol, no unsatisfied ionic bond , >2 unsatisfied h-bond donors, > 4 unsatisfied h-bond acceptors). 22, 49 The number of plausible solutions (7,712) being still important, a custom filter was finally applied to keep only poses anchored at both sides of the binding pocket (H-bond either Glu305 or Arg346, and to His475), as seen for all potent ERβ agonists (recall genistein X-ray pose, Fig. 3A ). The final hit list comprises 102 poses from 64 unique ligands (filter #7), including 54 benzoxazoles and 10 sulfonamides ( Fig. 3B, Supplementary Table S3) ranked by decreasing full IFP similarity to the reference ligand, then by decreasing polar IFP similarity, and last by increasing absolute binding free energy predicted by the HYDE scoring function. 48 Despite in minority in the initial space (0.57%), it is reassuring that the ground truth chemotype was considerably enriched (84 %) in the final hit list. Inspecting the structures and binding poses of the hits, we observed that SpaceDock was indeed able to recover, among the top-ranked hits, the ground truth ligand (rank #9), a known ERβ agonist ChEMBL187673 50 (IC 50 = 50 nM, rank #25) and 52 other 2-arylbenzoxazoles, with almost perfect binding modes (rmsd =1.15 Å for the ground-truth ligand, Fig. 3C ). About half of the hits (30 out 64; all from the benzoxazole space), were considered chemically similar to existing ERβ ligands ( Supplementary Fig. 5 ), evidencing that SpaceDock can propose both known ligands (or very close analogs thereof) and new chemical entities. However, only a lower number of compounds (17, out of which 10 share the sulfonamide space) strictly intersected Enamine REAL space ( Supplementary Fig. 5 ). This observation does not preclude for their synthesizability but just illustrates that these hits, despite the commercial availability of their starting building blocks, cannot be obtained within the scope of 167 parallel synthesis protocols defining REAL space. From this preliminary proof-of-concept, it appears that the herein presented method is able to perform a complex organic chemistry reaction (ring cyclisation) from suitably posed and chemically compatible chemical reagents, under the 3D constraints of the target's structure, to generate and prioritize fully enumerated ligands for meaningful reasons. We therefore decided to apply SpaceDock to a prospective screening of a much larger chemical space. Prospective chemical space docking of 670 million compounds for human dopamine D3 receptor antagonists. We next applied the method to a much larger chemical space of 670 million carboxamides targeting the human dopamine D3 receptor (DRD3). Since the only available high-resolution DRD3 receptor structure (PDB 3PBL) has been obtained in complex with the antagonist eticlopride ( Fig. 4A ), 51 the latter orthomethoxybenzamide (OMB) ligand was used as both reference and ground-truth ligand to recover. Commercially available carboxylic acids and primary/secondary amines ( Supplementary Table 2 ) were first filtered to remove reagents that, upon amide bond formation, would lead to non-drug-like ligands ( Supplementary Table 4 ), thereby keeping 19,887 acids and 33,726 amines (in 3D coordinates) to explore a chemical space of 670 million carboxamides ( Fig. 4B ). The resulting 53,613 chemical reagents were then docked to the eticlopride-free DRD3 structure using GOLD docking and PLP scoring, as previously described. Since 20 poses were saved for each reactant, a total of 268 billion (19,887*20*33,726*20) possible reactions were passed to the SpaceDock flowchart ( Fig. 4B ), removing first impossible amide bond formation according to geometrical criteria ( Supplementary Fig.6 ) while keeping only amine poses exhibiting the crucial ionic bond to the key Asp110 residue 51 (filter #1, Fig. 4B ), then retaining pair of reactant poses for which the IFP similarity to the reference ligand is higher than 0.60 for all interactions and 0.50 for polar interactions only (filter #2). 40 A total of 24,674,693 reactions were conducted in silico to generate the corresponding carboxamides inside the receptor pocket, that were later energy-minimized. Keeping only minimized poses that did not deviate much from the initial pose (rmsd < 1.0 Å) afforded 15,120,198 plausible solutions (filter #3, Fig. 4B ). At this stage, hits bearing a cis-amide bond or more than 2 chiral centers or more than 9 rotatable bonds were removed to keep only drug-like compounds. The resulting number of hits being still very high, we pruned the hit list by keeping only minimized poses with a high full IFP similarity to the reference ligand (IFP similarity > 0.60) while exhibiting a perfect IFP similarity to eticlopride (IFP =1) with respect to polar interactions (H-bond and ionic bond to Asp110). This filter (filter #4, Fig. 4B ) yielded to 518,306 SpaceDock poses (corresponding to 500,041 unique compounds) that had to be confirmed by full atomistic docking (GOLD docking, PLP scoring, 20 poses saved) of the corresponding ligands and comparison with the minimized SpaceDock poses. Only docking poses verifying the following three criteria (rmsd ≤2.0 Å & IFP_full ≥ 0.60 & IFP_polar =1) were retained, leaving 712,120 good docking poses (filter #5, Fig. 4B ) for sanity check (no strained torsion, local strain energy ≤ 4 kcal/mol, global strain energy ≤ 8 kcal/mol, no unsatisfied ionic bond, ≤ 2 unsatisfied h-bond donors, ≤ 4 unsatisfied h-bond acceptors, filter #6, Fig. 4B ). The number of remaining poses being still important (97,096), a custom filter (not implemented by default, Table 1 ) was added to remove poses for compounds with no aromatic ring (always present in known DRD3 antagonists), 52 exhibiting a predicted absolute binding free energy (HYDEscore) lower than 30 kJ/mol and further restricting the deviation to the original SpaceDock poses (rmsd ≤1.0 Å & IFP_full ≥ 0.75). A reasonable number of 757 docking poses from 315 unique ligands (filter #7, Fig. 4B ) defined the final hit list. Compounds were ranked by decreasing full IFP similarity to the reference ligand, then by decreasing polar IFP similarity, and last by increasing HYDE binding free energy ( Supplementary Table 5 ). As for the first attempt on ERβ ligands, we first check whether the ground-truth ligand and its corresponding OMB scaffold were present in the list. Indeed, 15 OMBs including eticlopride (rank 30) were part of the list with binding poses very similar to that observed for the reference ligand (rmsd of eticlopride = 0.73 Å, Fig. 4C ). Interestingly, 300 additional hits not sharing the OBM scaffold were prioritized with poses and protein-ligand interaction patterns quite close to that seen for eticlopride ( Fig. 4D ). Most ligands were scaffold hops for which the orthomethoxybenzamide has been replaced by a bicyclic heteroaryl-amide, connected by 2-3 carbon atoms to a basic amine. By comparison to the ERβ hit list, the DRD3 hits deviate more from known ChEMBL ligands (24% considered as chemically similar) but are more easily obtainable in REAL space (53% being directly purchasable, and additional 38% being very close to REAL space compounds; Supplementary Fig. 7 ). 16 chemically diverse and representative hits were directly purchased at Enamine, out of which 15 could be synthesized in six weeks (5 mg quantity, > 90% purity) and further tested for binding to human DRD3 ( Fig. 5 ). Out of the tested 15 compounds, ten exhibited detectable binding (> 20% inhibition) to the DRD3 receptor at the single concentration of 10 µM ( Fig. 5 ). The six strongest binders (#1, #25, #66, #107, #142, #161) were selected for dose-curve responses for inhibition constants (K i ) determination ( Fig. 5 , Supplementary Fig. 8 ). Three of them (#1, #66, #142) exhibited K i values in the 300-400 nM range, the three others at 1.4-1.6 µM. The remarkable hit rates (66% at 10 µM, 20% at 500 nM) are in line with previous observations from docking ultra-large libraries, 10, 11 and suggests that SpaceDock competes rather well with much more demanding full atomistic docking when screening large chemical spaces. Interestingly, novel heteroamatic-carboxamide scaffolds were disclosed for 4 of the strong binders (#66, #107, #142 and #161) that could not be found in any of 6,714 dopamine DRD2/DRD3 ligands from ChEMBL ( Table 2 ). SpaceDock proposals should still be considered as primary hits. As such, their potency is lower than that of the closest dopamine D2/D3 antagonists from ChEMBL, albeit with a higher ligand efficiency. Conclusion We herein describe a novel computational method (SpaceDock) to exhaustively browse ultra-large chemical spaces under specific constraints of a target protein and known binders. When applied to two nicely druggable targets (estrogen receptor β, dopamine D3 receptor) and chemical spaces up to 670 million compounds, it enabled the fast recovery of known ligands/scaffolds (in both cases) and the identification of novel and potent new chemical entities (dopamine D3 receptor). SpaceDock departs from existing methods 20-22 by two major differences: (i) fully unmodified chemical reagents and not synthons (scaffolds with chemistry-informed exit vectors) are used as primary sources of hits, (ii) most promising ligands are directly obtained within the protein binding site, by 3D in silico synthesis according to geometrical and chemical cross-compatibility of previously posed reagents pairs. Indeed, direct docking of chemical reagents has, to the best of our knowledge, never been reported. Interestingly, our preliminary benchmark demonstrates that docking chemical reagents is as accurate as docking low-molecular weight fragments 38 with ca. 75% of chemicals properly posed with respect to their corresponding substructures in full PDB ligands. Noteworthy, the docking accuracy is independent on the docking tool used, as well on the reactive moiety of the reactants and on the target protein family; therefore opening the method to any druggable target and set of commercial building blocks. To enable an easy synthetic access to most SpaceDock hits, the method relies on chemical reagents contributing to Enamine's REAL space, and generate hits in the binding site 3D space using a set of 36 robust two-component organic chemistry reactions. Given the 70% average docking accuracy of reactants, we therefore expect the likehood to properly couple two chemically compatible reactants into a fully enumerated and suitably posed ligand at ca. 50%. Docking the starting chemical reagents is clearly the most time-consuming step of the entire flowchart (ca 15 s/reagent), meaning that SpaceDock scales with the number of reactants and not the number of products defining the chemical space to be screened. To optimize the speed of the further processing, a series of filters of increasing complexity is applied, step to step, to a decreasing number of plausible solutions. Just checking the relative position of compatible reactants to be paired by fast distance/angles measures permits to remove 99.8% of possible solutions. Although not mandatory, we applied IFP similarity to a reference pose to remove topologically valid ligands not fulfilling expected interactions with key residues. This filter permits to reduce the number of full ligand poses to the third most time-consuming but necessary energy-minimization step (ca. 1s/ recombined pose), and remove local strains around the newly created bonds. We assume that a SpaceDock proposal is all the more interesting if it does not vary (in terms of rmsd and IFP similarity) upon energy minimization within the protein binding site, and if it can be recovered by full atomistic docking of the corresponding ligand. Although not necessary, we recommend this redocking step to ensure that SpaceDock and any state-of-the-art docking tool (we here used GOLD but other tools may be used as well) agrees on the final poses to be sent to the very important quality check. A particular importance is given to local and global strain energies (≤4 and 8 kcal/mol, respectively), as well as to the number of unsatisfied ionic bonds (none) and of unsatisfied hydrogen-bond donors/acceptors (≤ 2 and 4, respectively). In the DRD3 test case, omitting this step drastically enriched the final hit list in false positives which could not be confirmed experimentally (data not shown). The herein proposed chemical space docking approach could yield, at least for the present case of a G protein-coupled receptor, to experimentally-validated hits with a high hit rate and nanomolar potencies that agrees with tendencies already noticed upon full atomistic docking of ultra-large library virtual screens. 10, 11 SpaceDock remains a relatively light computational procedure since browsing a chemical space of 100 million compounds can be achieved within 2 days on a 16-core Intel (R) Xeon (R) Silver 4210 processor. Mining the entire 5.5 billion chemical space has been made possible for the 4 th international CACHE challenge 54 with still limited resources (1 week on 400 cores). Preliminary attempts to scan even larger chemical spaces (e.g. by adding three-component reactions) suggests that the method can be easily applied up to a trillion compounds. Methods Setting-up a library of chemical reagents from fragmented protein-bound ligands. 37,922 ligands from the sc-PDB database of druggable protein-ligand 3D structures, 31, 55 were fragmented using a set of 12 RECAP 56 -inspired retrosynthetic rules to yield 97,024 chemical reagents ( Supplementary Fig. 1 ) with standard topologies (bond length, angle bending, torsion angles) retrieved from the TRIPOS force-field. 57 The resulting building blocks were then filtered using the following rules: (i) IChem v.5.2. 8 45 detection of at least four non-covalent interactions (one of which being a ionic bond or an hydrogen-bond) with the original sc-PDB target protein, (ii) a total number of heavy atoms between 3 and 23, (iii) a total number of rotatable bonds inferior or equal to 6, (iv) a heteroatom to carbon ratio between 0.05 and 4.5, (v) no more than two fused cycles, (vi) a number of aromatic rings inferior to 3. The final library comprised 5,845 reagents (mol2 file format) derived from 4,656 unique sc-PDB ligands. Although the building blocks have not been explicitly crystallized with their target, the corresponding poses will be further annotated as "surrogate X-ray" pose. Docking sc-PDB building reagents to their cognate targets The above described reagents were docked to the sc-PDB target originally bound to the ligand they were derived of, after randomizing their initial orientation and dihedral angles with the Surflex 37 ran_archive routine, using 5 state-of-the-art docking tools (FlexX v.5.2.0, 33 GOLD v.2022, 34 PLANTS v1.2, 35 RDPSOVina v.2.0, 36 Surflex v.4.5.4.3 37 ) with almost standard parameters ( Supplementary Tables 6-8 ). Since the boron atom is not parametrized in some docking tools, it was replaced by either a dummy atom (FlexX, GOLD, PLANTS, Surflex) or a carbon (RDPSOvina) while keeping the trigonal planar geometry of the boronic acid unchanged. Up to 20 poses were preferentially saved in mol2 file format whenever possible (GOLD, PLANTS, Surflex), in sd file format (FlexX) or in pdbqt file format (RDPSOVina). For each docking pose, the root-mean-square deviation (rmsd) of heavy atoms to the corresponding surrogate X-ray pose was computed thanks to the Surflex rms routine when comparing mol2 files, or the ADFRsuite-1.0 58 obrms routine when comparing files of different formats (mol2 vs. pdbqt, mol2 vs. sd). In addition, we measured the similarity of protein-ligand interactions between docked and X-ray poses with the IFP module of the IChem v.5.2.8 package. 45 Preparation of bespoke chemical spaces encoded by 36 robust organic chemistry reactions The global stock of commercially available building blocks (250,355 compounds, sd file format, date: 2022-12-28) was downloaded from Enamine's website 59 and filtered by catalog identification number to retain 145,707 reagents contributing to the REAL space. 43 Building blocks were then filtered to remove unsuitable entries as previously described. 41 For each of 36 different one or two-steps organic chemistry reactions ( Supplementary Table 2 ), the corresponding reactants were retrieved using SMARTS strings 41 queries in PipelinePilot v.22.1.0.2935 60 ( Supplementary Figure 9 ). In order to avoid side reactions, building blocks need to be monofunctional for the reactive group of interest and free of any possible poisoning chemical function for the reaction of interest ( Supplementary Table 2 ). For each retained building block and possible reaction, an annotation triplet is provided: (i) reaction type, reactant role, reactive atoms. The final annotation table comprises 713,155 annotation triplets for 134,331 REAL building blocks. Selected building blocks were finally ionized at their most likely ionization state at pH 7.4 using PipelinePilot and converted into 3D atomic coordinates with Corina v.3.40, 61 allowing to generate up to 4 diastereoisomers by entry, in a single ready-to-dock mol2 file format. Docking of chemical reagents to human estrogen receptor beta The X-ray structure of the human estrogen receptor beta in complex with the agonist genistein 62 was downloaded from the Protein Data Bank (PDB 1QKM). Hydrogen atoms and simultaneous optimisation of protonation states of protein, water and ligand atoms was performed with Protoss v.4.0. 63 All water molecule and genistein were removed, keeping only remaining protein atoms of chain A which were saved in mol2 file format. The commercial building blocks selected for a possible benzoxazole ring or sulfonamide bond formation (145 aminophenols and 3,874 benzaldehydes; 1,275 sulfonyl chlorides and 76,758 amines) were docked to the ERβ atomic coordinates with GOLD using previously reported parameter settings ( Supplementary Table 7 ). The cavity was detected from X-ray atomic coordinates of genistein. Up to 20 poses, scored by the PLP scoring function, were retained for each building block. Docking of chemical reagents to the human dopamine D3 receptor (DRD3) The X-ray structure of the human dopamine D3 receptor in complex with the antagonist eticlopride 51 was downloaded from the Protein Data Bank (PDB 3PBL). Hydrogen atoms and simultaneous optimisation of protonation states of protein, water and ligand atoms was performed with Protoss v.4.0. 63 The inserted T4-lysozyme sequence (Asn1002-Tyr1161), all water molecule and eticlopride were removed, keeping only remaining protein atoms of chain A which were saved in mol2 file format. The commercial building blocks were initially filtered based on their capacity to form a drug-like molecule through an amide bond formation ( Supplementary Table 4 ) and their inclusion in the pool of reagents utilized in the REAL Space. The reagents selected for a possible amide bond formation (33,726 amines and 19,887 carboxylic acids) were docked to the DRD3 atomic coordinates with GOLD using previously reported parameter settings ( Supplementary Table 7 ). The cavity was detected from X-ray atomic coordinates of eticlopride. Up to 20 poses, scored by the PLP scoring function, were retained for each building block. To decrease the number of possible recombinations, only docking poses of amines exhibiting an ionic bond to the key residue Asp110, detected on the fly with IChem, were further retained for amide bond formation. Ligand enumeration by reagents coupling Given two poses of chemically compatible reagents, a ligand is generated within the protein binding site, according to their respective location and chemical compatibility. Reagent poses are initially loaded using an in-house mol2 parser and annotated for at least one reaction based on the tag table shown in Fig. 2 . Atomic coordinates of reactive atoms carbon and their immediate neighbors, are extracted and stored for subsequent calculations. This process is repeated for each reaction, following a similar workflow. A subsequent set of filters is applied to pairs of reagent poses, including the distance between their center of mass to promptly eliminate distant pairs, the distance between connectable atoms, examination of certain angles of the future formed bond/ring to ensure a suitable geometry, and consideration of clashes (≤4 between non-reacting atoms) to prevent overlapping substituents. If a pair satisfies all the rules, a bond is created between the connectable atoms. The hybridization of reacting atoms is then updated to reflect the newly created bonds and exit atoms (to be removed after the reaction) are deleted. The fully enumerated molecule is then saved into a single mol2 file. An optional step is available at this stage. If a reference ligand exists, the molecule is initially written to a temporary mol2 file to assess its IFP similarity (default values are ≥ 0.60 for all non-bonded interactions and ≥ 0.50 for polar interactions) to the reference pose using IChem v.5.2.8. If the similarity threshold is reached, the molecule is transferred to the final mol2 file. Detailed rules of these filters can be found in Supplementary Figs. 3, 4, 6 . The fully enumerated molecule, in presence of the target protein, is last energy-minimized in Szybki v2.4.0.0, 46 using standard settings and the MMFF94 force-field. 64 Comparisons to reference ligands Interaction fingerprint similarity search between any pose (before and after energy refinement) and a reference X-ray ligand was done using standard parameters of the IFP module implemented in the IChem v.5.2.8 package. 45 Likewise, root-mean square deviations were computed with the rms routine of Surflex-Dock v.4.5.4.3. 37 Redocking of SpaceDock poses The coupling of two reagent poses, followed by protein constraint refinement (referred to as the "SpaceDock" pose), was redocked into the target protein structure using GOLD. The scoring function employed was PLP, with 20 generated poses, and the same parameter file as described in Supplementary Table 7 . To eliminate structural biases, input ligand structures were converted to SMILES format using OEChem Toolkit v.3.4.0.1 46 and further transformed into 3D structures with Corina v.3.40. 61 Up to four diastereoisomers were generated in a single mol2 file. The resulting full atomistic docking pose, exhibiting a rmsd (computed with Surflex rms) below 2 Å, all non-bonded interactions IFP similarity ≥ 0.60, and precisely the same polar IFP as the corresponding SpaceDock pose, was considered as confirmation and retained for subsequent investigations. If multiple docking poses satisfy these rules for each SpaceDock pose, all of them are retained. Quality check of redocked poses The number of torsion strains in every redocking pose was estimated with TorsionAnalyzer v.2.0.0. 47 Any pose with at least one torsion annotated as 'strained' was discarded from further analysis. Local strain (distortion of the specific conformation from the nearest local minima) and global strain (energy required to select the specific conformation from the full conformational ensemble of the corresponding compound in water) energies were then computed with standard parameter of Freeform v.2.4.0.0. 46 Any pose with local and global strain energies higher than 4 and 8 kcal/mol, respectively, were discarded. Last, remaining poses were inspected, in their protein-bound state, for counting the number of unsatisfied ionic bonds, hydrogen-bond donors and acceptors. First, protein-ligand ionic and hydrogen-bonds were registered with IChem. Any charged atom or hydrogen-bond donor/acceptor atom of the ligand (according to IChem definitions) 40 not present in the above list was annotated as "unsatisfied" atom. Unsatisfied heavy atoms being both donors and acceptors (e.g. hydroxyl oxygen atom) were only counted once. Ligand atoms participating to intra-molecular hydrogen bonds were considered as satisfied. Altogether, ligand poses with more than 2 unsatisfied donors and 4 unsatisfied acceptors were removed from the final hit list. Similarity to ChEMBL and REAL Space ligands Known ligands of the human estrogen receptor beta (CHEMBL242) and human dopamine D2 (CHEMBL217) and D3 (CHEMBL234) receptors were retrieved from the ChEMBL database (release 33) 50 as SMILES strings for ligand entries fulfilling the following criteria: K i < 1 µM, Assay_type = B). Pairwise chemical similarity between SpaceDock hits and ChEMBL ligands was computed with PipelinePilot v.22.1.0.2935 60 from ECFP4 circular fingerprints and scored by the value of the Tanimoto coefficient. Maximum common substructure (MCS) similarity of SpaceDock hits (converted from mol2 to SMILES strings, thanks to Open Babel v.3.1.0) 65 to 36 billion REAL space ligands (version REALSpace_36bn_2023-03.space 12 ) was computed with SpaceMACS v.0.9.2, 15 to save the top 15 REAL space compounds ranked by decreasing MCS-Tanimito similarity value. Declarations Data availability List of reactants to build benzoxazole, sulfonamide and amide chemical spaces, docked poses of test reactants (ERβ, DRD3 test cases), annotation table of Enamine REAL reactants, IChem configuration files for IFP filtering. All data and SpaceDock processing scripts are available at https://github.com/litfsindt/LIT-SpaceDock Code availability Filter v.4.2.1.1, Szbyki v2.5.1.1, OEChem Toolkit v.3.4.0.1; Freeform v.2.5.1.1: OpenEye Scientific, Santa Fe, N.M., USA, https://www.eyesopen.com/ FlexX v.5.2.0, Hyde v.1.5.0, SpaceMACS v.0.9.2, REAL space in fragment space format: BioSolveIT GmbH, Sankt Augustin, Germany, www.biosolveit.de GOLD v.2022: CCDC Software Ltd., Cambridge CB2 1EZ, United Kingdom, www.ccdc.cam.ac.uk Open Babel v.3.1.0, https://github.com/openbabel/openbabel PLANTS v1.2: University of Konstanz, Germany, http://www.tcd.uni-konstanz.de/research/plants.php RDPSOVina v2.0: Jiangnan University, Jiangsu, China, https://github.com/li-jin-xing/RDPSOVina SpaceDock v.1.0.0: https://github.com/litfsindt/LIT-SpaceDock Surflex-Dock v4.5.4.3: BioPharmics LLC, https://www.biopharmics.com Acknowledgements We thank Guillaume Bret (Laboratoire d'innovation thérapeutique) for technical assistance, Michael Bossert and Yurii Moroz (Enamine Ltd.) for sharing the list of REAL space reagents, and the CC-IN2P3 calculation center (Villeurbanne, France) for allocation of computing time and excellent support. Contributions D.R conceived the study. A.S. performed the initial benchmarking study on chemical reagents. M.E. designed the rules to filter sc-PDB building blocks. F.S. encoded organic chemistry reactions in 3D space, performed the whole docking of Enamine chemical reagents and wrote the SpaceDock source code to enumerate full ligands. F.S. and D.R analyzed the data. All of the authors contributed to writing and editing the manuscript. Competing interests D.R. is co-founder and shareholder of BIODOL Therapeutics. M.E. is employee of Amgen References Bleicher KH, Bohm HJ, Muller K, Alanine AI. Hit and lead generation: beyond high-throughput screening. Nat Rev Drug Discov 2 , 369-378 (2003). Hughes JP, Rees S, Kalindjian SB, Philpott KL. Principles of early drug discovery. Br J Pharmacol 162 , 1239-1249 (2011). Lucas X, Gruning BA, Bleher S, Gunther S. The purchasable chemical space: a detailed picture. J Chem Inf Model 55 , 915-924 (2015). Tingle BI , et al. ZINC-22 horizontal line A Free Multi-Billion-Scale Database of Tangible Compounds for Ligand Discovery. J Chem Inf Model 63 , 1666-1776 (2023). Grygorenko OO, Radchenko DS, Dziuba I, Chuprina A, Gubina KE, Moroz YS. Generating Multibillion Chemical Space of Readily Accessible Screening Compounds. iScience 23 , 101681 (2020). Lyu J , et al. Ultra-large library docking for discovering new chemotypes. Nature 566 , 224-229 (2019). Sadybekov AA , et al. Structure-Based Virtual Screening of Ultra-Large Library Yields Potent Antagonists for a Lipid GPCR. Biomolecules 10 , (2020). Stein RM , et al. Virtual discovery of melatonin receptor ligands to modulate circadian rhythms. Nature 579 , 609-614 (2020). Alon A , et al. Structures of the sigma2 receptor enable docking for bioactive ligand discovery. Nature 600 , 759-764 (2021). Lyu J, Irwin JJ, Shoichet BK. Modeling the expansion of virtual screening libraries. Nat Chem Biol 19 , 712-718 (2023). Sadybekov AV, Katritch V. Computational approaches streamlining drug discovery. Nature 616 , 673-685 (2023). Readily-accessible on-demand chemical spaces, https://www.biosolveit.de/infiniSee (accessed 11-16-2023) Warr WA, Nicklaus MC, Nicolaou CA, Rarey M. Exploration of Ultralarge Compound Collections for Drug Discovery. J Chem Inf Model 62 , 2021-2034 (2022). Bellmann L, Penner P, Rarey M. Topological Similarity Search in Large Combinatorial Fragment Spaces. J Chem Inf Model 61 , 238-251 (2021). Schmidt R, Klein R, Rarey M. Maximum Common Substructure Searching in Combinatorial Make-on-Demand Compound Spaces. J Chem Inf Model 62 , 2133-2150 (2022). Meyenburg C, Dolfus U, Briem H, Rarey M. Galileo: Three-dimensional searching in large combinatorial fragment spaces on the example of pharmacophores. J Comput Aided Mol Des 37 , 1-16 (2023). Gentile F , et al. Deep Docking: A Deep Learning Platform for Augmentation of Structure Based Drug Discovery. ACS Cent Sci 6 , 939-949 (2020). Berenger F, Kumar A, Zhang KYJ, Yamanishi Y. Lean-Docking: Exploiting Ligands' Predicted Docking Scores to Accelerate Molecular Docking. J Chem Inf Model 61 , 2341-2352 (2021). Graff DE, Aldeghi M, Morrone JA, Jordan KE, Pyzer-Knapp EO, Coley CW. Self-Focusing Virtual Screening with Active Design Space Pruning. J Chem Inf Model 62 , 3854-3862 (2022). Sadybekov AA , et al. Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. Nature 601 , 452-459 (2022). Muller J , et al. Magnet for the Needle in Haystack: "Crystal Structure First" Fragment Hits Unlock Active Chemical Matter Using Targeted Exploration of Vast Chemical Spaces. J Med Chem 65 , 15663-15678 (2022). Beroza P , et al. Chemical space docking enables large-scale structure-based virtual screening to discover ROCK1 kinase inhibitors. Nat Commun 13 , 6447 (2022). Gorgulla C , et al. An open-source drug discovery platform enables ultra-large virtual screens. Nature 580 , 663-668 (2020). Gadioli D , et al. EXSCALATE: An extreme-scale in-silico virtual screening platform to evaluate 1 trillion compounds in 60 hours on 81 PFLOPS supercomputers. arXiv:211011644v1 , (2021). Neumann A, Marrison L, Klein R. Relevance of the Trillion-Sized Chemical Space "eXplore" as a Source for Drug Discovery. ACS Med Chem Lett 14 , 466-472 (2023). Hartenfeller M , et al. A collection of robust organic synthesis reactions for in silico molecule design. J Chem Inf Model 51 , 3093-3098 (2011). Penner P , et al. FastGrow: on-the-fly growing and its application to DYRK1A. J Comput Aided Mol Des 36 , 639-651 (2022). Sivula T, Yetukuri L, Kalliokoski T, Kasnanen H, Poso A, Pohner I. Machine Learning-Boosted Docking Enables the Efficient Structure-Based Virtual Screening of Giga-Scale Enumerated Chemical Libraries. J Chem Inf Model 63 , 5773-5783 (2023). Roggia M, Natale B, Amendola G, Di Maro S, Cosconati S. Streamlining Large Chemical Library Docking with Artificial Intelligence: the PyRMD2Dock Approach. J Chem Inf Model , https://doi.org/10.1021/acs.jcim.1023c00647 (2023). Gentile F , et al. Automated discovery of noncovalent inhibitors of SARS-CoV-2 main protease by consensus Deep Docking of 40 billion small molecules. Chem Sci 12 , 15960-15974 (2021). Desaphy J, Bret G, Rognan D, Kellenberger E. sc-PDB: a 3D-database of ligandable binding sites--10 years on. Nucleic Acids Res 43 , D399-404 (2015). Bostrom J, Brown DG, Young RJ, Keseru GM. Expanding the medicinal chemistry synthetic toolbox. Nat Rev Drug Discov 17 , 709-727 (2018). Rarey M, Kramer B, Lengauer T, Klebe G. A fast flexible docking method using an incremental construction algorithm. J Mol Biol 261 , 470-489 (1996). Jones G, Willett P, Glen RC, Leach AR, Taylor R. Development and validation of a genetic algorithm for flexible docking. J Mol Biol 267 , 727-748 (1997). Korb O, Stutzle T, Exner TE. Empirical scoring functions for advanced protein-ligand docking with PLANTS. J Chem Inf Model 49 , 84-96 (2009). Li J, Li C, Sun J, Palade V. RDPSOVina: the random drift particle swarm optimization for protein-ligand docking. J Comput Aided Mol Des 36 , 415-425 (2022). Jain AN. Surflex-Dock 2.1: robust performance from ligand energetic modeling, ring flexibility, and knowledge-based search. J Comput Aided Mol Des 21 , 281-306 (2007). Chachulski L, Windshugel B. LEADS-FRAG: A Benchmark Data Set for Assessment of Fragment Docking Performance. J Chem Inf Model 60 , 6544-6554 (2020). Verdonk ML, Giangreco I, Hall RJ, Korb O, Mortenson PN, Murray CW. Docking performance of fragments and druglike compounds. J Med Chem 54 , 5422-5431 (2011). Marcou G, Rognan D. Optimizing fragment and scaffold docking by use of molecular interaction fingerprints. J Chem Inf Model 47 , 195-207 (2007). Hartenfeller M , et al. DOGS: reaction-driven de novo design of bioactive compounds. PLoS Comput Biol 8 , e1002380 (2012). Sommer K, Flachsenberg F, Rarey M. NAOMInext - Synthetically feasible fragment growing in a structure-based design context. Eur J Med Chem 163 , 747-762 (2019). Moroz YS. 2022q3-4 REAL database reagents, personal communication. (ed^(eds) (2023). Malamas MS , et al. Design and synthesis of aryl diphenolic azoles as potent and selective estrogen receptor-beta ligands. J Med Chem 47 , 5021-5040 (2004). Da Silva F, Desaphy J, Rognan D. IChem: A Versatile Toolkit for Detecting, Comparing, and Predicting Protein-Ligand Interactions. ChemMedChem 13 , 507-510 (2018). OpenEye Scientific Software, Sante Fe, NM, U.S.A. (ed^(eds). Penner P, Guba W, Schmidt R, Meyder A, Stahl M, Rarey M. The Torsion Library: Semiautomated Improvement of Torsion Rules with SMARTScompare. J Chem Inf Model 62 , 1644-1653 (2022). Schneider N, Lange G, Hindle S, Klein R, Rarey M. A consistent description of HYdrogen bond and DEhydration energies in protein-ligand complexes: methods behind the HYDE scoring function. J Comput Aided Mol Des 27 , 15-29 (2013). Fischer A, Smiesko M, Sellner M, Lill MA. Decision Making in Structure-Based Drug Discovery: Visual Inspection of Docking Results. J Med Chem 64 , 2489-2500 (2021). https://www.ebi.ac.uk/chembl/ (accessed 11-16-2023). Chien EY , et al. Structure of the human dopamine D3 receptor in complex with a D2/D3 selective antagonist. Science 330 , 1091-1095 (2010). Maramai S , et al. Dopamine D3 Receptor Antagonists as Potential Therapeutics for the Treatment of Neurological Diseases. Front Neurosci-Switz 10 , 451 (2016). Hopkins AL, Groom CR, Alex A. Ligand efficiency: a useful metric for lead selection. Drug Discov Today 9 , 430-431 (2004). CACHE challenge: Critical assessment of computational hit-finding experiments. https://cache-challenge.org/ , (accessed 11-16-2023) sc-PDB: An Annotated Database of Druggable Binding Sites from the Protein DataBank. http://bioinfo-pharma.u-strasbg.fr/scPDB/, (accessed 11-16-2023) Lewell XQ, Judd DB, Watson SP, Hann MM. RECAP--retrosynthetic combinatorial analysis procedure: a powerful new technique for identifying privileged molecular fragments with useful applications in combinatorial chemistry. J Chem Inf Comput Sci 38 , 511-522 (1998). Clark M, Cramer RD, III., Van Opdenbosch N. Validation of the general purpose tripos 5.2 force field. J Comput Chem 10 , 982-1012 (1989). ADFR software suite downloads, https://ccsb.scripps.edu/adfr/downloads/ (accessed 11-16-2023). Enamine building blocks catalog. https://enamine.net/building-blocks/building-blocks-catalog (accessed 03-25-2023) Dassault Systèmes Biovia Corp., San Diego, CA. https://www.3ds.com/products-services/biovia/products/data-science/pipeline-pilot/ Molecular Networks GmbH, Nürnberg, Germany. https://mn-am.com/products/corina/ (accessed 03-25-2023) Pike AC , et al. Structure of the ligand-binding domain of oestrogen receptor beta in the presence of a partial agonist and a full antagonist. EMBO J 18 , 4608-4618 (1999). Bietz S, Urbaczek S, Schulz B, Rarey M. Protoss: a holistic approach to predict tautomers and protonation states in protein-ligand complexes. J Cheminform 6 , 12 (2014). Halgren TA. Merck molecular force field. I. Basis, form, scope, parameterization, and performance of MMFF94. J Comput Chem 17 , 490-519 (1996). O'Boyle NM, Banck M, James CA, Morley C, Vandermeersch T, Hutchison GR. Open Babel: An open chemical toolbox. J Cheminform 3 , 33 (2011). Table Table 2 is available in the Supplementary Files section. Supplementary Information Supplementary Figures and Supplementary Tables are not available with this version. Additional Declarations Yes there is potential Competing Interest. D.R. is co-founder and shareholder of BIODOL Therapeutics. M.E. is employee of Amgen. Supplementary Files Table2.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3687338","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":257163446,"identity":"7ca94271-80bc-4559-b5c9-815a6cf459c6","order_by":0,"name":"Didier Rognan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAyklEQVRIiWNgGAWjYPACGwYGZgiLn1gtaSAtjA1AlmQDkVoOgwgitZi3t1978HPH+cTt7OzPH1fUMEiYE9Ijc+ZMuWHvmduJO5t5DBvPHGOQkDlAQIuERE6aBG/b7cQNh3kYGxvYGOokCDkMpEXyb9s5oBb2h40N/4AChLWkH5PmbTsA1MJg2NjYRowWnjNs0rJtycZAhxnObOyTIEILe/szybdtdrIbzh9/8LHhmw1hLQwMPAYoRhDWwMDA/oAYVaNgFIyCUTCSAQC5mTzEOWyPXQAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-0577-641X","institution":"Laboratoire d'innovation therapeutique","correspondingAuthor":true,"prefix":"","firstName":"Didier","middleName":"","lastName":"Rognan","suffix":""},{"id":257163447,"identity":"91feb449-3e1b-49c6-a33a-225c93d39459","order_by":1,"name":"François Sindt","email":"","orcid":"","institution":"Laboratoire d'Innovation Thérapeutique, UMR7200 CNRS/Université de Strasbourg","correspondingAuthor":false,"prefix":"","firstName":"François","middleName":"","lastName":"Sindt","suffix":""},{"id":257163448,"identity":"30ace1eb-6647-4fe1-a4a6-50331678566c","order_by":2,"name":"Anthony SEYLLER","email":"","orcid":"","institution":"Laboratoire d'Innovation Thérapeutique, UMR7200 CNRS/Université de Strasbourg","correspondingAuthor":false,"prefix":"","firstName":"Anthony","middleName":"","lastName":"SEYLLER","suffix":""},{"id":257163449,"identity":"e615915f-782e-4e57-8400-0411d396add7","order_by":3,"name":"Merveille Eguida","email":"","orcid":"","institution":"Laboratoire d'Innovation Thérapeutique, UMR7200 CNRS/Université de Strasbourg","correspondingAuthor":false,"prefix":"","firstName":"Merveille","middleName":"","lastName":"Eguida","suffix":""}],"badges":[],"createdAt":"2023-11-30 14:10:35","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3687338/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3687338/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":49385793,"identity":"0933a591-a997-4996-8316-8bfdf42b5554","added_by":"auto","created_at":"2024-01-09 19:51:42","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":252251,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAccuracy of state-of-the-art docking tools to dock 5,845 sc-PDB reagents in their cognate targets.\u003c/strong\u003e \u003cstrong\u003e(A)\u003c/strong\u003eRoot-mean square deviation (rmsd) of the best pose (lowest rmsd, heavy atoms only) to the surrogate X-ray structure, \u003cstrong\u003e(B)\u003c/strong\u003e Similarity of protein-reagent interaction fingerprints between the best pose (highest interaction fingerprint similarity) and surrogate X-ray structures, measured by a Tanimoto coefficient. Fingerprints could not be measured for RDPSOVina poses in pdbqt format, \u003cstrong\u003e(C)\u003c/strong\u003eCumulative rmsd of the best pose (GOLD-PLP docking) for each of the 13 chemical functions. Numbers in brackets indicate the absolute number of each chemical function, \u003cstrong\u003e(D)\u003c/strong\u003e Cumulative rmsd of the best pose (GOLD-PLP docking) according to protein class. Numbers in brackets indicate the absolute number of samples from each protein family.\u003c/p\u003e","description":"","filename":"image1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/bd97a29ffefea0324c391a3c.jpeg"},{"id":49385789,"identity":"fdb034d8-4df7-4af3-89b1-b44f74632721","added_by":"auto","created_at":"2024-01-09 19:51:41","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":90184,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAnnotation of chemical reagents by reaction type, reactant role and reactive atoms\u003c/strong\u003e.\u003c/p\u003e","description":"","filename":"image2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/e3e5f31743862bf6cec6522b.jpeg"},{"id":49385792,"identity":"bf026686-debd-428e-8121-5be33ebfdaa3","added_by":"auto","created_at":"2024-01-09 19:51:42","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":283066,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSpace docking of benzoxazole and sulfonamide chemical spaces to the human estrogen receptor beta (ERβ). (A) \u003c/strong\u003eX-ray structure of human ERβ (tan ribbons, PDB 1QKM) in complex with the agonist genistein (blue sticks)\u003cstrong\u003e. \u003c/strong\u003eThe genistein binding site is\u003cstrong\u003e \u003c/strong\u003edelimited by ERβ residues displayed as tan sticks with main receptor-ligand hydrogen-bonds indicated by cyan broken lines. The known benzoxazole agonist (WAY-338) is taken as the ground truth ligand to recover. \u003cstrong\u003e(B)\u003c/strong\u003e SpaceDock flowchart affording 64 potential ERβ agonists according to a series of filters (\u003cstrong\u003eTable 1\u003c/strong\u003e). The custom filter (H-bond either Glu305 or Arg346, and to His475) is target-specific. \u003cstrong\u003e(C) \u003c/strong\u003eStructures and rank (#) of 4 representative benzoxazoles. The proposed binding poses are overlayed to the X-ray pose of the ground truth ligand (WAY-338, cyan), the protein being masked for sake of clarity.\u003c/p\u003e","description":"","filename":"image3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/cea4972a5b906cbf0ee255e3.jpeg"},{"id":49385795,"identity":"98259773-7601-48e0-9abf-46a937d86786","added_by":"auto","created_at":"2024-01-09 19:51:42","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":349111,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSpace docking of an amides chemical space to the human dopamine D3 receptor (DRD3). (A) \u003c/strong\u003eX-ray structure of human DRD3 (tan ribbons, PDB 3PBL) in complex with the antagonist eticlopride (blue sticks)\u003cstrong\u003e. \u003c/strong\u003eThe eticlopride binding site is\u003cstrong\u003e \u003c/strong\u003edelimited by DRD3 residues displayed as tan sticks with the main receptor-ligand ionic bond indicated by cyan broken lines. Eticlopride is taken as both the reference and the ground truth ligand to recover. \u003cstrong\u003e(B)\u003c/strong\u003eSpaceDock flowchart affording 315 potential DRD3 antagonists according to a series of filters (\u003cstrong\u003eTable 1\u003c/strong\u003e). The custom filter (IFP similarity to eticlopride X-ray pose) is target-specific. \u003cstrong\u003e(C) \u003c/strong\u003eStructures and rank of 4 representative orthomethoxybenzamides. The proposed binding poses are overlayed to the X-ray pose of the ground truth ligand (eticlopride, cyan), the protein being masked for sake of clarity. \u003cstrong\u003e(D) \u003c/strong\u003eStructure and binding poses of other hits, aligned to the X-ray pose of eticlopride.\u003c/p\u003e","description":"","filename":"image4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/7ea0f6a57b86bf5d24651ab9.jpeg"},{"id":49385790,"identity":"dcd718c2-83fa-414c-9e3d-40385fe32036","added_by":"auto","created_at":"2024-01-09 19:51:42","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":242875,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStructure and binding to human DRD3 of 15 SpaceDock hits from the amide space. \u003c/strong\u003eHits are labelled according to their SpaceDock rank, Enamine's catalog identifiers and purchased as racemates, unless specified. \u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eBinding affinities to human DRD3 are expressed as the percentage of inhibition of [\u003csup\u003e3\u003c/sup\u003eH]-methylspiperone binding to human recombinant DRD3 expressed in CHO cells (Eurofins Discovery assay #48) at a single concentration of 10 µM competitor (mean of two independent experiments). The inhibition constant (K\u003csub\u003ei\u003c/sub\u003e) was determined from dose-response curves for six strong binders (in green). Compound #123 could not be synthesized (n.s.)\u003c/p\u003e","description":"","filename":"image5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/5ad33408d068902b0cd89b85.jpeg"},{"id":49386751,"identity":"7ec25f9c-fd77-4bce-a7bb-9e13c3b65565","added_by":"auto","created_at":"2024-01-09 19:59:42","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1318471,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/84fe0208-616c-4d4a-b64f-57c9ddd5ff04.pdf"},{"id":49385791,"identity":"fca8c46e-3648-498c-a624-f577bd558130","added_by":"auto","created_at":"2024-01-09 19:51:42","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":61030,"visible":true,"origin":"","legend":"","description":"","filename":"Table2.docx","url":"https://assets-eu.researchsquare.com/files/rs-3687338/v1/d6da161e241d0e8e67a5cf38.docx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential Competing Interest.\nD.R. is co-founder and shareholder of BIODOL Therapeutics. \r\nM.E. is employee of Amgen.","formattedTitle":"Protein structure-based organic chemistry-driven ligand design from ultra-large chemical spaces","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIdentifying the first hit compounds able to target a macromolecule of interest is often achieved by screening experimentally or computationally a library of drug-like compounds,\u003csup\u003e1\u003c/sup\u003e thereby enabling a hit to lead follow-up using classical medicinal chemistry strategies.\u003csup\u003e2\u003c/sup\u003e Until recently, the commercially-available chemical space describing drug-like compounds amenable to screening has been restricted to 10-15 million compounds with a yearly growth of ca. half a million compounds.\u003csup\u003e3\u003c/sup\u003e On-demand compound libraries\u003csup\u003e4, 5\u003c/sup\u003e have completely changed this situation by proposing billions of compounds not yet available but easily synthesizable in a few steps and reproducible parallel synthesis. Early approaches to virtually screen subsets of ultra large chemical spaces led to spectacular successes,\u003csup\u003e6-9\u003c/sup\u003e notably unexpected high hit rates, very high potencies and fine selectivity.\u003csup\u003e10, 11\u003c/sup\u003e Today, ca. 70 billion compounds are accessible on-demand with fast delivery (6-8 weeks) and high-purity grade (\u0026gt; 95%).\u003csup\u003e12\u003c/sup\u003e Due to their huge size, compounds describing these ultra-large chemical spaces cannot be fully enumerated and requires dedicated computational tools for registration, storage and navigation.\u003csup\u003e13\u003c/sup\u003e Usually, large chemical spaces are described in a combinatorial manner from the building blocks and organic chemistry reactions required to synthesize them.\u003csup\u003e5\u003c/sup\u003e If ligand-based approaches are now available to efficiently query these large chemical spaces\u003csup\u003e14-16\u003c/sup\u003e, structure-based approaches including macromolecular target information (e.g. topology of a binding site) still need to be developed to exhaustively mine multibillion chemical spaces. Several computational methods have indeed been described for such a task\u003csup\u003e17-23\u003c/sup\u003e, albeit with moderate to severe restrictions. One the one hand, brute force docking of 1.4 billion compounds\u003csup\u003e23\u003c/sup\u003e has been successfully described with the help of costly dedicated platforms,\u003csup\u003e23, 24\u003c/sup\u003e but will soon reach its limits with next-to-come trillion-sized chemical spaces,\u003csup\u003e25\u003c/sup\u003e since full atomistic docking just scales linearly with the number of compounds to be screened. A workaround consists in the proper selection of seed fragments/scaffolds to screen a representative subset of the entire space. The seed fragment may originate from the early docking of fragment-based representative synthons,\u003csup\u003e20\u003c/sup\u003e X-ray diffraction screening data\u003csup\u003e21\u003c/sup\u003e or medicinal chemistry knowledge.\u003csup\u003e22\u003c/sup\u003e Once a seed fragment has been identified, scaffold-focused two-dimensional (2D) libraries, exploring the corresponding chemical space via a set of organic chemistry reactions,\u003csup\u003e26\u003c/sup\u003e can be enumerated, converted in three-dimensional (3D) atomic coordinates and physically docked to propose novel hits. This approach has been applied with success to a few targets\u003csup\u003e20-22, 27\u003c/sup\u003e but still requires hardware settings enabling docking a significant subset (a few million) of the entire chemical space. Last, fast machine learning approaches may be first trained on a set of representative ligand-annotated docking poses to simply predict docking scores,\u0026nbsp;\u003csup\u003e17-19, 28, 29\u003c/sup\u003e and next be applied to predict docking scores for the remaining space. Even if only a small fraction of the full space (1-5%) has to be docked at the atomic level, this strategy cannot be further applied to trillion- sized chemical spaces since it would require gathering first billion of docking scores on a single target. Moreover, this approach has led to very mitigated results with respect to hit rate and hit potencies,\u003csup\u003e30\u003c/sup\u003e and deserves further experimental validations.\u003c/p\u003e\n\u003cp\u003eHerein, we present a simple and fast computational approach (SpaceDock) avoiding the above-cited drawbacks. It first requires docking commercially available chemical reagents in the target of interest, in order to couple them according to standard organic chemistry reactions to propose multibillion compound libraries in one or two synthetic steps. When applied to two targets of pharmaceutical interest, the method was able to quickly retrieve hits that are chemically identical (or very close) to existing ligands, but also to propose chemically novel and potent ligands.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eSince the SpaceDock method heavily relies on the possibility to accurately dock chemical reagents, we first investigate the best docking prototols for the latter task by setting-up a dedicated benchmarking study. We then describe how chemical reagents are annotated by reactive groups and organic chemistry reactions, to define a chemical space of 53.5 billion synthesizable compounds. Last we present two concrete application of the SpaceDock workflow to two receptors of pharmaceutical interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSetting up the conditions for accurate docking of chemical reagents\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo evaluate the feasibility of the SpaceDock approach, we first needed to set-up an archive of reference 3D structures for protein-bound chemical reagents. Since experimental data for such a dataset are missing, we fragmented in 3D space drug-like ligands from known protein-ligand X-ray structures (sc-PDB dataset)\u003csup\u003e31\u003c/sup\u003e using a set of 12 common organic chemistry reactions, then added the 3D atomic coordinates of the missing reactive moieties (e.g. boronic acid, halide; \u003cstrong\u003eSupplementary Fig. 1\u003c/strong\u003e) and last created on-the-fly \u0026quot;surrogate X-ray poses\u0026quot; for the corresponding reagents expected to yield the parent ligands with the above-described reactions. The final archive of 5,845 reagents was selected after appropriate filtering \u003cstrong\u003e(Supplementary Table S1\u003c/strong\u003e) and exhibit 13 chemical functions with a prevalence of reactive groups (e.g. amines, aryl halides, boronic acids) reflecting the frequent usage of simple organic chemistry reactions in drug discovery.\u003csup\u003e32\u003c/sup\u003e\u0026nbsp; With a set of reference reagents in hand, we could next verify whether state-of-the-art docking algorithms were able to reproduce the surrogate X-ray poses. Five algorithms relying on different principles (FlexX:\u003csup\u003e33\u003c/sup\u003e incremental construction, GOLD:\u003csup\u003e34\u003c/sup\u003e genetic algorithm, PLANTS:\u003csup\u003e35\u003c/sup\u003e ant colony optimization, RDPSOVina:\u003csup\u003e36\u003c/sup\u003e random drift particle swarm optimization, \u0026nbsp;Surflex:\u003csup\u003e37\u003c/sup\u003e surface-based molecular \u0026nbsp;similarity) were used for that purpose. Since the SpaceDock strategy just need a single pair of complementary reagents to be properly docked to reconstitute a full ligand, the docking performance was measured by computing the root-mean square deviation (rmsd) of the pose found to be the closest (best pose) to that of the surrogate X-ray structure (\u003cstrong\u003eFig. 1\u003c/strong\u003e). All docking tools exhibit an excellent docking performance with 70-80% of chemical reagents being docked within 2 \u0026Aring; rmsd accuracy (\u003cstrong\u003eFig. 1A\u003c/strong\u003e). Up to 70% of very high-quality poses (rmsd \u0026lt; 1 \u0026Aring;) could be generated by the apparently best docking/scoring scheme (GOLD docking, PLP scoring; \u003cstrong\u003eFig. 1A\u003c/strong\u003e). The observed docking accuracy is therefore independent on the chosen docking algorithm, and remains in agreement with docking benchmarks on low molecular weight fragments.\u003csup\u003e38, 39\u003c/sup\u003e Since rmsd is a global measure that does not take into account whether key protein-reagent interactions are verified or not, we additionally computed the similarity of protein-reagent interaction fingerprints (IFPs)\u003csup\u003e40\u003c/sup\u003e between docked and surrogate X-ray poses. Again, an excellent performance could be noticed using this orthogonal quality descriptor, with 75-85 % of chemical reagents for which the IFP similarity to the X-ray pose is deemed acceptable (Tc-IFP \u0026gt; 0.60;\u003csup\u003e40\u003c/sup\u003e \u003cstrong\u003eFig. 1B\u003c/strong\u003e). To ascertain that all chemical functions are equally suitable for docking, the same analysis was repeated for each of the 13 chemical groups (\u003cstrong\u003eFig. 1C\u003c/strong\u003e) present in our library, focusing on the best docking strategy (GOLD docking, PLP scoring). Reassuringly, the docking performance appears to be relatively independent on the chemical function of the reagent (\u003cstrong\u003eFig. 1C\u003c/strong\u003e) as well as on the target protein family (\u003cstrong\u003eFig. 1D\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDefining a readily-accessible ultra-large chemical space from simple organic chemistry reactions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStarting from the pioneering work of Hartenfeller et al.,\u003csup\u003e26\u003c/sup\u003e we selected 36 robust, stereo- and regioselective organic chemistry reactions to define a chemical space of 5.5 billion compounds readily accessible in one or two synthesis steps (\u003cstrong\u003eSupplementary Table S2, Supplementary Fig. 2\u003c/strong\u003e). Contrarily to previous similar approaches\u003csup\u003e26, 41, 42\u003c/sup\u003e, chemical reagents were here carefully chosen from specific SMARTS strings in a list of 145,705 commercial chemical reagents contributing to Enamine\u0026apos;s REAL space\u003csup\u003e43\u003c/sup\u003e of 36 billion compounds. Moreover, possible side reactions affecting synthesis yields were minored by selecting reagents that are monofunctional for a particular chemical function (e.g. monocarboxylic acid), and lacking additional chemical functions (e.g. nucleophilic groups for an electrophilic reactant) that would decrease the reaction yield (\u003cstrong\u003eSupplementary Table S2\u003c/strong\u003e). Altogether, 134,331 commercial reactants could be unambiguously annotated by reaction type, reactant role and reactive atoms yielding a total of 713,155 atomic tags (\u003cstrong\u003eFig. 2\u003c/strong\u003e). Conversion in 3D atomic coordinates provided a total of 176,824 ready-to-dock unique reagents, ionized at pH 7, including stereoisomers for reactants bearing up to two undefined chiral centers.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRetrospective chemical space docking of 97 million compounds for human estrogen receptor beta agonists.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor a first proof-of-concept, we selected as a target the activated form of the human estrogen receptor beta (ER\u0026beta;) for the following two reasons: (i) the ligand-binding cavity is nicely druggable with a good hydrophobicity/hydrophilicity balance, (ii) the receptor has been co-crystallized with many high-affinity low molecular-weight agonists, notably compounds sharing a 2-aryl-benzoxazole scaffold\u003csup\u003e44\u003c/sup\u003e whose one-step synthesis from 2-aminophenols and benzaldehydes is one of the 36 reactions that we have encoded. To avoid a possible chemotype bias, we selected an X-ray receptor structure co-crystallized with genistein (PDB 1QKM), a non-benzoxazole high-affinity agonist used from hereon as \u0026quot;reference ligand\u0026quot; (\u003cstrong\u003eFig. 3A\u003c/strong\u003e) and asked whether we could recover a \u0026quot;ground truth\u0026quot; benzoxazole agonist (WAY-338, \u003cstrong\u003eFig. 3A\u003c/strong\u003e) or any close analog, by first docking the necessary reactants (2-aminophenols, benzaldehydes) and then enabling the benzoxazole ring formation within the protein binding site. To this end, 145 commercial 2-aminophenols and 3,874 benzaldehydes were generated in 3D and docked into the 1QKM structure, in order to explore a combinatorial space of 561,730 possible benzoxazoles. Since the later space is small, we additionally considered a much larger space of 97 million sulfonamide decoys synthesizable from 1,275 sulfonyl chlorides and 76,758 amines, thereby strongly minoring the benzoxazole space (0.57%) in the full chemical space to scan. After docking all reagents necessary to mine both chemical spaces according to the previously found best protocol (GOLD docking, PLP scoring), a series of filters of increasing complexity (\u003cstrong\u003eTable 1\u003c/strong\u003e) was iteratively passed to a decreasing number of possible solutions, first starting with pairs of potentially reacting reagent poses, then with successfully enumerated ligand poses, and last with quality checked redocking poses.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1 | Incremental series of filters applied to prioritize SpaceDock hits\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eFilter\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eType\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eCriteria\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eApplies to\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eSoftware used\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eGeometry\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eDistances, angles, clashes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003ePair of reactant poses\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003ethis work\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eInteraction\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eInteraction fingerprint similarity to reference\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003ePair of reactant poses\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eIChem\u003csup\u003e45\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eEnergy, geometry\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eRmsd of refined pose to non-refined pose\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eFully enumerated ligand\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eSzybki\u003csup\u003e46\u003c/sup\u003e\u003c/p\u003e\n \u003cp\u003eSurflex-Dock\u003csup\u003e37\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eInteraction, structure\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eInteraction fingerprint similarity (IFP) to reference\u003c/p\u003e\n \u003cp\u003eNumber of stereocenters Number of rotatable bonds\u003c/p\u003e\n \u003cp\u003eDrug-likeness\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eFully enumerated ligand\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eIChem\u003csup\u003e45\u003c/sup\u003e\u003c/p\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003cp\u003eFilter\u003csup\u003e46\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eRedocking\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eRmsd to energy-minimized SpaceDock pose\u0026nbsp;\u003c/p\u003e\n \u003cp\u003eIFP similarity to energy-minimized SpaceDock pose\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eDocking poses\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eGOLD\u003csup\u003e34\u003c/sup\u003e\u003c/p\u003e\n \u003cp\u003eSurflex-Dock\u003csup\u003e37\u003c/sup\u003e\u003c/p\u003e\n \u003cp\u003eIChem\u003csup\u003e45\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eQuality check\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eNumber of strained torsions Local and global strain energy\u003c/p\u003e\n \u003cp\u003eNumber of unsatisfied H-bond donors and acceptors, number of unsatisfied ionic bonds\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eDocking poses\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eTorsion_analyzer\u003csup\u003e47\u003c/sup\u003e Freeform\u003csup\u003e46\u003c/sup\u003e\u003c/p\u003e\n \u003cp\u003ethis work\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"7.781456953642384%\" valign=\"top\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.85430463576159%\" valign=\"top\"\u003e\n \u003cp\u003eFinal selection\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"30.29801324503311%\" valign=\"top\"\u003e\n \u003cp\u003eDuplicates removal\u003c/p\u003e\n \u003cp\u003eHYDEscore\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eDocking poses\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.033112582781456%\" valign=\"top\"\u003e\n \u003cp\u003eThis work\u003c/p\u003e\n \u003cp\u003eHydescorer\u003csup\u003e48\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe SpaceDock flowchart is displayed \u003cstrong\u003eFig. 3\u003c/strong\u003e. In a first step, pure chemical and topological filters (\u003cstrong\u003eSupplementary Figs. 3, 4\u003c/strong\u003e) are passed to all docking poses of possible reactant pairs to quickly remove impossible reactions (filter #1). To stay on a safe side, we only considered pairs of bound reactants exhibiting a total interaction fingerprint (IFP) similarity\u003csup\u003e40\u003c/sup\u003e to the genistein X-ray pose above an acceptable threshold\u003csup\u003e40\u003c/sup\u003e (IFP \u0026ge; 0.60 considering all non-bonded interactions, IFP \u0026ge; 0.50 considering polar interactions only; filter #2). The 821,702 remaining pairs of reactants were then converted, in the protein 3D space, into the corresponding benzoxazoles and sulfonamides, respectively and the fully enumerated ligands were quickly minimized in the protein binding site. Only 539,906 poses deviated less than 1.0 \u0026Aring; rmsd from the non-refined poses after energy refinement (filter #3). The remaining minimized poses were filtered again according to IFP similarity of the genistein X-ray pose (IFP \u0026ge; 0.60 considering all non-bonded interactions, IFP \u0026ge; 0.60 considering polar interactions only; filter #4). Compounds with more than 2 stereocenters and 8 rotatable bonds were removed at this stage, leaving 49,569 poses for further processing. To ensure that the selected SpaceDock poses might be recovered by classical docking, all remaining hits were redocked to the ER\u0026beta; structure, as previously done for the reagents. Only 121,470 poses close to the corresponding energy-minimized SpaceDock poses (rmsd \u0026le;2.0 \u0026Aring;; IFP \u0026ge; 0.60 considering all non-bonded interactions, IFP \u0026ge; 0.60 considering polar interactions only) were retained (filter #5). A quality check of remaining poses (filter #6) was next applied to remove unlikely solutions (\u0026ge; 1 strained torsion, local strain energy \u0026gt; 4 kcal/mol, global strain energy \u0026gt; 8 kcal/mol, no unsatisfied ionic bond , \u0026gt;2 unsatisfied h-bond donors, \u0026gt; 4 unsatisfied h-bond acceptors).\u003csup\u003e22, 49\u003c/sup\u003e The number of plausible solutions (7,712) being still important, a custom filter was finally applied to keep only poses anchored at both sides of the binding pocket (H-bond either Glu305 or Arg346, and to His475), as seen for all potent ER\u0026beta; agonists (recall genistein X-ray pose, \u003cstrong\u003eFig. 3A\u003c/strong\u003e). The final hit list comprises 102 poses from 64 unique ligands (filter #7), including 54 benzoxazoles and 10 sulfonamides (\u003cstrong\u003eFig. 3B, Supplementary Table S3)\u003c/strong\u003e ranked by decreasing full IFP similarity to the reference ligand, then by decreasing polar IFP similarity, and last by increasing absolute binding free energy predicted by the HYDE scoring function.\u003csup\u003e48\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eDespite in minority in the initial space (0.57%), it is reassuring that the ground truth chemotype was considerably enriched (84 %) in the final hit list. Inspecting the structures and binding poses of the hits, we observed that SpaceDock was indeed able to recover, among the top-ranked hits, the ground truth ligand (rank #9), a known ER\u0026beta;\u0026nbsp;agonist\u0026nbsp;ChEMBL187673\u003csup\u003e50\u003c/sup\u003e (IC\u003csub\u003e50\u003c/sub\u003e = 50 nM, rank #25) and 52 other 2-arylbenzoxazoles, with almost perfect binding modes (rmsd =1.15 \u0026Aring; for the ground-truth ligand, \u003cstrong\u003eFig. 3C\u003c/strong\u003e). About half of the hits (30 out 64; all from the benzoxazole space), were considered chemically similar to existing ER\u0026beta;\u0026nbsp;ligands (\u003cstrong\u003eSupplementary Fig. 5\u003c/strong\u003e), evidencing that SpaceDock can propose both known ligands (or very close analogs thereof) and new chemical entities. However, only a lower number of compounds (17, out of which 10 share the sulfonamide space) strictly intersected Enamine REAL space (\u003cstrong\u003eSupplementary Fig. 5\u003c/strong\u003e). This observation does not preclude for their synthesizability but just illustrates that these hits, despite the commercial availability of their starting building blocks, cannot be obtained within the scope of 167 parallel synthesis protocols defining REAL space.\u003c/p\u003e\n\u003cp\u003eFrom this preliminary proof-of-concept, it appears that the herein presented method is able to perform a complex organic chemistry reaction (ring cyclisation) from suitably posed and chemically compatible chemical reagents, under the 3D constraints of the target\u0026apos;s structure, to generate and prioritize fully enumerated ligands for meaningful reasons. We therefore decided to apply SpaceDock to a prospective screening of a much larger chemical space.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProspective chemical space docking of 670 million compounds for human dopamine D3 receptor antagonists.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe next applied the method to a much larger chemical space of 670 million carboxamides targeting the human dopamine D3 receptor (DRD3). Since the only available high-resolution DRD3 receptor structure (PDB 3PBL) has been obtained in complex with the antagonist eticlopride (\u003cstrong\u003eFig. 4A\u003c/strong\u003e),\u003csup\u003e51\u003c/sup\u003e the latter orthomethoxybenzamide (OMB) ligand was used as both reference and ground-truth ligand to recover. \u0026nbsp;Commercially available carboxylic acids and primary/secondary amines (\u003cstrong\u003eSupplementary Table 2\u003c/strong\u003e) were first filtered to remove reagents that, upon amide bond formation, would lead to non-drug-like ligands (\u003cstrong\u003eSupplementary Table 4\u003c/strong\u003e), thereby keeping 19,887 acids and 33,726 amines (in 3D coordinates) to explore a chemical space of 670 million carboxamides (\u003cstrong\u003eFig. 4B\u003c/strong\u003e). The resulting 53,613 chemical reagents were then docked to the eticlopride-free DRD3 structure using GOLD docking and PLP scoring, as previously described. Since 20 poses were saved for each reactant, a total of 268 billion (19,887*20*33,726*20) possible reactions were passed to the SpaceDock flowchart (\u003cstrong\u003eFig. 4B\u003c/strong\u003e), removing first impossible amide bond formation according to geometrical criteria (\u003cstrong\u003eSupplementary Fig.6\u003c/strong\u003e) while keeping only amine poses exhibiting the crucial ionic bond to the key Asp110 residue\u003csup\u003e51\u003c/sup\u003e (filter #1, \u003cstrong\u003eFig. 4B\u003c/strong\u003e), then retaining pair of reactant poses for which the IFP similarity to the reference ligand is higher than 0.60 for all interactions and 0.50 for polar interactions only (filter #2).\u003csup\u003e40\u003c/sup\u003e A total of 24,674,693 reactions were conducted \u003cem\u003ein silico\u003c/em\u003e to generate the corresponding carboxamides inside the receptor pocket, that were later energy-minimized. Keeping only minimized poses that did not deviate much from the initial pose (rmsd \u0026lt; 1.0 \u0026Aring;) afforded 15,120,198 plausible solutions (filter #3, \u003cstrong\u003eFig. 4B\u003c/strong\u003e). At this stage, hits bearing a cis-amide bond or more than 2 chiral centers or more than 9 rotatable bonds were removed to keep only drug-like compounds. The resulting number of hits being still very high, we pruned the hit list by keeping only minimized poses with a high full IFP similarity to the reference ligand (IFP similarity \u0026gt; 0.60) while exhibiting a perfect IFP similarity to eticlopride (IFP =1) with respect to polar interactions (H-bond and ionic bond to Asp110). This filter (filter #4, \u003cstrong\u003eFig. 4B\u003c/strong\u003e) yielded to 518,306 SpaceDock poses (corresponding to 500,041 unique compounds) that had to be confirmed by full atomistic docking (GOLD docking, PLP scoring, 20 poses saved) of the corresponding ligands and comparison with the minimized SpaceDock poses. Only docking poses verifying the following three criteria (rmsd \u0026le;2.0 \u0026Aring; \u0026amp; IFP_full \u0026ge; 0.60 \u0026amp; IFP_polar =1) were retained, leaving 712,120 good docking poses (filter #5, \u003cstrong\u003eFig. 4B\u003c/strong\u003e) for sanity check (no strained torsion, local strain energy \u0026le; 4 kcal/mol, global strain energy \u0026le; 8 kcal/mol, no unsatisfied ionic bond, \u0026le; 2 unsatisfied h-bond donors, \u0026le; 4 unsatisfied h-bond acceptors, filter #6, \u003cstrong\u003eFig. 4B\u003c/strong\u003e). The number of remaining poses being still important (97,096), a custom filter (not implemented by default, \u003cstrong\u003eTable 1\u003c/strong\u003e) was added to remove poses for compounds with no aromatic ring (always present in known DRD3 antagonists),\u003csup\u003e52\u003c/sup\u003e exhibiting a predicted absolute binding free energy (HYDEscore) lower than 30 kJ/mol and further restricting the deviation to the original SpaceDock poses (rmsd \u0026le;1.0 \u0026Aring; \u0026amp; IFP_full \u0026ge; 0.75). A reasonable number of 757 docking poses from 315 unique ligands (filter #7, \u003cstrong\u003eFig. 4B\u003c/strong\u003e) defined the final hit list. Compounds were ranked by\u0026nbsp;decreasing full IFP similarity to the reference ligand, then by decreasing polar IFP similarity, and last by increasing HYDE binding free energy\u0026nbsp;(\u003cstrong\u003eSupplementary Table 5\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eAs for the first attempt on ER\u0026beta; ligands, we first check whether the ground-truth ligand and its corresponding OMB scaffold were present in the list. Indeed, 15 OMBs including eticlopride (rank 30) were part of the list with binding poses very similar to that observed for the reference ligand (rmsd of eticlopride = 0.73 \u0026Aring;, \u003cstrong\u003eFig. 4C\u003c/strong\u003e). Interestingly, 300 additional hits not sharing the OBM scaffold were prioritized with poses and protein-ligand interaction patterns quite close to that seen for eticlopride (\u003cstrong\u003eFig. 4D\u003c/strong\u003e). Most ligands were scaffold hops for which the orthomethoxybenzamide has been replaced by a bicyclic heteroaryl-amide, connected by 2-3 carbon atoms to a basic amine. By comparison to the ER\u0026beta; hit list, the DRD3 hits deviate more from known ChEMBL ligands (24% considered as chemically similar) but are more easily obtainable in REAL space (53% being directly purchasable, and additional 38% being very close to REAL space compounds; \u003cstrong\u003eSupplementary Fig. 7\u003c/strong\u003e). 16 chemically diverse and representative hits were directly purchased at Enamine, out of which 15 could be synthesized in six weeks (5 mg quantity, \u0026gt; 90% purity) and further tested for binding to human DRD3 (\u003cstrong\u003eFig. 5\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eOut of the tested 15 compounds, ten exhibited detectable binding (\u0026gt; 20% inhibition) to the DRD3 receptor at the single concentration of 10\u0026nbsp;\u0026micro;M (\u003cstrong\u003eFig. 5\u003c/strong\u003e). The six strongest binders (#1, #25, #66, #107, #142, #161) were selected for dose-curve responses for inhibition constants (K\u003csub\u003ei\u003c/sub\u003e) determination (\u003cstrong\u003eFig. 5\u003c/strong\u003e, \u003cstrong\u003eSupplementary Fig. 8\u003c/strong\u003e). Three of them (#1, #66, #142) exhibited K\u003csub\u003ei\u003c/sub\u003e values in the 300-400 nM range, the three others at 1.4-1.6 \u0026micro;M. The remarkable hit rates (66% at 10 \u0026micro;M, 20% at 500 nM) are in line with previous observations from docking ultra-large libraries,\u003csup\u003e10, 11\u003c/sup\u003e and suggests that SpaceDock competes rather well with much more demanding full atomistic docking when screening large chemical spaces.\u003c/p\u003e\n\u003cp\u003eInterestingly, novel heteroamatic-carboxamide scaffolds were disclosed for 4 of the strong binders (#66, #107, #142 and #161) that could not be found in any of 6,714 dopamine DRD2/DRD3 ligands from ChEMBL (\u003cstrong\u003eTable 2\u003c/strong\u003e). SpaceDock proposals should still be considered as primary hits. As such, their potency is lower than that of the closest dopamine D2/D3 antagonists from ChEMBL, albeit with a higher ligand efficiency.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eWe herein describe a novel computational method (SpaceDock) to exhaustively browse ultra-large chemical spaces under specific constraints of a target protein and known binders. When applied to two nicely druggable targets (estrogen receptor \u0026beta;, dopamine D3 receptor) and chemical spaces up to 670 million compounds, it enabled the fast recovery of known ligands/scaffolds (in both cases) and the identification of novel and potent new chemical entities (dopamine D3 receptor).\u003c/p\u003e\n\u003cp\u003eSpaceDock departs from existing methods\u003csup\u003e20-22\u003c/sup\u003e by two major differences: (i) fully unmodified chemical reagents and not synthons (scaffolds with chemistry-informed exit vectors) are used as primary sources of hits, (ii) most promising ligands are directly obtained within the protein binding site, by 3D \u003cem\u003ein silico\u003c/em\u003e synthesis according to geometrical and chemical cross-compatibility of previously posed reagents pairs.\u003c/p\u003e\n\u003cp\u003eIndeed, direct docking of chemical reagents has, to the best of our knowledge, never been reported. Interestingly, our preliminary benchmark demonstrates that docking chemical reagents is as accurate as docking low-molecular weight fragments\u003csup\u003e38\u003c/sup\u003e with ca. 75% of chemicals properly posed with respect to their corresponding substructures in full PDB ligands. Noteworthy, the docking accuracy is independent on the docking tool used, as well on the reactive moiety of the reactants and on the target protein family; therefore opening the method to any druggable target and set of commercial building blocks. To enable an easy synthetic access to most SpaceDock hits, the method relies on chemical reagents contributing to Enamine\u0026apos;s REAL space, and generate hits in the binding site 3D space using a set of 36 robust two-component organic chemistry reactions. Given the 70% average docking accuracy of reactants, we therefore expect the likehood to properly couple two chemically compatible reactants into a fully enumerated and suitably posed ligand at ca. 50%. Docking the starting chemical reagents is clearly the most time-consuming step of the entire flowchart (ca 15 s/reagent), meaning that SpaceDock scales with the number of reactants and not the number of products defining the chemical space to be screened. To optimize the speed of the further processing, a series of filters of increasing complexity is applied, step to step, to a decreasing number of plausible solutions. Just checking the relative position of compatible reactants to be paired by fast distance/angles measures permits to remove 99.8% of possible solutions. Although not mandatory, we applied IFP similarity to a reference pose to remove topologically valid ligands not fulfilling expected interactions with key residues. This filter permits to reduce the number of full ligand poses to the third most time-consuming but necessary energy-minimization step (ca. 1s/ recombined pose), and remove local strains around the newly created bonds. We assume that a SpaceDock proposal is all the more interesting if it does not vary (in terms of rmsd and IFP similarity) upon energy minimization within the protein binding site, and if it can be recovered by full atomistic docking of the corresponding ligand. Although not necessary, we recommend this redocking step to ensure that SpaceDock and any state-of-the-art docking tool (we here used GOLD but other tools may be used as well) agrees on the final poses to be sent to the very important quality check. A particular importance is given to local and global strain energies (\u0026le;4 and 8 kcal/mol, respectively), as well as to the number of unsatisfied ionic bonds (none) and of unsatisfied hydrogen-bond donors/acceptors (\u0026le; 2 and 4, respectively). In the DRD3 test case, omitting this step drastically enriched the final hit list in false positives which could not be confirmed experimentally (data not shown). The herein proposed chemical space docking approach could yield, at least for the present case of a G protein-coupled receptor, to experimentally-validated hits with a high hit rate and nanomolar potencies that agrees with tendencies already noticed upon full atomistic docking of ultra-large library virtual screens. \u003csup\u003e10, 11\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eSpaceDock remains a relatively light computational procedure since browsing a chemical space of 100 million compounds can be achieved within 2 days on a 16-core Intel\u003csup\u003e(R)\u003c/sup\u003e Xeon\u003csup\u003e(R)\u003c/sup\u003e Silver 4210 processor. Mining the entire 5.5 billion chemical space has been made possible for the 4\u003csup\u003eth\u003c/sup\u003e international CACHE challenge\u003csup\u003e54\u003c/sup\u003e with still limited resources (1 week on 400 cores). Preliminary attempts to scan even larger chemical spaces (e.g. by adding three-component reactions) suggests that the method can be easily applied up to a trillion compounds.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eSetting-up a library of chemical reagents from fragmented protein-bound ligands.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e37,922 ligands from the sc-PDB database of druggable protein-ligand 3D structures,\u003csup\u003e31, 55\u003c/sup\u003e were fragmented using a set of 12 RECAP\u003csup\u003e56\u003c/sup\u003e-inspired retrosynthetic rules to yield 97,024 chemical reagents (\u003cstrong\u003eSupplementary\u003c/strong\u003e \u003cstrong\u003eFig. 1\u003c/strong\u003e) with standard topologies (bond length, angle bending, torsion angles) retrieved from the TRIPOS force-field.\u003csup\u003e57\u003c/sup\u003e The resulting building blocks were then filtered using the following rules: (i) IChem v.5.2. 8\u003csup\u003e45\u003c/sup\u003e detection of at least four non-covalent interactions (one of which being a ionic bond or an hydrogen-bond) with the original sc-PDB target protein, (ii) a total number of heavy atoms between 3 and 23, (iii) a total number of rotatable bonds inferior or equal to 6, (iv) a heteroatom to carbon ratio between 0.05 and 4.5, (v) no more than two fused cycles, (vi) a number of aromatic rings inferior to 3. The final library comprised 5,845 reagents (mol2 file format) derived from 4,656 unique sc-PDB ligands. Although the building blocks have not been explicitly crystallized with their target, the corresponding poses will be further annotated as \u0026quot;surrogate X-ray\u0026quot; pose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDocking sc-PDB building reagents to their cognate targets\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe above described reagents were docked to the sc-PDB target originally bound to the ligand they were derived of, after randomizing their initial orientation and dihedral angles with the Surflex\u003csup\u003e37\u003c/sup\u003e \u003cem\u003eran_archive\u003c/em\u003e routine, using 5 state-of-the-art docking tools (FlexX v.5.2.0,\u003csup\u003e33\u003c/sup\u003e GOLD v.2022,\u003csup\u003e34\u003c/sup\u003e PLANTS v1.2,\u003csup\u003e35\u003c/sup\u003e RDPSOVina v.2.0,\u003csup\u003e36\u003c/sup\u003e Surflex v.4.5.4.3\u003csup\u003e37\u003c/sup\u003e) with almost standard parameters (\u003cstrong\u003eSupplementary Tables 6-8\u003c/strong\u003e). Since the boron atom is not parametrized in some docking tools, it was replaced by either a dummy atom (FlexX, GOLD, PLANTS, Surflex) or a carbon (RDPSOvina) while keeping the trigonal planar geometry of the boronic acid unchanged. Up to 20 poses were preferentially saved in mol2 file format whenever possible (GOLD, PLANTS, Surflex), in sd file format (FlexX) or in pdbqt file format (RDPSOVina). For each docking pose, the root-mean-square deviation (rmsd) of heavy atoms to the corresponding surrogate X-ray pose was computed thanks to the Surflex \u003cem\u003erms\u003c/em\u003e routine when comparing mol2 files, or the ADFRsuite-1.0\u003csup\u003e58\u003c/sup\u003e \u003cem\u003eobrms\u003c/em\u003e routine when comparing files of different formats (mol2 vs. pdbqt, mol2 vs. sd). In addition, we measured the similarity of protein-ligand interactions between docked and X-ray poses with the IFP module of the IChem v.5.2.8 package.\u003csup\u003e45\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePreparation of bespoke chemical spaces encoded by 36 robust organic chemistry reactions\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe global stock of commercially available building blocks (250,355 compounds, sd file format, date: 2022-12-28) was downloaded from Enamine\u0026apos;s website\u003csup\u003e59\u003c/sup\u003e and filtered by catalog identification number to retain 145,707 reagents contributing to the REAL space.\u003csup\u003e43\u003c/sup\u003e Building blocks were then filtered to remove unsuitable entries as previously described.\u003csup\u003e41\u003c/sup\u003e For each of 36 different one or two-steps organic chemistry reactions (\u003cstrong\u003eSupplementary Table 2\u003c/strong\u003e), the corresponding reactants were retrieved using SMARTS strings\u003csup\u003e41\u003c/sup\u003e queries in PipelinePilot v.22.1.0.2935\u003csup\u003e60\u003c/sup\u003e (\u003cstrong\u003eSupplementary Figure 9\u003c/strong\u003e). In order to avoid side reactions, building blocks need to be monofunctional for the reactive group of interest and free of any possible poisoning chemical function for the reaction of interest (\u003cstrong\u003eSupplementary Table 2\u003c/strong\u003e). For each retained building block and possible reaction, an annotation triplet is provided: (i) reaction type, reactant role, reactive atoms. The final annotation table comprises 713,155 annotation triplets for 134,331 REAL building blocks. Selected building blocks were finally ionized at their most likely ionization state at pH 7.4 using PipelinePilot and converted into 3D atomic coordinates with Corina v.3.40,\u003csup\u003e61\u003c/sup\u003e allowing to generate up to 4 diastereoisomers by entry, in a single ready-to-dock mol2 file format.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDocking of chemical reagents to human estrogen receptor beta\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe X-ray structure of the human estrogen receptor beta in complex with the agonist genistein\u003csup\u003e62\u003c/sup\u003e was downloaded from the Protein Data Bank (PDB 1QKM). Hydrogen atoms and simultaneous optimisation of protonation states of protein, water and ligand atoms was performed with Protoss v.4.0.\u003csup\u003e63\u003c/sup\u003e All water molecule and genistein were removed, keeping only remaining protein atoms of chain A which were saved in mol2 file format. The commercial building blocks selected for a possible benzoxazole ring or sulfonamide bond formation (145 aminophenols and 3,874 benzaldehydes; 1,275 sulfonyl chlorides and 76,758 amines) were docked to the ER\u0026beta; atomic coordinates with GOLD using previously reported parameter settings (\u003cstrong\u003eSupplementary Table 7\u003c/strong\u003e). The cavity was detected from X-ray atomic coordinates of genistein. Up to 20 poses, scored by the PLP scoring function, were retained for each building block.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDocking of chemical reagents to the human dopamine D3 receptor (DRD3)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe X-ray structure of the human dopamine D3 receptor in complex with the antagonist eticlopride\u003csup\u003e51\u003c/sup\u003e was downloaded from the Protein Data Bank (PDB 3PBL). Hydrogen atoms and simultaneous optimisation of protonation states of protein, water and ligand atoms was performed with Protoss v.4.0.\u003csup\u003e63\u003c/sup\u003e The inserted T4-lysozyme sequence (Asn1002-Tyr1161), all water molecule and eticlopride were removed, keeping only remaining protein atoms of chain A which were saved in mol2 file format. The commercial building blocks were initially filtered based on their capacity to form a drug-like molecule through an amide bond formation (\u003cstrong\u003eSupplementary Table 4\u003c/strong\u003e) and their inclusion in the pool of reagents utilized in the REAL Space. The reagents selected for a possible amide bond formation (33,726 amines and 19,887 carboxylic acids) were docked to the DRD3 atomic coordinates with GOLD using previously reported parameter settings (\u003cstrong\u003eSupplementary Table 7\u003c/strong\u003e). The cavity was detected from X-ray atomic coordinates of eticlopride. Up to 20 poses, scored by the PLP scoring function, were retained for each building block. To decrease the number of possible recombinations, only docking poses of amines exhibiting an ionic bond to the key residue Asp110, detected on the fly with IChem, were further retained for amide bond formation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLigand enumeration by reagents coupling\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGiven two poses of chemically compatible reagents, a ligand is generated within the protein binding site, according to their respective location and chemical compatibility. Reagent poses are initially loaded using an in-house mol2 parser and annotated for at least one reaction based on the tag table shown in \u003cstrong\u003eFig. 2\u003c/strong\u003e. Atomic coordinates of reactive atoms carbon and their immediate neighbors, are extracted and stored for subsequent calculations. This process is repeated for each reaction, following a similar workflow. A subsequent set of filters is applied to pairs of reagent poses, including the distance between their center of mass to promptly eliminate distant pairs, the distance between connectable atoms, examination of certain angles of the future formed bond/ring to ensure a suitable geometry, and consideration of clashes (\u0026le;4 between non-reacting atoms) to prevent overlapping substituents. If a pair satisfies all the rules, a bond is created between the connectable atoms. The hybridization of reacting atoms is then updated to reflect the newly created bonds and exit atoms (to be removed after the reaction) are deleted. The fully enumerated molecule is then saved into a single mol2 file. An optional step is available at this stage. If a reference ligand exists, the molecule is initially written to a temporary mol2 file to assess its IFP similarity (default values are \u0026ge; 0.60 for all non-bonded interactions and \u0026ge; 0.50 for polar interactions) to the reference pose using IChem v.5.2.8. If the similarity threshold is reached, the molecule is transferred to the final mol2 file. Detailed rules of these filters can be found in \u003cstrong\u003eSupplementary Figs. 3, 4,\u003c/strong\u003e \u003cstrong\u003e6\u003c/strong\u003e. The fully enumerated molecule, in presence of the target protein, is last energy-minimized in Szybki v2.4.0.0,\u003csup\u003e46\u003c/sup\u003e using standard settings and the MMFF94 force-field.\u003csup\u003e64\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eComparisons to reference ligands\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eInteraction fingerprint similarity search between any pose (before and after energy refinement) and a reference X-ray ligand was done using standard parameters of the IFP module implemented in the IChem v.5.2.8 package.\u003csup\u003e45\u003c/sup\u003e Likewise, root-mean square deviations were computed with the \u003cem\u003erms\u0026nbsp;\u003c/em\u003eroutine of Surflex-Dock v.4.5.4.3.\u003csup\u003e37\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRedocking of SpaceDock poses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe coupling of two reagent poses, followed by protein constraint refinement (referred to as the \u0026quot;SpaceDock\u0026quot; pose), was redocked into the target protein structure using GOLD. The scoring function employed was PLP, with 20 generated poses, and the same parameter file as described in \u003cstrong\u003eSupplementary Table 7\u003c/strong\u003e. To eliminate structural biases, input ligand structures were converted to SMILES format using OEChem Toolkit v.3.4.0.1\u003csup\u003e46\u003c/sup\u003e and further transformed into 3D structures with Corina v.3.40.\u003csup\u003e61\u003c/sup\u003e Up to four diastereoisomers were generated in a single mol2 file. The resulting full atomistic docking pose, exhibiting a rmsd (computed with Surflex rms) below 2 \u0026Aring;, all non-bonded interactions IFP similarity \u0026ge; 0.60, and precisely the same polar IFP as the corresponding SpaceDock pose, was considered as confirmation and retained for subsequent investigations. If multiple docking poses satisfy these rules for each SpaceDock pose, all of them are retained.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eQuality check of redocked poses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe number of torsion strains in every redocking pose was estimated with TorsionAnalyzer v.2.0.0.\u003csup\u003e47\u003c/sup\u003e Any pose with at least one torsion annotated as \u0026apos;strained\u0026apos; was discarded from further analysis. Local strain (distortion of the specific conformation from the nearest local minima) and global strain (energy required to select the specific conformation from the full conformational ensemble of the corresponding compound in water) energies were then computed with standard parameter of Freeform v.2.4.0.0.\u003csup\u003e46\u003c/sup\u003e Any pose with local and global strain energies higher than 4 and 8 kcal/mol, respectively, were discarded.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLast, remaining poses were inspected, in their protein-bound state, for counting the number of unsatisfied ionic bonds, hydrogen-bond donors and acceptors. First, protein-ligand ionic and hydrogen-bonds were registered with IChem. Any charged atom or hydrogen-bond donor/acceptor atom of the ligand (according to IChem definitions)\u003csup\u003e40\u003c/sup\u003e not present in the above list was annotated as \u0026quot;unsatisfied\u0026quot; atom. Unsatisfied heavy atoms being both donors and acceptors (e.g. hydroxyl oxygen atom) were only counted once. Ligand atoms participating to intra-molecular hydrogen bonds were considered as satisfied. Altogether, ligand poses with more than 2 unsatisfied donors and 4 unsatisfied acceptors were removed from the final hit list.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSimilarity to ChEMBL and REAL Space ligands\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eKnown ligands of the human estrogen receptor beta (CHEMBL242) and human dopamine D2 (CHEMBL217) and D3 (CHEMBL234) receptors were retrieved from the ChEMBL database (release 33)\u003csup\u003e50\u003c/sup\u003e as SMILES strings for ligand entries fulfilling the following criteria: K\u003csub\u003ei\u003c/sub\u003e \u0026lt; 1\u0026nbsp;\u0026micro;M, Assay_type = B). Pairwise chemical similarity between SpaceDock hits and ChEMBL ligands was computed with PipelinePilot v.22.1.0.2935\u003csup\u003e60\u003c/sup\u003e from ECFP4 circular fingerprints and scored by the value of the Tanimoto coefficient.\u003c/p\u003e\n\u003cp\u003eMaximum common substructure (MCS) similarity of SpaceDock hits (converted from mol2 to SMILES strings, thanks to Open Babel v.3.1.0)\u003csup\u003e65\u003c/sup\u003e to 36 billion REAL space ligands (version REALSpace_36bn_2023-03.space\u003csup\u003e12\u003c/sup\u003e) was computed with SpaceMACS v.0.9.2,\u003csup\u003e15\u003c/sup\u003e to save the top 15 REAL space compounds ranked by decreasing MCS-Tanimito similarity value.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eList of reactants to build benzoxazole, sulfonamide and amide chemical spaces, docked poses of test reactants (ER\u0026beta;, DRD3 test cases), annotation table of Enamine REAL reactants, IChem configuration files for IFP filtering.\u003c/p\u003e\n\u003cp\u003eAll data and SpaceDock processing scripts are available at https://github.com/litfsindt/LIT-SpaceDock\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFilter v.4.2.1.1, Szbyki v2.5.1.1, OEChem Toolkit v.3.4.0.1; Freeform v.2.5.1.1: OpenEye Scientific, Santa Fe, N.M., USA, https://www.eyesopen.com/\u003c/p\u003e\n\u003cp\u003eFlexX v.5.2.0, Hyde v.1.5.0, SpaceMACS v.0.9.2, REAL space in fragment space format: BioSolveIT GmbH, Sankt Augustin, Germany, www.biosolveit.de\u003c/p\u003e\n\u003cp\u003eGOLD v.2022: CCDC Software Ltd., Cambridge CB2 1EZ, United Kingdom, www.ccdc.cam.ac.uk\u003c/p\u003e\n\u003cp\u003eOpen Babel v.3.1.0, https://github.com/openbabel/openbabel\u003c/p\u003e\n\u003cp\u003ePLANTS v1.2: University of Konstanz, Germany, http://www.tcd.uni-konstanz.de/research/plants.php\u003c/p\u003e\n\u003cp\u003eRDPSOVina v2.0: Jiangnan University, Jiangsu, China, https://github.com/li-jin-xing/RDPSOVina\u003c/p\u003e\n\u003cp\u003eSpaceDock v.1.0.0: https://github.com/litfsindt/LIT-SpaceDock\u003c/p\u003e\n\u003cp\u003eSurflex-Dock v4.5.4.3: BioPharmics LLC, https://www.biopharmics.com\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank Guillaume Bret (Laboratoire d\u0026apos;innovation th\u0026eacute;rapeutique) for technical assistance, Michael Bossert and Yurii Moroz (Enamine Ltd.) for sharing the list of REAL space reagents, and the CC-IN2P3 calculation center (Villeurbanne, France) for allocation of computing time and excellent support.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eContributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eD.R conceived the study. A.S. performed the initial benchmarking study on chemical reagents. M.E. designed the rules to filter sc-PDB building blocks. F.S. encoded organic chemistry reactions in 3D space, performed the whole docking of Enamine chemical reagents and wrote the SpaceDock source code to enumerate full ligands. F.S. and D.R analyzed the data. All of the authors contributed to writing and editing the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eD.R. is co-founder and shareholder of BIODOL Therapeutics. M.E. is employee of Amgen\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBleicher KH, Bohm HJ, Muller K, Alanine AI. Hit and lead generation: beyond high-throughput screening. \u003cem\u003eNat Rev Drug Discov\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 369-378 (2003).\u003c/li\u003e\n\u003cli\u003eHughes JP, Rees S, Kalindjian SB, Philpott KL. Principles of early drug discovery. \u003cem\u003eBr J Pharmacol\u003c/em\u003e \u003cstrong\u003e162\u003c/strong\u003e, 1239-1249 (2011).\u003c/li\u003e\n\u003cli\u003eLucas X, Gruning BA, Bleher S, Gunther S. The purchasable chemical space: a detailed picture. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e55\u003c/strong\u003e, 915-924 (2015).\u003c/li\u003e\n\u003cli\u003eTingle BI\u003cem\u003e, et al.\u003c/em\u003e ZINC-22 horizontal line A Free Multi-Billion-Scale Database of Tangible Compounds for Ligand Discovery. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e63\u003c/strong\u003e, 1666-1776 (2023).\u003c/li\u003e\n\u003cli\u003eGrygorenko OO, Radchenko DS, Dziuba I, Chuprina A, Gubina KE, Moroz YS. Generating Multibillion Chemical Space of Readily Accessible Screening Compounds. \u003cem\u003eiScience\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 101681 (2020).\u003c/li\u003e\n\u003cli\u003eLyu J\u003cem\u003e, et al.\u003c/em\u003e Ultra-large library docking for discovering new chemotypes. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e566\u003c/strong\u003e, 224-229 (2019).\u003c/li\u003e\n\u003cli\u003eSadybekov AA\u003cem\u003e, et al.\u003c/em\u003e Structure-Based Virtual Screening of Ultra-Large Library Yields Potent Antagonists for a Lipid GPCR. \u003cem\u003eBiomolecules\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, (2020).\u003c/li\u003e\n\u003cli\u003eStein RM\u003cem\u003e, et al.\u003c/em\u003e Virtual discovery of melatonin receptor ligands to modulate circadian rhythms. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e579\u003c/strong\u003e, 609-614 (2020).\u003c/li\u003e\n\u003cli\u003eAlon A\u003cem\u003e, et al.\u003c/em\u003e Structures of the sigma2 receptor enable docking for bioactive ligand discovery. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e600\u003c/strong\u003e, 759-764 (2021).\u003c/li\u003e\n\u003cli\u003eLyu J, Irwin JJ, Shoichet BK. Modeling the expansion of virtual screening libraries. \u003cem\u003eNat Chem Biol\u003c/em\u003e \u003cstrong\u003e19\u003c/strong\u003e, 712-718 (2023).\u003c/li\u003e\n\u003cli\u003eSadybekov AV, Katritch V. Computational approaches streamlining drug discovery. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e616\u003c/strong\u003e, 673-685 (2023).\u003c/li\u003e\n\u003cli\u003eReadily-accessible on-demand chemical spaces, https://www.biosolveit.de/infiniSee (accessed 11-16-2023)\u003c/li\u003e\n\u003cli\u003eWarr WA, Nicklaus MC, Nicolaou CA, Rarey M. Exploration of Ultralarge Compound Collections for Drug Discovery. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e62\u003c/strong\u003e, 2021-2034 (2022).\u003c/li\u003e\n\u003cli\u003eBellmann L, Penner P, Rarey M. Topological Similarity Search in Large Combinatorial Fragment Spaces. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e61\u003c/strong\u003e, 238-251 (2021).\u003c/li\u003e\n\u003cli\u003eSchmidt R, Klein R, Rarey M. Maximum Common Substructure Searching in Combinatorial Make-on-Demand Compound Spaces. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e62\u003c/strong\u003e, 2133-2150 (2022).\u003c/li\u003e\n\u003cli\u003eMeyenburg C, Dolfus U, Briem H, Rarey M. Galileo: Three-dimensional searching in large combinatorial fragment spaces on the example of pharmacophores. \u003cem\u003eJ Comput Aided Mol Des\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 1-16 (2023).\u003c/li\u003e\n\u003cli\u003eGentile F\u003cem\u003e, et al.\u003c/em\u003e Deep Docking: A Deep Learning Platform for Augmentation of Structure Based Drug Discovery. \u003cem\u003eACS Cent Sci\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 939-949 (2020).\u003c/li\u003e\n\u003cli\u003eBerenger F, Kumar A, Zhang KYJ, Yamanishi Y. Lean-Docking: Exploiting Ligands\u0026apos; Predicted Docking Scores to Accelerate Molecular Docking. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e61\u003c/strong\u003e, 2341-2352 (2021).\u003c/li\u003e\n\u003cli\u003eGraff DE, Aldeghi M, Morrone JA, Jordan KE, Pyzer-Knapp EO, Coley CW. Self-Focusing Virtual Screening with Active Design Space Pruning. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e62\u003c/strong\u003e, 3854-3862 (2022).\u003c/li\u003e\n\u003cli\u003eSadybekov AA\u003cem\u003e, et al.\u003c/em\u003e Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e601\u003c/strong\u003e, 452-459 (2022).\u003c/li\u003e\n\u003cli\u003eMuller J\u003cem\u003e, et al.\u003c/em\u003e Magnet for the Needle in Haystack: \u0026quot;Crystal Structure First\u0026quot; Fragment Hits Unlock Active Chemical Matter Using Targeted Exploration of Vast Chemical Spaces. \u003cem\u003eJ Med Chem\u003c/em\u003e \u003cstrong\u003e65\u003c/strong\u003e, 15663-15678 (2022).\u003c/li\u003e\n\u003cli\u003eBeroza P\u003cem\u003e, et al.\u003c/em\u003e Chemical space docking enables large-scale structure-based virtual screening to discover ROCK1 kinase inhibitors. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 6447 (2022).\u003c/li\u003e\n\u003cli\u003eGorgulla C\u003cem\u003e, et al.\u003c/em\u003e An open-source drug discovery platform enables ultra-large virtual screens. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e580\u003c/strong\u003e, 663-668 (2020).\u003c/li\u003e\n\u003cli\u003eGadioli D\u003cem\u003e, et al.\u003c/em\u003e EXSCALATE: An extreme-scale in-silico virtual screening platform to evaluate 1 trillion compounds in 60 hours on 81 PFLOPS supercomputers.\u003cem\u003e arXiv:211011644v1\u003c/em\u003e, (2021).\u003c/li\u003e\n\u003cli\u003eNeumann A, Marrison L, Klein R. Relevance of the Trillion-Sized Chemical Space \u0026quot;eXplore\u0026quot; as a Source for Drug Discovery. \u003cem\u003eACS Med Chem Lett\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 466-472 (2023).\u003c/li\u003e\n\u003cli\u003eHartenfeller M\u003cem\u003e, et al.\u003c/em\u003e A collection of robust organic synthesis reactions for in silico molecule design. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e51\u003c/strong\u003e, 3093-3098 (2011).\u003c/li\u003e\n\u003cli\u003ePenner P\u003cem\u003e, et al.\u003c/em\u003e FastGrow: on-the-fly growing and its application to DYRK1A. \u003cem\u003eJ Comput Aided Mol Des\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 639-651 (2022).\u003c/li\u003e\n\u003cli\u003eSivula T, Yetukuri L, Kalliokoski T, Kasnanen H, Poso A, Pohner I. Machine Learning-Boosted Docking Enables the Efficient Structure-Based Virtual Screening of Giga-Scale Enumerated Chemical Libraries. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e63\u003c/strong\u003e, 5773-5783 (2023).\u003c/li\u003e\n\u003cli\u003eRoggia M, Natale B, Amendola G, Di Maro S, Cosconati S. Streamlining Large Chemical Library Docking with Artificial Intelligence: the PyRMD2Dock Approach. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e, https://doi.org/10.1021/acs.jcim.1023c00647 (2023).\u003c/li\u003e\n\u003cli\u003eGentile F\u003cem\u003e, et al.\u003c/em\u003e Automated discovery of noncovalent inhibitors of SARS-CoV-2 main protease by consensus Deep Docking of 40 billion small molecules. \u003cem\u003eChem Sci\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 15960-15974 (2021).\u003c/li\u003e\n\u003cli\u003eDesaphy J, Bret G, Rognan D, Kellenberger E. sc-PDB: a 3D-database of ligandable binding sites--10 years on. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e43\u003c/strong\u003e, D399-404 (2015).\u003c/li\u003e\n\u003cli\u003eBostrom J, Brown DG, Young RJ, Keseru GM. Expanding the medicinal chemistry synthetic toolbox. \u003cem\u003eNat Rev Drug Discov\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 709-727 (2018).\u003c/li\u003e\n\u003cli\u003eRarey M, Kramer B, Lengauer T, Klebe G. A fast flexible docking method using an incremental construction algorithm. \u003cem\u003eJ Mol Biol\u003c/em\u003e \u003cstrong\u003e261\u003c/strong\u003e, 470-489 (1996).\u003c/li\u003e\n\u003cli\u003eJones G, Willett P, Glen RC, Leach AR, Taylor R. Development and validation of a genetic algorithm for flexible docking. \u003cem\u003eJ Mol Biol\u003c/em\u003e \u003cstrong\u003e267\u003c/strong\u003e, 727-748 (1997).\u003c/li\u003e\n\u003cli\u003eKorb O, Stutzle T, Exner TE. Empirical scoring functions for advanced protein-ligand docking with PLANTS. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, 84-96 (2009).\u003c/li\u003e\n\u003cli\u003eLi J, Li C, Sun J, Palade V. RDPSOVina: the random drift particle swarm optimization for protein-ligand docking. \u003cem\u003eJ Comput Aided Mol Des\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 415-425 (2022).\u003c/li\u003e\n\u003cli\u003eJain AN. Surflex-Dock 2.1: robust performance from ligand energetic modeling, ring flexibility, and knowledge-based search. \u003cem\u003eJ Comput Aided Mol Des\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 281-306 (2007).\u003c/li\u003e\n\u003cli\u003eChachulski L, Windshugel B. LEADS-FRAG: A Benchmark Data Set for Assessment of Fragment Docking Performance. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e60\u003c/strong\u003e, 6544-6554 (2020).\u003c/li\u003e\n\u003cli\u003eVerdonk ML, Giangreco I, Hall RJ, Korb O, Mortenson PN, Murray CW. Docking performance of fragments and druglike compounds. \u003cem\u003eJ Med Chem\u003c/em\u003e \u003cstrong\u003e54\u003c/strong\u003e, 5422-5431 (2011).\u003c/li\u003e\n\u003cli\u003eMarcou G, Rognan D. Optimizing fragment and scaffold docking by use of molecular interaction fingerprints. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e47\u003c/strong\u003e, 195-207 (2007).\u003c/li\u003e\n\u003cli\u003eHartenfeller M\u003cem\u003e, et al.\u003c/em\u003e DOGS: reaction-driven de novo design of bioactive compounds. \u003cem\u003ePLoS Comput Biol\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, e1002380 (2012).\u003c/li\u003e\n\u003cli\u003eSommer K, Flachsenberg F, Rarey M. NAOMInext - Synthetically feasible fragment growing in a structure-based design context. \u003cem\u003eEur J Med Chem\u003c/em\u003e \u003cstrong\u003e163\u003c/strong\u003e, 747-762 (2019).\u003c/li\u003e\n\u003cli\u003eMoroz YS. 2022q3-4 REAL database reagents, personal communication. (ed^(eds) (2023).\u003c/li\u003e\n\u003cli\u003eMalamas MS\u003cem\u003e, et al.\u003c/em\u003e Design and synthesis of aryl diphenolic azoles as potent and selective estrogen receptor-beta ligands. \u003cem\u003eJ Med Chem\u003c/em\u003e \u003cstrong\u003e47\u003c/strong\u003e, 5021-5040 (2004).\u003c/li\u003e\n\u003cli\u003eDa Silva F, Desaphy J, Rognan D. IChem: A Versatile Toolkit for Detecting, Comparing, and Predicting Protein-Ligand Interactions. \u003cem\u003eChemMedChem\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 507-510 (2018).\u003c/li\u003e\n\u003cli\u003eOpenEye Scientific Software, Sante Fe, NM, U.S.A. (ed^(eds).\u003c/li\u003e\n\u003cli\u003ePenner P, Guba W, Schmidt R, Meyder A, Stahl M, Rarey M. The Torsion Library: Semiautomated Improvement of Torsion Rules with SMARTScompare. \u003cem\u003eJ Chem Inf Model\u003c/em\u003e \u003cstrong\u003e62\u003c/strong\u003e, 1644-1653 (2022).\u003c/li\u003e\n\u003cli\u003eSchneider N, Lange G, Hindle S, Klein R, Rarey M. A consistent description of HYdrogen bond and DEhydration energies in protein-ligand complexes: methods behind the HYDE scoring function. \u003cem\u003eJ Comput Aided Mol Des\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 15-29 (2013).\u003c/li\u003e\n\u003cli\u003eFischer A, Smiesko M, Sellner M, Lill MA. Decision Making in Structure-Based Drug Discovery: Visual Inspection of Docking Results. \u003cem\u003eJ Med Chem\u003c/em\u003e \u003cstrong\u003e64\u003c/strong\u003e, 2489-2500 (2021).\u003c/li\u003e\n\u003cli\u003ehttps://www.ebi.ac.uk/chembl/ (accessed 11-16-2023).\u003c/li\u003e\n\u003cli\u003eChien EY\u003cem\u003e, et al.\u003c/em\u003e Structure of the human dopamine D3 receptor in complex with a D2/D3 selective antagonist. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e330\u003c/strong\u003e, 1091-1095 (2010).\u003c/li\u003e\n\u003cli\u003eMaramai S\u003cem\u003e, et al.\u003c/em\u003e Dopamine D3 Receptor Antagonists as Potential Therapeutics for the Treatment of Neurological Diseases. \u003cem\u003eFront Neurosci-Switz\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 451 (2016).\u003c/li\u003e\n\u003cli\u003eHopkins AL, Groom CR, Alex A. Ligand efficiency: a useful metric for lead selection. \u003cem\u003eDrug Discov Today\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 430-431 (2004).\u003c/li\u003e\n\u003cli\u003eCACHE challenge: Critical assessment of computational hit-finding experiments. https://cache-challenge.org/ , (accessed 11-16-2023)\u003c/li\u003e\n\u003cli\u003esc-PDB: An Annotated Database of Druggable Binding Sites from the Protein DataBank. http://bioinfo-pharma.u-strasbg.fr/scPDB/, (accessed 11-16-2023)\u003c/li\u003e\n\u003cli\u003eLewell XQ, Judd DB, Watson SP, Hann MM. RECAP--retrosynthetic combinatorial analysis procedure: a powerful new technique for identifying privileged molecular fragments with useful applications in combinatorial chemistry. \u003cem\u003eJ Chem Inf Comput Sci\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, 511-522 (1998).\u003c/li\u003e\n\u003cli\u003eClark M, Cramer RD, III., Van Opdenbosch N. Validation of the general purpose tripos 5.2 force field. \u003cem\u003eJ Comput Chem\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 982-1012 (1989).\u003c/li\u003e\n\u003cli\u003eADFR software suite downloads, https://ccsb.scripps.edu/adfr/downloads/ (accessed 11-16-2023).\u003c/li\u003e\n\u003cli\u003eEnamine building blocks catalog. https://enamine.net/building-blocks/building-blocks-catalog (accessed 03-25-2023)\u003c/li\u003e\n\u003cli\u003eDassault Syst\u0026egrave;mes Biovia Corp., San Diego, CA. https://www.3ds.com/products-services/biovia/products/data-science/pipeline-pilot/\u003c/li\u003e\n\u003cli\u003eMolecular Networks GmbH, N\u0026uuml;rnberg, Germany. https://mn-am.com/products/corina/ (accessed 03-25-2023)\u003c/li\u003e\n\u003cli\u003ePike AC\u003cem\u003e, et al.\u003c/em\u003e Structure of the ligand-binding domain of oestrogen receptor beta in the presence of a partial agonist and a full antagonist. \u003cem\u003eEMBO J\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 4608-4618 (1999).\u003c/li\u003e\n\u003cli\u003eBietz S, Urbaczek S, Schulz B, Rarey M. Protoss: a holistic approach to predict tautomers and protonation states in protein-ligand complexes. \u003cem\u003eJ Cheminform\u003c/em\u003e \u003cstrong\u003e6\u003c/strong\u003e, 12 (2014).\u003c/li\u003e\n\u003cli\u003eHalgren TA. Merck molecular force field. I. Basis, form, scope, parameterization, and performance of MMFF94. \u003cem\u003eJ Comput Chem\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 490-519 (1996).\u003c/li\u003e\n\u003cli\u003eO\u0026apos;Boyle NM, Banck M, James CA, Morley C, Vandermeersch T, Hutchison GR. Open Babel: An open chemical toolbox. \u003cem\u003eJ Cheminform\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 33 (2011).\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Table","content":"\u003cp\u003eTable 2 is available in the Supplementary Files section.\u003c/p\u003e "},{"header":"Supplementary Information","content":"\u003cp\u003eSupplementary Figures and Supplementary Tables are not available with this version.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3687338/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3687338/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eUltra-large chemical spaces describing several billion compounds are revolutionizing hit identification in early drug discovery. Because of their size, such chemical spaces cannot be fully enumerated and requires ad-hoc computational tools to navigate them and pick potentially interesting hits. We here propose a structure-based approach to ultra-large chemical space screening in which commercial chemical reagents are first docked to the target of interest and then directly connected according to organic chemistry and topological rules, to enumerate drug-like compounds under three-dimensional constraints of the target. When applied to bespoke chemical spaces of different sizes and chemical complexity targeting two receptors of pharmaceutical interest, the computational method was able to quickly enumerate hits that were either known ligands (or very close analogs) of targeted receptors as well as chemically novel candidates that could be experimentally confirmed by \u003cem\u003ein vitro \u003c/em\u003ebinding assays. The proposed approach is generic, can be applied to any docking algorithm and requires few computational resources to prioritize easily synthesizable hits from billion-sized chemical spaces.\u003c/p\u003e","manuscriptTitle":"Protein structure-based organic chemistry-driven ligand design from ultra-large chemical spaces","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-01-09 19:51:37","doi":"10.21203/rs.3.rs-3687338/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c992b067-365d-4c2b-a334-050c6236a450","owner":[],"postedDate":"January 9th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":27167092,"name":"Health sciences/Medical research/Drug development"},{"id":27167093,"name":"Biological sciences/Computational biology and bioinformatics/High-throughput screening"}],"tags":[],"updatedAt":"2024-01-09T19:51:37+00:00","versionOfRecord":[],"versionCreatedAt":"2024-01-09 19:51:37","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3687338","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3687338","identity":"rs-3687338","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.