A tailored enzyme cascade facilitates DNA-encoded library technology and gives access to a broad substrate scope

preprint OA: gold CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract DNA-encoded chemical library (DEL) technology has emerged as a powerful tool in early-stage drug discovery. Although the methodology is widely applied in industry and academia, major challenges persist in generating DELs with high quality and chemical diversity. Low yields in building-block incorporation, limited stereo-, regio-, and chemoselectivity, and, most importantly, DNA damage from harsh reaction conditions compromise library quality, reduce signal-to-noise in affinity selections, and ultimately hinder drug discovery. Here, we show that tailored enzymes can be harnessed for the effective construction of molecular diversity on DNA, opening new avenues for the generation of high-quality and diverse DELs under mild conditions. Targeting amide bond formation, we engineered a cascade of complementary CoA ligases and rationally tailored N-acyltransferases (NATs) to access a broad amide scope on-DNA (> 120 examples), identifying transferable structural motifs that optimize DNA-compatibility of the biocatalysts in the process. The successful integration of the biocatalytic steps with chemical synthesis led to the construction of a diverse DEL without damage to the DNA barcode and underscored the broad utility of the enzymatic cascade, enabling applications in both early scaffold construction and late-stage functionalization.
Full text 162,205 characters · extracted from preprint-html · click to expand
A tailored enzyme cascade facilitates DNA-encoded library technology and gives access to a broad substrate scope | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A tailored enzyme cascade facilitates DNA-encoded library technology and gives access to a broad substrate scope Rebecca Buller, Daniela Schaub, Alice Lessing, Fabian Meyer, Michael Eichenberger, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7598475/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract DNA-encoded chemical library (DEL) technology has emerged as a powerful tool in early-stage drug discovery. Although the methodology is widely applied in industry and academia, major challenges persist in generating DELs with high quality and chemical diversity. Low yields in building-block incorporation, limited stereo-, regio-, and chemoselectivity, and, most importantly, DNA damage from harsh reaction conditions compromise library quality, reduce signal-to-noise in affinity selections, and ultimately hinder drug discovery. Here, we show that tailored enzymes can be harnessed for the effective construction of molecular diversity on DNA, opening new avenues for the generation of high-quality and diverse DELs under mild conditions. Targeting amide bond formation, we engineered a cascade of complementary CoA ligases and rationally tailored N-acyltransferases (NATs) to access a broad amide scope on-DNA (> 120 examples), identifying transferable structural motifs that optimize DNA-compatibility of the biocatalysts in the process. The successful integration of the biocatalytic steps with chemical synthesis led to the construction of a diverse DEL without damage to the DNA barcode and underscored the broad utility of the enzymatic cascade, enabling applications in both early scaffold construction and late-stage functionalization. Biological sciences/Chemical biology/Biocatalysis Biological sciences/Chemical biology/Enzymes Biological sciences/Chemical biology/Chemical libraries/Combinatorial libraries Biocatalysis DNA-encoded library Protein engineering Amide-bond formation Chemoenzymatic cascade Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction The discovery and development of novel therapeutic modalities remains both time-consuming and costly. 1,2 While technologies for the generation of specific antibodies are well established, 3–5 the identification of high-affinity and specific small-molecule drugs continues to pose a major challenge. To address this bottleneck, barcoded compound collections, most notably DNA-encoded Chemical Libraries (DELs), 6 have emerged as powerful and cost-efficient discovery platforms. 7–10 Today, DEL technology is widely applied across academia and industry, including in nearly all major pharmaceutical companies, where it is fuelling clinical pipelines. 11–14 DELs are generated through iterative split-and-pool combinatorial synthesis, in which each chemical reaction step is recorded by ligating a unique DNA tag. 15–17 This strategy enables the creation of vast libraries, typically ranging from 10 6 - 10 9 library members. These libraries can be screened in parallel against protein targets of pharmaceutical interest by affinity-based selection ( Figure 1a ). 10,18–20 Enriched binders are then identified by PCR amplification of the DNA tags and next-generation DNA sequencing. 21 Since its introduction in the early 2000s, DNA-encoded library technology has been continuously improved and delivered numerous clinical drug candidates ( Figure 1b ). 11,22–24 Nevertheless, fundamental limitations persist. The unavoidable presence of DNA during library synthesis imposes strict constraints on reaction conditions, as base modifications or even complete degradation of the DNA tag can occur, leading to code alterations or loss of PCR amplifiability. 25,26 Consequently, chemical transformations must always be optimized for DNA compatibility and only a restricted subset of organic chemical conversions can, to date, be used for synthesizing productive DELs. 27 Recent innovations, such as solid-phase based synthesis to circumvent aqueous conditions or micellar reaction compartmentalization, have expanded the toolkit, 28,29 but the identification of high-yielding, diversity-generating, and DNA-compatible transformations remains the central challenge in DEL technology. 30 Currently, DEL construction is limited to a low number of split-and-pool cycles (typically 2-4) as otherwise the accumulation of truncated products 31,32 rapidly degrades library quality and complicates hit identification during affinity selections by reducing signal-to-noise ratios. 28 To address these limitations, we turned our attention to biocatalysis-aided DEL synthesis, as mild enzymatic transformations hold the promise to be DNA-preserving while simultaneously providing products with excellent chemo-, regio- and stereoselectivity and high yields. Intrigued by early examples describing the use of wildtype enzymes for on-DNA aldol reactions (threonine aldolases) 33 and carbohydrate couplings (glycosyltransferases) 34 we hypothesized that the development of engineered biocatalysts could broaden the hitherto moderate substrate spectrum and improve product yields, providing more effective tools for building molecular diversity in the presence of a DNA tag and leading to enhanced DEL quality and scope. In addition, we envisioned that the tailor-made enzymes could easily be combined with existing chemical methods and, theoretically, be employed in any of the consecutive split-and-pool cycles needed for library synthesis. As the value of a reaction for DEL synthesis is largely determined by its applicability across a wide range of substrates, 35 transformations that make use of commercially available building blocks are especially attractive. 36 Analyses of supplier catalogues show that amines and carboxylic acids are by far the most widely available building blocks, followed by less diverse and chemically less stable compound classes such as aryl halides, aldehydes and ketones. 37 Consequently, investigating the potential of enzymatic amide-bond formation on DNA was considered particularly valuable, given that this transformation represents a fundamental reaction in pharmaceutical chemistry and underlies the synthesis of a broad spectrum of small-molecule therapeutics. 38 While chemical procedures to couple carboxylic acids and amines have been established in DEL technology, these methods often require large molar excesses of building blocks and coupling reagents, lead to varying coupling efficiency, and potentially damage the DNA tag. 25,39 By contrast, enzymatic coupling of carboxylic acids and amines can generate amides with excellent selectivity from a diverse set of amines and carboxylic acids under mild reaction conditions. 40–43 When assessing the suitability (based on substrate and product scope as well as reaction conditions) of amide-bond forming enzymes for DEL technology (e.g., ATP-dependent amide bond forming enzymes 41,42,44 ; lipases 45 ; nitrile hydratases 46 ), we were particularly intrigued by the versatility of an enzymatic cascade consisting of CoA ligases (CLs) and N -acyltransferases (NATs). 40 We envisioned that by applying such a cascade, we could harness complementary wildtype CoA ligases to activate a broad set of carboxylic acids in an ATP-dependent manner, while only the N -acyltransferases, applied in the amide-forming step, would need to be specifically tailored for on-DNA chemistry ( Figure 1c ). 47,48 Here, we show how to implement this concept by designing an enzymatic cascade consisting of CoA ligases and engineered NATs, giving access to a wide portfolio of on-DNA amides (>120 examples). The enzymatic method affords excellent yields on-DNA (on average > 91% conversion in the enzymatically constructed DEL) while fully preserving the integrity of the encoding DNA tag. Applying a structure-guided engineering approach as well as a rational exchange of structural elements (loop swap), we identified key motifs important for on-DNA chemistry that could be transferred to distant enzyme scaffolds (27 % sequence identity) effectively improving on-DNA activity. Combining the enzymatic cascade with chemical steps, a DEL consisting of 72individual compounds was obtained, highlighting the versatility of our biocatalysis-aided DEL construction approach, which can be applied for initial scaffold building as well as for late-stage functionalization. Results Enzyme selection and assessment of substrate scope To enable access to broad molecular diversity in the presence of DNA, we selected seven wildtype CoA ligases and N -acyltransferases based on their reported substrate scope 40,49 and sequence diversity ( Table S1, Figure S1 ). For the NATs in particular, the selection encompassed several different families (BAHD family, GNAT family, and arylamine N -acyltransferases), each characterized by a distinct catalytic mechanism (BAHD family: HXXXD motif; GNAT family: catalytic tyrosine; arylamine N -acyltransferases: Cys-His-Asp motif) and protein architecture (CATh IDs 50 : Chloramphenicol acetyltransferase-like domain (BAHD): 3.30.559.10, GNAT: 3.40.630.30, N -arylamine acyltransferase: 3.30.2140.10). 48,51,52 Synthetic genes were cloned into a pET28b vector containing an N -terminal His-tag ( Table S2 ), transformed into Escherichia coli ( E. coli ) BL21 (DE3), and expressed to produce the corresponding enzyme ( Figure S2 ). In initial activity tests, crude enzyme preparations were used to evaluate whether modifications to reported amine substrates, i.e. the addition of synthetic handles intended for the subsequent coupling to the DNA barcodes (carboxylic acid or methyl ester moieties), would be tolerated by the N -acyltransferases ( Figure S3 ). In detail, we screened our NAT panel in the presence of CoA ligase ipfF from Sphingomonas Ibu-2 (UniProt: A1E027)for the amide bond formation between 4-fluorophenylacetic acid ( 1 )and freestanding amines A-D . Product formation was monitored using liquid chromatography coupled to mass spectrometry (LC-MS) by following the m/z ratios of the expected products ( Figure S3 ). For enzymatic cascades including N -acyltransferase 05PaAT from Pseudomonas aeruginosa (UniProt: Q9HUY3) or 42GmAT from Gibberella moniliformis (UniProt: B7SP66), successful formation of products 1A - 1D or 1A - 1C was observed, respectively, providing opportunities for further substrate modification. In addition, a side product with a mass corresponding to the acylated amine was detected in some reactions, likely resulting from a competing amide bond forming reaction with acetyl-CoA from the lysate. All lysate reactions yielding product formation were validated using assays with enzymes purified by Ni-NTA affinity chromatography ( Table S3, Figure S4 ). Building on thesuccessfulresults from the amine screening, on-DNA substrates were constructed. Examples from DEL 33,53 and PROTAC 54 literature highlight that the nature of the linker moiety can be decisive for ligand binding and recognition; particularly, length and hydrophobicity of the chosen tether can have an impact on affinity. 33 For example, ligands conjugated to a double stranded headpiece-DNA were shown to exhibit lower binding affinity against target proteins compared to the same ligands conjugated to a single-stranded DNA via a C6-linker. 53,55 To explore this effect and ensure to work with the most suitable linker architecture, molecular tethers of varying length and properties (C6, C12, and TEG) were installed at the 5'-end of a 14-mer oligonucleotide (GGA CGG GCG GCA CA-3’) and the resulting DNA tags were coupled to amines A , C and D ( Figure S5-S7, Table S4 ). The utilized oligomer was designed to prevent the formation of secondary structures such as hairpins or self-dimers and to form a stable duplex with a melting temperature of 59.4°C. To test the influence of double stranded DNA on enzyme performance, the single-stranded conjugates attached to molecule C were hybridized with the complementary strand (T m = 59.4°C, Figure S8 ). Proof-of-concept biocatalytic on-DNA reactions were conducted with the established enzymatic cascade consisting of purified CoA ligase ipfF (c = 20 µM) and N -acyltransferase 05PaAT or 42GmAT (c = 20 µM), using 0.1 mM carboxylic acid 1 and 0.05 mM DNA-conjugated amines ( a - l ) in aqueous Tris buffer (pH 8) at 25 °C. To closely mimic DEL synthesis conditions already from this early development stage, the reactions were carried out in a volume of 20 µL on a nanomolar scale. To our delight, successful product formation was observed for DNA-conjugated amines a-I and k , and, encouragingly, up to 95 % conversion was achieved for amides 1b and 1c - even under non-optimized reaction conditions ( Figure 1 d ). Further analysing the on-DNA experiments, we found that longer, hydrophobic linkers worked better in conjunction with the active N -acyltransferases. Reactions containing C12 mostly showed higher conversions than those with C6- or TEG-linked amines ( Figure 1d ), giving particularly good results with NAT 42GmAT. Importantly, we did not observe any DNA degradation in the recorded chromatograms underscoring the DNA-compatibility of the enzymatic reactions, while the negative control (without any enzyme) allowed us to exclude the occurrence of non-enzymatic background ( Figure S9 ). Notably, we found that the type of DNA oligomer used (double-stranded vs. single-stranded) had a clear effect on enzyme performance. For example, N -acyltransferase 42GmAT exhibited only 44 % relative activity for amines tethered to double-stranded DNA ( e , g , i ) compared to the identical substrates tagged with single stranded oligonucleotide ( d , f , h ) ( Figure 1 d ). We hypothesized that this effect might be due to interactions of the enzymes with unpaired bases within the single-stranded DNA tag, facilitating substrate positioning. To explore whether the utilized DNA sequence could be used to modulate enzyme activity, we replaced guanine - serving as the first base in the standard single-stranded 14mer oligo - with adenine, cytosine, and thymine, and coupled the resulting DNA strands via a C12 linker onto amine C . Cascade reactions with CoA ligase ipfF and NAT 42GmAT with this customized substrate set highlighted a preference of the enzyme for the purine bases guanine (100 % relative activity), followed by adenine (85.8 % relative activity) and the pyrimidine cytosine (83.8 % relative activity), while thymine was less well accepted (47.9 % relative activity) ( Extended Data Figure 1 ). Taken together, these results pointed to exploitable interactions between enzyme and DNA tag and identified the single-stranded 14mer oligo (GGA CGG GCG GCA CA-3’) coupled via its 5’ end to a C12 linker as the optimal set-up for all further enzyme engineering, substrate scoping and DEL construction experiments. Chimera Engineering Even though being consistently less active in the off-DNA amine screening ( Figure S4 ), our experimental data ( Figure 1d ) demonstrated that 42GmAT generally outperformed 05PaAT in on-DNA amide-coupling reactions, triggering us to conduct structural analyses to uncover underlying principles. Interestingly, both N -acyltransferases belong to the family of arylamine N -acyltransferases and are characterized by relatively large and accessible binding pockets. Sequence-independent protein structure comparison (TM-align 56 ) of the experimental structures PDB 7QI3 (42GmAT) and 1W4T (05PaAT) showed high structural agreement between the active NATs (TM-score = 0.73) despite low sequence similarity (27 %, Figure S1 ). Using pyKVFinder 57 , a python-based tool for the detection and characterization of protein cavities, we found, however, that 42GmAT featured a larger pocket volume (1538 Å 3 ) than05PaAT (1068 Å 3 ) and, concomitantly was characterized by a more spacious access tunnel (42GmAT: 1104 Å 2 vs 05PaAT: 811 Å 2 ) suggesting binding site accessibility as a potential determinant for 42GmAT’s improved on-DNA performance ( Figure 2a ). Figure 2 | Engineering N -acyltransferase 42GmAT. a , Representation of the loop swap between N -acyltransferase 05PaAT and 42GmAT. Structure and sequence of the exchanged protein motif is highlighted in purple. b, Fold improvement over parent (FIOP) of 05PaAT chimera and 42GmAT chimera when screened for the amide bond forming reaction between carboxylic acid 1 and DNA-conjugated amine f to the DNA-conjugated product 1f . c , Overall structure of NAT 42GmAT (PDB: 7QI3) and a zoom-in view of the active site highlighting the substrate entrance tunnel (pink mesh), which was identified using the bioinformatic tool CAVER. 58 Residues selected for sites-saturation mutagenesis are shown as coloured spheres, while the catalytic triad is shown as grey sticks (C110, H158, D173). d, Fold improvement over parent (FIOP) of 42GmAT variants screened for the amide-bond forming reaction between the screening substrate E (no DNA tag) and carboxylic acid 1 . Since the terminal ester group of screening substrate E was susceptible to hydrolysis in E. coli lysate (empty vector control: 90 % hydrolysis of substrate), only the major hydrolysed product was considered for calculations. Top performing variants are annotated. e, HPLC-chromatograms of amide-bond forming reactions between carboxylic acid 1 and DNA-conjugated amine f to the DNA-conjugated product 1f by best-performing 42GmAT variants and the wildtype enzyme. The accompanying barplot shows the corresponding enzymatic conversion to product 1f (c Enzyme = 10 µM). Variant data in c , e are the mean values ± s.d. of three technical replicates Analysing the structures further, we additionally found that 42GmAT’s binding pocket architecture is partially defined by a helix–loop-helix motif (residues 136–150) spanning the active site ( Figure 2 a ). This structural element is absent in the less DNA-accepting N -acyltransferase 05PaAT. Here, the active site is covered by a flexible loop (residues 99–105). Intriguingly, helix-turn-helix motifs are structural elements often found in DNA-binding domains, such as in transcription factors. 59 Thus, to explore this element’s role in on-DNA catalysis, we performed a loop swap between 42GmAT and 05PaAT. Both desired chimeric enzymes could be obtained in soluble form ( Figure S12 ) and were tested in on-DNA amide bond forming reactions. Strikingly, the 05PaAT-based chimera incorporating the helix–loop–helix motif from 42GmAT exhibited a 17-fold improvement over wildtype for the amide bond formation between carboxylic acid 1 and on-DNA amine f , whereas the activity of the 42GmAT-based chimera harbouring the 05PaAT loop dropped by 20-fold ( Figure 2b ). Applying Boltz-2, 60 a foundational model capable of co-folding protein structures with ligands such as DNA and RNA, we modelled the two wildtype NATs and their corresponding chimeras in the presence of the DNA-conjugated amine f ( Extended Data Figure 2 ). In the 42GmAT wildtype model, several stabilizing interactions between the protein and the first base of the 14-mer oligonucleotide were detected: Residue R146, which is part of the helix-loop-helix element, as well as residue R306 formed salt bridges with the DNA backbone, whereas amino acid Y229 engaged in p-p stacking with the initial guanine ( Extended Data Figure 2b ). Such stabilizing interactions were reduced in the 42GmAT chimera, where the salt bridge with a loop amino acid and the p-p interaction with the first base were no longer observed ( Extended Data Figure 2c ). Analogously, the binary model of wildtype 05PaAT did not show any interactions between its native loop and the DNA-bound substrate ( Extended Data Figure 2e ), while in the structure of the more active chimeric enzyme a salt bridge between residue R109 (analogous to R146 in 42GmAT) and the DNA backbone was detected ( Extended Data Figure 2f ), potentially facilitating on-DNA catalysis. 42GmAT engineering to improve on-DNA performance Considering the superior activity of NAT 42GmAT for DNA-tagged amines ( Figure 1d ), we opted to tailor this enzyme for improved activity for on-DNA amide bond formation. Guided by our structural analysis 57,58 and the experimental results highlighting the importance of linker-type and initial DNA base, we chose four positions (M139, P141, S178, M179) at the tunnel entrance potentially in contact with the linker/DNA as targets for full site-saturation ( Figure 2 c ). The libraries were constructed using NNK codon containing primers ( Table S5 ) and full genes were assembled by overlap extension polymerase chain reactions (PCRs). 61 To ensure sufficient oversampling of the site-saturation libraries, 62 each randomized position was screened on a 96-well microtiter plate, including the wildtype enzyme and appropriate negative controls. Due to the DNA-tagged substrates’ limited stability in E. coli lysates ( Figure S10 ), the screening work was carried out with substrate E , dubbed “screening substrate”, in which the target amine was linked to the C12 linker but not the DNA tag ( Figure 2 d and Figure S11 ). Following the biocatalytic reactions, fold improvement over parent (FIOP) was quantified for each variant by comparing relative amide product formation in individual wells to that of the wild-type enzyme. This analysis identified several improved variants from positions M139, M179, and P141, with FIOPs exceeding 5 (e.g. M139F, M139Y, and M179W), whereas more modest improvements (FIOPs ~ 2) were observed for variants M139L and P141I. Notably, none of the variants from the library S178NNK exhibited improved amide bond forming activity ( Figure 2 d ). In a next step, the plasmids harbouring the five hit variants (M139L, M139F, M139Y, P141I, M179W) were sequenced, retransformed into E. coli BL21 (DE3) and the corresponding enzymes were produced and purified ( Figure S12 ). When testing these enzyme preparations with on-DNA amine f (reduced enzyme load of 10 µM), improved activity of variants M139Y (FIOP 2.4), M139F (FIOP 2.3) and M179W (FIOP 2.4) was confirmed ( Figure 2 e ). A further combination of successful mutations at position 139 (L, F, Y) with M179W, however, did not result in any improvement beyond that observed for the single mutants ( Table S8 ). Carboxylic acid scope of the enzymatic cascade To supplement the improved on-DNA activity of tailored N -acyltransferase 42GmAT M139Y ( Table S9 ), we opted to explore the substrate scope of CoA ligases ipfF from Sphingomonas sp. Ibu-2 and complementary LcCL from Leptothrix cholodnii (UniProt: B1Y4L8) with the goal of introducing additional molecular diversity on DNA. Going beyond known examples, 40 we compiled a customized carboxylic acid collection containing a balanced mixture of acids with high (n=86) or low (n=50) structural similarity to the reported ipfF substrate scope 40 ( Extended Data Figure 3a and b ) as measured by the Tanimoto index. 63 Importantly, the collection also comprised carboxylic acids (n=100) which were previously reported as demanding building blocks for on-DNA amide bond formation, each giving less than 50 % yield in chemical coupling reactions using DMT-MM or EDC/HOAt/DIPEA. 64 This led to a library consisting of 236 acids ( Extended Data Figure 3a ), which we tested with purified ipfF and LcCL using a spectrophotometric pyrophosphate detection assay (Malachite Green assay), which allows to quantify enzyme activity by measuring the free phosphate concentration ( Figure 3 a, Extended Data Figure 3c ). The colorimetric screening highlighted the broad substrate scope and complementarity of the two chosen CoA ligases( Figure 3 b and Table S10 ). While ipfF accepted 46 carboxylic acids, LcCL was able to activate 27 compounds, including only one common structure ( 37 ). Notably, we identified 22 compoundsfrom the subset of acids with poor chemical reactivity that were activated when incubated with our selected CoA ligases, including 6-chloro-3-pyridineacetic acid ( 66 ), 3,4-dichlorophenylacetic acid( 46 ) and 3-(trifluoromethoxy)phenylacetic acid( 50 ). 64 Based on the Malachite Green assay results, 32 acids were selected for coupling with on-DNA amine f using the full CL/NAT cascade, including either wildtype 42GmAT or 42GmAT M139Y. Gratifyingly, the N -acyltransferase and its variant exhibited good compatibility with the structurally diverse carboxylic acids, yielding products in all cases except for compound 65 ( Table S11, c Carboxylic acid = 0.1 mM). Notably, the engineered enzyme outperformed the wildtype 42GmAT across nearly all DNA-conjugated substrates and carboxylic acids tested ( Extended Data Figure 3d, Table S9, Table S11 ) highlighting mutations M139Y role as a general enabler for on-DNA activity increase without negatively affecting substrate promiscuity. Applying optimized reaction conditions ( Table S12, c Carboxylic acid = 0.25 mM) to reactions with 37 different carboxylic acids (>75 % conversion Malachite Green Assay), the enzymatic cascade consisting of ipfF or LcCL with 42GmAT M139Y gave access to a wide portfolio of on-DNA amides with excellent conversions (up to 97 %) and purity ( Figure 3c, Figure S14 ). Enzymatic construction of a DNA-encoded library Expanding complexity, we next assayed a trifunctional scaffold installed on the C12-linked oligomer either bearing two primary amines ( n ) or including a Fmoc protecting group ( m ) with an enzymatic cascade containing NATs 05PaAT, 42GmAT, or 42GmAT M139Y ( Figure 4a ). While N -acyltransferase 05PaAT showed high specificity to form the mono-amidated product when assayed with amine n , 42GmAT M139Y could also modify the more hindered amine group in vicinity of the DNA tag ( Figure 4a ). Each enzymatic cascade afforded excellent conversion for both substrates (92 % - 99 %), including substrate m bearing the bulky Fmoc protecting group, opening the door for the enzymatic construction of complex DELs containing multiple building blocks. Based on this promising outcome, we set out to chemo-enzymatically construct a DNA-encoded library with a substrate set curated to consider both structural diversity and high conversion rates ( Figure 4b ). In detail, we chemically installed eight structurally diverse carboxylic acids (building block 1, BB1) on the alpha-amine of substrate o ( Figure S16-S17 , Table S13 ), before applying the best-performing enzymatic cascade (CoA ligase ipfF / NAT 42GmAT M139Y) for late-stage functionalization of the obtained DNA conjugates with a set of 9 carboxylic acids (building block 2, BB2) ( Figure 4b ). When evaluating library construction performance, we observed uniform incorporation and excellent conversions of all building blocks 2 (on average > 91 % yield), despite the structural variability and bulkiness of pre-installed fragments p - w ( Figure 4b, Table S14 ). DNA integrity Beyond the uniformity and yield of on-DNA conversions, DNA integrity is an additional key performance indicator when assessing methods applied in DEL construction. 25,28,65 Complementing our LC-MS analysis, which showed no evidence of DNA degradation, DNA-compatibility of the enzymatic cascade was further evaluated by quantitative PCR (qPCR). 25,28,65 Using an amine linked to an amplifiable 48mer DNA-tag ( aa ) and carboxylic acid 1 as substrates, we achieved almost quantitative conversion to amide product 1aa (93.9 %) when applying the tailored enzymatic cascade (ipfF /42GmAT M139Y) ( Extended Data Figure 4a and c ) underlining the independence of enzyme performance from DNA length and composition beyond the first base (utilized 48- and 14-oligomer only have the first three bases in common). Following the biocatalytic reactions, DNA content was quantified by qPCR indicating no decrease in amplifiability for any of the tested enzyme cascades compared to the no enzyme control ( Extended Data Figure 4b, Figure S19 ). In addition, to complete the use tests for enzymatic DEL construction, we investigated T4 DNA ligase-mediated ligation efficiency after conducting enzymatic reactions, which proceeded to completion ( Figure S20 ). Conclusion A critical factor in generating a successful DEL is that the chosen transformations must maintain the integrity of the DNA tag while leading to high conversions for a broad set of substrates. Unfortunately, reaction conditions routinely used in synthetic chemistry, such as low pH or the application of metal catalysts can lead to depurinations 66,67 and base transversions 26 or promote DNA cleavage. 25,68 Against this backdrop, enzymes are emerging as attractive additions to the DEL-reaction toolkit. 69,70 While, on first glance, the bulky DNA-tag might be perceived as an impediment for enzymatic substrate-acceptance, our increasing ability to discover, design and tailor proteins on-demand using different approaches 71,72 is helping to address this challenge, leading to highly promiscuous catalysts which can be applied in early but also late library generation ( Figure 4 ). In fact, we postulate that the DNA barcode may serve as an additional substrate recognition element and, in appropriately engineered enzymes, can facilitate catalysis ( Extended Data Figure 2 ). This insight is not only important for DEL synthesis but might also play a role in DEL screening as the linker/DNA construct potentially influences binding affinity. 53,55 Finally, our study on tailored enzyme cascades for DEL construction shows that harnessing customized enzymes for DEL synthesis allows to achieve excellent reaction yields for a broad substrate scope paving the way for employing biocatalysis also for other reaction types. Ultimately, the resulting diverse and high-quality DELs will accelerate our ability to discover new medicines. Materials and Methods Oligonucleotides were purchased from LGC Biosearch Technologies (Teddington, United Kingdom), Microsynth AG (Balgach, Switzerland), and Hitgen (Chengdu, China), analyzed by LC-MS, and quantified with a Nanodrop 2000c spectrophotometer. T4 DNA ligase, Q5 High-Fidelity DNA polymerase, restriction enzymes and dNTPs were purchased from New England Biolabs (Ipswich, USA). DNase I was purchased from Roche (Basel, Switzerland). Water was purified with a Millipore Milli-Q system (mQ, Merck, Darmstadt). Amicon TM Ultra 0.5 mL and 15 mL centrifugal filters (3 kDa and 10 kDa MW cutoff) were purchased from Merck Millipore (Darmstadt, Germany) and used according to the manufacturer’s recommendation. Inorganic pyrophosphatase from baker’s yeast for Malachite Green Assay was purchased from Sigma-Aldrich (St. Louis, USA), and the Malachite Green assay was obtained from ScienCell (Carlsbad, USA). All other reagents, solvents and building blocks were purchased from ABCR (Karlsruhe, Germany), Acros Organics (Geel, Belgium), Apollo Scientific Ltd (Cheshire, United Kingdom), Enamine (Kyiv, Ukraine), Fluorochem (Hadfield, United Kingdom), Iris Biotech (Marktredwitz, Germany), Merck (Darmstadt, Germany), Sigma-Aldrich (Burlington, United States), TCI chemical (Tokyo, Japan) and ThermoFischer Scientific (Waltham, United States) and used without further purification. Gene synthesis was performed by Twist Bioscience (California, USA). Sequencing service was provided by Microsynth AG (Balgach, Switzerland). Cloning, protein production and purification Genes encoding CoA ligases or N -acyltransferases containing an N-terminal His-tag were purchased in a pET28b(+) vector. The plasmids were transformed into E. coli BL21(DE3), and cells were plated on a Luria-Bertani (LB) agar plate containing 50 μg/mL kanamycin (kan). A single colony of freshly transformed cells was incubated at 37 °C overnight in 3 mL of LB medium containing 50 μg/mL kan. Glycerol stocks (30 % glycerol in LB) of BL21(DE3) harbouring the target genes were stored at -80 °C until further use. To produce the proteins of interest, 20 mL of LB medium containing 50 µg/mL kanamycin was inoculated from the corresponding glycerol stock in a 100 mL shaking flask. The cultures were incubated overnight using 120 rpm and 37 °C. Next, 10 mL of the overnight cultures were used to inoculate 190 mL Zym-5052 autoinduction medium containing 50 µg/mL kan in a 1 L shaking flask. The culture was incubated for at least 24 hrs at 20 °C and 120 rpm. The cell pellet was harvested by centrifugation (15’000 relative centrifugal force (rcf,) 15 mins) in a 50 mL falcon tube and was stored at -20 °C for at least 12 hrs (overnight) before lysis. For the lysis, the cell pellet was thawed on ice and resuspended in 5 mL buffer A (50 mM Tris-Cl, pH 7.4, 20 mM imidazole, 500 mM NaCl). Cells were lysed by sonication (2x 1 min, 2 second pulses, 40 % amplitude, on ice). The solid cell material was removed by centrifugation (15’000 rcf, 15 mins, 4°C) and the supernatant was filtered by passing it through a 0.45 µm polysulfone (PSE) filter membrane. Enzyme purification was conducted on an Äkta pure system (GE Healthcare) by applying the clarified lysate to a 5 mL equilibrated HisTrap-column (Cytiva Life Sciences). The column was washed with at least 5 column volumes (CV) of buffer A and 5 CV of 10 % buffer B (50 mM Tris-Cl, pH 7.4, 200 mM Imidazole and 500 mM NaCl) to separate any low-affinity binding proteins before eluting the protein of interest with 100 % buffer B. The target protein-containing fraction was concentrated by ultrafiltration (Amicon Ultra, cut-off 10 kDa) and desalted in 50 mM Tris buffer (pH 8) using a HiTrap desalting column (Cytiva Life Sciences) before snap-freezing in liquid nitrogen. The enzyme stocks were stored at -80 °C until further use. Synthesis of on-DNA substrates To synthesise on-DNA substrates a - aa ,correspondingactivation solutions were prepared by mixing 175µl DMSO, 32 µl amine A, C, D, F solution (200 mM in DMSO), 32 µl 1-ethyl-3-carbodiimide (EDC; 100 mM in DMSO), and 160 µl s-NHS solution (100 mM in 2:1 DMSO: water). The activation solution was shaken at 37°C for 30 minutes and added to a solution of 100 nmol 14-mer or 48-mer amino-modified oligonucleotide in 100 µl water and 100 µl MOPS buffer (pH = 8, 100 mM MOPS, 1 M NaCl). The reaction mixture was shaken overnight at 37°C before ethanol precipitation was carried out. For larger scales (up to 500 nmol), a linear scale-up was performed. For Fmoc deprotection,an aqueous solution of the Fmoc-protected oligonucleotide conjugates (100-200 µl) was added to an equal volume of 20% piperidine (v/v) in mQ water. The solution was shaken at 400 rpm at 30°C for 2h (Thermoshaker). Volatile components were evaporated in a vacuum concentrator; the residual solids were redissolved in water, and oligonucleotide conjugates were subjected to ethanol precipitation before further use. For azide reduction, 100 – 500 nmol of the relevantDNA conjugates were dissolved in 100 µl of water before 100 µl Tris(2-carboxyethyl)phosphine solution (TCEP; 200 mM in 500 mM Tris.HCl, pH 7.4) was added and the mixture was shaken at 400 rpm and 37°C for 5h (Thermoshaker). Following reaction, the samples were ethanol precipitated before further use. DNA conjugates were purified by preparative RP-HPLC and desalted by Amicon (MW cutoff 3 kDa) before biocatalysis. Double-stranded DNA conjugates were prepared by mixing DNA conjugates in an equimolar ratio with the complementary 14-mer oligomer. Hybridization was triggered by heating to 75°C for 10 min, followed by slow cool-down to room temperature. Analysis of successful hybridization was performed by agarose gel electrophoresis (4% agarose). Enzymatic on-DNA amide bond-formation Standard enzymatic reactions to form on-DNA amides contained 2.5 mM ATP, 0.025 mM CoA, 0.1 mM carboxylic acid, 0.05 mM on-DNA amine, 20 µM CoA ligase and 20 µM N -acyltransferase in 50 mM Tris buffer (pH 8). The final reaction volume was 20 µL. Biocatalytic reactions were carried out for 20 hours at 25 °C and 450 rpm (Eppendorf thermoshaker). Reactions were quenched by heat shock (90 °C, 5 min) and diluted with water: acetonitrile (ACN) and centrifuged (15’000 relative centrifugal force (rcf,) 30 mins) prior to LC-MS analysis. LC-MS analysis of oligonucleotides Oligonucleotide conjugates were analysed by ESI-ToF-MS on a Waters Xevo G2-XS QTof instrument after separation on a Waters Acquity UPLC H-Class System fitted with an Acquity Premier Oligonucleotide BEH C18 column (130Å 1.7µM, 2.1 x 50 mm) used for analysis of DNA conjugates and an Acquity UPLC Peptide BEH C18 column (300Å 1.7µM, 2.1 x 50 mm) used for analysis of biocatalysis samples, both from Waters (Milford, USA). LC separation was performed at 60°C and a flow rate of 0.5 ml/min. Eluent A consisted of 0.4 M hexafluoroisopropanol and 5 mM triethylamine in water; methanol was used as eluent B. Two gradients were used. The standard method started with 100% eluent A for 0.5 min, followed by a linear increase to 50% eluent B over 6 min and a further increase to 90% eluent B over 0.5 min, a return to 100% eluent A over 0.5 min, which was maintained for 0.5 min. A second method was used for analyzing more hydrophobic compounds: after starting with 100% eluent A for 0.5 min, a linear increase to 70% eluent B over 6.5 min and a further increase to 95% eluent B over 0.9 min followed, ending with a return to 100% eluent A over 0.1 min which was maintained for 0.9 min. Molecular analysis of 05PaAT and 42GmAT substrate binding sites To rationalize why 42GmAT’s activity toward on-DNA substrate was superior to 05PaAT, we investigated and compared substrate binding site characteristics of both enzymes. At first, we characterized the pocket size and volume using pyKVFinder Python package. 57 Crystal structures of 05PaAT and 42GmAT (PDB 1W4T and 7QI3, respectively) were utilized as input while default probe out and volume cutoffs (8 and 50 Å, respectively) were utilized for cavity detection. The detected pockets were manually verified, cavity dimensions were extracted, and the substrate binding pocket was visualized in mesh representation using PyMol. 73 Next, we opted to model the biomolecular complex of selected NATs with DNA-tagged substrate f utilizing Boltz-2. 60 For this purpose, yaml files containing protein sequences and substrate SMILES were prepared and the prediction was performed with 30 diffusion samples (--diffusion_samples 30) utilizing ColabFold 74 multiple sequence alignment server (--use_msa_server) and adding inference-time potentials to improve the physical plausibility (--use_potentials). The 30 resulting complexes were superposed clustered by their root mean square deviations (RMSDs) using MDTraj 75 and scikit-learn 76 Python libraries. Thereby, K-means algorithm was utilized (n_clusters=3 and random_state=42) and the centroid of the biggest cluster was utilized as representative complex for visualization and Protein-Ligand Interaction Profiler (PLIP) 77 analysis. Creation of 42GmAT site-saturation libraries NNK site saturation libraries of 42GmAT were constructed by overlap extension PCR as previously reported 78 using primers pET28(+), pET28(-) and the corresponding NNK primer pairs (M139NNK, P141NNK, S178NNK, M179NNK, Table S5 ). The libraries were subcloned into a pET28b vector containing the negative selection marker sacB using restriction sites for Xho I and Nco I and ligated with T4 ligase according to the supplier’s instruction. Subsequently, the libraries were transformed into E. coli NEB10-β before plating on LB-Agar (50 µg/mL kanamycin, 5 % sucrose). Library quality was confirmed by Sanger sequencing (Microsynth). Following library plasmid isolation with NucleoSpin Plasmid Easy Pure Kit, E. coli BL21(DE3) cells were transformed with the genetic material. Per library, 89 single colonies were picked to inoculate 1 mL of LB media (50 µg/mL kan) in a 96-deep well plate. Three wells on the plate were inoculated with cells carrying the empty vector only (no gene of interest) and three wells with cells harbouring wildtype NAT 42GmAT. Additionally, one position was carried forward without any cell material (no enzyme control). In parallel to NAT library production, a plate uniformly inoculated with CoA ligase ipfF was generated. The plates were incubated overnight at 37 °C and 300 rpm (Duetz system, 50 mm shaking amplitude). As a next step, 50 µL of each well from the preculture plate was used to inoculate 950 µL Zym-5052 on the main culture plate, which was incubated at 20 °C for 24 hours. Plates were centrifuged (3’428 rcf, 20 minutes at 4°C) and the supernatant was discarded before the main culture plate was being stored for at least 12 hours (overnight) at -20 °C. For library screening, plates were thawed at room temperature and lysis was initiated by adding 50 µL/well lysis buffer (50 mM Tris (pH 8), 1 mg/mL lysozyme, 0.5 mg/mL polymyxin B, and DNAseI) to the site-saturation plates and 200 µL/well lysis buffer to the CoA ligase plate. Lysis was carried out for 20 minutes at 25 °C and 300 rpm. After combining 50 µL lysate of each NNK library variant with 50 µL ipfF lysate, the biocatalytic reaction was initiated by adding 100 µL of 50 mM Tris buffer (pH 8) containing ATP, CoA, amine E , and carboxylic acid 1 yielding the following final concentrations: 0.625 mM amine E ,1.25 mM carboxylic acid 1 , 2.5 mM ATP and 0.025 mM CoA. After 20 hrs of incubation at 300 rpm and 25 °C, the reaction was quenched by the addition of 800 µL acetonitrile (5 min incubation, 25 °C). Reaction plates were centrifuged (15 mins, 3’428 rcf) and each well’s supernatant was analysed with reverse phase chromatography on a 1290 Infinity II Agilent system fitted with a single quadrupole (ESI). A Poroshell 120 EC-C18 column (2.7 µm, 2.1 x 50 mm) from Agilent was used for analysis. LC separation was performed at 40°C and a flow rate of 0.8 mL/min. Eluent A consisted of 5 % acetonitrile (ACN) with 0.2 % formic acid. Eluent B consisted of 100 % ACN with 0.2 % formic acid. The method started with 100% eluent A for 0.15 min, followed by a linear increase to 95 % eluent B over 2 min , which was maintained for 0.35 min, followed by a return to 100 % eluent A over 0.5 min. Substrates and products were detected by MS scan (m/z 300- 510) and MSD SIM ([M+H]+), targeting the expected m/z for products and substrates. Construction of combinatorial 42GmAT libraries Site-directed mutagenesis polymerase chain reaction (PCR) was used to combine mutations M139L, M139F, and M139Y with M179W. To this end, pET28b plasmids harbouring 42GmAT genes with L, F or Y at position 139 were used as templates and amplified with primers containing the M179W mutation (M179 (+) and (-), ( Table S5 ) using the following PCR conditions: 10 µL Q5 reaction buffer, 3 µL dNTPs (10 mM), 2.5 µL (+) primer (10 µM), 2.5 µL (-) primer (10 µM), 0.5 µL template, 0.5 µL Q5 polymerase, with the addition of mQ water to reach a final volume of 50 µL. The PCR was performed according to manufacturer’s recommendation using an annealing temperature of 61°C and 25 cycles ( Table S6 ). Following Dpn I digest, PCR products were transformed into E. coli NEB10-β. Sequences were validated by Sanger sequencing. Molecular analysis of 05PaAT and 42GmAT To rationalize why 42GmAT’s activity toward on-DNA substrate was superior to 05PaAT, we investigated and compared substrate binding site characteristics of both enzymes. At first, we characterized the pocket size and volume using pyKVFinder Python package. 57 Crystal structures of 05PaAT and 42GmAT (PDB 1W4T and 7QI3, respectively) were utilized as input while default probe out and volume cutoffs (8 and 50 Å, respectively) were utilized for cavity detection. The detected pockets were manually verified, cavity dimensions were extracted, and the substrate binding pocket was visualized in mesh representation using PyMol. 73 Next, we opted to model the biomolecular complex of the NATs with DNA-tagged substrate f utilizing Boltz-2. 60 For this purpose, yaml files containing protein sequences and substrate SMILES were prepared and the prediction was performed with 30 diffusion samples (--diffusion_samples 30) utilizing ColabFold 74 multiple sequence alignment server (--use_msa_server) and adding inference-time potentials to improve the physical plausibility (--use_potentials). The 30 resulting complexes were superposed clustered by their root mean square deviations (RMSDs) using MDTraj 75 and scikit-learn 76 Python libraries. Thereby, K-means algorithm was utilized (n_clusters=3 and random_state=42) and the centroid of the biggest cluster was utilized as representative complex for visualization and Protein-Ligand Interaction Profiler (PLIP) 77 analysis. 42GmAT and 05PaAT chimera engineering FastCloning PCR 79 was used to generate 42GmAT and 05PaAT chimera enzymes (loop swap) with primer pairs 42GmAT_05PaAT or 05PaAT_42GmAT ( Table S5 ) and pET28b plasmids harbouring 42GmAT or 05PaAT as templates. Following PCR conditions were used: 10 µL Q5 reaction buffer, 4 µL dNTPs (10 mM), 2.5 µL (+) primer (10 µM), 2.5 µL (-) primer (10 µM), 0.5 µL template, 0.5 µL Q5 polymerase, with the addition of mQ water to reach a final volume of 50 µL. The PCR was performed according to the manufacturer’s recommendation using an annealing temperature of 60°C and 30 cycles ( Table S7 ). Following Dpn I digest, PCR products were transformed into E. coli NEB10-β. Sequences were validated by Sanger sequencing. Malachite green assay The activity of CoA ligases ipfF and LcCL for the set of 236 carboxylic acids was determined via the Malachite Green assay. 80 Reactions were carried out in a 96-well plate, with each well containing 2 µM CoA ligase, 0.4 U/mL pyrophosphatase, 0.1 mM carboxylic acid, 1 mM ATP, 1 mM CoA and 5 mM MgCl 2 in 25 µL 50 mM Tris (pH 8). The reactions were incubated at 25 °C for 90 minutes. Following reaction, 2.5 µL of the reaction mix was diluted with 50 mM Tris (pH 8), to a volume of 50 µL in a clear F-bottom Greiner plate. The malachite green assay was carried out according to the supplier’s instruction: 10 µL Malachite Green Reagent A was added to the 50 µL sample and shaken for 10 minutes at RT before adding 10 µL Malachite Green Reagent B. The plate was loaded on the Tecan Spark plate reader and shaken for 15 seconds to mix components before incubating without shaking for 20 minutes at 30 °C. Absorbance was measured at a wavelength of 630 nm with 15 flashes per well at 30 °C. All reactions were performed as three technical replicates and corrected against a general blank reaction containing no CoA ligase. The phosphate calibration curve ( Figure S13 ) was generated according to the supplier's instruction and was used to quantify phosphate concentration which serves as a measure for carboxylic acid activation. qPCR experiments For qPCR experiments,7.5 µL 50 mM Tris buffer (pH 8) containing enzyme (ipfF; ipfF/05PaAT; ipfF/42GmAT; ipfF/42GmAT M139Y) as well as no enzyme control (7.5 µL 50 mM Tris buffer (pH 8)) preincubated at 25°C were added to 7.5 µL 50 mM Tris buffer (pH 8) containing ATP, CoA, carboxylic acid 1 and amine aa (48-mer oligonucleotide, Table S4 entry 13). Final reaction concentrations were; 2.5 mM ATP, 0.025 mM CoA, 0.05 mM on-DNA amine aa and 0.1 mM carboxylic acid 1 . The reactions were incubated for 20 hours at 25 °C on a thermocycler and quenched by heat shock (90 °C, 1 min). Solutions were diluted to a concentration of 1 µM and directly used for the qPCR experiments. All reactions were prepared in triplicates. Quantitative PCR (qPCR) of the control samples was performed using PowerTrack TM SYBR TM Green Master Mix (Applied Biosystems) on a QuantStudio TM 7 Flex Real-Time PCR System (Applied Biosystems) in 10 µl sample volume and according to the manufacturer’s instructions with an annealing temperature of 60°C using 40 cycles ( Tables S15 and S16 ). A standard curve based on a serial dilution (10 -3 µM - 10 -6 µM) of the substrate was prepared. All qPCR reactions were prepared and measured in quadruplicates for each of the enzymatic reaction samples. Declarations Data availability Amino acid sequences of enzymes can be found in the Supplementary Information. The crystal structures used in bioinformatic experiments can be accessed via PDB ID 7QI3 (42GmAT) and 1W4T (05PaAT). All source data are provided with this manuscript. Code availability All Python scripts utilized for the bioinformatic analysis were uploaded to a GitHub repository (https://github.com/Buller-Lab/EnzyDEL_NATs). Acknowledgements This work was created as part of a SINERGIA Project funded by the Swiss National Science Foundation (Grant number CRSII5_198673 to J.S. and R.B.). In addition, we would like to thank the Vorholt Group for providing the negative selection marker sacB . Author contributions J.S. and R.B. initiated and designed the project. D.S., A.L., F. M., J.S., and R.B. designed the experiments. D.S., A.L., M.E., G.v.H. and P.S. carried out the experiments and D.S., A.L., A.G., P.S., J.S. and R.B. analysed the data. P.S. and D.S. carried out the bioinformatic analyses. D.S., A.L., J.S., and R.B. wrote the manuscript with feedback from all authors. J.S. and R.B. supervised the project. Competing interests All authors declare no competing interests. References Wouters, O. J., McKee, M. & Luyten, J. Estimated Research and Development Investment Needed to Bring a New Medicine to Market, 2009-2018. JAMA 323, 844 (2020). OECD. Health at a Glance 2023: OECD Indicators . (OECD, 2023). doi:10.1787/7a7afb35-en. Köhler, G. & Milstein, C. Continuous cultures of fused cells secreting antibody of predefined specificity. Nature 256, 495–497 (1975). Clackson, T., Hoogenboom, H. R., Griffiths, A. D. & Winter, G. Making antibody fragments using phage display libraries. Nature 352, 624–628 (1991). Boder, E. T. & Wittrup, K. D. Yeast surface display for screening combinatorial polypeptide libraries. Nat. Biotechnol. 15, 553–557 (1997). Brenner, S. & Lerner, R. A. Encoded combinatorial chemistry. Proc. Natl. Acad. Sci. 89, 5381–5383 (1992). Gura, T. DNA helps build molecular libraries for drug testing. Science 350, 1139–1140 (2015). Mullard, A. DNA-encoded drug libraries come of age. Nat. Biotechnol. 34, 450–451 (2016). Goodnow, R. A., Dumelin, C. E. & Keefe, A. D. DNA-encoded chemistry: enabling the deeper sampling of chemical space. Nat. Rev. Drug Discov. 16, 131–147 (2017). Satz, A. L. et al. DNA-encoded chemical libraries. Nat. Rev. Methods Primer 2, 1–17 (2022). Peterson, A. A. & Liu, D. R. Small-molecule discovery through DNA-encoded libraries. Nat. Rev. Drug Discov. 22, 699–722 (2023). Gloger, A. & Scheuermann, J. DNA-encoded chemical libraries on stage. Nat. Chem. Biol. 21, 20–21 (2025). Thalji, R. K. et al. Discovery of 1-(1,3,5-triazin-2-yl)piperidine-4-carboxamides as inhibitors of soluble epoxide hydrolase. Bioorg. Med. Chem. Lett. 23, 3584–3588 (2013). Keller, M., Schira, K. & Scheuermann, J. Impact of DNA-Encoded Chemical Library Technology on Drug Discovery. CHIMIA 76, 388–395 (2022). Neri, D. & Lerner, R. A. DNA-Encoded Chemical Libraries: A Selection System Based on Endowing Organic Compounds with Amplifiable Information. Annu. Rev. Biochem. 87, 479–502 (2018). Mannocci, L. et al. High-throughput sequencing allows the identification of binding molecules isolated from DNA-encoded chemical libraries. Proc. Natl. Acad. Sci. 105, 17670–17675 (2008). Clark, M. A. et al. Design, synthesis and selection of DNA-encoded small-molecule libraries. Nat. Chem. Biol. 5, 647–654 (2009). DNA-Encoded Libraries . b000000342 (Georg Thieme Verlag KG, Stuttgart, 2024). doi:10.1055/b000000342. Huang, Y., Li, Y. & Li, X. Strategies for developing DNA-encoded libraries beyond binding assays. Nat. Chem. 14, 129–140 (2022). Vummidi, B. R. et al. A mating mechanism to generate diversity for the Darwinian selection of DNA-encoded synthetic molecules. Nat. Chem. 14, 141–152 (2022). Decurtins, W. et al. Automated screening for small organic ligands using DNA-encoded chemical libraries. Nat. Protoc. 11, 764–780 (2016). Ding, Y. et al. Discovery of soluble epoxide hydrolase inhibitors through DNA-encoded library technology (ELT). Bioorg. Med. Chem. 41, 116216 (2021). Cuozzo, J. W. et al. Novel Autotaxin Inhibitor for the Treatment of Idiopathic Pulmonary Fibrosis: A Clinical Candidate Discovered Using DNA-Encoded Chemistry. J. Med. Chem. 63, 7840–7856 (2020). Harris, P. A. et al. DNA-Encoded Library Screening Identifies Benzo[ b ][1,4]oxazepin-4-ones as Highly Potent and Monoselective Receptor Interacting Protein 1 Kinase Inhibitors. J. Med. Chem. 59, 2163–2178 (2016). Malone, M. L. & Paegel, B. M. What is a “DNA-Compatible” Reaction? ACS Comb. Sci. 18, 182–187 (2016). Sauter, B., Schneider, L., Stress, C. & Gillingham, D. An assessment of the mutational load caused by various reactions used in DNA encoded libraries. Bioorg. Med. Chem. 52, 116508 (2021). Fitzgerald, P. R. & Paegel, B. M. DNA-Encoded Chemistry: Drug Discovery from a Few Good Reactions. Chem. Rev. 121, 7155–7177 (2021). Keller, M. et al. Highly pure DNA-encoded chemical libraries by dual-linker solid-phase synthesis. Science 384, 1259–1265 (2024). Hunter, J. H. et al. Functional Group Tolerance of a Micellar on-DNA Suzuki–Miyaura Cross-Coupling Reaction for DNA-Encoded Library Design. J. Org. Chem. 86, 17930–17935 (2021). Wang, X., Li, L., Shen, X. & Lu, X. Rational Design Strategies in DNA‐Encoded Libraries for Drug Discovery. Angew. Chem. Int. Ed. 64, e202511839 (2025). Su, W. et al. Triaging of DNA-Encoded Library Selection Results by High-Throughput Resynthesis of DNA–Conjugate and Affinity Selection Mass Spectrometry. Bioconjug. Chem. 32, 1001–1007 (2021). Ratnayake, A. S. et al. Toward the assembly and characterization of an encoded library hit confirmation platform: Bead-Assisted Ligand Isolation Mass Spectrometry (BALI-MS). Bioorg. Med. Chem. 41, 116205 (2021). Chai, J., Lu, X., Arico-Muendel, C. C., Ding, Y. & Pollastri, M. P. Application of l -Threonine Aldolase to on-DNA Reactions. Bioconjug. Chem. 32, 1973–1978 (2021). Thomas, B. et al. Application of Biocatalysis to on-DNA Carbohydrate Library Synthesis. ChemBioChem 18, 858–863 (2017). Fair, R. J., Walsh, R. T. & Hupp, C. D. The expanding reaction toolkit for DNA-encoded libraries. Bioorg. Med. Chem. Lett. 51, 128339 (2021). Fitzgerald, P. R., Dixit, A., Zhang, C., Mobley, D. L. & Paegel, B. M. Building Block-Centric Approach to DNA-Encoded Library Design. J. Chem. Inf. Model. 64, 4661–4672 (2024). Building Blocks Catalog. Enamine https://enamine.net/building-blocks/building-blocks-catalog. McGrath, N. A., Brichacek, M. & Njardarson, J. T. A Graphical Journey of Innovative Organic Architectures That Have Improved Our Lives. J. Chem. Educ. 87, 1348–1349 (2010). Franzini, R. M. et al. Systematic Evaluation and Optimization of Modification Reactions of Oligonucleotides with Amines and Carboxylic Acids for the Synthesis of DNA-Encoded Chemical Libraries. Bioconjug. Chem. 25, 1453–1461 (2014). Philpott, H. K., Thomas, P. J., Tew, D., Fuerst, D. E. & Lovelock, S. L. A versatile biosynthetic approach to amide bond formation. Green Chem. 20, 3426–3431 (2018). Petchey, M. et al. The Broad Aryl Acid Specificity of the Amide Bond Synthetase McbA Suggests Potential for the Biocatalytic Synthesis of Amides. Angew. Chem. Int. Ed. 57, 11584–11588 (2018). Winn, M., Richardson, S. M., Campopiano, D. J. & Micklefield, J. Harnessing and engineering amide bond forming ligases for the synthesis of amides. Curr. Opin. Chem. Biol. 55, 77–85 (2020). Torri, D. et al. Enzymatic Cascades for Stereoselective and Regioselective Amide Bond Assembly. Angew. Chem. Int. Ed. e202422185 (2025) doi:10.1002/anie.202422185. Tang, Q. et al. Broad Spectrum Enantioselective Amide Bond Synthetase from Streptoalloteichus hindustanus . ACS Catal. 1021–1029 (2024) doi:10.1021/acscatal.3c05656. Lima, R. N., Dos Anjos, C. S., Orozco, E. V. M. & Porto, A. L. M. Versatility of Candida antarctica lipase in the amide bond formation applied in organic synthesis and biotechnological processes. Mol. Catal. 466, 75–105 (2019). Bering, L., Craven, E. J., Sowerby Thomas, S. A., Shepherd, S. A. & Micklefield, J. Merging enzymes with chemocatalysis for amide bond synthesis. Nat. Commun. 13, (2022). Petchey, M. R. & Grogan, G. Enzyme‐Catalysed Synthesis of Secondary and Tertiary Amides. Adv. Synth. Catal. 361, 3895–3914 (2019). Sim, E., Walters, K. & Boukouvala, S. Arylamine N-acetyltransferases: From Structure to Function. Drug Metab. Rev. 40, 479–510 (2008). Lelièvre, C. M., Balandras, M., Petit, J., Vergne‐Vaxelaire, C. & Zaparucha, A. ATP Regeneration System in Chemoenzymatic Amide Bond Formation with Thermophilic CoA Ligase. ChemCatChem 12, 1184–1189 (2020). Orengo, C. et al. CATH – a hierarchic classification of protein domain structures. Structure 5, 1093–1109 (1997). Xu, D., Wang, Z., Zhuang, W., Wang, T. & Xie, Y. Family characteristics, phylogenetic reconstruction, and potential applications of the plant BAHD acyltransferase family. Front. Plant Sci. 14, 1218914 (2023). Burckhardt, R. M. & Escalante-Semerena, J. C. Small-Molecule Acetylation by GCN5-Related N -Acetyltransferases in Bacteria. Microbiol. Mol. Biol. Rev. 84, (2020). Bittner, P. et al. Native Mass Spectrometry Facilitates Hit Validation in DNA‐Encoded Library Technology. Angew. Chem. Int. Ed. e202504470 (2025) doi:10.1002/anie.202504470. Cyrus, K. et al. Impact of linker length on the activity of PROTACs. Mol BioSyst 7, 359–364 (2011). Bittner, P. et al. The Influence of Single-Stranded or Double-Stranded DNA Tags on Ligand Binding Affinity in DNA-Encoded Libraries. Anal. Chem. (2025) doi:10.1021/acs.analchem.5c03540. Bittrich, S., Segura, J., Duarte, J. M., Burley, S. K. & Rose, Y. RCSB protein Data Bank: exploring protein 3D similarities via comprehensive structural alignments. Bioinformatics 40, btae370 (2024). Guerra, J. V. D. S. et al. pyKVFinder: an efficient and integrable Python package for biomolecular cavity detection and characterization in data science. BMC Bioinformatics 22, 607 (2021). Stourac, J. et al. Caver Web 1.0: identification of tunnels and channels in proteins and analysis of ligand transport. Nucleic Acids Res. 47, W414–W422 (2019). Corbella, M. et al. The N-terminal Helix-Turn-Helix Motif of Transcription Factors MarA and Rob Drives DNA Recognition. J. Phys. Chem. B 125, 6791–6806 (2021). Passaro, S. et al. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. Preprint at https://doi.org/10.1101/2025.06.14.659707 (2025). Horton, R. M., Cai, Z., Ho, S. N. & Pease, L. R. Gene Splicing by Overlap Extension: Tailor-Made Genes Using the Polymerase Chain Reaction. BioTechniques 54, 129–133 (2013). Privett, H. K. et al. Iterative approach to computational enzyme design. Proc. Natl. Acad. Sci. 109, 3790–3795 (2012). Tanimoto, T. T. An Elementary Mathematical Theory of Classification and Prediction . (International Business Machines Corporation, 1958). Li, Y. et al. Optimized Reaction Conditions for Amide Bond Formation in DNA-Encoded Combinatorial Libraries. ACS Comb. Sci. 18, 438–443 (2016). Li, Y., Zimmermann, G., Scheuermann, J. & Neri, D. Quantitative PCR is a Valuable Tool to Monitor the Performance of DNA‐Encoded Chemical Library Selections. ChemBioChem 18, 848–852 (2017). Zoltewicz, J. A., Clark, D. F., Sharpless, T. W. & Grahe, G. Kinetics and mechanism of the acid-catalyzed hydrolysis of some purine nucleosides. J. Am. Chem. Soc. 92, 1741–1750 (1970). Potowski, M. et al. Screening of metal ions and organocatalysts on solid support-coupled DNA oligonucleotides guides design of DNA-encoded reactions. Chem. Sci. 10, 10481–10492 (2019). Kim, G. A., Kang, S., Win, M. N. & Han, M. S. Simple and High-Throughput Fluorescence Assay Method for DNA Damage Analysis in Single-Stranded DNA-Encoded Library Synthesis. Bioconjug. Chem. (2025) doi:10.1021/acs.bioconjchem.5c00219. Götte, K., Chines, S. & Brunschweiger, A. Reaction development for DNA-encoded library technology: From evolution to revolution? Tetrahedron Lett. 61, 151889 (2020). Song, M. & Hwang, G. T. DNA-Encoded Library Screening as Core Platform Technology in Drug Discovery: Its Synthetic Method Development and Applications in DEL Synthesis. J. Med. Chem. 63, 6578–6599 (2020). Buller, R., Damborsky, J., Hilvert, D. & Bornscheuer, U. T. Structure Prediction and Computational Protein Design for Efficient Biocatalysts and Bioactive Proteins. Angew. Chem. 137, (2025). Buller, R. et al. From nature to industry: Harnessing enzymes for biocatalysis. Science 382, (2023). Schrödinger, LLC. The PyMOL Molecular Graphics System, Version 2.5. (2015). Mirdita, M. et al. ColabFold: making protein folding accessible to all. Nat. Methods 19, 679–682 (2022). McGibbon, R. T. et al. MDTraj: A Modern Open Library for the Analysis of Molecular Dynamics Trajectories. Biophys. J. 109, 1528–1532 (2015). Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 12, 2825–2830 (2011). Salentin, S., Schreiber, S., Haupt, V. J., Adasme, M. F. & Schroeder, M. PLIP: fully automated protein–ligand interaction profiler. Nucleic Acids Res. 43, W443–W447 (2015). Eichenberger, M. et al. Asymmetric Cation‐Olefin Monocyclization by Engineered Squalene–Hopene Cyclases. Angew. Chem. Int. Ed. 60, 26080–26086 (2021). Li, C. et al. FastCloning: a highly simplified, purification-free, sequence- and ligation-independent PCR cloning method. BMC Biotechnol. 11, 92 (2011). Vardakou, M., Salmon, M., Faraldos, J. A. & O’Maille, P. E. Comparative analysis and validation of the malachite green assay for the high throughput biochemical characterization of terpene synthases. MethodsX 1, 187–196 (2014). Morgan, H. L. The Generation of a Unique Machine Description for Chemical Structures-A Technique Developed at Chemical Abstracts Service. J. Chem. Doc. 5, 107–113 (1965). Additional Declarations There is NO Competing Interest. Supplementary Files ExtendedDataFigures.docx SIEnzyDEL.pdf Supllementary Information Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7598475","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":519215051,"identity":"10b3001a-d4ec-4f3c-b8e8-3c7e0ff41b97","order_by":0,"name":"Rebecca Buller","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8UlEQVRIiWNgGAWjYFACHghlwMDceICBwYaBgRkscoAYLYwNQGVppGs5DBPBrcW8vffggx81dgzmEokNBz7uOZ+4nZ334AOGmjs4tcicOZds2HMsmcFyRmLDwRnPbifubOZLNmA49gynFgmJHDNpxgZmBoMbiQ2HeQ7cTtxwmMdMgrHhMG4t8m/MfzM21EO0/DlwDqTF/AdeLRI8ZswgBWAtDAcOgG1hwKuFJy9ZsufYcR7LnocNB3sOJBtvOMyXLJFwDI8W9rMHP/yoqZYzZ08GBt0BO9kN54EiH2pwa4EBHlRuAkENo2AUjIJRMArwAQCN+Vz5N4KUDQAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-5997-1616","institution":"University of Bern","correspondingAuthor":true,"prefix":"","firstName":"Rebecca","middleName":"","lastName":"Buller","suffix":""},{"id":519215052,"identity":"5c0d64a4-2254-4787-a70c-c8fe0e1b478c","order_by":1,"name":"Daniela Schaub","email":"","orcid":"https://orcid.org/0000-0001-9967-984X","institution":"Zurich University of Applied Sciences","correspondingAuthor":false,"prefix":"","firstName":"Daniela","middleName":"","lastName":"Schaub","suffix":""},{"id":519215053,"identity":"8d71b7b0-e49e-4ca0-91da-02b26c689819","order_by":2,"name":"Alice Lessing","email":"","orcid":"","institution":"ETH Zurich","correspondingAuthor":false,"prefix":"","firstName":"Alice","middleName":"","lastName":"Lessing","suffix":""},{"id":519215054,"identity":"f730b327-1f5b-4aa5-8eec-e1fdc1a2d617","order_by":3,"name":"Fabian Meyer","email":"","orcid":"","institution":"Zurich University of Applied Sciences","correspondingAuthor":false,"prefix":"","firstName":"Fabian","middleName":"","lastName":"Meyer","suffix":""},{"id":519215055,"identity":"307a9632-2f9f-4048-84bf-529806733400","order_by":4,"name":"Michael Eichenberger","email":"","orcid":"","institution":"Zurich University of Applied Sciences","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"","lastName":"Eichenberger","suffix":""},{"id":519215056,"identity":"e3d0d240-f8db-4d4f-8c40-d7ff781d16f6","order_by":5,"name":"Gerlis von Haugwitz","email":"","orcid":"","institution":"Zurich University of Applied Sciences","correspondingAuthor":false,"prefix":"","firstName":"Gerlis","middleName":"","lastName":"von Haugwitz","suffix":""},{"id":519215057,"identity":"f2ebb087-0af7-4d55-b141-fcd35cdc898c","order_by":6,"name":"Peter Stockinger","email":"","orcid":"","institution":"Zurich University of Applied Sciences","correspondingAuthor":false,"prefix":"","firstName":"Peter","middleName":"","lastName":"Stockinger","suffix":""},{"id":519215058,"identity":"58412ac7-d884-4be6-8268-08a710137082","order_by":7,"name":"Andreas Gloger","email":"","orcid":"","institution":"ETH Zurich","correspondingAuthor":false,"prefix":"","firstName":"Andreas","middleName":"","lastName":"Gloger","suffix":""},{"id":519215059,"identity":"a22423ce-9080-45d2-9c57-84a9fe844e7b","order_by":8,"name":"Jörg Scheuermann","email":"","orcid":"https://orcid.org/0000-0001-7746-4717","institution":"ETH Zurich","correspondingAuthor":false,"prefix":"","firstName":"Jörg","middleName":"","lastName":"Scheuermann","suffix":""}],"badges":[],"createdAt":"2025-09-12 09:05:35","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7598475/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7598475/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":92045575,"identity":"b0693202-24fa-4f5d-94c9-03542d5a9fd6","added_by":"auto","created_at":"2025-09-24 04:32:35","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":478662,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConcept of\u003c/strong\u003e \u003cstrong\u003eDNA-encoded libraries and\u003c/strong\u003e \u003cstrong\u003eenzymatic on-DNA amide-bond formation\u003c/strong\u003e. \u003cstrong\u003ea\u003c/strong\u003e, Schematic representation of a DEL structure (left), affinity-based selection against pharmaceutically relevant protein targets (middle) and advantages of biocatalysis-aided DEL synthesis (right). \u003cstrong\u003eb\u003c/strong\u003e, Optimised clinical candidates (bottom) derived from screening amide-bond containing DELs (top).\u003csup\u003e22–24\u003c/sup\u003e \u003cstrong\u003ec\u003c/strong\u003e, Enzymatic cascade consisting of a CoA ligase (CL) and an N-acyltransferase (NAT) for on-DNA amide bond formation. \u003cstrong\u003ed\u003c/strong\u003e, Enzymatic amide-coupling reactions between carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and a set of amines (\u003cstrong\u003ea\u003c/strong\u003e-\u003cstrong\u003el\u003c/strong\u003e) carrying DNA 14-mers (single- or double-stranded) conjugated via a set of strategically chosen linkers (C6, C12, TEG) using CoA ligase ipfF and NATs 05PaAT or 42GmAT, respectively. Data in \u003cstrong\u003ed\u003c/strong\u003e are the mean values ± s.d. of three technical replicates.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/eba09c4e22084753d67e1467.png"},{"id":92046007,"identity":"c5054afd-c7a9-40a7-9008-3ed28c05e76e","added_by":"auto","created_at":"2025-09-24 04:40:34","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":795556,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEngineering \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eN\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e-acyltransferase 42GmAT.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, Representation of the loop swap between \u003cem\u003eN\u003c/em\u003e-acyltransferase 05PaAT and 42GmAT. Structure and sequence of the exchanged protein motif is highlighted in purple. \u003cstrong\u003eb,\u003c/strong\u003e Fold improvement over parent (FIOP) of 05PaAT chimera and 42GmAT chimera when screened for the amide bond forming reaction between carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and DNA-conjugated amine \u003cstrong\u003ef\u003c/strong\u003e to the DNA-conjugated product \u003cstrong\u003e1f\u003c/strong\u003e. \u003cstrong\u003ec\u003c/strong\u003e, Overall structure of NAT 42GmAT (PDB: 7QI3) and a zoom-in view of the active site highlighting the substrate entrance tunnel (pink mesh), which was identified using the bioinformatic tool CAVER.\u003csup\u003e58\u003c/sup\u003e Residues selected for sites-saturation mutagenesis are shown as coloured spheres, while the catalytic triad is shown as grey sticks (C110, H158, D173). \u003cstrong\u003ed,\u003c/strong\u003e Fold improvement over parent (FIOP) of 42GmAT variants screened for the amide-bond forming reaction between the screening substrate \u003cstrong\u003eE\u003c/strong\u003e (no DNA tag) and carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e. Since the terminal ester group of screening substrate \u003cstrong\u003eE\u003c/strong\u003e was susceptible to hydrolysis in \u003cem\u003eE. coli\u003c/em\u003e lysate (empty vector control: 90 % hydrolysis of substrate), only the major hydrolysed product was considered for calculations. Top performing variants are annotated. \u003cstrong\u003ee,\u003c/strong\u003e HPLC-chromatograms of amide-bond forming reactions between carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and DNA-conjugated amine \u003cstrong\u003ef\u003c/strong\u003e to the DNA-conjugated product \u003cstrong\u003e1f\u003c/strong\u003e by best-performing 42GmAT variants and the wildtype enzyme. The accompanying barplot shows the corresponding enzymatic conversion to product \u003cstrong\u003e1f\u003c/strong\u003e (c\u003csub\u003eEnzyme\u003c/sub\u003e = 10 µM). Variant data in \u003cstrong\u003ec\u003c/strong\u003e,\u003cstrong\u003ee\u003c/strong\u003e are the mean values ± s.d. of three technical replicates\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/8afd4f5a793a129fe273ec4d.png"},{"id":92045576,"identity":"f25fe690-706b-4502-af45-6ace5f34cd07","added_by":"auto","created_at":"2025-09-24 04:32:35","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":600045,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCarboxylic acid substrate scope\u003c/strong\u003e. \u003cstrong\u003ea\u003c/strong\u003e, Reaction steps of enzymatic amide bond formation including carboxylic acid activation by CoA ligases, followed by coupling of the activated carboxylic acid to DNA-conjugated amine \u003cstrong\u003ef\u003c/strong\u003e by N-acyltransferase variant 42GmAT M139Y \u003cstrong\u003eb,\u003c/strong\u003e Heat map showing conversions of carboxylic acid activation by CoA ligases ipfF and LcCL as quantified by malachite green assay (t = 90 min, 2 µM CoA ligase, 0.4 U/mL pyrophosphatase, 0.1 mM carboxylic acid, 1 mM ATP, 1 mM CoA, and 5 mM MgCl\u003csub\u003e2 \u003c/sub\u003ein 50 mM Tris (pH 8)). \u003cstrong\u003ec,\u003c/strong\u003e Carboxylic acid substrate scope of CoA ligases ipfF and LcCL. For carboxylic acids \u003cstrong\u003e37 \u003c/strong\u003e– \u003cstrong\u003e73\u003c/strong\u003e (highlighted in grey) the full enzymatic cascade was tested for on-DNA amide bond formation (reaction scheme as depicted in \u003cstrong\u003ea,\u003c/strong\u003e c\u003csub\u003eCarboxylic acid\u003c/sub\u003e= 0.25 mM). Conversions to on-DNA amides are given below the structures. Data in \u003cstrong\u003ec\u003c/strong\u003e are the mean values ± s.d. of three technical replicates.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/034ce28fa98c3b14aea24892.png"},{"id":92045577,"identity":"3b733fe9-40c0-4c89-af62-e6dfa7ca9e45","added_by":"auto","created_at":"2025-09-24 04:32:35","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":390931,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eEnzymatic construction of a DNA-encoded library\u003c/strong\u003e. \u003cstrong\u003ea\u003c/strong\u003e, Enzymatic modification of trifunctional scaffolds \u003cstrong\u003em\u003c/strong\u003e (with Fmoc protecting group) and\u003cstrong\u003e n\u003c/strong\u003e (without Fmoc protecting group) with carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e using enzymatic cascades consisting of CoA ligase ipFf and NATs 05PaAT, 42GmAT, or 42GmAT M139Y, respectively. The accompanying table shows the conversion yields obtained for each enzymatic cascade. \u003cstrong\u003eb,\u003c/strong\u003e Construction scheme of the chemo-enzymatic DEL: Following chemical coupling of building blocks 1 (\u003cstrong\u003eP-W\u003c/strong\u003e) to amine \u003cstrong\u003eo\u003c/strong\u003e, nine carboxylic acids (shown in green) were used for the late-stage functionalization of the obtained DNA conjugates (\u003cstrong\u003ep\u003c/strong\u003e-\u003cstrong\u003ew\u003c/strong\u003e) by applying the best-performing enzymatic cascade (CoA ligase ipfF / NAT 42GmAT M139Y, c\u003csub\u003ecarboxylic acid\u003c/sub\u003e= 0.25 mM). Chromatograms of selected biocatalytic reactions highlight the purity of the obtained DNA-coupled products while observed conversions are indicated in the accompanying grid chart (on average \u0026gt; 91 %). The asterisk denotes coeluting substrate and product peaks preventing accurate calculation of conversion values. Data in \u003cstrong\u003ea\u003c/strong\u003e are the mean values ± s.d. of three technical replicates.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/c339b9a7df31b46a264dd5a0.png"},{"id":92046147,"identity":"0ef5c337-4c9a-4ea5-9d12-3b7abf7e0516","added_by":"auto","created_at":"2025-09-24 04:48:37","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3800400,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/b95a6613-9f14-47c2-b8eb-7f800d60ea00.pdf"},{"id":92045573,"identity":"c75e763b-9ce9-4a99-8a58-fd6eadb92306","added_by":"auto","created_at":"2025-09-24 04:32:34","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2065398,"visible":true,"origin":"","legend":"","description":"","filename":"ExtendedDataFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/2d27ded2319fc48fd102a8ad.docx"},{"id":92045578,"identity":"42272dc6-9a83-4668-923c-f75211319b8c","added_by":"auto","created_at":"2025-09-24 04:32:35","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":5224124,"visible":true,"origin":"","legend":"Supllementary Information","description":"","filename":"SIEnzyDEL.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7598475/v1/02276a3a1b3c6ff691332f2b.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"A tailored enzyme cascade facilitates DNA-encoded library technology and gives access to a broad substrate scope","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe discovery and development of novel therapeutic modalities remains both time-consuming and costly.\u003csup\u003e1,2\u003c/sup\u003e While technologies for the generation of specific antibodies are well established,\u003csup\u003e3–5\u003c/sup\u003e the identification of high-affinity and specific small-molecule drugs continues to pose a major challenge. To address this bottleneck, barcoded compound collections, most notably DNA-encoded Chemical Libraries (DELs),\u003csup\u003e6\u003c/sup\u003e have emerged as powerful and cost-efficient discovery platforms.\u003csup\u003e7–10\u003c/sup\u003e Today, DEL technology is widely applied across academia and industry, including in nearly all major pharmaceutical companies, where it is fuelling clinical pipelines.\u003csup\u003e11–14\u003c/sup\u003e DELs are generated through iterative split-and-pool combinatorial synthesis, in which each chemical reaction step is recorded by ligating a unique DNA tag.\u003csup\u003e15–17\u003c/sup\u003e This strategy enables the creation of vast libraries, typically ranging from 10\u003csup\u003e6\u0026nbsp;\u003c/sup\u003e- 10\u003csup\u003e9\u003c/sup\u003e library members. These libraries can be screened in parallel against protein targets of pharmaceutical interest by affinity-based selection (\u003cstrong\u003eFigure 1a\u003c/strong\u003e).\u003csup\u003e10,18–20\u003c/sup\u003e Enriched binders are then identified by PCR amplification of the DNA tags and next-generation DNA sequencing.\u003csup\u003e21\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eSince its introduction in the early 2000s, DNA-encoded library technology has been continuously improved and delivered numerous clinical drug candidates (\u003cstrong\u003eFigure 1b\u003c/strong\u003e).\u003csup\u003e11,22–24\u003c/sup\u003e Nevertheless, fundamental limitations persist. The unavoidable presence of DNA during library synthesis imposes strict constraints on reaction conditions, as base modifications or even complete degradation of the DNA tag can occur, leading to code alterations or loss of PCR amplifiability.\u003csup\u003e25,26\u003c/sup\u003e Consequently, chemical transformations must always be optimized for DNA compatibility and only a restricted subset of organic chemical conversions can, to date, be used for synthesizing productive DELs.\u003csup\u003e27\u003c/sup\u003e Recent innovations, such as solid-phase based synthesis to circumvent aqueous conditions or micellar reaction compartmentalization, have expanded the toolkit,\u003csup\u003e28,29\u003c/sup\u003e but the identification of high-yielding, diversity-generating, and DNA-compatible transformations remains the central challenge in DEL technology.\u003csup\u003e30\u003c/sup\u003e Currently, DEL construction is limited to a low number of split-and-pool cycles (typically 2-4) as otherwise the accumulation of truncated products\u003csup\u003e31,32\u003c/sup\u003e rapidly degrades library quality and complicates hit identification during affinity selections by reducing signal-to-noise ratios.\u003csup\u003e28\u003c/sup\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo address these limitations, we turned our attention to biocatalysis-aided DEL synthesis, as mild enzymatic transformations hold the promise to be DNA-preserving while simultaneously providing products with excellent chemo-, regio- and stereoselectivity and high yields. Intrigued by early examples describing the use of wildtype enzymes for on-DNA aldol reactions (threonine aldolases)\u003csup\u003e33\u003c/sup\u003e and carbohydrate couplings (glycosyltransferases)\u003csup\u003e34\u003c/sup\u003e we hypothesized that the development of engineered biocatalysts\u0026nbsp;could broaden the hitherto moderate substrate spectrum and improve product yields, providing\u0026nbsp;more effective tools for building molecular diversity in the presence of a DNA tag and leading to enhanced DEL quality and scope. In addition, we envisioned that the tailor-made enzymes could easily be combined with existing chemical methods and, theoretically, be employed in any of the consecutive split-and-pool cycles needed for library synthesis.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAs the value of a reaction for DEL synthesis is largely determined by its applicability across a wide range of substrates,\u003csup\u003e35\u003c/sup\u003e transformations that make use of commercially available building blocks are especially attractive.\u003csup\u003e36\u003c/sup\u003e Analyses of supplier catalogues show that amines and carboxylic acids are by far the most widely available building blocks, followed by less diverse and chemically less stable compound classes such as aryl halides, aldehydes and ketones.\u003csup\u003e37\u003c/sup\u003e Consequently, investigating the potential of enzymatic amide-bond formation on DNA was considered particularly valuable, given that this transformation represents a fundamental reaction in pharmaceutical chemistry and underlies the synthesis of a broad spectrum of small-molecule therapeutics.\u003csup\u003e38\u003c/sup\u003e While chemical procedures to couple carboxylic acids and amines have been established in DEL technology, these methods often require large molar excesses of building blocks and coupling reagents, lead to varying coupling efficiency, and potentially damage the DNA tag.\u003csup\u003e25,39\u003c/sup\u003e By contrast, enzymatic coupling of carboxylic acids and amines can generate amides with excellent selectivity from a diverse set of amines and carboxylic acids under mild reaction conditions.\u003csup\u003e40–43\u003c/sup\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWhen assessing the suitability (based on substrate and product scope as well as reaction conditions) of amide-bond forming enzymes for DEL technology (e.g., ATP-dependent amide bond forming enzymes\u003csup\u003e41,42,44\u003c/sup\u003e; lipases\u003csup\u003e45\u003c/sup\u003e; nitrile hydratases\u003csup\u003e46\u003c/sup\u003e), we were particularly intrigued by the versatility of an enzymatic cascade consisting of CoA ligases (CLs) and \u003cem\u003eN\u003c/em\u003e-acyltransferases (NATs).\u003csup\u003e40\u003c/sup\u003e We envisioned that by applying such a cascade, we could harness complementary wildtype CoA ligases to activate a broad set of carboxylic acids in an ATP-dependent manner, while only the \u003cem\u003eN\u003c/em\u003e-acyltransferases, applied in the amide-forming step, would need to be specifically tailored for on-DNA chemistry (\u003cstrong\u003eFigure 1c\u003c/strong\u003e).\u003csup\u003e47,48\u003c/sup\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHere, we show how to implement this concept by designing an enzymatic cascade consisting of CoA ligases and engineered NATs, giving access to a wide portfolio of on-DNA amides (\u0026gt;120 examples). The enzymatic method affords excellent yields on-DNA (on average \u0026gt; 91% conversion in the enzymatically constructed DEL) while fully preserving the integrity of the encoding DNA tag. Applying a structure-guided engineering approach as well as a rational exchange of structural elements (loop swap), we identified key\u0026nbsp;motifs important for on-DNA chemistry that could be transferred to distant enzyme scaffolds (27 % sequence identity) effectively improving on-DNA activity.\u0026nbsp;Combining the enzymatic cascade with chemical steps, a DEL consisting of 72individual compounds was obtained, highlighting the versatility of our biocatalysis-aided DEL construction approach, which can be applied for initial scaffold building as well as for late-stage functionalization.\u0026nbsp;\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eEnzyme selection and assessment of substrate scope\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo enable access to broad molecular diversity in the presence of DNA, we selected seven wildtype CoA ligases and \u003cem\u003eN\u003c/em\u003e-acyltransferases based on their reported substrate scope\u003csup\u003e40,49\u003c/sup\u003e and sequence diversity (\u003cstrong\u003eTable S1, Figure S1\u003c/strong\u003e). For the NATs in particular, the selection encompassed several different families (BAHD family, GNAT family, and arylamine \u003cem\u003eN\u003c/em\u003e-acyltransferases), each characterized by a distinct catalytic mechanism (BAHD family: HXXXD motif; GNAT family: catalytic tyrosine; arylamine \u003cem\u003eN\u003c/em\u003e-acyltransferases: Cys-His-Asp motif) and protein architecture (CATh IDs\u003csup\u003e50\u003c/sup\u003e: Chloramphenicol acetyltransferase-like domain (BAHD): 3.30.559.10, GNAT: 3.40.630.30, \u003cem\u003eN\u003c/em\u003e-arylamine acyltransferase: 3.30.2140.10).\u003csup\u003e48,51,52\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eSynthetic genes were cloned into a pET28b vector containing an \u003cem\u003eN\u003c/em\u003e-terminal His-tag (\u003cstrong\u003eTable S2\u003c/strong\u003e), transformed into \u003cem\u003eEscherichia coli\u003c/em\u003e (\u003cem\u003eE. coli\u003c/em\u003e) BL21 (DE3), and expressed to produce the corresponding enzyme (\u003cstrong\u003eFigure S2\u003c/strong\u003e). In initial activity tests, crude enzyme preparations were used to evaluate whether modifications to reported amine substrates, i.e. the addition of synthetic handles intended for the subsequent coupling to the DNA barcodes (carboxylic acid or methyl ester moieties), would be tolerated by the \u003cem\u003eN\u003c/em\u003e-acyltransferases (\u003cstrong\u003eFigure S3\u003c/strong\u003e). In detail, we screened our NAT panel in the presence of CoA ligase ipfF from \u003cem\u003eSphingomonas Ibu-2\u0026nbsp;\u003c/em\u003e(UniProt:\u0026nbsp;A1E027)for the amide bond formation between 4-fluorophenylacetic acid (\u003cstrong\u003e1\u003c/strong\u003e)and freestanding amines \u003cstrong\u003eA-D\u003c/strong\u003e. Product formation was monitored using liquid chromatography coupled to mass spectrometry (LC-MS) by following the m/z ratios of the expected products (\u003cstrong\u003eFigure S3\u003c/strong\u003e). For enzymatic cascades including \u003cem\u003eN\u003c/em\u003e-acyltransferase 05PaAT from \u003cem\u003ePseudomonas aeruginosa\u0026nbsp;\u003c/em\u003e(UniProt: Q9HUY3) or 42GmAT from \u003cem\u003eGibberella moniliformis\u0026nbsp;\u003c/em\u003e(UniProt: B7SP66), successful formation of products \u003cstrong\u003e1A\u003c/strong\u003e-\u003cstrong\u003e1D\u003c/strong\u003e or \u003cstrong\u003e1A\u003c/strong\u003e-\u003cstrong\u003e1C\u0026nbsp;\u003c/strong\u003ewas observed, respectively, providing opportunities for further substrate modification. In addition, a side product with a mass corresponding to the acylated amine was detected in some reactions, likely resulting from a competing amide bond forming reaction with acetyl-CoA from the lysate. All lysate reactions yielding product formation were validated using assays with enzymes purified by Ni-NTA affinity chromatography (\u003cstrong\u003eTable S3, Figure S4\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBuilding on thesuccessfulresults from the amine screening, on-DNA substrates were constructed. Examples from DEL\u003csup\u003e33,53\u003c/sup\u003e and PROTAC\u003csup\u003e54\u003c/sup\u003e literature highlight that the nature of the linker moiety can be decisive for ligand binding and recognition; particularly, length and hydrophobicity of the chosen tether can have an impact on affinity.\u003csup\u003e33\u003c/sup\u003e For example, ligands conjugated to a double stranded headpiece-DNA were shown to exhibit lower binding affinity against target proteins compared to the same ligands conjugated to a single-stranded DNA via a C6-linker.\u003csup\u003e53,55\u003c/sup\u003e To explore this effect and ensure to work with the most suitable linker architecture, molecular tethers of varying length and properties (C6, C12, and TEG) were installed at the 5\u0026apos;-end of a 14-mer oligonucleotide (GGA CGG GCG GCA CA-3\u0026rsquo;) and the resulting DNA tags were coupled to amines \u003cstrong\u003eA\u003c/strong\u003e, \u003cstrong\u003eC\u0026nbsp;\u003c/strong\u003eand\u003cstrong\u003e\u0026nbsp;D\u0026nbsp;\u003c/strong\u003e(\u003cstrong\u003eFigure S5-S7, Table S4\u003c/strong\u003e). The utilized oligomer was designed to prevent the formation of secondary structures such as hairpins or self-dimers and to form a stable duplex with a melting temperature of 59.4\u0026deg;C. To test the influence of double stranded DNA on enzyme performance, the single-stranded conjugates attached to molecule \u003cstrong\u003eC\u003c/strong\u003e were hybridized with the complementary strand (T\u003csub\u003em\u0026nbsp;\u003c/sub\u003e= 59.4\u0026deg;C, \u003cstrong\u003eFigure S8\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eProof-of-concept biocatalytic on-DNA reactions were conducted with the established enzymatic cascade consisting of purified CoA ligase ipfF (c = 20 \u0026micro;M) and \u003cem\u003eN\u003c/em\u003e-acyltransferase 05PaAT or 42GmAT (c = 20 \u0026micro;M), using 0.1 mM carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and 0.05 mM DNA-conjugated amines (\u003cstrong\u003ea\u003c/strong\u003e-\u003cstrong\u003el\u003c/strong\u003e) in aqueous Tris buffer (pH 8) at 25 \u0026deg;C. To closely mimic DEL synthesis conditions already from this early development stage, the reactions were carried out in a volume of 20 \u0026micro;L on a nanomolar scale. To our delight, successful product formation was observed for DNA-conjugated amines \u003cstrong\u003ea-I\u0026nbsp;\u003c/strong\u003eand\u003cstrong\u003e\u0026nbsp;k\u003c/strong\u003e, and, encouragingly, up to 95 % conversion was achieved for amides\u003cstrong\u003e\u0026nbsp;1b\u0026nbsp;\u003c/strong\u003eand\u003cstrong\u003e\u0026nbsp;1c\u003c/strong\u003e - even under non-optimized reaction conditions (\u003cstrong\u003eFigure 1\u003c/strong\u003e\u003cstrong\u003ed\u003c/strong\u003e). Further analysing the on-DNA experiments, we found that longer, hydrophobic linkers worked better in conjunction with the active \u003cem\u003eN\u003c/em\u003e-acyltransferases. Reactions containing C12 mostly showed higher conversions than those with C6- or TEG-linked amines (\u003cstrong\u003eFigure 1d\u003c/strong\u003e), giving particularly good results with NAT 42GmAT. Importantly, we did not observe any DNA degradation in the recorded chromatograms underscoring the DNA-compatibility of the enzymatic reactions, while the negative control (without any enzyme) allowed us to exclude the occurrence of non-enzymatic background (\u003cstrong\u003eFigure S9\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eNotably, we found that the type of DNA oligomer used (double-stranded vs. single-stranded) had a clear effect on enzyme performance. For example, \u003cem\u003eN\u003c/em\u003e-acyltransferase 42GmAT exhibited only 44 % relative activity for amines tethered to double-stranded DNA (\u003cstrong\u003ee\u003c/strong\u003e, \u003cstrong\u003eg\u003c/strong\u003e, \u003cstrong\u003ei\u003c/strong\u003e) compared to the identical substrates tagged with single stranded oligonucleotide (\u003cstrong\u003ed\u003c/strong\u003e, \u003cstrong\u003ef\u003c/strong\u003e, \u003cstrong\u003eh\u003c/strong\u003e) (\u003cstrong\u003eFigure 1\u003c/strong\u003e\u003cstrong\u003ed\u003c/strong\u003e). We hypothesized that this effect might be due to interactions of the enzymes with unpaired bases within the single-stranded DNA tag, facilitating substrate positioning. To explore whether the utilized DNA sequence could be used to modulate enzyme activity, we replaced guanine - serving as the first base in the standard single-stranded 14mer oligo - with adenine, cytosine, and thymine, and coupled the resulting DNA strands via a C12 linker onto amine \u003cstrong\u003eC\u003c/strong\u003e. Cascade reactions with CoA ligase ipfF and NAT 42GmAT with this customized substrate set highlighted a preference of the enzyme for the purine bases guanine (100 % relative activity), followed by adenine (85.8 % relative activity) and the pyrimidine cytosine (83.8 % relative activity), while thymine was less well accepted (47.9 % relative activity) (\u003cstrong\u003eExtended Data Figure 1\u003c/strong\u003e). Taken together, these results pointed to exploitable interactions between enzyme and DNA tag and identified the single-stranded 14mer oligo (GGA CGG GCG GCA CA-3\u0026rsquo;) coupled via its 5\u0026rsquo; end to a C12 linker as the optimal set-up for all further enzyme engineering, substrate scoping and DEL construction experiments.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eChimera Engineering\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEven though being consistently less active in the off-DNA amine screening (\u003cstrong\u003eFigure S4\u003c/strong\u003e), our experimental data (\u003cstrong\u003eFigure 1d\u003c/strong\u003e) demonstrated that 42GmAT generally outperformed 05PaAT in on-DNA amide-coupling reactions, triggering us to conduct structural analyses to uncover underlying principles. Interestingly, both \u003cem\u003eN\u003c/em\u003e-acyltransferases belong to the family of arylamine \u003cem\u003eN\u003c/em\u003e-acyltransferases and are characterized by relatively large and accessible binding pockets. Sequence-independent protein structure comparison (TM-align\u003csup\u003e56\u003c/sup\u003e) of the experimental structures PDB 7QI3 (42GmAT) and 1W4T (05PaAT) showed high structural agreement between the active NATs (TM-score = 0.73) despite low sequence similarity (27 %, \u003cstrong\u003eFigure S1\u003c/strong\u003e). Using pyKVFinder\u003csup\u003e57\u003c/sup\u003e, a python-based tool for the detection and characterization of protein cavities, we found, however, that 42GmAT featured a larger pocket volume (1538 \u0026Aring;\u003csup\u003e3\u003c/sup\u003e) than05PaAT (1068 \u0026Aring;\u003csup\u003e3\u003c/sup\u003e) and, concomitantly was characterized by a more spacious access tunnel (42GmAT: 1104 \u0026Aring;\u003csup\u003e2\u003c/sup\u003e vs 05PaAT: 811 \u0026Aring;\u003csup\u003e2\u003c/sup\u003e) suggesting binding site accessibility as a potential determinant for 42GmAT\u0026rsquo;s improved on-DNA performance (\u003cstrong\u003eFigure 2a\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003e2\u003c/strong\u003e| \u003cstrong\u003eEngineering \u003cem\u003eN\u003c/em\u003e-acyltransferase 42GmAT.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, Representation of the loop swap between \u003cem\u003eN\u003c/em\u003e-acyltransferase 05PaAT and 42GmAT. Structure and sequence of the exchanged protein motif is highlighted in purple. \u003cstrong\u003eb,\u003c/strong\u003e Fold improvement over parent (FIOP) of 05PaAT chimera and 42GmAT chimera when screened for the amide bond forming reaction between carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and DNA-conjugated amine \u003cstrong\u003ef\u003c/strong\u003e to the DNA-conjugated product \u003cstrong\u003e1f\u003c/strong\u003e. \u003cstrong\u003ec\u003c/strong\u003e, Overall structure of NAT 42GmAT (PDB: 7QI3) and a zoom-in view of the active site highlighting the substrate entrance tunnel (pink mesh), which was identified using the bioinformatic tool CAVER.\u003csup\u003e58\u003c/sup\u003e Residues selected for sites-saturation mutagenesis are shown as coloured spheres, while the catalytic triad is shown as grey sticks (C110, H158, D173). \u003cstrong\u003ed,\u003c/strong\u003e Fold improvement over parent (FIOP) of 42GmAT variants screened for the amide-bond forming reaction between the screening substrate \u003cstrong\u003eE\u003c/strong\u003e (no DNA tag) and carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e. Since the terminal ester group of screening substrate \u003cstrong\u003eE\u003c/strong\u003e was susceptible to hydrolysis in \u003cem\u003eE. coli\u003c/em\u003e lysate (empty vector control: 90 % hydrolysis of substrate), only the major hydrolysed product was considered for calculations. Top performing variants are annotated. \u003cstrong\u003ee,\u003c/strong\u003e HPLC-chromatograms of amide-bond forming reactions between carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and DNA-conjugated amine \u003cstrong\u003ef\u003c/strong\u003e to the DNA-conjugated product \u003cstrong\u003e1f\u003c/strong\u003e by best-performing 42GmAT variants and the wildtype enzyme. The accompanying barplot shows the corresponding enzymatic conversion to product \u003cstrong\u003e1f\u003c/strong\u003e (c\u003csub\u003eEnzyme\u003c/sub\u003e = 10 \u0026micro;M). Variant data in \u003cstrong\u003ec\u003c/strong\u003e,\u003cstrong\u003ee\u003c/strong\u003e are the mean values \u0026plusmn; s.d. of three technical replicates\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAnalysing the structures further, we additionally found that 42GmAT\u0026rsquo;s binding pocket architecture is partially defined by a helix\u0026ndash;loop-helix motif (residues 136\u0026ndash;150) spanning the active site (\u003cstrong\u003eFigure 2\u003c/strong\u003e\u003cstrong\u003ea\u003c/strong\u003e). This structural element is absent in the less DNA-accepting \u003cem\u003eN\u003c/em\u003e-acyltransferase 05PaAT. Here, the active site is covered by a flexible loop (residues 99\u0026ndash;105). Intriguingly, helix-turn-helix motifs are structural elements often found in DNA-binding domains, such as in transcription factors.\u003csup\u003e59\u003c/sup\u003e Thus, to explore this element\u0026rsquo;s role in on-DNA catalysis, we performed a loop swap between 42GmAT and 05PaAT. Both desired chimeric enzymes could be obtained in soluble form (\u003cstrong\u003eFigure S12\u003c/strong\u003e) and were tested in on-DNA amide bond forming reactions. Strikingly, the 05PaAT-based chimera incorporating the helix\u0026ndash;loop\u0026ndash;helix motif from 42GmAT exhibited a 17-fold improvement over wildtype for the amide bond formation between carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and on-DNA amine \u003cstrong\u003ef\u003c/strong\u003e, whereas the activity of the 42GmAT-based chimera harbouring the 05PaAT loop dropped by 20-fold (\u003cstrong\u003eFigure 2b\u003c/strong\u003e). Applying Boltz-2,\u003csup\u003e60\u003c/sup\u003e a foundational model capable of co-folding protein structures with ligands such as DNA and RNA, we modelled the two wildtype NATs and their corresponding chimeras in the presence of the DNA-conjugated amine \u003cstrong\u003ef\u003c/strong\u003e (\u003cstrong\u003eExtended Data Figure 2\u003c/strong\u003e). In the 42GmAT wildtype model, several stabilizing interactions between the protein and the first base of the 14-mer oligonucleotide were detected: Residue R146, which is part of the helix-loop-helix element, as well as residue R306 formed salt bridges with the DNA backbone, whereas amino acid Y229 engaged in\u0026nbsp;p-p\u0026nbsp;stacking with the initial guanine (\u003cstrong\u003eExtended Data Figure 2b\u003c/strong\u003e). Such stabilizing interactions were reduced in the 42GmAT chimera, where the salt bridge with a loop amino acid and the\u0026nbsp;p-p\u0026nbsp;interaction with the first base were no longer observed (\u003cstrong\u003eExtended Data Figure 2c\u003c/strong\u003e). Analogously, the binary model of wildtype 05PaAT did not show any interactions between its native loop and the DNA-bound substrate (\u003cstrong\u003eExtended Data Figure 2e\u003c/strong\u003e), while in the structure of the more active chimeric enzyme a salt bridge between residue R109 (analogous to R146 in 42GmAT) and the DNA backbone was detected (\u003cstrong\u003eExtended Data Figure 2f\u003c/strong\u003e), potentially facilitating on-DNA catalysis.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e42GmAT engineering to improve on-DNA performance\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConsidering the superior activity of NAT 42GmAT for DNA-tagged amines (\u003cstrong\u003eFigure 1d\u003c/strong\u003e), we opted to tailor this enzyme for improved activity for on-DNA amide bond formation. Guided by our structural analysis\u003csup\u003e57,58\u003c/sup\u003e and the experimental results highlighting the importance of linker-type and initial DNA base, we chose four positions (M139, P141, S178, M179) at the tunnel entrance potentially in contact with the linker/DNA as targets for full site-saturation (\u003cstrong\u003eFigure 2\u003c/strong\u003e\u003cstrong\u003ec\u003c/strong\u003e). The libraries were constructed using NNK codon containing primers (\u003cstrong\u003eTable S5\u003c/strong\u003e) and full genes were assembled by overlap extension polymerase chain reactions (PCRs).\u003csup\u003e61\u003c/sup\u003e To ensure sufficient oversampling of the site-saturation libraries,\u003csup\u003e62\u003c/sup\u003e each randomized position was screened on a 96-well microtiter plate, including the wildtype enzyme and appropriate negative controls. Due to the DNA-tagged substrates\u0026rsquo; limited stability in \u003cem\u003eE. coli\u0026nbsp;\u003c/em\u003elysates (\u003cstrong\u003eFigure S10\u003c/strong\u003e), the screening work was carried out with substrate \u003cstrong\u003eE\u003c/strong\u003e, dubbed \u0026ldquo;screening substrate\u0026rdquo;, in which the target amine was linked to the C12 linker but not the DNA tag (\u003cstrong\u003eFigure 2\u003c/strong\u003e\u003cstrong\u003ed\u0026nbsp;\u003c/strong\u003eand\u003cstrong\u003e\u0026nbsp;Figure S11\u003c/strong\u003e). Following the biocatalytic reactions, fold improvement over parent (FIOP) was quantified for each variant by comparing relative amide product formation in individual wells to that of the wild-type enzyme. This analysis identified several improved variants from positions M139, M179, and P141, with FIOPs exceeding 5 (e.g. M139F, M139Y, and M179W), whereas more modest improvements (FIOPs ~ 2) were observed for variants M139L and P141I. Notably, none of the variants from the library S178NNK exhibited improved amide bond forming activity (\u003cstrong\u003eFigure 2\u003c/strong\u003e\u003cstrong\u003ed\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn a next step, the plasmids harbouring the five hit variants (M139L, M139F, M139Y, P141I, M179W) were sequenced, retransformed into \u003cem\u003eE. coli\u003c/em\u003e BL21 (DE3) and the corresponding enzymes were produced and purified (\u003cstrong\u003eFigure S12\u003c/strong\u003e). When testing these enzyme preparations with on-DNA amine \u003cstrong\u003ef\u0026nbsp;\u003c/strong\u003e(reduced enzyme load of 10 \u0026micro;M), improved activity of variants M139Y (FIOP 2.4), M139F (FIOP 2.3) and M179W (FIOP 2.4) was confirmed (\u003cstrong\u003eFigure 2\u003c/strong\u003e\u003cstrong\u003ee\u003c/strong\u003e). A further combination of successful mutations at position 139 (L, F, Y) with M179W, however, did not result in any improvement beyond that observed for the single mutants (\u003cstrong\u003eTable S8\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCarboxylic acid scope of the enzymatic cascade\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo supplement the improved on-DNA activity of tailored \u003cem\u003eN\u003c/em\u003e-acyltransferase 42GmAT M139Y (\u003cstrong\u003eTable S9\u003c/strong\u003e), we opted to explore the substrate scope of CoA ligases ipfF from \u003cem\u003eSphingomonas sp. Ibu-2\u003c/em\u003e and complementary LcCL from \u003cem\u003eLeptothrix cholodnii\u003c/em\u003e (UniProt: B1Y4L8) with the goal of introducing additional molecular diversity on DNA. Going beyond known examples,\u003csup\u003e40\u003c/sup\u003e we compiled a customized carboxylic acid collection containing a balanced mixture of acids with high (n=86) or low (n=50) structural similarity to the reported ipfF substrate scope\u003csup\u003e40\u003c/sup\u003e (\u003cstrong\u003eExtended Data Figure 3a\u0026nbsp;\u003c/strong\u003eand\u003cstrong\u003e\u0026nbsp;b\u003c/strong\u003e) as measured by the Tanimoto index.\u003csup\u003e63\u003c/sup\u003e Importantly, the collection also comprised carboxylic acids (n=100) which were previously reported as demanding building blocks for on-DNA amide bond formation, each giving less than 50 % yield in chemical coupling reactions using DMT-MM or EDC/HOAt/DIPEA.\u003csup\u003e64\u003c/sup\u003e This led to a library consisting of 236 acids (\u003cstrong\u003eExtended Data Figure 3a\u003c/strong\u003e), which we tested with purified ipfF and LcCL using a spectrophotometric pyrophosphate detection assay (Malachite Green assay), which allows to quantify enzyme activity by measuring the free phosphate concentration (\u003cstrong\u003eFigure 3\u003c/strong\u003e\u003cstrong\u003ea, Extended Data Figure 3c\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe colorimetric screening highlighted the broad substrate scope and complementarity of the two chosen CoA ligases(\u003cstrong\u003eFigure 3\u003c/strong\u003e\u003cstrong\u003eb\u0026nbsp;\u003c/strong\u003eand\u003cstrong\u003e\u0026nbsp;Table S10\u003c/strong\u003e). While ipfF accepted 46 carboxylic acids, LcCL was able to activate 27 compounds, including only one common structure (\u003cstrong\u003e37\u003c/strong\u003e). Notably, we identified 22 compoundsfrom the subset of acids with poor chemical reactivity that were activated when incubated with our selected CoA ligases, including 6-chloro-3-pyridineacetic acid (\u003cstrong\u003e66\u003c/strong\u003e), 3,4-dichlorophenylacetic acid(\u003cstrong\u003e46\u003c/strong\u003e) and 3-(trifluoromethoxy)phenylacetic acid(\u003cstrong\u003e50\u003c/strong\u003e).\u003csup\u003e64\u003c/sup\u003e Based on the Malachite Green assay results, 32 acids were selected for coupling with on-DNA amine \u003cstrong\u003ef\u0026nbsp;\u003c/strong\u003eusing the full CL/NAT cascade, including either wildtype 42GmAT or 42GmAT M139Y. Gratifyingly, the \u003cem\u003eN\u003c/em\u003e-acyltransferase and its variant exhibited good compatibility with the structurally diverse carboxylic acids, yielding products in all cases except for compound \u003cstrong\u003e65\u003c/strong\u003e (\u003cstrong\u003eTable S11,\u0026nbsp;\u003c/strong\u003ec\u003csub\u003eCarboxylic acid\u003c/sub\u003e= 0.1 mM). Notably, the engineered enzyme outperformed the wildtype 42GmAT across nearly all DNA-conjugated substrates and carboxylic acids tested (\u003cstrong\u003eExtended Data Figure 3d, Table S9, Table S11\u003c/strong\u003e)\u0026nbsp;highlighting mutations M139Y role as a general enabler for on-DNA activity increase without negatively affecting substrate promiscuity. Applying optimized reaction conditions (\u003cstrong\u003eTable S12,\u0026nbsp;\u003c/strong\u003ec\u003csub\u003eCarboxylic acid\u003c/sub\u003e= 0.25 mM) to reactions with 37 different carboxylic acids (\u0026gt;75 % conversion Malachite Green Assay), the enzymatic cascade consisting of ipfF or LcCL with 42GmAT M139Y gave access to a wide portfolio of on-DNA amides with excellent conversions (up to 97 %) and purity (\u003cstrong\u003eFigure 3c, Figure S14\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEnzymatic construction of a DNA-encoded library\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eExpanding complexity, we next assayed a trifunctional scaffold installed on the C12-linked oligomer either bearing two primary amines (\u003cstrong\u003en\u003c/strong\u003e) or including a Fmoc protecting group (\u003cstrong\u003em\u003c/strong\u003e) with an enzymatic cascade containing NATs 05PaAT, 42GmAT, or 42GmAT M139Y (\u003cstrong\u003eFigure 4a\u003c/strong\u003e). While \u003cem\u003eN\u003c/em\u003e-acyltransferase 05PaAT showed high specificity to form the mono-amidated product when assayed with amine \u003cstrong\u003en\u003c/strong\u003e, 42GmAT M139Y could also modify the more hindered amine group in vicinity of the DNA tag (\u003cstrong\u003eFigure 4a\u003c/strong\u003e). Each enzymatic cascade afforded excellent conversion for both substrates (92 % - 99 %), including substrate \u003cstrong\u003em\u003c/strong\u003e bearing the bulky Fmoc protecting group, opening the door for the enzymatic construction of complex DELs containing multiple building blocks.\u003c/p\u003e\n\u003cp\u003eBased on this promising outcome, we set out to chemo-enzymatically construct a DNA-encoded library with a substrate set curated to consider both structural diversity and high conversion rates (\u003cstrong\u003eFigure 4b\u003c/strong\u003e). In detail, we chemically installed eight structurally diverse carboxylic acids (building block 1, BB1) on the alpha-amine of substrate \u003cstrong\u003eo\u0026nbsp;\u003c/strong\u003e(\u003cstrong\u003eFigure S16-S17\u003c/strong\u003e, \u003cstrong\u003eTable S13\u003c/strong\u003e), before applying the best-performing enzymatic cascade (CoA ligase ipfF / NAT 42GmAT M139Y) for late-stage functionalization of the obtained DNA conjugates with a set of 9 carboxylic acids (building block 2, BB2) (\u003cstrong\u003eFigure 4b\u003c/strong\u003e). When evaluating library construction performance, we observed uniform incorporation and excellent conversions of all building blocks 2 (on average \u0026gt; 91 % yield), despite the structural variability and bulkiness of pre-installed fragments \u003cstrong\u003ep\u003c/strong\u003e-\u003cstrong\u003ew\u0026nbsp;\u003c/strong\u003e(\u003cstrong\u003eFigure 4b, Table S14\u003c/strong\u003e).\u003cstrong\u003e\u003cbr\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDNA integrity\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBeyond the uniformity and yield of on-DNA conversions, DNA integrity is an additional key performance indicator when assessing methods applied in DEL construction.\u003csup\u003e25,28,65\u003c/sup\u003e Complementing our LC-MS analysis, which showed no evidence of DNA degradation, DNA-compatibility of the enzymatic cascade was further evaluated by quantitative PCR (qPCR).\u003csup\u003e25,28,65\u003c/sup\u003e Using an amine linked to an amplifiable 48mer DNA-tag (\u003cstrong\u003eaa\u003c/strong\u003e) and carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e as substrates, we achieved almost quantitative conversion to amide product \u003cstrong\u003e1aa\u003c/strong\u003e (93.9 %) when applying the tailored enzymatic cascade (ipfF /42GmAT M139Y) (\u003cstrong\u003eExtended Data Figure 4a\u0026nbsp;\u003c/strong\u003eand \u003cstrong\u003ec\u003c/strong\u003e) underlining the independence of enzyme performance from DNA length and composition beyond the first base (utilized 48- and 14-oligomer only have the first three bases in common). Following the biocatalytic reactions, DNA content was quantified by qPCR indicating no decrease in amplifiability for any of the tested enzyme cascades compared to the no enzyme control (\u003cstrong\u003eExtended Data Figure 4b, Figure S19\u003c/strong\u003e). In addition, to complete the use tests for enzymatic DEL construction, we investigated T4 DNA ligase-mediated ligation efficiency after conducting enzymatic reactions, which proceeded to completion (\u003cstrong\u003eFigure S20\u003c/strong\u003e).\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eA critical factor in generating a successful DEL is that the chosen transformations must maintain the integrity of the DNA tag while leading to high conversions for a broad set of substrates. Unfortunately, reaction conditions routinely used in synthetic chemistry, such as low pH or the application of metal catalysts can lead to depurinations\u003csup\u003e66,67\u003c/sup\u003e and base transversions\u003csup\u003e26\u003c/sup\u003e or promote DNA cleavage.\u003csup\u003e25,68\u003c/sup\u003e Against this backdrop, enzymes are emerging as attractive additions to the DEL-reaction toolkit.\u003csup\u003e69,70\u003c/sup\u003e While, on first glance, the bulky DNA-tag might be perceived as an impediment for enzymatic substrate-acceptance, our increasing ability to discover, design and tailor proteins on-demand using different approaches\u003csup\u003e71,72\u003c/sup\u003e is helping to address this challenge, leading to highly promiscuous catalysts which can be applied in early but also late library generation (\u003cstrong\u003eFigure 4\u003c/strong\u003e). In fact, we postulate that the DNA barcode may serve as an additional substrate recognition element and, in appropriately engineered enzymes, can facilitate catalysis (\u003cstrong\u003eExtended Data Figure 2\u003c/strong\u003e). This insight is not only important for DEL synthesis but might also play a role in DEL screening as the linker/DNA construct potentially influences binding affinity.\u003csup\u003e53,55\u003c/sup\u003e Finally, our study on tailored enzyme cascades for DEL construction shows that harnessing customized enzymes for DEL synthesis allows to achieve excellent reaction yields for a broad substrate scope paving the way for employing biocatalysis also for other reaction types. Ultimately, the resulting diverse and high-quality DELs will accelerate our ability to discover new medicines.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003eOligonucleotides were purchased from LGC Biosearch Technologies (Teddington, United Kingdom), Microsynth AG (Balgach, Switzerland), and Hitgen (Chengdu, China), analyzed by LC-MS, and quantified with a Nanodrop 2000c spectrophotometer. T4 DNA ligase, Q5 High-Fidelity DNA polymerase, restriction enzymes and dNTPs were purchased from New England Biolabs (Ipswich, USA). DNase I was purchased from Roche (Basel, Switzerland). Water was purified with a Millipore Milli-Q system (mQ, Merck, Darmstadt). Amicon\u003csup\u003eTM\u003c/sup\u003e Ultra 0.5 mL and 15 mL centrifugal filters (3 kDa and 10 kDa MW cutoff) were purchased from Merck Millipore (Darmstadt, Germany) and used according to the manufacturer’s recommendation. Inorganic pyrophosphatase from baker’s yeast for Malachite Green Assay was purchased from Sigma-Aldrich (St. Louis, USA), and the Malachite Green assay was obtained from ScienCell (Carlsbad, USA). All other reagents, solvents and building blocks were purchased from ABCR (Karlsruhe, Germany), Acros Organics (Geel, Belgium), Apollo Scientific Ltd (Cheshire, United Kingdom), Enamine (Kyiv, Ukraine), Fluorochem (Hadfield, United Kingdom), Iris Biotech (Marktredwitz, Germany), Merck (Darmstadt, Germany), Sigma-Aldrich (Burlington, United States), TCI chemical (Tokyo, Japan) and ThermoFischer Scientific (Waltham, United States) and used without further purification. Gene synthesis was performed by Twist Bioscience (California, USA). Sequencing service was provided by Microsynth AG (Balgach, Switzerland).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCloning, protein production and purification\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGenes encoding CoA ligases or \u003cem\u003eN\u003c/em\u003e-acyltransferases containing an N-terminal His-tag were purchased in a pET28b(+) vector. The plasmids were transformed into \u003cem\u003eE. coli\u003c/em\u003e BL21(DE3), and cells were plated on a Luria-Bertani (LB) agar plate containing 50 μg/mL kanamycin (kan). A single colony of freshly transformed cells was incubated at 37 °C overnight in 3 mL of LB medium containing 50 μg/mL kan. Glycerol stocks (30 % glycerol in LB) of BL21(DE3) harbouring the target genes were stored at -80 °C until further use.\u003c/p\u003e\n\u003cp\u003eTo produce the proteins of interest, 20 mL of LB medium containing 50 µg/mL kanamycin was inoculated from the corresponding glycerol stock in a 100 mL shaking flask. The cultures were incubated overnight using 120 rpm and 37 °C. Next, 10 mL of the overnight cultures were used to inoculate 190 mL Zym-5052 autoinduction medium containing 50 µg/mL kan in a 1 L shaking flask. The culture was incubated for at least 24 hrs at 20 °C and 120 rpm. The cell pellet was harvested by centrifugation (15’000 relative centrifugal force (rcf,) 15 mins) in a 50 mL falcon tube and was stored at -20 °C for at least 12 hrs (overnight) before lysis. For the lysis, the cell pellet was thawed on ice and resuspended in 5 mL buffer A (50 mM Tris-Cl, pH 7.4, 20 mM imidazole, 500 mM NaCl). Cells were lysed by sonication (2x 1 min, 2 second pulses, 40 % amplitude, on ice). The solid cell material was removed by centrifugation (15’000 rcf, 15 mins, 4°C) and the supernatant was filtered by passing it through a 0.45 µm polysulfone (PSE) filter membrane. Enzyme purification was conducted on an Äkta pure system (GE Healthcare) by applying the clarified lysate to a 5 mL equilibrated HisTrap-column (Cytiva Life Sciences). The column was washed with at least 5 column volumes (CV) of buffer A and 5 CV of 10 % buffer B (50 mM Tris-Cl, pH 7.4, 200 mM Imidazole and 500 mM NaCl) to separate any low-affinity binding proteins before eluting the protein of interest with 100 % buffer B. The target protein-containing fraction was concentrated by ultrafiltration (Amicon Ultra, cut-off 10 kDa) and desalted in 50 mM Tris buffer (pH 8) using a HiTrap desalting column (Cytiva Life Sciences) before snap-freezing in liquid nitrogen. The enzyme stocks were stored at -80 °C until further use.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSynthesis of on-DNA substrates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo synthesise on-DNA substrates \u003cstrong\u003ea\u003c/strong\u003e-\u003cstrong\u003eaa\u003c/strong\u003e,correspondingactivation solutions were prepared by mixing 175µl DMSO, 32 µl amine A, C, D, F solution (200 mM in DMSO), 32 µl 1-ethyl-3-carbodiimide (EDC; 100 mM in DMSO), and 160 µl s-NHS solution (100 mM in 2:1 DMSO: water). The activation solution was shaken at 37°C for 30 minutes and added to a solution of 100 nmol 14-mer or 48-mer amino-modified oligonucleotide in 100 µl water and 100 µl MOPS buffer (pH = 8, 100 mM MOPS, 1 M NaCl). The reaction mixture was shaken overnight at 37°C before ethanol precipitation was carried out. For larger scales (up to 500 nmol), a linear scale-up was performed.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor Fmoc deprotection,an aqueous solution of the Fmoc-protected oligonucleotide conjugates (100-200 µl) was added to an equal volume of 20% piperidine (v/v) in mQ water. The solution was shaken at 400 rpm at 30°C for 2h (Thermoshaker). Volatile components were evaporated in a vacuum concentrator; the residual solids were redissolved in water, and oligonucleotide conjugates were subjected to ethanol precipitation before further use.\u003c/p\u003e\n\u003cp\u003eFor azide reduction, 100 – 500 nmol of the relevantDNA conjugates were dissolved in 100 µl of water before 100 µl Tris(2-carboxyethyl)phosphine solution (TCEP; 200 mM in 500 mM Tris.HCl, pH 7.4) was added and the mixture was shaken at 400 rpm and 37°C for 5h (Thermoshaker). Following reaction, the samples were ethanol precipitated before further use.\u003c/p\u003e\n\u003cp\u003eDNA conjugates were purified by preparative RP-HPLC and desalted by Amicon (MW cutoff 3 kDa) before biocatalysis.\u003c/p\u003e\n\u003cp\u003eDouble-stranded DNA conjugates were prepared by mixing DNA conjugates in an equimolar ratio with the complementary 14-mer oligomer. Hybridization was triggered by heating to 75°C for 10 min, followed by slow cool-down to room temperature. Analysis of successful hybridization was performed by agarose gel electrophoresis (4% agarose).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEnzymatic on-DNA amide bond-formation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStandard enzymatic reactions to form on-DNA amides contained 2.5 mM ATP, 0.025 mM CoA, 0.1 mM carboxylic acid, 0.05 mM on-DNA amine, 20 µM CoA ligase and 20 µM \u003cem\u003eN\u003c/em\u003e-acyltransferase in 50 mM Tris buffer (pH 8). The final reaction volume was 20 µL. Biocatalytic reactions were carried out for 20 hours at 25 °C and 450 rpm (Eppendorf thermoshaker). Reactions were quenched by heat shock (90 °C, 5 min) and diluted with water: acetonitrile (ACN) and centrifuged (15’000 relative centrifugal force (rcf,) 30 mins) prior to LC-MS analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLC-MS analysis of oligonucleotides\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOligonucleotide conjugates were analysed by ESI-ToF-MS on a Waters Xevo G2-XS QTof instrument after separation on a Waters Acquity UPLC H-Class System fitted with an Acquity Premier Oligonucleotide BEH C18 column (130Å 1.7µM, 2.1 x 50 mm) used for analysis of DNA conjugates and an Acquity UPLC Peptide BEH C18 column (300Å 1.7µM, 2.1 x 50 mm) used for analysis of biocatalysis samples, both from Waters (Milford, USA). LC separation was performed at 60°C and a flow rate of 0.5 ml/min. Eluent A consisted of 0.4 M hexafluoroisopropanol and 5 mM triethylamine in water; methanol was used as eluent B. Two gradients were used.\u003cbr\u003e\u0026nbsp;The standard method started with 100% eluent A for 0.5 min, followed by a linear increase to 50% eluent B over 6 min and a further increase to 90% eluent B over 0.5 min, a return to 100% eluent A over 0.5 min, which was maintained for 0.5 min. A second method was used for analyzing more hydrophobic compounds: after starting with 100% eluent A for 0.5 min, a linear increase to 70% eluent B over 6.5 min and a further increase to 95% eluent B over 0.9 min followed, ending with a return to 100% eluent A over 0.1 min which was maintained for 0.9 min.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMolecular analysis of 05PaAT and 42GmAT substrate binding sites\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo rationalize why 42GmAT’s activity toward on-DNA substrate was superior to 05PaAT, we investigated and compared substrate binding site characteristics of both enzymes. At first, we characterized the pocket size and volume using pyKVFinder Python package.\u003csup\u003e57\u003c/sup\u003e Crystal structures of 05PaAT and 42GmAT (PDB 1W4T and 7QI3, respectively) were utilized as input while default probe out and volume cutoffs (8 and 50 Å, respectively) were utilized for cavity detection. The detected pockets were manually verified, cavity dimensions were extracted, and the substrate binding pocket was visualized in mesh representation using PyMol.\u003csup\u003e73\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eNext, we opted to model the biomolecular complex of selected NATs with DNA-tagged substrate \u003cstrong\u003ef\u003c/strong\u003e utilizing Boltz-2.\u003csup\u003e60\u003c/sup\u003e For this purpose, yaml files containing protein sequences and substrate SMILES were prepared and the prediction was performed with 30 diffusion samples (--diffusion_samples 30) utilizing ColabFold\u003csup\u003e74\u003c/sup\u003e multiple sequence alignment server (--use_msa_server) and adding inference-time potentials to improve the physical plausibility (--use_potentials). The 30 resulting complexes were superposed clustered by their root mean square deviations (RMSDs) using MDTraj\u003csup\u003e75\u003c/sup\u003e and scikit-learn\u003csup\u003e76\u003c/sup\u003e Python libraries. Thereby, K-means algorithm was utilized (n_clusters=3 and random_state=42) and the centroid of the biggest cluster was utilized as representative complex for visualization and Protein-Ligand Interaction Profiler (PLIP)\u003csup\u003e77\u003c/sup\u003e analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCreation of 42GmAT site-saturation libraries\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNNK site saturation libraries of 42GmAT were constructed by overlap extension PCR as previously reported\u003csup\u003e78\u003c/sup\u003e using primers pET28(+), pET28(-) and the corresponding NNK primer pairs (M139NNK, P141NNK, S178NNK, M179NNK,\u0026nbsp;\u003cstrong\u003eTable S5\u003c/strong\u003e). The libraries were subcloned into a pET28b vector containing the negative selection marker \u003cem\u003esacB\u003c/em\u003e using restriction sites for \u003cem\u003eXho\u003c/em\u003eI and \u003cem\u003eNco\u003c/em\u003eI and ligated with T4 ligase according to the supplier’s instruction. Subsequently, the libraries were transformed into \u003cem\u003eE. coli\u003c/em\u003e NEB10-β before plating on LB-Agar (50 µg/mL kanamycin, 5 % sucrose). Library quality was confirmed by Sanger sequencing (Microsynth). Following library plasmid isolation with NucleoSpin Plasmid Easy Pure Kit, \u003cem\u003eE. coli\u003c/em\u003e BL21(DE3) cells were transformed with the genetic material. Per library, 89 single colonies were picked to inoculate 1 mL of LB media (50 µg/mL kan) in a 96-deep well plate. Three wells on the plate were inoculated with cells carrying the empty vector only (no gene of interest) and three wells with cells harbouring wildtype NAT 42GmAT. Additionally, one position was carried forward without any cell material (no enzyme control). In parallel to NAT library production, a plate uniformly inoculated with CoA ligase ipfF was generated. The plates were incubated overnight at 37 °C and 300 rpm (Duetz system, 50 mm shaking amplitude). As a next step, 50 µL of each well from the preculture plate was used to inoculate 950 µL Zym-5052 on the main culture plate, which was incubated at 20 °C for 24 hours. Plates were centrifuged (3’428 rcf, 20 minutes at 4°C) and the supernatant was discarded before the main culture plate was being stored for at least 12 hours (overnight) at -20 °C.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor library screening, plates were thawed at room temperature and lysis was initiated by adding 50 µL/well lysis buffer (50 mM Tris (pH 8), 1 mg/mL lysozyme, 0.5 mg/mL polymyxin B, and DNAseI) to the site-saturation plates and 200 µL/well lysis buffer to the CoA ligase plate. Lysis was carried out for 20 minutes at 25 °C and 300 rpm. After combining 50 µL lysate of each NNK library variant with 50 µL ipfF lysate, the biocatalytic reaction was initiated by adding 100 µL of 50 mM Tris buffer (pH 8) containing ATP, CoA, amine \u003cstrong\u003eE\u003c/strong\u003e, and carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e yielding the following final concentrations: 0.625 mM amine \u003cstrong\u003eE\u003c/strong\u003e,1.25 mM carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e, 2.5 mM ATP and 0.025 mM CoA. After 20 hrs of incubation at 300 rpm and 25 °C, the reaction was quenched by the addition of 800 µL acetonitrile (5 min incubation, 25 °C). Reaction plates were centrifuged (15 mins, 3’428 rcf) and each well’s supernatant was analysed with reverse phase chromatography on a 1290 Infinity II Agilent system fitted with a single quadrupole (ESI). A Poroshell 120 EC-C18 column (2.7 µm, 2.1 x 50 mm) from Agilent was used for analysis. LC separation was performed at 40°C and a flow rate of 0.8 mL/min. Eluent A consisted of 5 % acetonitrile (ACN) with 0.2 % formic acid. Eluent B consisted of 100 % ACN with 0.2 % formic acid. The method started with 100% eluent A for 0.15 min, followed by a linear increase to 95 % eluent B over 2 min , which was maintained for 0.35 min, followed by a return to 100 % eluent A over 0.5 min. Substrates and products were detected by MS scan (m/z 300- 510) and MSD SIM ([M+H]+), targeting the expected m/z for products and substrates.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConstruction of combinatorial 42GmAT libraries\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSite-directed mutagenesis polymerase chain reaction (PCR) was used to combine mutations M139L, M139F, and M139Y with M179W. To this end, pET28b plasmids harbouring 42GmAT genes with L, F or Y at position 139 were used as templates and amplified with primers containing the M179W mutation (M179 (+) and (-), (\u003cstrong\u003eTable S5\u003c/strong\u003e) using the following PCR conditions: 10 µL Q5 reaction buffer, 3 µL dNTPs (10 mM), 2.5 µL (+) primer (10 µM), 2.5 µL (-) primer (10 µM), 0.5 µL template, 0.5 µL Q5 polymerase, with the addition of mQ water to reach a final volume of 50 µL. The PCR was performed according to manufacturer’s recommendation using an annealing temperature of 61°C and 25 cycles (\u003cstrong\u003eTable S6\u003c/strong\u003e). Following \u003cem\u003eDpn\u003c/em\u003eI digest, PCR products were transformed into \u003cem\u003eE. coli\u003c/em\u003e NEB10-β. Sequences were validated by Sanger sequencing.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMolecular analysis of 05PaAT and 42GmAT\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo rationalize why 42GmAT’s activity toward on-DNA substrate was superior to 05PaAT, we investigated and compared substrate binding site characteristics of both enzymes. At first, we characterized the pocket size and volume using pyKVFinder Python package.\u003csup\u003e57\u003c/sup\u003e Crystal structures of 05PaAT and 42GmAT (PDB 1W4T and 7QI3, respectively) were utilized as input while default probe out and volume cutoffs (8 and 50 Å, respectively) were utilized for cavity detection. The detected pockets were manually verified, cavity dimensions were extracted, and the substrate binding pocket was visualized in mesh representation using PyMol.\u003csup\u003e73\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eNext, we opted to model the biomolecular complex of the NATs with DNA-tagged substrate \u003cstrong\u003ef\u003c/strong\u003e utilizing Boltz-2.\u003csup\u003e60\u003c/sup\u003e For this purpose, yaml files containing protein sequences and substrate SMILES were prepared and the prediction was performed with 30 diffusion samples (--diffusion_samples 30) utilizing ColabFold\u003csup\u003e74\u003c/sup\u003e multiple sequence alignment server (--use_msa_server) and adding inference-time potentials to improve the physical plausibility (--use_potentials). The 30 resulting complexes were superposed clustered by their root mean square deviations (RMSDs) using MDTraj\u003csup\u003e75\u003c/sup\u003e and scikit-learn\u003csup\u003e76\u003c/sup\u003e Python libraries. Thereby, K-means algorithm was utilized (n_clusters=3 and random_state=42) and the centroid of the biggest cluster was utilized as representative complex for visualization and Protein-Ligand Interaction Profiler (PLIP)\u003csup\u003e77\u003c/sup\u003e analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e42GmAT and 05PaAT chimera engineering\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFastCloning PCR\u003csup\u003e79\u003c/sup\u003e was used to generate 42GmAT and 05PaAT chimera enzymes (loop swap) with primer pairs 42GmAT_05PaAT or 05PaAT_42GmAT (\u003cstrong\u003eTable S5\u003c/strong\u003e) and pET28b plasmids harbouring 42GmAT or 05PaAT as templates. Following PCR conditions were used: 10 µL Q5 reaction buffer, 4 µL dNTPs (10 mM), 2.5 µL (+) primer (10 µM), 2.5 µL (-) primer (10 µM), 0.5 µL template, 0.5 µL Q5 polymerase, with the addition of mQ water to reach a final volume of 50 µL. The PCR was performed according to the manufacturer’s recommendation using an annealing temperature of 60°C and 30 cycles (\u003cstrong\u003eTable S7\u003c/strong\u003e). Following \u003cem\u003eDpn\u003c/em\u003eI digest, PCR products were transformed into \u003cem\u003eE. coli\u003c/em\u003e NEB10-β. Sequences were validated by Sanger sequencing.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMalachite green assay\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe activity of CoA ligases ipfF and LcCL for the set of 236 carboxylic acids was determined via the Malachite Green assay.\u003csup\u003e80\u003c/sup\u003e Reactions were carried out in a 96-well plate, with each well containing 2 µM CoA ligase, 0.4 U/mL pyrophosphatase, 0.1 mM carboxylic acid, 1 mM ATP, 1 mM CoA and 5 mM MgCl\u003csub\u003e2\u0026nbsp;\u003c/sub\u003ein 25 µL 50 mM Tris (pH 8). The reactions were incubated at 25 °C for 90 minutes. Following reaction, 2.5 µL of the reaction mix was diluted with 50 mM Tris (pH 8), to a volume of 50 µL in a clear F-bottom Greiner plate. The malachite green assay was carried out according to the supplier’s instruction: 10 µL Malachite Green Reagent A was added to the 50 µL sample and shaken for 10 minutes at RT before adding 10 µL Malachite Green Reagent B. The plate was loaded on the Tecan Spark plate reader and shaken for 15 seconds to mix components before incubating without shaking for 20 minutes at 30 °C. Absorbance was measured at a wavelength of 630 nm with 15 flashes per well at 30 °C. All reactions were performed as three technical replicates and corrected against a general blank reaction containing no CoA ligase. The phosphate calibration curve (\u003cstrong\u003eFigure S13\u003c/strong\u003e) was generated according to the supplier's instruction and was used to quantify phosphate concentration which serves as a measure for carboxylic acid activation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eqPCR experiments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor qPCR experiments,7.5 µL 50 mM Tris buffer (pH 8) containing enzyme (ipfF; ipfF/05PaAT; ipfF/42GmAT; ipfF/42GmAT M139Y) as well as no enzyme control (7.5 µL 50 mM Tris buffer (pH 8)) preincubated at 25°C were added to 7.5 µL 50 mM Tris buffer (pH 8) containing ATP, CoA, carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e and amine \u003cstrong\u003eaa\u0026nbsp;\u003c/strong\u003e(48-mer oligonucleotide,\u0026nbsp;\u003cstrong\u003eTable S4\u003c/strong\u003e entry 13). Final reaction concentrations were; 2.5 mM ATP, 0.025 mM CoA, 0.05 mM on-DNA amine \u003cstrong\u003eaa\u003c/strong\u003e and 0.1 mM carboxylic acid \u003cstrong\u003e1\u003c/strong\u003e. The reactions were incubated for 20 hours at 25 °C on a thermocycler and quenched by heat shock (90 °C, 1 min). Solutions were diluted to a concentration of 1 µM and directly used for the qPCR experiments. All reactions were prepared in triplicates.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eQuantitative PCR (qPCR) of the control samples was performed using PowerTrack\u003csup\u003eTM\u0026nbsp;\u003c/sup\u003eSYBR\u003csup\u003eTM\u0026nbsp;\u003c/sup\u003eGreen Master Mix (Applied Biosystems) on a QuantStudio\u003csup\u003eTM\u0026nbsp;\u003c/sup\u003e7 Flex Real-Time PCR System (Applied Biosystems) in 10 µl sample volume and according to the manufacturer’s instructions with an annealing temperature of 60°C using 40 cycles (\u003cstrong\u003eTables S15\u003c/strong\u003e and\u0026nbsp;\u003cstrong\u003eS16\u003c/strong\u003e). A standard curve based on a serial dilution (10\u003csup\u003e-3\u0026nbsp;\u003c/sup\u003eµM - 10\u003csup\u003e-6\u0026nbsp;\u003c/sup\u003eµM) of the substrate was prepared. All qPCR reactions were prepared and measured in quadruplicates for each of the enzymatic reaction samples.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAmino acid sequences of enzymes can be found in the Supplementary Information. The crystal structures used in bioinformatic experiments can be accessed via PDB ID 7QI3 (42GmAT) and 1W4T (05PaAT). All source data are provided with this manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll Python scripts utilized for the bioinformatic analysis were uploaded to a GitHub repository (https://github.com/Buller-Lab/EnzyDEL_NATs).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was created as part of a SINERGIA Project funded by the Swiss National Science Foundation (Grant number CRSII5_198673 to J.S. and R.B.). In addition, we would like to thank the Vorholt Group for providing the negative selection marker \u003cem\u003esacB\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eJ.S. and R.B. initiated and designed the project. D.S., A.L., F. M., J.S., and R.B. designed the experiments. D.S., A.L., M.E., G.v.H. and P.S. carried out the experiments and D.S., A.L., A.G., P.S., J.S. and R.B. analysed the data. P.S. and D.S. carried out the bioinformatic analyses. D.S., A.L., J.S., and R.B. wrote the manuscript with feedback from all authors. J.S. and R.B. supervised the project.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eWouters, O. J., McKee, M. \u0026amp; Luyten, J. Estimated Research and Development Investment Needed to Bring a New Medicine to Market, 2009-2018. \u003cem\u003eJAMA\u003c/em\u003e 323, 844 (2020).\u003c/li\u003e\n\u003cli\u003eOECD. \u003cem\u003eHealth at a Glance 2023: OECD Indicators\u003c/em\u003e. (OECD, 2023). doi:10.1787/7a7afb35-en.\u003c/li\u003e\n\u003cli\u003eK\u0026ouml;hler, G. \u0026amp; Milstein, C. Continuous cultures of fused cells secreting antibody of predefined specificity. \u003cem\u003eNature\u003c/em\u003e 256, 495\u0026ndash;497 (1975).\u003c/li\u003e\n\u003cli\u003eClackson, T., Hoogenboom, H. R., Griffiths, A. D. \u0026amp; Winter, G. Making antibody fragments using phage display libraries. \u003cem\u003eNature\u003c/em\u003e 352, 624\u0026ndash;628 (1991).\u003c/li\u003e\n\u003cli\u003eBoder, E. T. \u0026amp; Wittrup, K. D. Yeast surface display for screening combinatorial polypeptide libraries. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e 15, 553\u0026ndash;557 (1997).\u003c/li\u003e\n\u003cli\u003eBrenner, S. \u0026amp; Lerner, R. A. Encoded combinatorial chemistry. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e 89, 5381\u0026ndash;5383 (1992).\u003c/li\u003e\n\u003cli\u003eGura, T. DNA helps build molecular libraries for drug testing. \u003cem\u003eScience\u003c/em\u003e 350, 1139\u0026ndash;1140 (2015).\u003c/li\u003e\n\u003cli\u003eMullard, A. DNA-encoded drug libraries come of age. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e 34, 450\u0026ndash;451 (2016).\u003c/li\u003e\n\u003cli\u003eGoodnow, R. A., Dumelin, C. E. \u0026amp; Keefe, A. D. DNA-encoded chemistry: enabling the deeper sampling of chemical space. \u003cem\u003eNat. Rev. Drug Discov.\u003c/em\u003e 16, 131\u0026ndash;147 (2017).\u003c/li\u003e\n\u003cli\u003e Satz, A. L. \u003cem\u003eet al.\u003c/em\u003e DNA-encoded chemical libraries. \u003cem\u003eNat. Rev. Methods Primer\u003c/em\u003e 2, 1\u0026ndash;17 (2022).\u003c/li\u003e\n\u003cli\u003e Peterson, A. A. \u0026amp; Liu, D. R. Small-molecule discovery through DNA-encoded libraries. \u003cem\u003eNat. Rev. Drug Discov.\u003c/em\u003e 22, 699\u0026ndash;722 (2023).\u003c/li\u003e\n\u003cli\u003e Gloger, A. \u0026amp; Scheuermann, J. DNA-encoded chemical libraries on stage. \u003cem\u003eNat. Chem. Biol.\u003c/em\u003e 21, 20\u0026ndash;21 (2025).\u003c/li\u003e\n\u003cli\u003e Thalji, R. K. \u003cem\u003eet al.\u003c/em\u003e Discovery of 1-(1,3,5-triazin-2-yl)piperidine-4-carboxamides as inhibitors of soluble epoxide hydrolase. \u003cem\u003eBioorg. Med. Chem. Lett.\u003c/em\u003e 23, 3584\u0026ndash;3588 (2013).\u003c/li\u003e\n\u003cli\u003e Keller, M., Schira, K. \u0026amp; Scheuermann, J. Impact of DNA-Encoded Chemical Library Technology on Drug Discovery. \u003cem\u003eCHIMIA\u003c/em\u003e 76, 388\u0026ndash;395 (2022).\u003c/li\u003e\n\u003cli\u003e Neri, D. \u0026amp; Lerner, R. A. DNA-Encoded Chemical Libraries: A Selection System Based on Endowing Organic Compounds with Amplifiable Information. \u003cem\u003eAnnu. Rev. Biochem.\u003c/em\u003e 87, 479\u0026ndash;502 (2018).\u003c/li\u003e\n\u003cli\u003e Mannocci, L. \u003cem\u003eet al.\u003c/em\u003e High-throughput sequencing allows the identification of binding molecules isolated from DNA-encoded chemical libraries. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e 105, 17670\u0026ndash;17675 (2008).\u003c/li\u003e\n\u003cli\u003e Clark, M. A. \u003cem\u003eet al.\u003c/em\u003e Design, synthesis and selection of DNA-encoded small-molecule libraries. \u003cem\u003eNat. Chem. Biol.\u003c/em\u003e 5, 647\u0026ndash;654 (2009).\u003c/li\u003e\n\u003cli\u003e \u003cem\u003eDNA-Encoded Libraries\u003c/em\u003e. b000000342 (Georg Thieme Verlag KG, Stuttgart, 2024). doi:10.1055/b000000342.\u003c/li\u003e\n\u003cli\u003e Huang, Y., Li, Y. \u0026amp; Li, X. Strategies for developing DNA-encoded libraries beyond binding assays. \u003cem\u003eNat. Chem.\u003c/em\u003e 14, 129\u0026ndash;140 (2022).\u003c/li\u003e\n\u003cli\u003e Vummidi, B. R. \u003cem\u003eet al.\u003c/em\u003e A mating mechanism to generate diversity for the Darwinian selection of DNA-encoded synthetic molecules. \u003cem\u003eNat. Chem.\u003c/em\u003e 14, 141\u0026ndash;152 (2022).\u003c/li\u003e\n\u003cli\u003e Decurtins, W. \u003cem\u003eet al.\u003c/em\u003e Automated screening for small organic ligands using DNA-encoded chemical libraries. \u003cem\u003eNat. Protoc.\u003c/em\u003e 11, 764\u0026ndash;780 (2016).\u003c/li\u003e\n\u003cli\u003e Ding, Y. \u003cem\u003eet al.\u003c/em\u003e Discovery of soluble epoxide hydrolase inhibitors through DNA-encoded library technology (ELT). \u003cem\u003eBioorg. Med. Chem.\u003c/em\u003e 41, 116216 (2021).\u003c/li\u003e\n\u003cli\u003e Cuozzo, J. W. \u003cem\u003eet al.\u003c/em\u003e Novel Autotaxin Inhibitor for the Treatment of Idiopathic Pulmonary Fibrosis: A Clinical Candidate Discovered Using DNA-Encoded Chemistry. \u003cem\u003eJ. Med. Chem.\u003c/em\u003e 63, 7840\u0026ndash;7856 (2020).\u003c/li\u003e\n\u003cli\u003e Harris, P. A. \u003cem\u003eet al.\u003c/em\u003e DNA-Encoded Library Screening Identifies Benzo[ \u003cem\u003eb\u003c/em\u003e ][1,4]oxazepin-4-ones as Highly Potent and Monoselective Receptor Interacting Protein 1 Kinase Inhibitors. \u003cem\u003eJ. Med. Chem.\u003c/em\u003e 59, 2163\u0026ndash;2178 (2016).\u003c/li\u003e\n\u003cli\u003e Malone, M. L. \u0026amp; Paegel, B. M. What is a \u0026ldquo;DNA-Compatible\u0026rdquo; Reaction? \u003cem\u003eACS Comb. Sci.\u003c/em\u003e 18, 182\u0026ndash;187 (2016).\u003c/li\u003e\n\u003cli\u003e Sauter, B., Schneider, L., Stress, C. \u0026amp; Gillingham, D. An assessment of the mutational load caused by various reactions used in DNA encoded libraries. \u003cem\u003eBioorg. Med. Chem.\u003c/em\u003e 52, 116508 (2021).\u003c/li\u003e\n\u003cli\u003e Fitzgerald, P. R. \u0026amp; Paegel, B. M. DNA-Encoded Chemistry: Drug Discovery from a Few Good Reactions. \u003cem\u003eChem. Rev.\u003c/em\u003e 121, 7155\u0026ndash;7177 (2021).\u003c/li\u003e\n\u003cli\u003e Keller, M. \u003cem\u003eet al.\u003c/em\u003e Highly pure DNA-encoded chemical libraries by dual-linker solid-phase synthesis. \u003cem\u003eScience\u003c/em\u003e 384, 1259\u0026ndash;1265 (2024).\u003c/li\u003e\n\u003cli\u003e Hunter, J. H. \u003cem\u003eet al.\u003c/em\u003e Functional Group Tolerance of a Micellar on-DNA Suzuki\u0026ndash;Miyaura Cross-Coupling Reaction for DNA-Encoded Library Design. \u003cem\u003eJ. Org. Chem.\u003c/em\u003e 86, 17930\u0026ndash;17935 (2021).\u003c/li\u003e\n\u003cli\u003e Wang, X., Li, L., Shen, X. \u0026amp; Lu, X. Rational Design Strategies in DNA‐Encoded Libraries for Drug Discovery. \u003cem\u003eAngew. Chem. Int. Ed.\u003c/em\u003e 64, e202511839 (2025).\u003c/li\u003e\n\u003cli\u003e Su, W. \u003cem\u003eet al.\u003c/em\u003e Triaging of DNA-Encoded Library Selection Results by High-Throughput Resynthesis of DNA\u0026ndash;Conjugate and Affinity Selection Mass Spectrometry. \u003cem\u003eBioconjug. Chem.\u003c/em\u003e 32, 1001\u0026ndash;1007 (2021).\u003c/li\u003e\n\u003cli\u003e Ratnayake, A. S. \u003cem\u003eet al.\u003c/em\u003e Toward the assembly and characterization of an encoded library hit confirmation platform: Bead-Assisted Ligand Isolation Mass Spectrometry (BALI-MS). \u003cem\u003eBioorg. Med. Chem.\u003c/em\u003e 41, 116205 (2021).\u003c/li\u003e\n\u003cli\u003e Chai, J., Lu, X., Arico-Muendel, C. C., Ding, Y. \u0026amp; Pollastri, M. P. Application of l -Threonine Aldolase to on-DNA Reactions. \u003cem\u003eBioconjug. Chem.\u003c/em\u003e 32, 1973\u0026ndash;1978 (2021).\u003c/li\u003e\n\u003cli\u003e Thomas, B. \u003cem\u003eet al.\u003c/em\u003e Application of Biocatalysis to on-DNA Carbohydrate Library Synthesis. \u003cem\u003eChemBioChem\u003c/em\u003e 18, 858\u0026ndash;863 (2017).\u003c/li\u003e\n\u003cli\u003e Fair, R. J., Walsh, R. T. \u0026amp; Hupp, C. D. The expanding reaction toolkit for DNA-encoded libraries. \u003cem\u003eBioorg. Med. Chem. Lett.\u003c/em\u003e 51, 128339 (2021).\u003c/li\u003e\n\u003cli\u003e Fitzgerald, P. R., Dixit, A., Zhang, C., Mobley, D. L. \u0026amp; Paegel, B. M. Building Block-Centric Approach to DNA-Encoded Library Design. \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e 64, 4661\u0026ndash;4672 (2024).\u003c/li\u003e\n\u003cli\u003e Building Blocks Catalog. \u003cem\u003eEnamine\u003c/em\u003e https://enamine.net/building-blocks/building-blocks-catalog.\u003c/li\u003e\n\u003cli\u003e McGrath, N. A., Brichacek, M. \u0026amp; Njardarson, J. T. A Graphical Journey of Innovative Organic Architectures That Have Improved Our Lives. \u003cem\u003eJ. Chem. Educ.\u003c/em\u003e 87, 1348\u0026ndash;1349 (2010).\u003c/li\u003e\n\u003cli\u003e Franzini, R. M. \u003cem\u003eet al.\u003c/em\u003e Systematic Evaluation and Optimization of Modification Reactions of Oligonucleotides with Amines and Carboxylic Acids for the Synthesis of DNA-Encoded Chemical Libraries. \u003cem\u003eBioconjug. Chem.\u003c/em\u003e 25, 1453\u0026ndash;1461 (2014).\u003c/li\u003e\n\u003cli\u003e Philpott, H. K., Thomas, P. J., Tew, D., Fuerst, D. E. \u0026amp; Lovelock, S. L. A versatile biosynthetic approach to amide bond formation. \u003cem\u003eGreen Chem.\u003c/em\u003e 20, 3426\u0026ndash;3431 (2018).\u003c/li\u003e\n\u003cli\u003e Petchey, M. \u003cem\u003eet al.\u003c/em\u003e The Broad Aryl Acid Specificity of the Amide Bond Synthetase McbA Suggests Potential for the Biocatalytic Synthesis of Amides. \u003cem\u003eAngew. Chem. Int. Ed.\u003c/em\u003e 57, 11584\u0026ndash;11588 (2018).\u003c/li\u003e\n\u003cli\u003e Winn, M., Richardson, S. M., Campopiano, D. J. \u0026amp; Micklefield, J. Harnessing and engineering amide bond forming ligases for the synthesis of amides. \u003cem\u003eCurr. Opin. Chem. Biol.\u003c/em\u003e 55, 77\u0026ndash;85 (2020).\u003c/li\u003e\n\u003cli\u003e Torri, D. \u003cem\u003eet al.\u003c/em\u003e Enzymatic Cascades for Stereoselective and Regioselective Amide Bond Assembly. \u003cem\u003eAngew. Chem. Int. Ed.\u003c/em\u003e e202422185 (2025) doi:10.1002/anie.202422185.\u003c/li\u003e\n\u003cli\u003e Tang, Q. \u003cem\u003eet al.\u003c/em\u003e Broad Spectrum Enantioselective Amide Bond Synthetase from \u003cem\u003eStreptoalloteichus hindustanus\u003c/em\u003e. \u003cem\u003eACS Catal.\u003c/em\u003e 1021\u0026ndash;1029 (2024) doi:10.1021/acscatal.3c05656.\u003c/li\u003e\n\u003cli\u003e Lima, R. N., Dos Anjos, C. S., Orozco, E. V. M. \u0026amp; Porto, A. L. M. Versatility of Candida antarctica lipase in the amide bond formation applied in organic synthesis and biotechnological processes. \u003cem\u003eMol. Catal.\u003c/em\u003e 466, 75\u0026ndash;105 (2019).\u003c/li\u003e\n\u003cli\u003e Bering, L., Craven, E. J., Sowerby Thomas, S. A., Shepherd, S. A. \u0026amp; Micklefield, J. Merging enzymes with chemocatalysis for amide bond synthesis. \u003cem\u003eNat. Commun.\u003c/em\u003e 13, (2022).\u003c/li\u003e\n\u003cli\u003e Petchey, M. R. \u0026amp; Grogan, G. Enzyme‐Catalysed Synthesis of Secondary and Tertiary Amides. \u003cem\u003eAdv. Synth. Catal.\u003c/em\u003e 361, 3895\u0026ndash;3914 (2019).\u003c/li\u003e\n\u003cli\u003e Sim, E., Walters, K. \u0026amp; Boukouvala, S. Arylamine N-acetyltransferases: From Structure to Function. \u003cem\u003eDrug Metab. Rev.\u003c/em\u003e 40, 479\u0026ndash;510 (2008).\u003c/li\u003e\n\u003cli\u003e Leli\u0026egrave;vre, C. M., Balandras, M., Petit, J., Vergne‐Vaxelaire, C. \u0026amp; Zaparucha, A. ATP Regeneration System in Chemoenzymatic Amide Bond Formation with Thermophilic CoA Ligase. \u003cem\u003eChemCatChem\u003c/em\u003e 12, 1184\u0026ndash;1189 (2020).\u003c/li\u003e\n\u003cli\u003e Orengo, C. \u003cem\u003eet al.\u003c/em\u003e CATH \u0026ndash; a hierarchic classification of protein domain structures. \u003cem\u003eStructure\u003c/em\u003e 5, 1093\u0026ndash;1109 (1997).\u003c/li\u003e\n\u003cli\u003e Xu, D., Wang, Z., Zhuang, W., Wang, T. \u0026amp; Xie, Y. Family characteristics, phylogenetic reconstruction, and potential applications of the plant BAHD acyltransferase family. \u003cem\u003eFront. Plant Sci.\u003c/em\u003e 14, 1218914 (2023).\u003c/li\u003e\n\u003cli\u003e Burckhardt, R. M. \u0026amp; Escalante-Semerena, J. C. Small-Molecule Acetylation by GCN5-Related \u003cem\u003eN\u003c/em\u003e -Acetyltransferases in Bacteria. \u003cem\u003eMicrobiol. Mol. Biol. Rev.\u003c/em\u003e 84, (2020).\u003c/li\u003e\n\u003cli\u003e Bittner, P. \u003cem\u003eet al.\u003c/em\u003e Native Mass Spectrometry Facilitates Hit Validation in DNA‐Encoded Library Technology. \u003cem\u003eAngew. Chem. Int. Ed.\u003c/em\u003e e202504470 (2025) doi:10.1002/anie.202504470.\u003c/li\u003e\n\u003cli\u003e Cyrus, K. \u003cem\u003eet al.\u003c/em\u003e Impact of linker length on the activity of PROTACs. \u003cem\u003eMol BioSyst\u003c/em\u003e 7, 359\u0026ndash;364 (2011).\u003c/li\u003e\n\u003cli\u003e Bittner, P. \u003cem\u003eet al.\u003c/em\u003e The Influence of Single-Stranded or Double-Stranded DNA Tags on Ligand Binding Affinity in DNA-Encoded Libraries. \u003cem\u003eAnal. Chem.\u003c/em\u003e (2025) doi:10.1021/acs.analchem.5c03540.\u003c/li\u003e\n\u003cli\u003e Bittrich, S., Segura, J., Duarte, J. M., Burley, S. K. \u0026amp; Rose, Y. RCSB protein Data Bank: exploring protein 3D similarities via comprehensive structural alignments. \u003cem\u003eBioinformatics\u003c/em\u003e 40, btae370 (2024).\u003c/li\u003e\n\u003cli\u003e Guerra, J. V. D. S. \u003cem\u003eet al.\u003c/em\u003e pyKVFinder: an efficient and integrable Python package for biomolecular cavity detection and characterization in data science. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e 22, 607 (2021).\u003c/li\u003e\n\u003cli\u003e Stourac, J. \u003cem\u003eet al.\u003c/em\u003e Caver Web 1.0: identification of tunnels and channels in proteins and analysis of ligand transport. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e 47, W414\u0026ndash;W422 (2019).\u003c/li\u003e\n\u003cli\u003e Corbella, M. \u003cem\u003eet al.\u003c/em\u003e The N-terminal Helix-Turn-Helix Motif of Transcription Factors MarA and Rob Drives DNA Recognition. \u003cem\u003eJ. Phys. Chem. B\u003c/em\u003e 125, 6791\u0026ndash;6806 (2021).\u003c/li\u003e\n\u003cli\u003e Passaro, S. \u003cem\u003eet al.\u003c/em\u003e Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. Preprint at https://doi.org/10.1101/2025.06.14.659707 (2025).\u003c/li\u003e\n\u003cli\u003e Horton, R. M., Cai, Z., Ho, S. N. \u0026amp; Pease, L. R. Gene Splicing by Overlap Extension: Tailor-Made Genes Using the Polymerase Chain Reaction. \u003cem\u003eBioTechniques\u003c/em\u003e 54, 129\u0026ndash;133 (2013).\u003c/li\u003e\n\u003cli\u003e Privett, H. K. \u003cem\u003eet al.\u003c/em\u003e Iterative approach to computational enzyme design. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e 109, 3790\u0026ndash;3795 (2012).\u003c/li\u003e\n\u003cli\u003e Tanimoto, T. T. \u003cem\u003eAn Elementary Mathematical Theory of Classification and Prediction\u003c/em\u003e. (International Business Machines Corporation, 1958).\u003c/li\u003e\n\u003cli\u003e Li, Y. \u003cem\u003eet al.\u003c/em\u003e Optimized Reaction Conditions for Amide Bond Formation in DNA-Encoded Combinatorial Libraries. \u003cem\u003eACS Comb. Sci.\u003c/em\u003e 18, 438\u0026ndash;443 (2016).\u003c/li\u003e\n\u003cli\u003e Li, Y., Zimmermann, G., Scheuermann, J. \u0026amp; Neri, D. Quantitative PCR is a Valuable Tool to Monitor the Performance of DNA‐Encoded Chemical Library Selections. \u003cem\u003eChemBioChem\u003c/em\u003e 18, 848\u0026ndash;852 (2017).\u003c/li\u003e\n\u003cli\u003e Zoltewicz, J. A., Clark, D. F., Sharpless, T. W. \u0026amp; Grahe, G. Kinetics and mechanism of the acid-catalyzed hydrolysis of some purine nucleosides. \u003cem\u003eJ. Am. Chem. Soc.\u003c/em\u003e 92, 1741\u0026ndash;1750 (1970).\u003c/li\u003e\n\u003cli\u003e Potowski, M. \u003cem\u003eet al.\u003c/em\u003e Screening of metal ions and organocatalysts on solid support-coupled DNA oligonucleotides guides design of DNA-encoded reactions. \u003cem\u003eChem. Sci.\u003c/em\u003e 10, 10481\u0026ndash;10492 (2019).\u003c/li\u003e\n\u003cli\u003e Kim, G. A., Kang, S., Win, M. N. \u0026amp; Han, M. S. Simple and High-Throughput Fluorescence Assay Method for DNA Damage Analysis in Single-Stranded DNA-Encoded Library Synthesis. \u003cem\u003eBioconjug. Chem.\u003c/em\u003e (2025) doi:10.1021/acs.bioconjchem.5c00219.\u003c/li\u003e\n\u003cli\u003e G\u0026ouml;tte, K., Chines, S. \u0026amp; Brunschweiger, A. Reaction development for DNA-encoded library technology: From evolution to revolution? \u003cem\u003eTetrahedron Lett.\u003c/em\u003e 61, 151889 (2020).\u003c/li\u003e\n\u003cli\u003e Song, M. \u0026amp; Hwang, G. T. DNA-Encoded Library Screening as Core Platform Technology in Drug Discovery: Its Synthetic Method Development and Applications in DEL Synthesis. \u003cem\u003eJ. Med. Chem.\u003c/em\u003e 63, 6578\u0026ndash;6599 (2020).\u003c/li\u003e\n\u003cli\u003e Buller, R., Damborsky, J., Hilvert, D. \u0026amp; Bornscheuer, U. T. Structure Prediction and Computational Protein Design for Efficient Biocatalysts and Bioactive Proteins. \u003cem\u003eAngew. Chem.\u003c/em\u003e 137, (2025).\u003c/li\u003e\n\u003cli\u003e Buller, R. \u003cem\u003eet al.\u003c/em\u003e From nature to industry: Harnessing enzymes for biocatalysis. \u003cem\u003eScience\u003c/em\u003e 382, (2023).\u003c/li\u003e\n\u003cli\u003e Schr\u0026ouml;dinger, LLC. The PyMOL Molecular Graphics System, Version 2.5. (2015).\u003c/li\u003e\n\u003cli\u003e Mirdita, M. \u003cem\u003eet al.\u003c/em\u003e ColabFold: making protein folding accessible to all. \u003cem\u003eNat. Methods\u003c/em\u003e 19, 679\u0026ndash;682 (2022).\u003c/li\u003e\n\u003cli\u003e McGibbon, R. T. \u003cem\u003eet al.\u003c/em\u003e MDTraj: A Modern Open Library for the Analysis of Molecular Dynamics Trajectories. \u003cem\u003eBiophys. J.\u003c/em\u003e 109, 1528\u0026ndash;1532 (2015).\u003c/li\u003e\n\u003cli\u003e Pedregosa, F. \u003cem\u003eet al.\u003c/em\u003e Scikit-learn: Machine Learning in Python. \u003cem\u003eJ. Mach. Learn. Res.\u003c/em\u003e 12, 2825\u0026ndash;2830 (2011).\u003c/li\u003e\n\u003cli\u003e Salentin, S., Schreiber, S., Haupt, V. J., Adasme, M. F. \u0026amp; Schroeder, M. PLIP: fully automated protein\u0026ndash;ligand interaction profiler. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e 43, W443\u0026ndash;W447 (2015).\u003c/li\u003e\n\u003cli\u003e Eichenberger, M. \u003cem\u003eet al.\u003c/em\u003e Asymmetric Cation‐Olefin Monocyclization by Engineered Squalene\u0026ndash;Hopene Cyclases. \u003cem\u003eAngew. Chem. Int. Ed.\u003c/em\u003e 60, 26080\u0026ndash;26086 (2021).\u003c/li\u003e\n\u003cli\u003e Li, C. \u003cem\u003eet al.\u003c/em\u003e FastCloning: a highly simplified, purification-free, sequence- and ligation-independent PCR cloning method. \u003cem\u003eBMC Biotechnol.\u003c/em\u003e 11, 92 (2011).\u003c/li\u003e\n\u003cli\u003e Vardakou, M., Salmon, M., Faraldos, J. A. \u0026amp; O\u0026rsquo;Maille, P. E. Comparative analysis and validation of the malachite green assay for the high throughput biochemical characterization of terpene synthases. \u003cem\u003eMethodsX\u003c/em\u003e 1, 187\u0026ndash;196 (2014).\u003c/li\u003e\n\u003cli\u003e Morgan, H. L. The Generation of a Unique Machine Description for Chemical Structures-A Technique Developed at Chemical Abstracts Service. \u003cem\u003eJ. Chem. Doc.\u003c/em\u003e 5, 107\u0026ndash;113 (1965).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Biocatalysis, DNA-encoded library, Protein engineering, Amide-bond formation, Chemoenzymatic cascade","lastPublishedDoi":"10.21203/rs.3.rs-7598475/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7598475/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"DNA-encoded chemical library (DEL) technology has emerged as a powerful tool in early-stage drug discovery. Although the methodology is widely applied in industry and academia, major challenges persist in generating DELs with high quality and chemical diversity. Low yields in building-block incorporation, limited stereo-, regio-, and chemoselectivity, and, most importantly, DNA damage from harsh reaction conditions compromise library quality, reduce signal-to-noise in affinity selections, and ultimately hinder drug discovery. Here, we show that tailored enzymes can be harnessed for the effective construction of molecular diversity on DNA, opening new avenues for the generation of high-quality and diverse DELs under mild conditions. Targeting amide bond formation, we engineered a cascade of complementary CoA ligases and rationally tailored N-acyltransferases (NATs) to access a broad amide scope on-DNA (\u003e 120 examples), identifying transferable structural motifs that optimize DNA-compatibility of the biocatalysts in the process. The successful integration of the biocatalytic steps with chemical synthesis led to the construction of a diverse DEL without damage to the DNA barcode and underscored the broad utility of the enzymatic cascade, enabling applications in both early scaffold construction and late-stage functionalization.","manuscriptTitle":"A tailored enzyme cascade facilitates DNA-encoded library technology and gives access to a broad substrate scope","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-24 04:32:30","doi":"10.21203/rs.3.rs-7598475/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-catalysis","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"natcatal","sideBox":"Learn more about [Nature Catalysis](http://www.nature.com/natcatal/)","snPcode":"","submissionUrl":"","title":"Nature Catalysis","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"962f4920-138c-41fd-9b93-e547384b725a","owner":[],"postedDate":"September 24th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":55170610,"name":"Biological sciences/Chemical biology/Biocatalysis"},{"id":55170611,"name":"Biological sciences/Chemical biology/Enzymes"},{"id":55170612,"name":"Biological sciences/Chemical biology/Chemical libraries/Combinatorial libraries"}],"tags":[],"updatedAt":"2026-05-06T08:52:05+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-24 04:32:30","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7598475","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7598475","identity":"rs-7598475","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0