Protein-mediated folding of the genome is essential for site-specific integration of foreign DNA into CRISPR loci | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Protein-mediated folding of the genome is essential for site-specific integration of foreign DNA into CRISPR loci Andrew Santiago-Frangos, William Henriques, Tanner Wiegand, Colin Gauvin, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2982802/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Sep, 2023 Read the published version in Nature Structural & Molecular Biology → Version 1 posted You are reading this latest preprint version Abstract Bacteria and archaea acquire resistance to viruses and plasmids by integrating fragments of foreign DNA into the first repeat of a CRISPR array. However, the mechanism of site-specific integration remains poorly understood. Here, we determine a 560 kDa integration complex structure that explains how Cas (Cas1-2/3) and non-Cas proteins (IHF) fold 150 base-pairs of host DNA into a U-shaped bend and a loop that protrude from Cas1-2/3 at right angles. The U-shaped bend traps foreign DNA on one face of the Cas1-2/3 integrase, while the loop places the first CRISPR repeat in the Cas1 active site. Both Cas3s rotate 100-degrees to expose DNA binding sites on either side of the Cas2 homodimer, that each bind an inverted repeat motif in the leader. Leader sequence motifs direct Cas1-2/3-mediated integration to diverse repeat sequences that have a 5’-GT. Biological sciences/Structural biology/Electron microscopy/Cryoelectron microscopy Biological sciences/Biochemistry/Proteins CRISPR-Cas integration CRISPR adaptation Cas1-2/3 IHF DNA recording devices transposition transposase DNA bending Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Full Text Vertebrates, bacteria, and archaea have domesticated transposases (e.g., RAG1 and Cas1) for adaptive immunity 1,2 . Integrases, transposases and recombinases often co-opt additional DNA-bending proteins (e.g., IHF, HU, H-NS, or HMGB1) that facilitate DNA integration and excision 3–7 . However, the structural role of DNA folding during this mobilization of DNA remains largely enigmatic. C lustered R egularly I nterspaced S hort P alindromic R epeats (CRISPRs) are essential components of an adaptive immune system that stores DNA-based molecular memories of past infections 8 . C RISPR- as sociated proteins, Cas1 and Cas2, integrate fragments of foreign DNA ("spacers") into CRISPRs. Integration duplicates a repeat sequence, which thereby maintains the characteristic repeat-spacer-repeat architecture ( Fig. 1a ). Cas1 and Cas2 form a heterohexameric complex that consists of two Cas1 homodimers (Cas1a-a* and Cas1b-b*) flanking a Cas2 homodimer ( Fig. 1a-b ) 9–11 . Foreign DNA fragments bind across one face of the Cas2 homodimer, which positions the 3’-ends into Cas1 active sites on either end of the complex (i.e., Cas1a* and Cas1b*) 8–10,12 . The CRISPR repeat sequence wraps around the opposing face of Cas2, sandwiching the Cas2 homodimer between the foreign and repeat DNA duplexes. Opposing Cas1 subunits (Cas1a* and Cas1b*) catalyze two successive strand-transfer reactions, linking the 3'-ends of the foreign DNA to opposite ends of the repeat 5,11 . CRISPR integration complexes sense a 2-5 bp 3’ overhang called a protospacer-adjacent motif (PAM) in the foreign DNA to determine the integration orientation 9 . Correct spacer orientation is necessary to produce a functional CRISPR RNA that guides the CRISPR interference machinery (i.e., Cascade) to complementary targets 8,13 . Integration occurs in a stepwise manner. First, the non-PAM end of the foreign DNA is integrated at the leader-side of the repeat 14 . Second, the PAM is cleaved by Cas or non-Cas nucleases and the trimmed 3’ end is integrated at the spacer-side of the repeat 14–16 . These integration events tie a non-covalent knot around the Cas2 homodimer (foreign DNA on one side and repeat DNA on the other) that is held together by complementary base-pairing in the foreign DNA. New foreign DNA is preferentially integrated at the first repeat in a CRISPR locus, ensuring efficient transcription and processing of CRISPR RNAs that target the most recently encountered genetic parasites ( Fig. 1a ) 17,18 . Cas1-2 is thought to recognize a palindromic sequence within the CRISPR repeat, similar to target site recognition by many DNA transposases 5,11,19–21 . However, Cas1-2 recognition of the palidromic repeat does not explain how the first repeat is differentiated from downstream repeat sequences in a CRISPR. Thus, polarized integration often relies on additional proteins and DNA sequence motifs upstream of the CRISPR (i.e., leader) 3–5,11,18,22–27 . I ntegration H ost F actor (IHF) facilitates polarized integration in the type I-E CRISPR system from Escherichia coli 3 . A structure of the I-E integration complex revealed that IHF bends the leader DNA to bring an upstream sequence motif into contact with Cas1, and IHF further stabilizes the Cas1-2 integrase at the first repeat through direct Cas1-IHF interactions 3,5 . Cas1 and Cas2 are conserved components of CRISPR-mediated immune systems. However, the type I-F CRISPR system has a unique fusion of the Cas2 subunit to the Cas3-nuclease/helicase found in many type I systems ( Fig. 1a-b ) 28,29 . Cas3 degrades Cascade-bound DNA into fragments with PAM-containing termini, that are captured by Cas1-2 and integrated into the CRISPR locus, in a process called “primed acquisition” 4,30–37 . However, the structural mechanism for primed acquisition is unclear. Additionally, in contrast to the E. coli I-E system, which relies on one IHF to bend DNA and recruit a DNA motif found ~50 bp upstream of the CRISPR repeat, most I-F, I-C, and some I-E CRISPR leaders, contain multiple IHF binding sites and multiple subtype-specific DNA motifs found up to 100-200 bp upstream of the CRISPR repeat ( Fig. 1a and Extended Data Fig. 1a ) 22 . To determine how DNA sequence motifs in the CRISPR leader regulate the Cas1-2/3 integrase, we determined the structure of a ~560 kDa CRISPR integration complex from Pseudomonas aeruginosa . The structure reveals that Cas1-2/3 cannot interact with the leader without first undergoing a large conformational change that may be induced by foreign DNA binding. Further, the structure explains how the I-F leader and IHF proteins guide Cas1-2/3 to deliver and integrate foreign DNA at the first repeat in the CRISPR array ( Fig. 1 ). Cas1-2/3 and IHF interact with all five DNA sequence motifs (i.e., two Inverted Repeats, two IHF binding sites, and the CRISPR repeat) primarily through a shape-based readout 38,39 . The shape of the folded I-F CRISPR leader is similar to that of the λ-phage excision complex, suggesting that DNA is often used as a flexible scaffold to regulate DNA mobilization 3–7,40–53 . The structure suggests that site-specific integration relies on protein-induced folding of the upstream DNA rather than sequence-specific recognition of the repeat. To test this idea, we perform a series of integration reactions demonstrating that efficient integration relies on conserved sequences in the leader and a 5’ GT dinucleotide in the repeat. We show that 5’ GT dinucleotides are broadly conserved in repeats derived from different CRISPR types, suggesting that they play a conserved role in integration across diverse CRISPR systems. In addition, the I-F CRISPR integration complex suggests a structural mechanism for interactions of the Cas1-2/3 integrase with the Cascade surveillance complex, that may be neccesary for rapid adaptation to phage escape mutants 4,30–37 . Cryo-EM structure of type I-F CRISPR integration complex To understand how the Cas1-2/3 integrase cooperates with IHF and CRISPR leader motifs to integrate foreign DNA at the first CRISPR repeat, we purified the heterohexameric Cas1-2/3 integrase and the IHFα-β heterodimer, incubated these proteins with a DNA substrate representing a half-site integration intermediate, and isolated the assembled complex using S ize E xclusion C hromatography (SEC) ( Extended Data Fig. 1a-f ). The purified I-F integration complex was applied to cryo-EM grids and vitrified. We recorded 10,740 movies and picked 366,794 particles to determine a ~3.48 Å-resolution structure of the integration complex. The reconstructed density was of sufficient to model 90.7% of the 10 polypeptides and 88.4% of the 396 nucleotides of DNA ( Fig. 1b-d, Extended Data Fig. 1g-k, Table 1 ). The model explains how the Cas1-2/3 subunits cooperate with two IHF heterodimers to kink and twist ~150 base-pairs of host DNA into a structure that precisely positions foreign DNA for integration at the first repeat of the CRISPR ( Fig. 1b-d ). The Cas1 and Cas2 subunits adopt a familiar quaternary arrangement that binds a foreign DNA on one face of the Cas2 homodimer and CRISPR repeat DNA on the other face ( Fig. 1d, 2a and 3a ) 8 . A previously determined structure of Cas1-2/3 alone revealed that the Cas3 and Cas1 domains surround the central Cas2 homodimer like petals of a closed flower (Cas1 4 :Cas2/3 2 ) ( Fig. 1b ) 54 . While this structure explained how Cas1 regulates the Cas3 nuclease, the role of Cas3 during integration remained unclear 54 . Here, we show that the addition of DNA drives a series of conformational changes in both the DNA and proteins. The Cas3 domains rotate ~100° to align in a planar configuration with Cas2, simulating the motion of a bloomed flower, and exposing equivalent surfaces on opposite sides of the Cas2 homodimer that recognize an inverted repeat (IR) that is conserved in I-F leaders ( Fig. 1d, Extended Data Fig. 2a and Supplementary Video 1 ) 54 . Thus, the new planar conformation of Cas1-2/3 enables the simultaneous coordination of four DNA helices (IR distal , IR proximal , foreign DNA, and CRISPR repeat) around the central Cas2 homodimer ( Fig. 1d and 3 ). Further, this Cas3 rotation flips the nuclease domain from an interaction with Cas1 that suppresses the Cas3 nuclease activity, to the opposite side of the complex, where the back of the Cas3 nuclease domain docks onto a groove created at the Cas1-Cas1 interaface ( Fig. 1b,d and Extended Data Fig. 2a,b ). The structure reveals two prominent DNA bends that protrude at right angles from Cas1-2/3 ( Fig. 1c,d ). An IHF heterodimer is wedged at the apex of each DNA bend, consistent with IHF's well-defined role in DNA bending 38 . These two DNA protrussions extend ~75 Å from the Cas1-2 core. Flexablity of these DNA extentions limits the resolution of the regions to 4-8 Å ( Fig. 1c,d and Supplementary Video 2 ). IHF-mediated bending of the IHF distal site positions the flanking IR sequences as symmetrical DNA pillars, which are recognized by equivalent surfaces on opposite sides of the Cas2 homodimer ( Fig. 3 and Extended Data Fig. 3b and 4c ) 22 . Cas2 binding to these DNA pillars traps foreign DNA on one face of the Cas1-2/3 integrase. Further, Cas2 bends the IRs and steers downstream DNA away from Cas1-2/3, which would project the downstream CRISPR repeat away from the Cas1-2/3 integrase ( Fig. 1d ). However, Cas1-2/3 and IHF cooperate to constrict the DNA around the IHF proximal site, forming a loop that places the CRISPR repeat into the Cas1a* active site ( Fig. 1d and 3 ). Foreign DNA constrains the Cas2/3 linker against conserved Cas1 surfaces The type I-F Cas2 and Cas3 subunits are connected by a 20 amino acid disordered linker (residues 90-110) 4,28,29,55 . The structure explains how foreign DNA constrains the Cas2/3 linker against conserved surfaces of Cas1, which suggests that foreign DNA-binding either initiates, or stabilizes the Cas3 rotation ( Fig. 2a ) 54,55 . The constrained Cas2/3 linker positions the HD nuclease domain of Cas3 (residues 111-374) against the Cas1-Cas1 interface, and facilitates Cas3 interactions with the IRs ( Fig. 2a and Extended Data Fig. 2 and 4c ). The foreign DNA and amino acids in the Cas2/3 linker contact conserved residues in type I-F Cas1 proteins ( Fig. 2d,e and Extended Data Fig. 2c ). Polar residues in the Cas2/3 linker may assist the binding or splaying of the foreign DNA duplex at the conserved histidine wedge (H25) in Cas1 ( Fig. 2b,c and Extended Data Fig. 3 ). Mutation of the histidine wedge (Cas1 H25A ) decreases Cas1-2/3 integration activity of foreign DNA that has either fully complementary or splayed DNA ends ( Extended Data Fig. 5 and 6a,b ). The integration defect on substrates with splayed ends suggest that H25 is more than a simple wedge that pries apart the ends for foreign DNA 10 . The histidine steers the 3’-ends down a positively charged channel that positions each 3’-hydroxyl into Cas1 active sites on opposite ends of the complex ( Fig. 2 and Fig. 4a,d ), whereas the 5’-ends of the protospacer DNA are directed towards the back face of the Cas3 HD domain ( Fig. 2 ) 9,10,56 . Cas2 homodimers recognize and bend inverted repeat sequences The structure of the type I-F CRISPR integration complex reveals that Cas2 is the homodimer that binds the IRs ( Fig. 3a ). Mutations that scramble the order of nucleotides in either the IR distal or IR proximal motifs limit Cas1-2/3-mediated integration 22 . While Cas2 doesn’t make extensive sequence-specific contacts with nucleobases of the IR, a single residue (Cas2 R55 ) intercalates in the minor groove, and may participate in recognizing two conserved bases in the 10 bp-long motif ( Fig. 3a and Extended Data Fig. 4c,e ). However, there is insufficient density for the R55 sidechain to confidently assign contacts. Other conserved Cas2 residues (i.e., K11, R12, and N56) form additional hydrogen bonds with the phosphate backbone of one DNA strand in each IR ( Fig. 3a and Extended Data Fig. 4c ). Mutation of these Cas2 residues (Cas2 K11D,R12E , Cas2 R55E,N56D , Cas2 K11D,R12E,R55E,N56D ) prevents Cas1-2/3-mediated DNA integration ( Extended Data Fig. 5 and 6c,d ). Cas2 acts as a wedge that induces a 25-35° bend in the DNA upstream of IR distal and downstream of IR proximal ( Fig. 3a ). These flared IRs lean against basic residues (K381, R393, K397) on the back surface of Cas3 ( Fig. 3 and Extended Data Fig. 3 and 4c ). In sum, these observations reveal that the IR DNA sequences are primarily recognized by Cas1-2/3 through shape readout rather than base readout 39 . Cas1-2/3 accommodates four dsDNA helices that surround the Cas2 homodimer Cas2 is a cube-shaped homodimer at the center of the Cas-integrase. The Cas2 cube is flanked by Cas1 homodimers to form an elongated DNA binding platform that interacts with the CRISPR repeat on one face and the foreign DNA on the other ( Fig. 3 ). Unique to the type I-F Cas1-2/3 integration complex, the IRs occupy the last two accessible surfaces of the Cas2 cube ( Fig. 3a ). Positively charged surfaces on Cas1-2/3 bind and shield negatively charged DNA, which enables the packing of four DNA helices around the small Cas2 homodimer ( Fig. 3b and Extended Data Fig. 3 ). The foreign DNA-binding face of Cas2 has two electronegative pillars of leader DNA that straddle the foreign DNA, such that major grooves of the leader DNA pillars are clamped against major grooves of the foreign DNA. The two DNA pillars continue past Cas2 to flank the Cas1 active sites ( Fig. 3b ). At the IHF proximal loop, Cas3 packs the leader against the Cas1-bound repeat, decreasing the phosphate-to-phosphate distances between these helices to ~11-12 Å. Although the latter two-thirds of the CRISPR repeat could not be resolved, the repeat’s trajectory suggests it will follow a path that threads between the distal leader DNA duplex and the 3’-hydroxyl of the foreign DNA that rests in the Cas1b* active site ( Fig. 3b ). Collectively, Cas1, the Cas2/3 linker, and Cas3 accommodate four DNA helices (IR distal , IR proximal , foreign DNA, and repeat) around the central Cas2 homodimer to facilitate site-specific integration. IHF and the leader sequence facilitate integration into diverse repeat sequences The structure suggests that Cas1-2/3 is guided to the first repeat of the CRISPR by IHF-mediated folding of the I-F leader, rather than direct recognition of the repeat sequence ( Fig. 1d and Extended Data Fig. 4b,d ). To determine if or how the repeat sequence impacts integration, we measured the efficiency of Cas1-2/3-catalyzed integration into DNAs containing either a I-F, I-E, I-C or II-A repeat downstream of a I-F leader ( Fig. 4a-c ). The type I-F leader supports Cas1-2/3-catalyzed leader-side integration at repeats derived from I-E, I-C, and II-A CRISPR loci ( Fig. 4a-c and Extended Data Fig. 7 and 8 ). Integration efficiency at non-native repeats is not correlated with sequence similarity to the I-F repeat or with GC-content ( Extended Data Fig. 8b ). Instead, integration efficiency is correlated to the length of the repeat. I-F and I-E repeats are similar in length (28 and 29 bps, respectively), whereas the I-C and II-A repeats are 0.5 to 1 full DNA turns longer than the I-F repeat (33 and 37 bp, respectively) ( Fig. 4b and Extended Data Fig. 9a,b ). While leader-side integration is robust with different repeats, spacer-side integration is ~3.5-fold slower for the I-E repeat, and undetectable for the longer repeats ( Extended Data Fig. 7 and 9a,b) . As expected, Cas1-2/3 does not catalyze integration at a I-F repeat downstream of a scrambled I-F leader, nor does Cas1-2/3 does catalyze integration at I-E, I-C or II-A repeats downstream of their respective leaders 22 ( Extended Data Fig. 7 ). Cas1-2 sequences are diverse, such that leader-interacting residues are only conserved in a subset of proteins within a given CRISPR subtype 22 . For example, the I-F Cas1 protein lacks residues required to interact with the I-E leader 22 . Further, the I-E, I-C and I-F leaders have distinct nucleotide spacings between the leader motifs and the repeat. These nucleotide spacings impart a unique shape to the DNA, which is critical for integration. These experiments were performed with foreign DNA substrates either with or without a PAM ( Extended Data Fig. 7 ). Cas1-2/3 integration of PAM-containing DNA is more specific, but the conclusions are otherwise consistent between the two substrates. The PAM must be trimmed by an ancillary nuclease before Cas1-2/3 can catalyze spacer-side integration, therefore we focused our discussion on results from the trimmed foreign DNA to compare differences in spacer-side integration ( Extended Data Fig. 7a-d ) 14–16 . Collectively, these integration experiments indicate that leader sequences and host factors dictate site-specific integation of foreign DNA at diverse DNA target sites, providing insight into the evolution of new CRISPR systems in nature. 5’ GT dinucleotides in CRISPR repeat sequences are critical for Cas1-mediated integration Repeat sequences are strongly conserved within CRISPR subtypes, but vary in sequence and length between subtypes 57,58 . However, in the small subset of repeats tested above, we noticed that the 5’ GT is conserved. To determine if conservation of the 5’ GT is a coincidence or a more widely conserved feature of repeats, we performed a bioinformatic analysis consisting of 24,940 CRISPRs. This bioinformatic analysis reveals that a 5’ GT dinucleotide is broadly conserved at the leader-side of the repeat, and conserved in some CRISPR systems at the spacer-side of the repeat ( Extended Data Fig. 4g ). Therefore we hypothesized that the 5’ GT dinucleotide is a base-specific determinant for leader-side integration. To test this hypothesis, we mutated the 5’ GT (G1A, T2A) and repeated the integration assays. The 5’ GT to AA mutation ablates both leader- and spacer-side integration, indicating that the 5’ GT is essential and that leader-side integration is a prerequisite for spacer-side integration ( Fig. 4d ). Since Cas1 requires a 5’ GT at the leader-side of the repeat, we hypothesized that introducing a 5’ GT at the spacer-side of the repeat would increase spacer-side integration efficiency. To test this hypothesis, we replaced adenosine 28 of the I-F repeat with cytosine (A28C) and repeated the integration assays. The A28C mutation increases the rate and amount of spacer-side integration by ~2-fold, relative to the WT I-F repeat ( Fig. 4d ). We do not detect integration into a I-F repeat that lacks a 5’ GT at the leader-side even if the repeat contains a 5’ GT at the spacer-side. This result further supports our conclusion that leader-side integration is a prerequisite for spacer-side integration ( Fig. 4d ). We examined the structure of the Cas1 active site to determine whether the 5’ G is directly recognized by protein contacts. The Cas1 residue E184 is within 4 Å of the 5’ G ( Extended Data Fig. 4b,d ). However, a Cas1 E184A mutation destabilizes the complex, decreasing the amount of Cas1 subunits per Cas1-2/3 complex ( Extended Data Fig. 5 ). Therefore, the decrease in integration activity of the Cas1 E184A -2/3 complex cannot be solely attributed to a decrease in recognition of the repeat ( Extended Data Fig. 9d,f ). Collectively, these data reveal that the 5’ GT is a conserved feature necessary for integration in most CRISPR systems, though no available structure provides a mechanism for direct recognition of the 5’ GT of the repeat 5,11,15,59 . Conclusions Here we demonstrate that Cas1-2/3 and IHF fold DNA into a structure that is necessary for site-specific integration of foreign DNA into CRISPRs. IHF proteins are highly expressed and most IHF binding sitesare thought to be occupied in vivo 60 . Therefore, IHF may pre-fold the CRISPR leader into a "landing pad" that recruits foreign DNA-bound Cas1-2/3 ( Fig. 4e ) 22,38 . The Cas3 and Cas1 domains of the Cas1-2/3 complex are arranged like petals of a closed flower around the central Cas2 homodimer, such that the Cas3 domains occlude two of the four DNA binding surfaces on Cas2, which precludes interactions with the leader ( Fig. 1b,d ). Foreign DNA binding to Cas1-2/3 physically constrains the Cas2/3 linker against the Cas1 homodimer, pulling the Cas3 HD domain against the Cas1-Cas1 interface ( Fig. 2, Extended Data Fig. 2 and Supplementary Video 1 ). The 100-degree rotation of each Cas3 simulates the motion of a bloomed flower and exposes DNA binding sites on Cas2 that interact with each of the inverted repeat (IR) motifs in the leader ( Fig. 1, 3, Extended Data Fig. 2,3 ). Cas1-2/3 and IHF proteins fold DNA around the leader-repeat junction into a 260° loop that docks the first CRISPR repeat into the Cas1 active site under tension ( Fig. 1, 4, Extended Data Fig. 3, 4 ). The structure suggests that Cas1-mediated strand-transfer releases tension in this DNA loop, which may prevent disintegration of an otherwise isoenergetic strand-transfer reaction, and thereby favour complete integration ( Extended Data Fig. 4f ). A similar mechanism has been proposed to favour complete integration in other systems, where both strand transfer events occur simultaneously 61,62 . We show that Cas1-2/3, IHF and the I-F leader facilitate leader-side integration at four different repeat sequences ( Fig. 4b-c ). These repeats are diverse in sequence identity, length, palindrome and GC-content, but they share a 5’ GT ( Fig. 4b and Extended Data Fig. 8b ). To determine whether the 5’ GT is a universal feature of CRISPR repeats we analyzed 24,940 CRISPRs. This analysis reveals that CRISPR repeats contain a strongly conserved 5’ GT at the leader-end, and that a 5’ GT is also conserved at the spacer-end of repeats from several CRISPR systems ( Extended Data Fig. 4g ). We demonstrate the 5’ GT is critical for leader-side integration and that introducing a 5’ GT at the spacer-end of the repeat increases spacer-side integration. The broad conservation of 5’ GT is consistent with previous reports that type I-A, II-A, and I-E systems require a 5’ G for integration 11,63 . Collectively, these data suggest that Cas1 proteins retain a shared sequence preference for a 5' G or a 5' GT, and lack strict sequence requirements for the central body of the repeat ( Fig. 4b,c ) 64 . The lack of strict sequence requirements may be advantageous because the CRISPR repeat is at the nexus of foreign DNA integration, processing of the transcribed CRISPR, and loading the mature crRNA into the surveillance complexes (e.g., Cascade, Cas9). Genetic parasites commonly escape CRISPR-based immunity through point mutations 65 . To counter escape mutants, many CRISPR-Cas systems use existing spacer sequences to enhance the acquisition of new spacers from the same foreign genetic element via “primed” acquisition 4,30–37 . In some examples of primed acquisition, the Cas3 nuclease/helicase degrades CRISPR-targeted DNAs into ssDNA fragments enriched in PAM-containing termini 66 . Cas1-2 has been proposed to anneal complementary ssDNA fragments and integrate these into CRISPRs 14,30,31,66 . Single-molecule co-localization and bulk immunoprecipitation suggest that the type I-E Cas1-2 integrase is recruited to a Cas3-Cascade-target DNA complex to facilitate primed acquisition 34,67 . The structure of the type I-F integration complex reveals conformational changes in Cas3 that may enable interactions with DNA-bound Cascade ( Fig. 5a ) 36,68 . Cascade improves integration efficiency and fidelity in vivo 68,69 , and the structure suggests a model for the formation of a primed acquisition complex (Cas1-2/3-Cascade-target DNA) that transfers new foreign DNA fragments to the integrase. Additional structures will be necessary to clarify the mechanism(s) of primed adaptation. A comparison of the I-E and I-F integration complexes with a structure of the λ-phage excision complex, reveals structural similarities and differences. In I-E systems, the IHF protein folds the leader DNA to present an upstream motif to a lobe of Cas1. In contrast, the I-F structure highlights extensive cooperation between IHF and Cas1-2/3 in bending the leader into an energetically strained conformation that may increase the specificity of CRISPR recognition ( Fig. 6a,b ). This cooperation includes Cas1-2/3 kinking the IR DNA motifs presented as parrallel DNA pillars by IHF, and Cas1-2/3 constricting the IHF proximal 180° bend ito a 260° ( Fig. 6b ). Further, the sequestration of the IR-binding surface of Cas2 suggests a unique structural mechanism that prevents Cas1-2/3 interactions with the leader until foreign DNA binding induces a rotation of Cas3 ( Fig. 4e, Supplementary Video 1 and 3 ) 54 . Many CRISPR leaders contain multiple IHF binding sites and subtype-specific motifs that are reminiscent of motifs found at the ends of transposons, phages, and plasmids, which facilitate integration, excision, or recombination complex assembly ( Extended Data Fig. 8a ) 3–7,22,40–53 . For example, the left and right ends of the λ-phage genome contain three IHF binding sites, five copies of a second DNA motif (P1-2,1’-3’), and three copies of third DNA motif (X1,1.5,2), that recruit λ-phage (Int, Xis) and non-λ-phage (IHF, Fis) proteins ( Fig. 6c ) 7 . These proteins use DNA as a flexible scaffold that organizes enzyme active sites and DNA substrates to facilitate either the integration, or excision, of λ-phage DNA from the bacterial genome ( Fig. 6c ). Similarly, the I-F CRISPR leader motifs recruit Cas (Cas1-2/3) and non-Cas (IHF) proteins, that bend DNA into a flexible scaffold to organize enzyme active sites and DNA substrates in space, facilitating the integration of foreign DNA into the bacterialgenome ( Fig. 6b ). Diverse systems thus use DNA as a flexible scaffold to regulate the isoenergetic mobilization of DNA ( Fig. 6a-c ).” The CRISPR adaptive immune system has been proposed to have originated through the domestication of a transposon that used a Cas1 homolog to catalyze integration (Casposon) 1 . The pathway for the domestication of proteins from different operons to form the Cas components of CRISPR-Cas system has been previously proposed 1,70 . However, the domestication or evolution of DNA elements that gave rise to the CRISPR repeat and leader has remained unclear 1,70 . The similarities in motif architectures between CRISPR leaders and transposon DNA ends suggests that the Capsoson DNA end may have been domesticated to form the primordial CRISPR “leader” ( Fig. 6d ). Similarly, the target DNA site that the ancient Casposase integrated at may have been co-opted to form the first CRISPR “repeat” ( Fig. 6d ). Originally, the transposon end directed Casposase to integrate Casposon (foreign) DNA into the target site. Now, the domesticated leader directs Cas1-2 to integrate foreign DNA into the domesticated repeat ( Fig. 6d ). Diverse DNA mobilizing enzymes across the tree of life co-opt DNA folding to regulate DNA mobilization 3–7 . In sum, these data provide a mechanistic understanding for the role of DNA as a flexible scaffold that controls DNA mobilization. These insights are critical to developing applications of DNA-mobilizing enzymes in gene therapy, genetic engineering, and chronological DNA recordings 6,53,71–75 . Methods Nucleic acid preparation Four single-stranded DNAs ( Supplementary Table S1 ) were synthesized (IDT) and resuspended in 1x TE buffer (10 mM Tris-HCl pH 8, 1 mM EDTA) before being used to assemble the structure of the type I-F integration complex, the assembly is detailed in a below section. The splayed foreign DNAs used in integration assays ( Supplementary Table S1 ) were synthesized and resuspended in 1x TE buffer before use. To make 32 P-labeled CRISPR integration substrates, the sequences consisting of the leader and CRISPR arrays were first synthesized and cloned into pUC57 (Genscript). These plasmids have been made available on Addgene ( Supplementary Table S1 ). These plasmids were transformed into chemically competent E. coli DH5α cells and the transformed cells were plated onto LB agar plates containing 100 µg/mL Ampicillin. These cells were cultured in LB media and plasmids were purified using ZymoPURE II Plasmid Midiprep kit (Zymo Research). Each plasmid was then digested with EcoRI-HF and BamHI-HF (NEB) restriction enzymes, and the 294-383 bp inserts of interest were separated from the vector backbone by agarose gel electrophoresis. The gel segments containing the DNA inserts of interest were excised and DNA was purified using a Zymoclean Gel DNA Recovery kit (Zymo Research) ( Supplementary Table S1 ). The 5' ends of the CRISPR leader and array fragments were dephosphorylated using Quick Calf Intestinal alkaline Phosphatase (NEB), and the DNAs were purified away from protein using a DNA Clean & Concentrator kit (Zymo Research). Both 5' ends of 1 pmole of the CRISPR leader and array fragments were then labelled with 32 P, by incubation with 4 pmoles of [γ- 32 P]ATP (PerkinElmer) by polynucleotide kinase (NEB) in 1x PNK buffer at 37°C for 45 minutes. PNK was heat-denatured by incubation at 65°C for 20 minutes. Spin column purification (G-25, GE Healthcare) was used to remove unincorporated radioactive nucleotides and to buffer exchange DNAs into 1x TE buffer. Cas1 & Cas2/3 Mutagenesis The plasmid used to express Cas1 and Cas2/3 (Addgene #89240) was PCR amplified with mutagenic primer pairs using Q5 polymerase (NEB) ( Supplementary Table S1 ). The parental template plasmid was digested with DpnI (NEB) and the PCR amplified products were purified using the DNA Clean and Concentrate kit (Zymo). The purified DNA was 5’ phosphorylated with T4 Polynucleotide Kinase (NEB) and ligated with homemade T4 DNA ligase. The ligation reactions were transformed into DH5α cells. The Cas2 K11D, R12E, R55E, N56D mutant was prepared using the Cas2 K11D, R12E mutant as the template with the PCR primers for Cas2 R55E, N56D in the mutagenic PCR reaction. Plasmids for expressing the mutant Cas1-2/3 complexes, Cas1H25A-2/3, Cas1E184A-2/3, Cas1-2K11D,R12E/3, Cas1-2R55E,N56D/3, and Cas1-2K11D,R12E,R55E,N56D/3 have been deposited at Addgene (#200213, #200214, #200216, #200217, #200218). Protein purification P. aeruginosa IHF heterodimer was purified as previously described 22 . Briefly, 6xHis-tagged IHFα and StrepII-tagged IHFβ were co-expressed in E. coli BL21(DE3). Cell pellets were lysed by sonication in IHF lysis buffer (25 mM HEPES-NaOH pH 7.5, 500 mM NaCl, 10 mM Imidazole, 1 mM TCEP, 5% Glycerol), supplemented with 0.3x Halt Protease Inhibitor Cocktail (ThermoFisher), at 4°C. Lysate was clarified by two rounds of centrifugation at 12,000 rpm for 15 minutes, at 4°C. His-tagged IHF was captured on HisTrap HP resin (Cytiva), and eluted with 500 mM Imidazole. Affinity tags were cleaved using PreScision protease, and the PreScision protease and remaining 6x-His-IHFα were removed by affinity chromatography using HisTrap HP resin (Cytiva). Untagged IHF heterodimer was then further purified on Heparin Sepharose (Cytiva) and eluted with a linear gradient to a buffer containing 2 M NaCl. Fractions containing IHF heterodimer were concentrated and further purified by SEC ( S ize E xclusion C hromatography) on a Superdex 75 column (Cytiva) equilibrated in IHF buffer (25 mM HEPES-NaOH pH 7.5, 200 mM NaCl, 5 % Glycerol) ( Extended Data Fig. 1a ). IHF overexpression plasmids have been previously described and deposited at Addgene (#149384, #149385) 22 . P. aeruginosa Cas1-2/3 heterohexamer complexes (wildtype and mutants) were purified as previously described 22 . Briefly, StrepII-tagged Cas1-Cas2/3 was overexpressed in E. coli BL21(DE3). Cell pellets were lysed via sonication in Cas1-2/3 lysis buffer (50 mM HEPES pH 7.5, 500 mM KCl, 10% Glycerol, 1 mM DTT), supplemented with 0.3x Halt Protease Inhibitor Cocktail (ThermoFisher), at 4°C. Lysate was clarifed as above. StrepII-tagged Cas1-Cas2/3 complexes were affinity purified on StrepTrap HP resin (GE Healthcare) and eluted with Cas1-2/3 Lysis Buffer containing 3 mM desthiobiotin (Sigma-Aldrich). Eluate was concentrated at 4°C (Corning Spin-X concentrators), before purification over a Superdex 200 size-exclusion column (Cytiva) equilibrated in 10 mM HEPES pH 7.5, 500 mM Potassium Glutamate, and 10% Glycerol ( Extended Data Fig. 1c and 5e ). In vitro integration assays End-point integration reactions were performed in triplicate using 300 nM of splayed foreign DNA fragments, containing or lacking a PAM (IDT), 200 nM of Cas1-2/3, 300 nM of IHF heterodimer and roughly 1 nM of a given 32 P-labelled CRISPR variant fragment in Integration Buffer (20 mM HEPES pH 7.5, 150 mM potassium glutamate, 5 mM MnCl 2 , 1 mM TCEP, 1% Glycerol) ( Supplementary Table S1) . Reactions were assembled on ice and then incubated at 35°C for 20 minutes.Timecourse integration reactions were performed in triplicate using 300 nM of splayed, or fully complementary, foreign DNA fragments lacking a PAM (IDT), 200 nM of Cas1-2/3, 300 nM of IHF heterodimer and roughly 1 nM of a given 32 P-labelled CRISPR variant fragment in modified Integration Buffer (20 mM HEPES pH 7.5, 150 mM potassium glutamate, 5 mM MnCl 2 , 10 mM TCEP, 1% Glycerol) ( Supplementary Table S1) . We noticed that IHF and Cas1-2/3 protein stocks exhibited a propensity to precipitate when diluted into pre-chilled buffer to form working dilution stocks. Therefore all working dilutions of IHF and Cas1-2/3 prepared for timecourse assays were made by first diluting the proteins into room-temperature buffer, mixed, and then chilled on ice. Reactions were assembled on ice and then incubated at 20 minutes. Timepoints were taken at 0, 1, 2, 4 and 8 minutes. Reactions were stopped by the addition of phenol. The aqueous (nucleic acid-containing) layer was mixed 1:1 with 2x formamide loading buffer (95% formamide, 20 mM EDTA, 0.05% bromophenol blue, 0.05% xylene cyanol) and then denatured at 95°C for 5 minutes, before resolving the 32 P-labeled CRISPR substrates and integration products on a 7% (w/v) (29:1 mono:bis) polyacrylamide Urea gel in 1x TBE (100 mM Tris-Borate pH 8.3, 2 mM EDTA). Gels were dried and quantified using a Typhoon phosphorimager (GE Healthcare). The intensity of full-length CRISPR variant, leader-side integration fragments and spacer-side integration fragments were quantified with Multi Gauge v3 (Fujifilm). These readings were then used to calculate leader- and spacer-side integration events as percentages of all events. Images of all gels that resolved integration reactions are shown Extended Data Fig. 6,7 and 9 ), and additional control gels show that Cas1-2/3 is required for integration, and show how the custom 32 P-labelled ladder was generated by restriction enzyme digestion of 32 P-labelled CRISPR variant DNAs ( Extended Data Fig. 8 ). Timecourse integration data was fit to a plateau followed by one phase association (GraphPad Prism). Assembly and purification of I-F integration complex A total of four single-stranded DNAs (ssDNAs) synthesized to mimic a half-site integration intermediate were annealed in a step-wise manner. 2 nanomoles of ssDNAs mostly corresponding to the sense and anti-sense strands of the CRISPR leader ("strand_1" and "strand_2" were denatured at 100°C and then slow annealed using a PCR program that cooled the samples to 25°C in 6°C over an hour, in 100 µL of hybridization buffer (20 mM Tris-HCl pH 7.5, 100 mM monopotassium glutamate, 5 mM EDTA, 1 mM TCEP). 2 nanomoles of ssDNAs mostly corresponding to the corresponding to the sense and anti-sense of the strands of the foreign DNA ("strand_3" and "strand_4") were slow annealed using the same protocol ( Extended Data Fig. 1a and Supplementary Table S1 ). The two sets of annealed DNAs (tube 1: "Strand_1" and "Strand_2"; tube 2: "Strand_3" and "Strand_4") were mixed together, heated to 80°C and then slow annealed using a PCR program that cooled the samples to 25°C in 6°C over an hour, to anneal the complementary sense and anti-sense regions of the CRISPR repeat included in Strand_2 and Strand_3 together. Next, 6 nanomoles of IHF heterodimer in 50 µL of hybridization buffer was warmed to 25°C and then mixed and incubated with the annealed DNAs at 25°C for 10 minutes. Next, 3 nanomoles of Cas1-2/3 in 250 µL of hybridization buffer was warmed to 25°C and mixed with the prepared DNA and IHF mixture, and incubated at 25°C for 10 minutes. The total concentration of monopotassium glutamate in the mixture at this stage was ~200 mM, due to carryover from the stored protein stocks. This sample was centrifuged at 22,000 g, 4°C for 20 minutes to remove precipitates. The type I-F CRISPR integration complex was then purified on a Superdex 200 10/300 column (Cytiva) equilibrated in SEC buffer (20 mM Tris-HCl pH 7.5, 200 mM monopotassium glutamate, 5 mM EDTA, 1 mM TCEP, 2 % Glycerol). 0.5 ml fractions were individually concentrated and stored. The sixth SEC fraction contained all DNAs and proteins of interest and was further analyzed by cryo-EM ( Extended Data Fig. 1d-f ). Cryo-EM sample preparation and data acquisition Purified integration complex was diluted to a concentration of 1 µM in SEC buffer lacking glycerol (20 mM Tris-HCl pH 7.5, 200 mM monopotassium glutamate, 5 mM EDTA, 1 mM TCEP), such that the final glycerol concentration was 0.2% within 1 hour of freezing. Sample was applied to Quantifoil R2/2 Cu 200 mesh grids that were glow discharged using 15 mA for 15 seconds with a 10 second hold (easiGlow, Pelco). 4 µl of diluted integration complex was applied to the grids, and then the grids were blotted for 5-6 seconds using Vitrobot TM Filter paper (Electron Microscopy Sciences) with a blot force of 6, at 100% humidity, 8°C, followed by plunge freezing into liquid ethane using a Vitrobot (Mk. IV, ThermoFisher Scientific). A preliminary dataset of 230 movies was collected on Montana State University's Talos Arctica transmission electron microscope (ThermoFisher Scientific), with a field emission gun operating at an accelerating voltage of 200 kV using parallel illumination conditions 76 . Movies were acquired using a Gatan K3 direct electron detector, operated in electron counting mode applying a total electron exposure of 50 e-/Å 2 over 50 frames (3.995 s exposure, 0.08 s frame time). The SerialEM data collection software was used to collect micrographs at 36,000-fold nominal magnification (1.152 Å/pixel at the specimen level) with a nominal defocus set to 0.5 µm - 2.0 µm 77 . Stage movement was used to target the center of four 2.0 µm holes for focusing, and image shift was used to acquire high magnification images in the center of each of the holes. An preliminary reconstruction was determined from a curated set of 160 images that had CTF fits less than 9 Å and a full-frame motion less than 40 pixels. Briefly, a round of blob picking (150-270 Å) followed by 2D classification was used to identify 2D classes used as templates for template picking in cryoSPARC 78 . Template picking identified 33,403 initial particles from the above 160 images. 2D classification of these 33,403 particles into 50 classes was used to identify 5 classes with strong structural features containing 4,002 particles. Non-uniform refinement of these 4,002 particles resulted in a ~14.7Å reconstruction that appeared to contain a complete integration complex, therefore new grids were prepared as described above and shipped to National Center for CryoEM Access and Training (NCCAT) and the Simons Electron Microscopy Center located at the New York Structural Biology Center (NYSBC) for additional data collection ( Table 1 ). At NCCAT, grids were imaged using a 300 kV Titan Krios G3i (Thermo Fisher Scientific) equipped with a GIF BioQuantum and K3 camera (Gatan). 10,740 images were recorded with Leginon 79 (Suloway et al., 2005) with a calibrated pixel size of 0.5335 Å/px (micrograph dimension of 11520 x 8184 px) over a nominal defocus range of −0.7 μm to −2.1 μm and 20 eV slit. Movies were recorded in “super-resolution mode” (native K3 camera binning 1) with subframes of 50 ms over a 2.5 s exposure (50 frames) to give a total exposure of ∼69 e-/Å 2 ( Table 1 ). Cryo-EM image processing Patch motion correction and patch CTF correction were performed in cryoSPARC 78 . 3,792 of 10,740 total images ( CTF < 8Å, Full-frame motion < 30Å ) were processed first to build an initial template. Blob picking was used to pick particles with diameters ranging from 120-280 Å. These ~1.8 million particles were extracted, Fourier-binned 2x2, and then subjected to 2D classification ( Custom parameters: Initial classification uncertainty factor = 3; Number of online-EM iterations = 30; Batchsize per class = 200 ) ( Extended Data Fig. 1g ). Particles from 82 of the 200 2D classes were selected for an initial round of ab initio reconstruction and heterogeneous refinement. Particles from 1 of 5 of these classes were selected for a second round of ab initio reconstruction and heterogeneous refinement. Particles from 1 of 3 of these classes (174,000 particles) were selected for Non-uniform refinement ( Custom parameters: Optimize per-particle defocus = true; Optimize per-group CTF params = true ) to create an initial reconstruction with a resolution of ~3.7Å, that was used to calculate templates 80 . These templates were used to choose particles from 9,858 of 10,740 total images (CTF < 8Å cutoff). These ~5.85 million particles were classified into a total of 6 classes by Heterogeneous refinement, that were seeded with 1 good volumes and 5 junk volumes taken from the above Heterogeneous refinement analyses. The ~1.31 million particles in the single selected class, were passed through a round of 2D classification ( Custom parameters: Batchsize per class = 200 ) ( Extended Data Fig. 1h ). ~1.29 million particles from 49 of the 50 2D classes were selected for a round of ab initio reconstruction followed by heterogeneous refinement into two classes. ~1.1 million particles from one of these two classes were subjected to 3D classification into 4 classes ( Custom parameters: Batchsize per class = 20,000; Initialization mode = PCA; Target resolution = 2Å; Particles per reconstruction = 500; Class similarity = 0.3 ) , followed by separate Non-uniform refinements of particles from each of these four classes ( Custom parameters: Optimize per-particle defocus = true; Optimize per-group CTF params = true ) 80 . The final set of 366,794 particles were re-extracted and re-centered, Fourier-binned 2x2, and subjected to Non-uniform refinement to generate a final reconstruction refined to a global resolution of 3.48 Å based on the 0.143 FSC cutoff ( Extended Data Fig. 1h-k and Table 1 ) 81 . The 3D FSC was calculated using webserver (3dfsc.salk.edu) 82 . Model building and validation The map was sharpened from two half-maps using the local anisotropic sharpening job in Phenix 83 . The published structure of the P. aeruginosa Cas1 homodimer was used as starting models 56 , because Colabfold consistently failed to predict the alternative fold that one Cas1 subunit adopts to form the assymetric homodimer interface 84 , even when provided template structures. Whereas the Colabfold-predicted models for the P. aeruginosa IHF heterodimer, and the P. aeruginosa Cas2/3 subunit were used as starting models. The conformation of the DNA sequences within the E. coli IHF-DNA co-crystal structure (PDB: 1IHF) 38 was used as starting models for DNA segments within the IHF distal and IHF proximal DNA bends. For all other double stranded DNA segments, B-form DNA was used as a starting model. Single-stranded DNA segments were built in de novo . Protein and DNA segments were individually rigid-body fitted into the EM density map. The relative orientation of the Cas2, Cas2/3 linker and Cas3 domains were corrected by real-space refinement into the EM density map in WinCoot 85 . The ReadySet job in Phenix was used to generate hydrogens on all proteins and nucleic acids and prepare the model for further refinement. Then protein and DNA segments were real-space refined in WinCoot 85 , restrained to ideal geometry, secondary structure and German McClure distance restraints generated in ProSMART from the input models 86 . The models were iteratively real-space refined in WinCoot and in Phenix using Ramachandran and secondary structure restraints 83,85 . The starting model was used as a reference model, and harmonic restraints on the starting coordinates were enabled. MolProbity 87 and the PDB validation service server (https://validate-rcsb-1.wwpdb.org/) were used to identify problem regions subsequently corrected in WinCoot 85 . For regions of the reconstruction where side chains are not visible (resolution >4.0Å) the atomic model was truncated to the peptide backbone. For regions of the reconstruction where the backbone was ambigous the sections of the peptide or DNA model were removed. Contacts and hydrogen bonds between residues were identified by ChimeraX v1.4 using the “contacts” and “hbonds” commands respectively, with default parameters 88,89 . The DNAproDB webserver ( https://dnaprodb.usc.edu/ ) was further used to analyze DNA-protein contacts ( Extended Data Fig. d,e ) 90 . Structure-guided mutagenesis was used to further validate key Cas1-2/3-DNA contacts in the above biochemical assays. Cas1, Cas2/3 and repeat conservation analysis To build a list of type I-F Cas1 sequences, CRISPRDetect v2.4 with default parameters was used to identify CRISPR arrays within a total of 18,225 bacterial and 376 archaeal complete genomes accessed from the NCBI Assembly database on June 10 th of 2019 as previously described 22,91 . 15,274 high-confidence CRISPR arrays were classified with a CRISPR subtype by CRISPRDetect v2.4 (by matching to a list of repeats with known subtype annotations), and by genetic proximity to subtype-specific cas genes (within 20,000 bp). To identify cas genes, the 20,000 bp flanking the CRISPR were submitted to PRODIGAL v2.6.3 (default parameters) to predict all potential o pen r eading f rames (ORFs) 92 . This ORF database was then used as input to search for cas gene clusters with MacsyFinder v1.0.5 93 . The following parameters were used: “ macsyfinder --sequence-db --db-type gembase -d -p -w 50 -vv all ”. HMM profiles and classification definitions used in MacsyFinder were acquired from the local version of CRISPRCasFinder v4.2.20 94 . Next, the first repeat and 200 nucleotides upstream of CRISPR arrays (leader) which were classified as Type I-F (1,683 arrays) were collected. A non-redundant list of I-F CRISPR leaders (536 leaders) was generated using CD-HIT v4.8.1 with a 95% identity cutoff 95 . A local copy of FIMO was used to find significant matches to the position weight matrix representing I-F IHF binging site as previously described 22,96 . I-F CRISPR arrays that possess more than one IHF site (IHF proximal and/or IHF distal ) in the leader sequences were extracted for downstream analyses. Cas1 homologs were identified within the 20,000 base-pair flanking regions of extracted 444 I-F CRISPR arrays by using PRODIGAL and MacsyFinder with the same parameters described above. 371 Cas1 homologs associated with Type I-F CRISPRs and possessing at least one IHF site in the leader sequences were identified. A non-redundant list of Cas1 sequences was generated CD-HIT v4.8.1 with a 95% identity cutoff, resulting in 222 sequences 95 . Sequences smaller than 200 residues and larger than 500 residues were removed, and the remaining 205 sequences were further curated with MaxAlign, which selected a list of 144 unique type I-F Cas1 sequences 97 . The P. aeruginosa PA14Cas1 sequence was then added to a final list of 145 type I-F Cas1 sequences. To build a list of type I-F Cas2/3 sequences, the P. aeruginosa PA14 Cas2/3 sequence was used as an input for HHMER for a search for homologs using 3 iterations, an E-value cutoff of 0.0001, against the UNIREF-90 database 98,99 . A list of 500 representative sequences was further curated with MaxAlign, to generate a final list of 458 unique Cas2/3 sequences. Type I-F Cas1 and Cas2/3 sequences were aligned using the MAFFT webserver with the E-INS-I iterative refinment methods to result in alignments with the highest number of gap-free sites 100 . To build an updated list of CRISPR repeat sequences, CRISPRDetect v3.0 with default parameters was used to identify CRISPR arrays within a total of 25,502 bacterial and 398 archaeal complete genomes and chromosomes accessed from the NCBI RefSeq Assembly database (accessed on June 10 th , 2021) 91 . This search identified CRISPR loci within 58,864 genomic and plasmid sequences, resulting in 24,940 high-confidence CRISPR loci predictions (array quality score >3). Similar to above, CRISPRDetect annotated the subtype of 14,446 of these CRISPR loci, based on the sequence similarity of the repeats in these loci to known CRISPR repeats. The subtypes of the remaining 10,494 CRISPR loci were determined by their proximity to subtype-specific cas genes as described above. 5,321 of the 10,494 unclassified CRISPR loci were assigned a subtype using this protocol, such that 5,173 CRISPR loci remained unclassified. The consensus repeat for each of the 24,940 CRISPR loci, as reported by CRISPRDetect, were used for downstream analyses. To ensure the repeats were arranged in the correct orientation, the 24,940 repeats were grouped by subtype, and each group was individually aligned by MAFFT using the “ --adjustdirection ” parameter. Sequence logos of the first and last three bps of CRISPR repeats were made using Weblogo v3.7.1 for CRISPR subtypes and across all subtypes 101,102 ( Extended Data Fig. 4g ). Reporting summary Further information on research design is available in the Nature Research Reporting Summary linked to this paper. Declarations Data availability The data that support the findings of this study are available from the corresponding author Blake Wiedenheft upon request. Cryo-EM maps were deposited in the Electron Microscopy Data Bank under accession number EMD-29280. The atomic model of the type I-F integration complex was deposited in the PDB under accession number 8FLJ. Plasmids generated in this study are available from Addgene. Code availability Code will be made available upon request and without restriction. Acknowledgements Thanks to members of the B.W. laboratory for feedback and discussions. We thank Dr. Mariusz Matyszewski and Dr. Jeliazko Jeliazkov for helpful discussions. Thanks to Coltran Hophan-Nichols for computational support. A.S-F. is a postdoctoral fellow of the Life Science Research Foundation that is supported by the Simons Foundation. A.S-F. is supported by the Postdoctoral Enrichment Program Award from the Burroughs Wellcome Fund. This work was supported by National Institutes of Health, United States grant 1K99GM147842 (A.S-F.). L.T., A.B.G. is supported by Montana State University’s Undergraduate Scholars Program, and by the NIH NIGMS IDeA program (P20GM103474). This work was performed using the cryo-EM Facility at Montana State University (NSF 1828765 and the M.J. Murdock Charitable Trust). Microscopy was also performed at the National Center for CryoEM Access and Training (NCCAT) and the Simons Electron Microscopy Center located at the New York Structural Biology Center, supported by the NIH Common Fund Transformative High Resolution Cryo-Electron Microscopy program (U24 GM129539), and by grants from the Simons Foundation (SF349247) and NY State Assembly. Research in the Wiedenheft lab is supported by the NIH (R35GM134867), the M.J. Murdock Charitable Trust, a young investigator award from Amgen, and the Montana State University Agricultural Experimental Station (USDA NIFA), and a sponsored research agreement from VIRIS Detection Systems. Molecular graphics and analyses performed with UCSF ChimeraX, developed by the Resource for Biocomputing, Visualization, and Informatics at the University of California, San Francisco, with support from National Institutes of Health R01-GM129325 and the Office of Cyber Infrastructure and Computational Biology, National Institute of Allergy and Infectious Diseases. Funders had no role in designing, performing, interpreting, or submitting the work. Author Contributions A.S.-F.: Conceptualization, Data Curation, Formal Analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualization, Writing – original draft. W.S.H., T.W., M.B.: Data curation, Investigation, Methodology, Visualization, Writing – review & editing. A.B.G., R.A.W.: Investigation, Methodology. L.T.: Visualization. C.C.G.: Software, Resources, Writing – review & editing. K.N. and E.E.: Investigation, Resources. G.C.L.: Methodology, Supervision, Visualization, Writing – review & editing. B.W.: Funding acquisition, Project administration, Resources, Supervision, Visualization, Writing – review & editing. Competing Interests B.W. is the founder of SurGene and VIRIS Detection Systems. B.W. and A.S.-F. are inventors on patent applications related to CRISPR-Cas systems and applications thereof. References Koonin, E. V. & Krupovic, M. Evolution of adaptive immunity from transposable elements combined with innate immune systems. Nat. Rev. Genet. 16 , 184–192 (2015). McCLINTOCK, B. The origin and behavior of mutable loci in maize. Proc. Natl. Acad. Sci. U. S. A. 36 , 344–355 (1950). Nuñez, J. K., Bai, L., Harrington, L. B., Hinder, T. L. & Doudna, J. A. CRISPR Immunological Memory Requires a Host Factor for Specificity. Mol. Cell 62 , 824–833 (2016). Fagerlund, R. D. et al. Spacer capture and integration by a type I-F Cas1–Cas2-3 CRISPR adaptation complex. Proc. Natl. Acad. Sci. 114 , 201618421 (2017). Wright, A. V. et al. Structures of the CRISPR genome integration complex. Science 357 , 1113–1118 (2017). Hickman, A. B. & Dyda, F. Mechanisms of DNA transposition. Mob. DNA III 529–553 (2015) doi:10.1128/9781555819217.ch25. Laxmikanthan, G. et al. Structure of a holliday junction complex reveals mechanisms governing a highly regulated DNA transaction. eLife 5 , 1–23 (2016). Lee, H. & Sashital, D. G. Creating memories: molecular mechanisms of CRISPR adaptation. Trends Biochem. Sci. 1–13 (2022) doi:10.1016/j.tibs.2022.02.004. Wang, J. et al. Structural and Mechanistic Basis of PAM-Dependent Spacer Acquisition in CRISPR-Cas Systems. Cell 163 , 840–853 (2015). Nuñez, J. K., Harrington, L. B., Kranzusch, P. J., Engelman, A. N. & Doudna, J. A. Foreign DNA capture during CRISPR-Cas adaptive immunity. Nature 527 , 535–538 (2015). Xiao, Y., Ng, S., Nam, K. H. & Ke, A. How type II CRISPR–Cas establish immunity through Cas1–Cas2-mediated spacer integration. Nature 550 , 137–141 (2017). Jackson, S. A. et al. CRISPR-Cas: Adapting to change. Science 356 , eaal5056 (2017). Mojica, F. J. M., Díez-Villaseñor, C., García-Martínez, J. & Almendros, C. Short motif sequences determine the targets of the prokaryotic CRISPR defence system. Microbiology 155 , 733–740 (2009). Kim, S. et al. Selective loading and processing of prespacers for precise CRISPR adaptation. Nature 579 , 141–145 (2020). Hu, C. et al. Mechanism for Cas4-assisted directional spacer acquisition in CRISPR–Cas. Nature 598 , 515–520 (2021). Ramachandran, A., Summerville, L., Learn, B. A., DeBell, L. & Bailey, S. Processing and integration of functionally oriented prespacers in the Escherichia coli CRISPR system depends on bacterial host exonucleases. J. Biol. Chem. 295 , 3403–3414 (2020). Liao, C. et al. Spacer prioritization in CRISPR–Cas9 immunity is enabled by the leader RNA. Nat. Microbiol. (2022) doi:10.1038/s41564-022-01074-3. McGinn, J. & Marraffini, L. A. CRISPR-Cas Systems Optimize Their Immune Response by Specifying the Site of Spacer Integration. Mol. Cell 64 , 616–623 (2016). Wang, R., Li, M., Gong, L., Hu, S. & Xiang, H. DNA motifs determining the accuracy of repeat duplication during CRISPR adaptation in Haloarcula hispanica. Nucleic Acids Res. 44 , 4266–4277 (2016). Goren, M. G. et al. Repeat Size Determination by Two Molecular Rulers in the Type I-E CRISPR Array. Cell Rep. 16 , 2811–2818 (2016). Linheiro, R. S. & Bergman, C. M. Testing the palindromic target site model for DNA transposon insertion using the Drosophila melanogaster P-element. Nucleic Acids Res. 36 , 6199–6208 (2008). Santiago-Frangos, A., Buyukyoruk, M., Wiegand, T., Krishna, P. & Wiedenheft, B. Distribution and phasing of sequence motifs that facilitate CRISPR adaptation. Curr. Biol. 1–10 (2021) doi:10.1016/j.cub.2021.05.068. Kieper, S. N., Almendros, C. & Brouns, S. J. J. Conserved motifs in the CRISPR leader sequence control spacer acquisition levels in Type I-D CRISPR-Cas systems. FEMS Microbiol. Lett. 366 , 2016–2020 (2019). Rollie, C., Graham, S., Rouillon, C. & White, M. F. Prespacer processing and specific integration in a Type I-A CRISPR system. Nucleic Acids Res. 46 , 1007–1020 (2018). Yosef, I., Goren, M. G. & Qimron, U. Proteins and DNA elements essential for the CRISPR adaptation process in Escherichia coli. Nucleic Acids Res. 40 , 5569–5576 (2012). Wei, Y., Chesne, M. T., Terns, R. M. & Terns, M. P. Sequences spanning the leader-repeat junction mediate CRISPR adaptation to phage in Streptococcus thermophilus. Nucleic Acids Res. 43 , 1749–1758 (2015). Wright, A. V. & Doudna, J. A. Protecting genome integrity during CRISPR immune adaptation. Nat. Struct. Mol. Biol. 23 , 876–883 (2016). Westra, E. R. et al. Parasite Exposure Drives Selective Evolution of Constitutive versus Inducible Defense. Curr. Biol. 25 , 1043–1049 (2015). Makarova, K. S. et al. Evolutionary classification of CRISPR–Cas systems: a burst of class 2 and derived variants. Nat. Rev. Microbiol. 18 , 67–83 (2020). Richter, C. et al. Priming in the Type I-F CRISPR-Cas system triggers strand-independent spacer acquisition, bi-directionally from the primed protospacer. Nucleic Acids Res. 42 , 8516–8526 (2014). Datsenko, K. A. et al. Molecular memory of prior infections activates the CRISPR/Cas adaptive bacterial immunity system. Nat. Commun. 3 , 945 (2012). Xiao, Y. et al. Structure basis for RNA-guided DNA degradation by Cascade and Cas3. 0839 , 1–12 (2018). Nicholson, T. J. et al. Bioinformatic evidence of widespread priming in type I and II CRISPR-Cas systems. RNA Biol. 16 , 566–576 (2019). Brown, M. W. et al. Assembly and translocation of a CRISPR-Cas primed acquisition complex. bioRxiv 41 , 1–11 (2017). Li, M., Wang, R., Zhao, D. & Xiang, H. Adaptation of the Haloarcula hispanica CRISPR-Cas system to a purified virus strictly requires a priming process. Nucleic Acids Res. 42 , 2483–2492 (2014). Semenova, E. et al. Highly efficient primed spacer acquisition from targets destroyed by the Escherichia coli type I-E CRISPR-Cas interfering complex. Proc. Natl. Acad. Sci. 113 , 7626–7631 (2016). Fineran, P. C. et al. Degenerate target sites mediate rapid primed CRISPR adaptation. Proc. Natl. Acad. Sci. U. S. A. 111 , (2014). Rice, P. A., Yang, S., Mizuuchi, K. & Nash, H. A. Crystal Structure of an IHF-DNA Complex: A Protein-Induced DNA U-Turn. Cell 87 , 1295–1306 (1996). Rohs, R. et al. Origins of specificity in protein-DNA recognition. Annu. Rev. Biochem. 79 , 233–269 (2010). Zayed, H. The DNA-bending protein HMGB1 is a cellular cofactor of Sleeping Beauty transposition. Nucleic Acids Res. 31 , 2313–2322 (2003). Little, A. J., Corbett, E., Ortega, F. & Schatz, D. G. Cooperative recruitment of HMGB1 during V(D)J recombination through interactions with RAG1 and DNA. Nucleic Acids Res. 41 , 3289–3301 (2013). Nash, H. A. & Robertson, C. A. Purification and properties of the Escherichia coli protein factor required for lambda integrative recombination. J. Biol. Chem. 256 , 9246–9253 (1981). Lavoie, B. D. & Chaconas, G. Site-specific HU binding in the Mu transpososome: conversion of a sequence-independent DNA-binding protein into a chemical nuclease. Genes Dev. 7 , 2510–2519 (1993). Chalmers, R., Guhathakurta, A., Benjamin, H. & Kleckner, N. IHF Modulation of Tn10 Transposition: Sensory Transduction of Supercoiling Status via a Proposed Protein/DNA Molecular Spring. Cell 93 , 897–908 (1998). Haniford, D. B. Transpososome Dynamics and Regulation in Tn10 Transposition. Crit. Rev. Biochem. Mol. Biol. 41 , 407–424 (2006). Whitfield, C. R., Wardle, S. J. & Haniford, D. B. The global bacterial regulator H-NS promotes transpososome formation and transposition in the Tn5 system. Nucleic Acids Res. 37 , 309–321 (2009). Liu, D., Haniford, D. B. & Chalmers, R. M. H-NS mediates the dissociation of a refractory protein-DNA complex during Tn10/IS10 transposition. Nucleic Acids Res. 39 , 6660–6668 (2011). van Gent, D. C., Hiom, K., Paull, T. T. & Gellert, M. Stimulation of V(D)J cleavage by high mobility group proteins. EMBO J. 16 , 2665–2670 (1997). Rowland, S.-J., Stark, W. M. & Boocock, M. R. Sin recombinase from Staphylococcus aureus: synaptic complex architecture and transposon targeting: Sin recombinase. Mol. Microbiol. 44 , 607–619 (2002). Alonso, J. C., Weise, F. & Rojo, F. The Bacillus subtilis Histone-like Protein Hbsu Is Required for DNA Resolution and DNA Inversion Mediated by the β Recombinase of Plasmid pSM19035. J. Biol. Chem. 270 , 2938–2945 (1995). Petit, M.-A., Ehrlich, D. & Jannière, L. pAMβ1 resolvase has an atypical recombination site and requires a histone-like protein HU. Mol. Microbiol. 18 , 271–282 (1995). Rojo, F. & Alonso, J. C. The β recombinase of plasmid pSM19035 binds to two adjacent sites, making different contacts at each of them. Nucleic Acids Res. 23 , 3181–3188 (1995). Walker, M. W. G., Klompe, S. E., Zhang, D. J. & Sternberg, S. H. Transposon mutagenesis libraries reveal novel molecular requirements during CRISPR RNA-guided DNA integration . http://biorxiv.org/lookup/doi/10.1101/2023.01.19.524723 (2023) doi:10.1101/2023.01.19.524723. Rollins, M. F. et al. Cas1 and the Csy complex are opposing regulators of Cas2/3 nuclease activity. Proc. Natl. Acad. Sci. 114 , 201616395 (2017). Wang, X. et al. Structural basis of Cas3 inhibition by the bacteriophage protein AcrF3. Nat. Struct. Mol. Biol. 23 , 868–870 (2016). Wiedenheft, B. et al. Structural Basis for DNase Activity of a Conserved Protein Implicated in CRISPR-Mediated Genome Defense. Structure 17 , 904–912 (2009). Kunin, V., Sorek, R. & Hugenholtz, P. Evolutionary conservation of sequence and secondary structures in CRISPR repeats. Genome Biol. 8 , R61 (2007). Nethery, M. A. et al. CRISPRclassify: Repeat-Based Classification of CRISPR Loci. CRISPR J. 4 , 558–574 (2021). Dhingra, Y., Suresh, S. K., Juneja, P. & Sashital, D. G. PAM binding ensures orientational integration during Cas4-Cas1-Cas2-mediated CRISPR adaptation. Mol. Cell 82 , 4353-4367.e6 (2022). Ali Azam, T., Iwata, A., Nishimura, A., Ueda, S. & Ishihama, A. Growth Phase-Dependent Variation in Protein Composition of the Escherichia coli Nucleoid. J. Bacteriol. 181 , 6361–6370 (1999). Montaño, S. P., Pigli, Y. Z. & Rice, P. A. The Mu transpososome structure sheds light on DDE recombinase evolution. Nature 491 , 413–417 (2012). Maertens, G. N., Hare, S. & Cherepanov, P. The mechanism of retroviral integration from X-ray structures of its key intermediates. Nature 468 , 326–329 (2010). Rollie, C., Schneider, S., Brinkmann, A. S., Bolt, E. L. & White, M. F. Intrinsic sequence specificity of the Cas1 integrase directs new spacer acquisition. eLife 4 , 1–19 (2015). Makarova, K. S., Wolf, Y. I. & Koonin, E. V. Classification and Nomenclature of CRISPR-Cas Systems: Where from Here? CRISPR J. 1 , 325–336 (2018). Deveau, H. et al. Phage response to CRISPR-encoded resistance in Streptococcus thermophilus. J. Bacteriol. 190 , 1390–1400 (2008). Künne, T. et al. Cas3-Derived Target DNA Degradation Fragments Fuel Primed CRISPR Adaptation. Mol. Cell 63 , 852–864 (2016). Musharova, O. et al. Prespacers formed during primed adaptation associate with the Cas1–Cas2 adaptation complex and the Cas3 interference nuclease–helicase. Proc. Natl. Acad. Sci. 118 , e2021291118 (2021). Wiegand, T. et al. Reproducible Antigen Recognition by the Type I-F CRISPR-Cas System. CRISPR J. 3 , 378–387 (2020). Vorontsova, D. et al. Foreign DNA acquisition by the I-F CRISPR–Cas system requires all components of the interference machinery. Nucleic Acids Res. 43 , 10848–10860 (2015). Koonin, E. V. & Makarova, K. S. Evolutionary plasticity and functional versatility of CRISPR systems. PLOS Biol. 20 , e3001481 (2022). Cavazzana-Calvo, M. et al. Gene Therapy of Human Severe Combined Immunodeficiency (SCID)-X1 Disease. Science 288 , 669–672 (2000). Strecker, J. et al. RNA-guided DNA insertion with CRISPR-associated transposases. Science 365 , 48–53 (2019). Klompe, S. E., Vo, P. L. H., Halpin-Healy, T. S. & Sternberg, S. H. Transposon-encoded CRISPR–Cas systems direct RNA-guided DNA integration. Nature 571 , 219–225 (2019). Shipman, S. L., Nivala, J., Macklis, J. D. & Church, G. M. Molecular recordings by directed CRISPR spacer acquisition. Science 353 , aaf1175 (2016). Schmidt, F., Cherepkova, M. Y. & Platt, R. J. Transcriptional recording by CRISPR spacer acquisition from RNA. Nature 562 , 380–385 (2018). Herzik, M. A., Wu, M. & Lander, G. C. High-resolution structure determination of sub-100 kDa complexes using conventional cryo-EM. Nat. Commun. 10 , 1–9 (2019). Mastronarde, D. N. Automated electron microscope tomography using robust prediction of specimen movements. J. Struct. Biol. 152 , 36–51 (2005). Punjani, A., Rubinstein, J. L., Fleet, D. J. & Brubaker, M. A. CryoSPARC: Algorithms for rapid unsupervised cryo-EM structure determination. Nat. Methods 14 , 290–296 (2017). Suloway, C. et al. Automated molecular microscopy: The new Leginon system. J. Struct. Biol. 151 , 41–60 (2005). Punjani, A., Zhang, H. & Fleet, D. J. Non-uniform refinement: adaptive regularization improves single-particle cryo-EM reconstruction. Nat. Methods 17 , 1214–1221 (2020). Scheres, S. H. W. & Chen, S. Prevention of overfitting in cryo-EM structure determination. Nat. Methods 9 , 853–854 (2012). Tan, Y. Z. et al. Addressing preferred specimen orientation in single-particle cryo-EM through tilting. Nat. Methods 14 , 793–796 (2017). Liebschner, D. et al. Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in Phenix . Acta Crystallogr. Sect. Struct. Biol. 75 , 861–877 (2019). Mirdita, M. et al. ColabFold: making protein folding accessible to all. Nat. Methods 19 , 679–682 (2022). Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr. Sect. D 66 , 486–501 (2010). Nicholls, R. A. Conformation-independent comparison of protein structures. (2011). Williams, C. J. et al. MolProbity: More and better reference data for improved all-atom structure validation: PROTEIN SCIENCE.ORG. Protein Sci. 27 , 293–315 (2018). Goddard, T. D. et al. UCSF ChimeraX: Meeting modern challenges in visualization and analysis. Protein Sci. 27 , 14–25 (2018). Pettersen, E. F. et al. UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci. 30 , 70–82 (2021). Sagendorf, J. M., Markarian, N., Berman, H. M. & Rohs, R. DNAproDB: an expanded database and web-based tool for structural analysis of DNA–protein complexes. Nucleic Acids Res. gkz889 (2019) doi:10.1093/nar/gkz889. Biswas, A., Staals, R. H. J., Morales, S. E., Fineran, P. C. & Brown, C. M. CRISPRDetect: A flexible algorithm to define CRISPR arrays. BMC Genomics 17 , 356 (2016). Hyatt, D. et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11 , 119 (2010). Abby, S. S., Néron, B., Ménager, H., Touchon, M. & Rocha, E. P. C. MacSyFinder: A program to mine genomes for molecular systems with an application to CRISPR-Cas systems. PLoS ONE (2014) doi:10.1371/journal.pone.0110726. Couvin, D. et al. CRISPRCasFinder, an update of CRISRFinder, includes a portable version, enhanced performance and integrates search for Cas proteins. Nucleic Acids Res. 46 , W246–W251 (2018). Li, W. & Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics 22 , 1658–1659 (2006). Grant, C. E., Bailey, T. L. & Noble, W. S. FIMO: scanning for occurrences of a given motif. Bioinformatics 27 , 1017–1018 (2011). Gouveia-Oliveira, R., Sackett, P. W. & Pedersen, A. G. MaxAlign: maximizing usable data in an alignment. BMC Bioinformatics 8 , 312 (2007). Finn, R. D., Clements, J. & Eddy, S. R. HMMER web server: interactive sequence similarity searching. Nucleic Acids Res. 39 , W29–W37 (2011). Suzek, B. E. et al. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics 31 , 926–932 (2015). Katoh, K. & Standley, D. M. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. Mol. Biol. Evol. 30 , 772–780 (2013). Schneider, T. D. & Stephens, R. M. Sequence logos: A new way to display consensus sequences. Nucleic Acids Res. 18 , 6097–6100 (1990). Crooks, G. E., Hon, G., Chandonia, J.-M. & Brenner, S. E. WebLogo: a sequence logo generator. Genome Res. 14 , 1188–90 (2004). Table Table 1. Cryo-EM data collection, refinement, and validation statistics I-F Integration Complex (EMD- 29280 ) (PDB 8FLJ ) Data collection and processing Microscope Krios Voltage (keV) 300 Detector Gatan K3 Magnification (nominal/calibrated) 81,000x/46,860x Exposure navigation Beam image-shift Data acquisition software Leginon Electron exposure (e - /Å 2 ) 69.09 Exposure rate (e - /Å 2 /s) 27.64 Frame length (ms) 50 Number of frames per micrograph 50 Energy filter width (eV) 20 Pixel size (Å) 1.067 Defocus range (μm) 0.7 – 2.1 Stage Tilt (°) 0 Micrographs collected (no.) 10,740 Reconstruction Image-processing package cryoSPARC Total extracted particles (no.) 5,846,923 Refined particles (no.) 1,103,036 Final particles (no.) 366,794 Symmetry imposed C 1 Global resolution (Å) FSC 0.143 (unmasked/masked) 3.88/3.47 FSC 0.5 (unmasked/masked) 4.22/3.77 Resolution range (local) (Å) 2.3 – 9.9 3DFSC sphericity 0.985 out of 1 Model sharpening B factor (Å 2 ) -55 Refinement Refinement package Phenix Model composition Non-hydrogen atoms 31,794 Protein residues 3,549 DNA nucleotides 350 Ligands - CC (volume/mask) 0.76/0.76 R.m.s deviations Bond lengths (Å) 0.005 Bond angles (°) 0.835 Validation Ramachandran plot Outliers (%) 0 Allowed (%) 4.28 Favored (%) 95.72 MolProbity score 1.54 Poor rotamers (%) 0.26 Clashscore (all atoms) 4.67 C-beta deviations (%) 0.09 CaBLAM Outliers (%) 1.70 EMRinger score 2.55 Additional Declarations Yes there is potential Competing Interest. B.W. is the founder of SurGene and VIRIS Detection Systems. B.W. and A.S.-F. are inventors on patent applications related to CRISPR-Cas systems and applications thereof. Supplementary Files IntegrationPaperSupplementaryTableS1.docx ASFCas123EDatatosubmit.docx ReportingSummaryWiedenheft.pdf IntegrationPaperSupplementaryVideo1.mp4 Cas1-2/3 undergoes a large conformational change to unveil Cas2-leader binding sites. IntegrationPaperSupplementaryVideo2.mp4 Overview of how the I-F CRISPR leader and IHF guide Cas1-2/3-mediated integration of foreign DNA at the first CRISPR repeat. IntegrationPaperSupplementaryVideo3.mp4 The strained IHF-mediated DNA bends sway in relation to the Cas1-2/3 integrase. Cite Share Download PDF Status: Published Journal Publication published 14 Sep, 2023 Read the published version in Nature Structural & Molecular Biology → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2982802","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":209093937,"identity":"988fb6e3-2713-4ff8-a43b-01536d5f0171","order_by":0,"name":"Andrew Santiago-Frangos","email":"","orcid":"https://orcid.org/0000-0001-9615-065X","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Andrew","middleName":"","lastName":"Santiago-Frangos","suffix":""},{"id":209093938,"identity":"b3f0477f-8adc-4a2a-973c-538fd7cab823","order_by":1,"name":"William Henriques","email":"","orcid":"","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"William","middleName":"","lastName":"Henriques","suffix":""},{"id":209093939,"identity":"73a9efb4-6d6e-4795-b80d-fdd01f7cc586","order_by":2,"name":"Tanner Wiegand","email":"","orcid":"https://orcid.org/0000-0002-0528-268X","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tanner","middleName":"","lastName":"Wiegand","suffix":""},{"id":209093940,"identity":"0b2cf92e-f187-4e23-b062-ad6fa0a3c717","order_by":3,"name":"Colin Gauvin","email":"","orcid":"https://orcid.org/0000-0001-7171-552X","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Colin","middleName":"","lastName":"Gauvin","suffix":""},{"id":209093941,"identity":"227326fd-ae03-4f08-bdc3-71406e80059e","order_by":4,"name":"Murat Buyukyoruk","email":"","orcid":"","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Murat","middleName":"","lastName":"Buyukyoruk","suffix":""},{"id":209093942,"identity":"9573fb70-009d-4930-b3f7-735f8efcb852","order_by":5,"name":"Kasahun Neselu","email":"","orcid":"","institution":"Simons Electron Microscopy Center","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kasahun","middleName":"","lastName":"Neselu","suffix":""},{"id":209093943,"identity":"a451d493-8109-4ab4-b809-9f0221ed2462","order_by":6,"name":"Edward Eng","email":"","orcid":"https://orcid.org/0000-0002-8014-7269","institution":"New York Structural Biology Center","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Edward","middleName":"","lastName":"Eng","suffix":""},{"id":209093944,"identity":"f0603fb3-de1f-45fe-95e6-84e7d728910a","order_by":7,"name":"Gabriel Lander","email":"","orcid":"https://orcid.org/0000-0003-4921-1135","institution":"Scripps Research","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Gabriel","middleName":"","lastName":"Lander","suffix":""},{"id":209093945,"identity":"8a3f611d-02d8-4702-a222-6ee9e8e8f3cb","order_by":8,"name":"Royce Wilkinson","email":"","orcid":"","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Royce","middleName":"","lastName":"Wilkinson","suffix":""},{"id":209093946,"identity":"e11cab1d-3a4f-48eb-8aca-ca32dbb393a1","order_by":9,"name":"Ava Graham","email":"","orcid":"","institution":"Montana State University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ava","middleName":"","lastName":"Graham","suffix":""},{"id":209093947,"identity":"30c4fc18-5a2c-4416-b4b6-d4d91a03dc2c","order_by":10,"name":"Blake Wiedenheft","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7klEQVRIiWNgGAWjYBACAwbGBgYeBoYEBuYDjA+AAjx8xGthS2Y2AGlhI6wFpAyihU0CxCGoxVwiufHBm4o7efxs/Mcqv+bYybAxMD98dAOPFssZic2Gc848K5ZsY2a7LbstGegwNmPjHHwOu53YJs3bdjhxw/1mttuS25iBWnjYpAloaf/N++9w4v5jzGzFktvqidLSxszbALSFjZmN8eO2w4S1WM5/2Cw559jhxBnHmI2lGbcd5wFqxe8Xc57jDz+8qTmc2N/G+PDjz23V9vzszQ8f49OCAph5wCSxykGA8QcpqkfBKBgFo2DEAAATv0jkCrgs9QAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0001-9297-5304","institution":"Montana State University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Blake","middleName":"","lastName":"Wiedenheft","suffix":""}],"badges":[],"createdAt":"2023-05-25 20:45:53","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2982802/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2982802/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41594-023-01097-2","type":"published","date":"2023-09-14T04:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":38561838,"identity":"13a112fb-19cd-4fec-b75a-0cab4ad1bfa3","added_by":"auto","created_at":"2023-06-14 20:24:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":660264,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCryo-EM structure of the type I-F CRISPR integration complex. a\u003c/strong\u003e, Scheme of the I-F CRISPR system of \u003cem\u003ePseudomonas aeruginosa \u003c/em\u003ePA14. A CRISPR is composed of repeated DNA sequences (diamonds) interspersed with unique spacer sequences (black squares). The CRISPR is adjacent to six \u003cem\u003ecas\u003c/em\u003e genes (arrows). Four Cas1 and two Cas2/3 proteins assemble into a heterohexamer in which the Cas3 and Cas1 subunits surround the central Cas2 homodimer like petals of a closed flower (Cas1\u003csub\u003e4\u003c/sub\u003e-Cas2/3\u003csub\u003e2\u003c/sub\u003e). Cas1-2/3 and integration host factor (IHF) proteins cooperate with DNA upstream of the CRISPR (leader) to integrate foreign DNA at the first repeat. The leader sequence contains two IHF binding sites and two Inverted Repeats (IRs) that are necessary for integration of foreign DNA at the leader-repeat junction. In addition to playing a central role in integration, the Cas2/3 fusion is recruited to DNA-bound Cascade CRISPR surveillance complex to degrade foreign genetic parasites. \u003cstrong\u003eb\u003c/strong\u003e, The Cas1-2/3 heterohexamer and IHF proteins were mixed with a half-site DNA integration intermediate consisting of a foreign DNA linked to one strand of the CRISPR DNA at the leader-repeat junction. \u003cstrong\u003ec\u003c/strong\u003e, Cryo-EM density map of the type I-F CRISPR integration complex at ~3.5 Å resolution (\u003cstrong\u003eExtended Data Fig. 1g-k, Table 1\u003c/strong\u003e). \u003cstrong\u003ed\u003c/strong\u003e, Atomic model of the type I-F CRISPR integration complex. Cas1-2/3 proteins alone (left) are shown in cartoon representation. Cas3 domains rotate by 100° simulating the motion of a bloomed flower and exposing DNA binding sites on Cas2 that interact with each of the IRs (\u003cstrong\u003eExtended Data Fig. 2a and Supplementary Video 1\u003c/strong\u003e). DNA alone is shown in the middle (surface representation). Proteins (cartoon representations) and DNAs of the integration complex are shown on the right.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/2b954f493314b351dba68724.png"},{"id":38561837,"identity":"64654c0a-6ed1-485f-a2d8-24cb5287baba","added_by":"auto","created_at":"2023-06-14 20:24:04","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":849956,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eForeign DNA constrains the Cas2/3 linker against conserved Cas1 residues. a\u003c/strong\u003e, View of the foreign DNA-bound face of Cas1-2/3. The foreign DNA, Cas2 subunits, Cas2/3 linker and Cas1 beta hairpins that contact the start and end of the Cas2/3 linker are shown in solid while other parts of the complex are shown at 40% transparency for clarity. Insets outline locations of close-up views shown in panels b-e. \u003cstrong\u003eb\u003c/strong\u003e,\u003cstrong\u003ec\u003c/strong\u003e, The foreign DNA constrains the Cas2/3 linker against each Cas1 subunit (\u003cstrong\u003eExtended Data Fig. 2b\u003c/strong\u003e). Cas2, the Cas2/3 linker and, Cas1 cooperate to bind the foreign DNA body and to splay the ends of the foreign DNA. Histidine wedges in Cas1 measure out a central foreign DNA duplex of 22 base-pairs. Most DNA-binding residues are conserved or undergo conservative mutations (\u003cstrong\u003eExtended Data Fig. 3\u003c/strong\u003e).\u003cstrong\u003e d\u003c/strong\u003e,\u003cstrong\u003ee\u003c/strong\u003e, Conserved Cas2/3 linker residues (blue, sticks) contact residues conserved in Cas1 proteins from type I-F CRISPR systems (mauve, surface) (\u003cstrong\u003eExtended Data Fig. 2b,c\u003c/strong\u003e). Cas2 and Cas3 domains are shown at 90% transparency for clarity. Inset shows the Cas1 sequence conservation color key.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/71d9d67c6c6038d0ad756eaa.png"},{"id":38562215,"identity":"bf843cf9-1245-4b11-a3b4-c91fed728117","added_by":"auto","created_at":"2023-06-14 20:32:04","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":758881,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe Cas2 homodimer simultaneously coordinates four dsDNA helices critical to CRISPR integration. a\u003c/strong\u003e, The Cas2 homodimer (pink surface) is flanked by DNA on four sides. Prior structures have shown that the CRISPR repeat (yellow) and foreign DNA (red) are bound to opposite faces of Cas2. Here, we show that symmetrical surfaces on Cas2 also bind inverted repeat (IR, left and right) motifs in the leader. Surface representations of the IHF heterodimers, Cas1 homodimers, Cas2/3 linkers, Cas3 domains, and the 3’ overhang of the foreign DNA are shown in 100% transparency for clarity. Each Cas2 inserts an arginine (R55) into the center of the IRs, which stack between deoxyribose sugars, and additional polar residues (R54, N56, R12, and K11) contact the DNA backbone. Cas2 induces 25-35° bends in the DNA (\u003cstrong\u003eExtended Data Fig. 4c,e\u003c/strong\u003e). The sequence logos of the type I-F IR\u003csub\u003eproximal\u003c/sub\u003e (left) and IR\u003csub\u003edistal\u003c/sub\u003e motifs (right), and the IR sequences present in the \u003cem\u003eP. aeruginosa\u003c/em\u003e PA14 CRISPR leader are shown. \u003cstrong\u003ed\u003c/strong\u003e, Views of the foreign DNA- (left) and repeat-bound (right) faces of Cas1-2/3 are shown in surface representation and colored by columbic potential. For clarity the highly electronegative DNA is shown in cartoon representation. Labels highlight highly basic and conserved surfaces of each Cas1-2/3 subunit that accommodate the packing of four dsDNA helices in proximity around the Cas2 homodimer (\u003cstrong\u003eExtended Data Fig. 3\u003c/strong\u003e). IHF heterodimers are shown in 100% transparency for clarity. The phosphate-to-phosphate distances of DNA helices packed around Cas2 are noted.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/cc1ab53dc6516be45ca13947.png"},{"id":38561275,"identity":"6e96fcec-9325-4566-acb7-fc21daf70bea","added_by":"auto","created_at":"2023-06-14 20:16:04","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":192596,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSequence motifs in the leader and IHF proteins facilitate Cas1-2/3-based integration into diverse repeat sequences\u003c/strong\u003e. \u003cstrong\u003ea\u003c/strong\u003e, Scheme of reactants and products of \u003cem\u003ein vitro\u003c/em\u003e CRISPR integration assays (\u003cstrong\u003eExtended Data Fig. 6,7 and 9\u003c/strong\u003e).\u003cstrong\u003e b\u003c/strong\u003e, Four CRISPR repeats used in the integration assays. A gapped sequence alignment highlights two identical (asterisks) and six similar (dots) positions. An un-gapped sequence alignment reveals nine identical nucleotide positions between the I-F and I-E repeats. All four repeats have different internal palindromes and GC-content (\u003cstrong\u003eExtended Data Fig. 8b\u003c/strong\u003e). \u003cstrong\u003ec\u003c/strong\u003e, Endpoint integration reactions with CRISPR repeat-swapped mutants, resolved on denaturing polyacrylamide gels. One of three representative gel images is shown (\u003cstrong\u003eExtended Data Fig. 7\u003c/strong\u003e). Quantification of leader- (grey circles) or spacer-side (white circles) integration events from all three replicate gels (\u003cstrong\u003eExtended Data Fig. 7\u003c/strong\u003e). The reactions were performed in triplicate, each dot represents one reaction, and some dots overlap. \u003cstrong\u003ed\u003c/strong\u003e, The 4-minute time point of time-course integration reactions with I-F repeat mutants, resolved on denaturing polyacrylamide gels. One of three representative images is shown (\u003cstrong\u003eExtended Data Fig. 9\u003c/strong\u003e). Quantification of leader- (grey circles) or spacer-side (white circles) integration events from all three replicate gels (\u003cstrong\u003eExtended Data Fig. 9\u003c/strong\u003e). \u003cstrong\u003ee\u003c/strong\u003e, CRISPR integration model. IHF-mediated folding of the genome presents IRs as symmetric DNA pillars that recruit foreign DNA-bound Cas1-2/3. Cas3 domains of Cas1-2/3 must rotate away from Cas2 to expose IR binding sites on Cas2. Cas1-2/3 and IHF cooperate to fold DNA into a loop, docking the leader-repeat junction at the Cas1 active site. Foreign DNA integration at the leader-repeat junction nicks the DNA duplex, releasing tension in the DNA duplex and inhibiting the reverse disintegration reaction (\u003cstrong\u003eExtended Data Fig. 4\u003c/strong\u003e) \u003csup\u003e33,34\u003c/sup\u003e. 5’ GT dinucleotides are required for efficient leader- and spacer-side integration, but no strict sequence requirements are necessary in the rest of the repeat.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/502e14b54c3f1b5a9808541b.png"},{"id":38561277,"identity":"e4d3806a-0409-4f34-ae94-3b15df5f364f","added_by":"auto","created_at":"2023-06-14 20:16:04","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":218216,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eI-F CRISPR integration complex suggests a mechanism for primed acquisition and Cascade impact on integration\u003c/strong\u003e. \u003cstrong\u003ea\u003c/strong\u003e, Cascade (grey) bound to foreign DNA (red) displays a Cas8 helical bundle (turquoise) that recruits Cas2/3. Cascade-recruited Cas3 degrades foreign DNA into fragments that are captured by the Cas1-2 integrase for subsequent integration \u003csup\u003e37,38,45,46,79–83\u003c/sup\u003e. The Cas1-2/3 integration complex can be docked onto dsDNA-bound Cascade with minimal clashing and suggests a model for the formation of a primed acquisition complex that facilitates rapid adaptation to genetic parasite variants \u003csup\u003e41,47,81\u003c/sup\u003e. \u003cstrong\u003eb\u003c/strong\u003e, Recruitment of the Cascade-Cas1-2/3 complex to IHF-folded CRISPR leader DNA is not predicted to form any new clashes, suggesting a mechanism for the role of Cascade in facilitating integration \u003cem\u003ein vivo\u003c/em\u003e \u003csup\u003e72,73\u003c/sup\u003e. Note that a total of two Cascade complexes can be docked onto the Cas1-2/3 integration complex (one per each Cas3) without introducing additional clashing, a single is shown above for clarity.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/81115ce3de3e4dbdd0539b04.png"},{"id":38561840,"identity":"07a1b47f-4218-4a91-b263-0ab5083d4ef4","added_by":"auto","created_at":"2023-06-14 20:24:04","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":358396,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDNA is a flexible scaffold that controls DNA mobilization\u003c/strong\u003e. Structures for the I-E CRISPR integration complex\u003cstrong\u003e (a\u003c/strong\u003e), I-F CRISPR integration complex (\u003cstrong\u003eb\u003c/strong\u003e), and Lambda phage excision complex (\u003cstrong\u003ec\u003c/strong\u003e). DNA shown as a surface, IHF (purple) and all other proteins shown as transparent cartoons. Integration and excision sites, along with DNA motifs that regulate DNA mobilization, are labeled and colored according to the schematic (bottom). \u003cstrong\u003ed\u003c/strong\u003e, A revised evolutionary scenario for CRISPR-Cas systems arising from a transposon that uses Cas1 as an integrase (“Casposon”) is proposed \u003csup\u003e51\u003c/sup\u003e. An ancestral Casposon encoding Cas1, Cas4 and Cas2 is shown. Transposon ends with DNA motifs that regulate transposition (grey) facilitate Casposon integration into a DNA target site near an operon encoding a generic type III Cascade complex (light green). Loss of one transposon end and target site duplication immobilizes the Casposon genes. The remaining transposon end (Leader) was repurposed for foreign DNA integration into the target site (repeat), forming CRISPR arrays with a repeat-spacer-repeat architecture.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/51fae7a574519f4a722095e1.png"},{"id":43886070,"identity":"6ef3c376-922c-456d-9727-e84ffca89157","added_by":"auto","created_at":"2023-09-29 15:34:17","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3938138,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/57685e2a-e88a-47a0-a9c7-ff6cea5f6d37.pdf"},{"id":38561273,"identity":"47cf31b3-e752-4c90-b3a6-b07972e059ab","added_by":"auto","created_at":"2023-06-14 20:16:04","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":18066,"visible":true,"origin":"","legend":"","description":"","filename":"IntegrationPaperSupplementaryTableS1.docx","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/4227ef3731468732503c310d.docx"},{"id":38561283,"identity":"8e80408c-4e93-48f1-9d35-9b2be533bc68","added_by":"auto","created_at":"2023-06-14 20:16:06","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":46776136,"visible":true,"origin":"","legend":"","description":"","filename":"ASFCas123EDatatosubmit.docx","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/5618e916f43a07c8f6355878.docx"},{"id":38561278,"identity":"8895875d-c703-48e5-97f3-d374ba3c5d5e","added_by":"auto","created_at":"2023-06-14 20:16:04","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":1260825,"visible":true,"origin":"","legend":"","description":"","filename":"ReportingSummaryWiedenheft.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/c46b1ee8d1b376d75111f571.pdf"},{"id":38561281,"identity":"9b47f3d3-b23c-4fef-a59e-c8f216efe8f0","added_by":"auto","created_at":"2023-06-14 20:16:04","extension":"mp4","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":12399609,"visible":true,"origin":"","legend":"Cas1-2/3 undergoes a large conformational change to unveil Cas2-leader binding sites.","description":"","filename":"IntegrationPaperSupplementaryVideo1.mp4","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/106bbc9d5c1c9c84165ca5ef.mp4"},{"id":38561284,"identity":"48bb8a29-ad73-4ea5-bfa9-ba58f27c5956","added_by":"auto","created_at":"2023-06-14 20:16:07","extension":"mp4","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":114041232,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of how the I-F CRISPR leader and IHF guide Cas1-2/3-mediated integration of foreign DNA at the first CRISPR repeat.\u003c/p\u003e","description":"","filename":"IntegrationPaperSupplementaryVideo2.mp4","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/280e89e7a4ba7ae915d04e8c.mp4"},{"id":38561282,"identity":"37de19e9-92f0-415e-a99f-3a32d670309c","added_by":"auto","created_at":"2023-06-14 20:16:05","extension":"mp4","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":39363903,"visible":true,"origin":"","legend":"\u003cp\u003eThe strained IHF-mediated DNA bends sway in relation to the Cas1-2/3 integrase.\u003c/p\u003e","description":"","filename":"IntegrationPaperSupplementaryVideo3.mp4","url":"https://assets-eu.researchsquare.com/files/rs-2982802/v1/331c2d595904e6f8630b2578.mp4"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential Competing Interest.\nB.W. is the founder of SurGene and VIRIS Detection Systems. B.W. and A.S.-F. are inventors on patent applications related to CRISPR-Cas systems and applications thereof.","formattedTitle":"Protein-mediated folding of the genome is essential for site-specific integration of foreign DNA into CRISPR loci","fulltext":[{"header":"Full Text","content":"\u003cp\u003eVertebrates, bacteria, and archaea have domesticated transposases (e.g., RAG1 and Cas1) for adaptive immunity\u0026nbsp;\u003csup\u003e1,2\u003c/sup\u003e. Integrases, transposases and recombinases often co-opt additional DNA-bending proteins (e.g., IHF, HU, H-NS, or HMGB1) that facilitate DNA integration and excision\u0026nbsp;\u003csup\u003e3–7\u003c/sup\u003e.\u0026nbsp;However, the structural role of DNA folding during this mobilization of DNA remains largely enigmatic.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eC\u003c/u\u003elustered \u003cu\u003eR\u003c/u\u003eegularly \u003cu\u003eI\u003c/u\u003enterspaced \u003cu\u003eS\u003c/u\u003ehort \u003cu\u003eP\u003c/u\u003ealindromic \u003cu\u003eR\u003c/u\u003eepeats (CRISPRs) are essential components of an adaptive immune system that stores DNA-based molecular memories of past infections\u0026nbsp;\u003csup\u003e8\u003c/sup\u003e. \u003cu\u003eC\u003c/u\u003eRISPR-\u003cu\u003eas\u003c/u\u003esociated proteins, Cas1 and Cas2, integrate fragments of foreign DNA (\"spacers\") into CRISPRs. Integration duplicates a repeat sequence, which thereby maintains the characteristic repeat-spacer-repeat architecture (\u003cstrong\u003eFig. 1a\u003c/strong\u003e). Cas1 and Cas2 form a heterohexameric complex that consists of two Cas1 homodimers (Cas1a-a* and Cas1b-b*) flanking a Cas2 homodimer (\u003cstrong\u003eFig. 1a-b\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e9–11\u003c/sup\u003e. Foreign DNA fragments bind across one face of the Cas2 homodimer, which positions the 3’-ends into Cas1 active sites on either end of the complex (i.e., Cas1a* and Cas1b*)\u0026nbsp;\u003csup\u003e8–10,12\u003c/sup\u003e. The CRISPR repeat sequence wraps around the opposing face of Cas2, sandwiching the Cas2 homodimer between the foreign and repeat DNA duplexes. Opposing Cas1 subunits (Cas1a* and Cas1b*) catalyze two successive strand-transfer reactions, linking the 3'-ends of the foreign DNA to opposite ends of the repeat\u0026nbsp;\u003csup\u003e5,11\u003c/sup\u003e. CRISPR integration complexes sense a 2-5 bp 3’ overhang called a protospacer-adjacent motif (PAM) in the foreign DNA to determine the integration orientation\u0026nbsp;\u003csup\u003e9\u003c/sup\u003e. Correct spacer orientation is necessary to produce a functional CRISPR RNA that guides the CRISPR interference machinery (i.e., Cascade) to complementary targets\u0026nbsp;\u003csup\u003e8,13\u003c/sup\u003e. Integration occurs in a stepwise manner. First, the non-PAM end of the foreign DNA is integrated at the leader-side of the repeat\u0026nbsp;\u003csup\u003e14\u003c/sup\u003e. Second, the PAM is cleaved by Cas or non-Cas nucleases and the trimmed 3’ end is integrated at the spacer-side of the repeat\u0026nbsp;\u003csup\u003e14–16\u003c/sup\u003e.\u0026nbsp;These integration events tie a non-covalent knot around the Cas2 homodimer (foreign DNA on one side and repeat DNA on the other) that is held together by complementary base-pairing in the foreign DNA.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eNew foreign DNA is preferentially integrated at the first repeat in a CRISPR locus, ensuring efficient transcription and processing of CRISPR RNAs that target the most recently encountered genetic parasites (\u003cstrong\u003eFig. 1a\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e17,18\u003c/sup\u003e. Cas1-2 is thought to recognize a palindromic sequence within the CRISPR repeat, similar to target site recognition by many DNA transposases\u0026nbsp;\u003csup\u003e5,11,19–21\u003c/sup\u003e. However, Cas1-2 recognition of the palidromic repeat does not explain how the first repeat is differentiated from downstream repeat sequences in a CRISPR. Thus, polarized integration often relies on additional proteins and DNA sequence motifs upstream of the CRISPR (i.e., leader)\u0026nbsp;\u003csup\u003e3–5,11,18,22–27\u003c/sup\u003e. \u003cu\u003eI\u003c/u\u003entegration \u003cu\u003eH\u003c/u\u003eost \u003cu\u003eF\u003c/u\u003eactor (IHF) facilitates polarized integration in the type I-E CRISPR system from \u003cem\u003eEscherichia coli\u0026nbsp;\u003c/em\u003e\u003csup\u003e3\u003c/sup\u003e. A structure of the I-E integration complex revealed that IHF bends the leader DNA to bring an upstream sequence motif into contact with Cas1, and IHF further stabilizes the Cas1-2 integrase at the first repeat through direct Cas1-IHF interactions\u0026nbsp;\u003csup\u003e3,5\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eCas1 and Cas2 are conserved components of CRISPR-mediated immune systems. However, the type I-F CRISPR system has a unique fusion of the Cas2 subunit to the Cas3-nuclease/helicase found in many type I systems (\u003cstrong\u003eFig. 1a-b\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e28,29\u003c/sup\u003e. Cas3 degrades Cascade-bound DNA into fragments with PAM-containing termini, that are captured by Cas1-2 and integrated into the CRISPR locus, in a process called “primed acquisition”\u0026nbsp;\u003csup\u003e4,30–37\u003c/sup\u003e. However, the structural mechanism for primed acquisition is unclear. Additionally, in contrast to the \u003cem\u003eE. coli\u003c/em\u003e I-E system, which relies on one IHF to bend DNA and recruit a DNA motif found ~50 bp upstream of the CRISPR repeat, most I-F, I-C, and some I-E CRISPR leaders, contain multiple IHF binding sites and multiple subtype-specific DNA motifs found up to 100-200 bp upstream of the CRISPR repeat (\u003cstrong\u003eFig. 1a and Extended Data Fig. 1a\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo determine how DNA sequence motifs in the CRISPR leader regulate the Cas1-2/3 integrase, we determined the structure of a ~560 kDa CRISPR integration complex from \u003cem\u003ePseudomonas aeruginosa\u003c/em\u003e.\u0026nbsp;The structure reveals that Cas1-2/3 cannot interact with the leader without first undergoing a large conformational change that may be induced by foreign DNA binding. Further, the structure explains how the I-F leader and IHF proteins guide Cas1-2/3 to deliver and integrate foreign DNA at the first repeat in the CRISPR array (\u003cstrong\u003eFig. 1\u003c/strong\u003e). Cas1-2/3 and IHF interact with all five DNA sequence motifs (i.e., two Inverted Repeats, two IHF binding sites, and the CRISPR repeat) primarily through a shape-based readout\u0026nbsp;\u003csup\u003e38,39\u003c/sup\u003e. The shape of the folded I-F CRISPR leader is similar to that of the λ-phage excision complex, suggesting that DNA is often used as a flexible scaffold to regulate DNA mobilization\u0026nbsp;\u003csup\u003e3–7,40–53\u003c/sup\u003e. The structure suggests that site-specific integration relies on protein-induced folding of the upstream DNA rather than sequence-specific recognition of the repeat. To test this idea, we perform a series of integration reactions demonstrating that efficient integration relies on conserved sequences in the leader and a 5’ GT dinucleotide in the repeat. We show that 5’ GT dinucleotides are broadly conserved in repeats derived from different CRISPR types, suggesting that they play a conserved role in integration across diverse CRISPR systems.\u0026nbsp;In addition, the I-F CRISPR integration complex suggests a structural mechanism for interactions of the Cas1-2/3 integrase with the Cascade surveillance complex, that may be neccesary for rapid adaptation to phage escape mutants\u003csup\u003e4,30–37\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCryo-EM structure of type I-F CRISPR integration complex\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo understand how the Cas1-2/3 integrase cooperates with IHF and CRISPR leader motifs to integrate foreign DNA at the first CRISPR repeat, we purified the heterohexameric Cas1-2/3 integrase and the IHFα-β heterodimer, incubated these proteins with a DNA substrate representing a half-site integration intermediate, and isolated the assembled complex using \u003cu\u003eS\u003c/u\u003eize \u003cu\u003eE\u003c/u\u003exclusion \u003cu\u003eC\u003c/u\u003ehromatography (SEC) (\u003cstrong\u003eExtended Data Fig. 1a-f\u003c/strong\u003e). The purified I-F integration complex was applied to cryo-EM grids and vitrified. We recorded 10,740 movies and picked 366,794 particles to determine a ~3.48 Å-resolution structure\u0026nbsp;of the integration complex. The reconstructed density was of sufficient to model 90.7% of the 10 polypeptides and 88.4% of the 396 nucleotides of DNA (\u003cstrong\u003eFig. 1b-d, Extended Data Fig. 1g-k, Table 1\u003c/strong\u003e). The model explains how the Cas1-2/3 subunits cooperate with two IHF heterodimers to kink and twist ~150 base-pairs of host DNA into a structure that precisely positions foreign DNA for integration at the first repeat of the CRISPR (\u003cstrong\u003eFig. 1b-d\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eThe Cas1 and Cas2 subunits adopt a familiar quaternary arrangement that binds a foreign DNA on one face of the Cas2 homodimer and CRISPR repeat DNA on the other face (\u003cstrong\u003eFig. 1d, 2a and 3a\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e8\u003c/sup\u003e. A previously determined\u0026nbsp;structure of Cas1-2/3 alone revealed that the Cas3 and Cas1 domains surround the central Cas2 homodimer like petals of a closed flower (Cas1\u003csub\u003e4\u003c/sub\u003e:Cas2/3\u003csub\u003e2\u003c/sub\u003e)\u0026nbsp;(\u003cstrong\u003eFig. 1b\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e54\u003c/sup\u003e. While this structure explained how Cas1 regulates the Cas3 nuclease, the role of Cas3 during integration remained unclear\u0026nbsp;\u003csup\u003e54\u003c/sup\u003e. Here, we show that the addition of DNA drives a series of conformational changes in both the DNA and proteins. The Cas3 domains rotate ~100° to align in a planar configuration with Cas2, simulating the motion of a bloomed flower, and exposing equivalent surfaces on opposite sides of the Cas2 homodimer that recognize an inverted repeat (IR) that is conserved in I-F leaders (\u003cstrong\u003eFig. 1d, Extended Data Fig. 2a and Supplementary Video 1\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e54\u003c/sup\u003e. Thus, the new planar conformation of Cas1-2/3 enables the simultaneous coordination of four DNA helices (IR\u003csub\u003edistal\u003c/sub\u003e, IR\u003csub\u003eproximal\u003c/sub\u003e, foreign DNA, and CRISPR repeat) around the central Cas2 homodimer (\u003cstrong\u003eFig. 1d and 3\u003c/strong\u003e). Further, this Cas3 rotation flips the nuclease domain from an interaction with Cas1 that suppresses the Cas3 nuclease activity, to the opposite side of the complex, where the back of the Cas3 nuclease domain docks onto a groove created at the Cas1-Cas1 interaface\u0026nbsp;(\u003cstrong\u003eFig. 1b,d and Extended Data Fig. 2a,b\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eThe structure reveals two prominent DNA bends that protrude at right angles from Cas1-2/3 (\u003cstrong\u003eFig. 1c,d\u003c/strong\u003e). An IHF heterodimer is wedged at the apex of each DNA bend, consistent with IHF's well-defined role in DNA bending\u0026nbsp;\u003csup\u003e38\u003c/sup\u003e. These two DNA protrussions extend ~75 Å from the Cas1-2 core. Flexablity of these DNA extentions limits the resolution of the regions to 4-8 Å (\u003cstrong\u003eFig. 1c,d and Supplementary Video 2\u003c/strong\u003e).\u0026nbsp;IHF-mediated bending of the IHF\u003csub\u003edistal\u003c/sub\u003e site positions the flanking IR sequences as symmetrical DNA pillars, which are recognized by equivalent surfaces on opposite sides of the Cas2 homodimer (\u003cstrong\u003eFig. 3 and\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eExtended Data Fig. 3b and 4c\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e. Cas2 binding to these DNA pillars traps foreign DNA on one face of the Cas1-2/3 integrase. Further, Cas2 bends the IRs and steers downstream DNA away from Cas1-2/3, which would project the downstream CRISPR repeat away from the Cas1-2/3 integrase (\u003cstrong\u003eFig. 1d\u003c/strong\u003e). However, Cas1-2/3 and IHF cooperate to constrict the DNA around the IHF\u003csub\u003eproximal\u003c/sub\u003e site, forming a loop that places the CRISPR repeat into the Cas1a* active site (\u003cstrong\u003eFig. 1d and 3\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eForeign DNA constrains the Cas2/3 linker against conserved Cas1 surfaces\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe type I-F Cas2 and Cas3 subunits are connected by a 20 amino acid disordered linker (residues 90-110)\u0026nbsp;\u003csup\u003e4,28,29,55\u003c/sup\u003e. The structure explains how foreign DNA constrains the Cas2/3 linker against conserved surfaces of Cas1, which suggests that foreign DNA-binding either initiates, or stabilizes the Cas3 rotation (\u003cstrong\u003eFig. 2a\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e54,55\u003c/sup\u003e.\u0026nbsp;The constrained Cas2/3 linker positions the HD nuclease domain of Cas3 (residues 111-374) against the Cas1-Cas1 interface, and facilitates Cas3 interactions with the IRs (\u003cstrong\u003eFig. 2a\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eand Extended Data Fig. 2 and 4c\u003c/strong\u003e).\u0026nbsp;The foreign DNA and amino acids in the Cas2/3 linker contact conserved residues in type I-F Cas1 proteins (\u003cstrong\u003eFig. 2d,e and Extended Data Fig. 2c\u003c/strong\u003e). Polar residues in the Cas2/3 linker may assist\u0026nbsp;the binding or splaying of the foreign DNA duplex at the conserved histidine wedge (H25) in Cas1 (\u003cstrong\u003eFig. 2b,c and Extended Data Fig. 3\u003c/strong\u003e). Mutation of the\u0026nbsp;histidine wedge (Cas1\u003csup\u003eH25A\u003c/sup\u003e) decreases Cas1-2/3 integration activity of foreign DNA that has either fully complementary or splayed DNA ends (\u003cstrong\u003eExtended Data Fig. 5 and 6a,b\u003c/strong\u003e). The integration defect on substrates with splayed ends suggest that H25 is more than a simple wedge that pries apart the ends for foreign DNA\u0026nbsp;\u003csup\u003e10\u003c/sup\u003e. The histidine steers the 3’-ends down a positively charged channel that positions each 3’-hydroxyl into Cas1 active sites on opposite ends of the complex (\u003cstrong\u003eFig. 2 and Fig. 4a,d\u003c/strong\u003e), whereas the 5’-ends of the protospacer DNA are directed towards the back face of the Cas3 HD domain (\u003cstrong\u003eFig. 2\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e9,10,56\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCas2 homodimers recognize and bend inverted repeat sequences\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe structure of the type I-F CRISPR integration complex reveals that Cas2 is the homodimer that binds the IRs (\u003cstrong\u003eFig. 3a\u003c/strong\u003e). Mutations that scramble the order of nucleotides in either the IR\u003csub\u003edistal\u003c/sub\u003e or IR\u003csub\u003eproximal\u003c/sub\u003e motifs limit Cas1-2/3-mediated integration\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e.\u0026nbsp;While Cas2 doesn’t make extensive sequence-specific contacts with nucleobases of the IR, a single residue (Cas2\u003csup\u003eR55\u003c/sup\u003e) intercalates in the minor groove, and may participate in recognizing two conserved bases in the 10 bp-long motif (\u003cstrong\u003eFig. 3a and Extended Data Fig. 4c,e\u003c/strong\u003e). However, there is insufficient density for the R55 sidechain to confidently assign contacts.\u0026nbsp;Other conserved Cas2 residues (i.e., K11, R12, and N56) form additional hydrogen bonds with the phosphate backbone of one DNA strand in each IR (\u003cstrong\u003eFig. 3a and Extended Data Fig. 4c\u003c/strong\u003e). Mutation of these Cas2 residues (Cas2\u003csup\u003eK11D,R12E\u003c/sup\u003e, Cas2\u003csup\u003eR55E,N56D\u003c/sup\u003e, Cas2\u003csup\u003eK11D,R12E,R55E,N56D\u003c/sup\u003e) prevents Cas1-2/3-mediated DNA integration (\u003cstrong\u003eExtended Data Fig. 5 and 6c,d\u003c/strong\u003e). Cas2 acts as a wedge that induces a 25-35° bend in the DNA upstream of IR\u003csub\u003edistal\u003c/sub\u003e and downstream of IR\u003csub\u003eproximal\u003c/sub\u003e (\u003cstrong\u003eFig. 3a\u003c/strong\u003e). These flared IRs lean against basic residues (K381, R393, K397) on the back surface of Cas3 (\u003cstrong\u003eFig. 3 and\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;Extended Data Fig. 3 and 4c\u003c/strong\u003e). In sum, these observations reveal that the IR DNA sequences are primarily recognized by Cas1-2/3 through shape readout rather than base readout\u0026nbsp;\u003csup\u003e39\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCas1-2/3 accommodates four dsDNA helices that surround the Cas2 homodimer\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCas2 is a cube-shaped homodimer at the center of the Cas-integrase. The Cas2 cube is flanked by Cas1 homodimers to form an elongated DNA binding platform that interacts with the CRISPR repeat on one face and the foreign DNA on the other (\u003cstrong\u003eFig. 3\u003c/strong\u003e). Unique to the type I-F Cas1-2/3 integration complex, the IRs occupy the last two accessible surfaces of the Cas2 cube (\u003cstrong\u003eFig. 3a\u003c/strong\u003e). Positively charged surfaces on Cas1-2/3 bind and shield negatively charged DNA, which enables the packing of four DNA helices around the small Cas2 homodimer (\u003cstrong\u003eFig. 3b and Extended Data Fig. 3\u003c/strong\u003e). The foreign DNA-binding face of Cas2 has two electronegative pillars of leader DNA that straddle the foreign DNA, such that major grooves of the leader DNA pillars are clamped against major grooves of the foreign DNA. The two DNA pillars continue past Cas2 to flank the Cas1 active sites (\u003cstrong\u003eFig. 3b\u003c/strong\u003e). At the IHF\u003csub\u003eproximal\u003c/sub\u003e loop, Cas3 packs the leader against the Cas1-bound repeat, decreasing the phosphate-to-phosphate distances between these helices to ~11-12 Å. Although the latter two-thirds of the CRISPR repeat could not be resolved, the repeat’s trajectory suggests it will follow a path that threads between the distal leader DNA duplex and the 3’-hydroxyl of the foreign DNA that rests in the Cas1b* active site (\u003cstrong\u003eFig. 3b\u003c/strong\u003e). Collectively, Cas1, the Cas2/3 linker, and Cas3 accommodate four DNA helices (IR\u003csub\u003edistal\u003c/sub\u003e, IR\u003csub\u003eproximal\u003c/sub\u003e, foreign DNA, and repeat) around the central Cas2 homodimer to facilitate site-specific integration.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIHF and the leader sequence facilitate integration into diverse repeat sequences\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe structure suggests that Cas1-2/3 is guided to the first repeat of the CRISPR by IHF-mediated folding of the I-F leader, rather than direct recognition of the repeat sequence (\u003cstrong\u003eFig. 1d and Extended Data Fig. 4b,d\u003c/strong\u003e). To determine if or how the repeat sequence impacts integration, we measured the efficiency of Cas1-2/3-catalyzed integration into DNAs containing either a I-F, I-E, I-C or II-A repeat downstream of a I-F leader (\u003cstrong\u003eFig. 4a-c\u003c/strong\u003e). The type I-F leader supports Cas1-2/3-catalyzed leader-side integration at repeats derived from I-E, I-C, and II-A CRISPR loci (\u003cstrong\u003eFig. 4a-c and Extended Data Fig. 7 and 8\u003c/strong\u003e). Integration efficiency at non-native repeats is not correlated with sequence similarity to the I-F repeat or with GC-content (\u003cstrong\u003eExtended Data\u003c/strong\u003e \u003cstrong\u003eFig. 8b\u003c/strong\u003e). Instead, integration efficiency is correlated to the length of the repeat. I-F and I-E repeats are similar in length (28 and 29 bps, respectively), whereas the I-C and II-A repeats are 0.5 to 1 full DNA turns longer than the I-F repeat (33 and 37 bp, respectively) (\u003cstrong\u003eFig. 4b\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;and Extended Data Fig. 9a,b\u003c/strong\u003e). While leader-side integration is robust with different repeats, spacer-side integration is ~3.5-fold slower for the I-E repeat, and undetectable for the longer repeats (\u003cstrong\u003eExtended Data Fig. 7 and 9a,b)\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eAs expected, Cas1-2/3 does not catalyze integration at a I-F repeat downstream of a scrambled I-F leader, nor does Cas1-2/3 does catalyze integration at I-E, I-C or II-A repeats downstream of their respective leaders\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e (\u003cstrong\u003eExtended Data Fig. 7\u003c/strong\u003e). Cas1-2 sequences are diverse, such that leader-interacting residues are only conserved in a subset of proteins within a given CRISPR subtype\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e. For example, the I-F Cas1 protein lacks residues required to interact with the I-E leader\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e.\u0026nbsp;Further, the I-E, I-C and I-F leaders have distinct nucleotide spacings between the leader motifs and the repeat. These nucleotide spacings impart a unique shape to the DNA, which is critical for integration.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThese experiments were performed with foreign DNA substrates either with or without a PAM (\u003cstrong\u003eExtended Data Fig. 7\u003c/strong\u003e). Cas1-2/3 integration of PAM-containing DNA is more specific, but the conclusions are otherwise consistent between the two substrates. The PAM must be trimmed by an ancillary nuclease before Cas1-2/3 can catalyze spacer-side integration, therefore we focused our discussion on results from the trimmed foreign DNA to compare differences in spacer-side integration (\u003cstrong\u003eExtended Data Fig. 7a-d\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e14–16\u003c/sup\u003e. Collectively, these integration experiments indicate that leader sequences and host factors dictate site-specific integation of foreign DNA at diverse DNA target sites, providing insight into the evolution of new CRISPR systems in nature.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5’ GT dinucleotides in CRISPR repeat sequences are critical for Cas1-mediated integration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRepeat sequences are strongly conserved within CRISPR subtypes, but vary in sequence and length between subtypes\u0026nbsp;\u003csup\u003e57,58\u003c/sup\u003e. However, in the small subset of repeats tested above, we noticed that the 5’ GT is conserved. To determine if conservation of the 5’ GT is a coincidence or a more widely conserved feature of repeats, we performed a bioinformatic analysis consisting of 24,940 CRISPRs. This bioinformatic analysis reveals that a 5’ GT dinucleotide is broadly conserved at the leader-side of the repeat, and conserved in some CRISPR systems at the spacer-side of the repeat (\u003cstrong\u003eExtended Data Fig. 4g\u003c/strong\u003e). Therefore we hypothesized that the 5’ GT dinucleotide is a base-specific determinant for leader-side integration. To test this hypothesis, we mutated the 5’ GT (G1A, T2A) and repeated the integration assays. The 5’ GT to AA mutation ablates both leader- and spacer-side integration, indicating that the 5’ GT is essential and that leader-side integration is a prerequisite for spacer-side integration (\u003cstrong\u003eFig. 4d\u003c/strong\u003e). Since Cas1 requires a 5’ GT at the leader-side of the repeat, we hypothesized that introducing a 5’ GT at the spacer-side of the repeat would increase spacer-side integration efficiency. To test this hypothesis, we replaced adenosine 28 of the I-F repeat with cytosine (A28C) and repeated the integration assays. The A28C mutation increases the rate and amount of spacer-side integration by ~2-fold, relative to the WT I-F repeat (\u003cstrong\u003eFig. 4d\u003c/strong\u003e). We do not detect integration into a I-F repeat that lacks a 5’ GT at the leader-side even if the repeat contains a 5’ GT at the spacer-side. This result further supports our conclusion that leader-side integration is a prerequisite for spacer-side integration (\u003cstrong\u003eFig. 4d\u003c/strong\u003e). We examined the structure of the Cas1 active site to determine whether the 5’ G is directly recognized by protein contacts. The Cas1 residue E184 is within 4 Å of the 5’ G\u0026nbsp;(\u003cstrong\u003eExtended Data Fig. 4b,d\u003c/strong\u003e). However, a Cas1\u003csup\u003eE184A\u003c/sup\u003e mutation destabilizes the complex, decreasing the amount of Cas1 subunits per Cas1-2/3 complex (\u003cstrong\u003eExtended Data Fig. 5\u003c/strong\u003e). Therefore, the decrease in integration activity of the Cas1\u003csup\u003e\u0026nbsp;E184A\u003c/sup\u003e-2/3 complex cannot be solely attributed to a decrease in recognition of the repeat (\u003cstrong\u003eExtended Data Fig. 9d,f\u003c/strong\u003e). Collectively, these data reveal that the 5’ GT is a conserved feature necessary for integration in most CRISPR systems, though no available structure provides a mechanism for direct recognition of the 5’ GT of the repeat\u0026nbsp;\u003csup\u003e5,11,15,59\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eHere we demonstrate that Cas1-2/3 and IHF fold DNA into a structure that is necessary for site-specific integration of foreign DNA into CRISPRs.\u0026nbsp;IHF proteins are highly expressed and most IHF binding sitesare thought to be occupied \u003cem\u003ein vivo\u003c/em\u003e \u003csup\u003e60\u003c/sup\u003e. Therefore, IHF may pre-fold the CRISPR leader into a \"landing pad\" that recruits foreign DNA-bound Cas1-2/3 (\u003cstrong\u003eFig. 4e\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e22,38\u003c/sup\u003e. The Cas3 and Cas1 domains of the Cas1-2/3 complex are arranged like petals of a closed flower around the central Cas2 homodimer, such that the Cas3 domains occlude two of the four DNA binding surfaces on Cas2, which precludes interactions with the leader (\u003cstrong\u003eFig. 1b,d\u003c/strong\u003e). Foreign DNA binding to Cas1-2/3 physically constrains the Cas2/3 linker against the Cas1 homodimer, pulling the Cas3 HD domain against the Cas1-Cas1 interface (\u003cstrong\u003eFig. 2, Extended Data Fig. 2 and Supplementary Video 1\u003c/strong\u003e). The 100-degree rotation of each Cas3 simulates the motion of a bloomed flower and exposes DNA binding sites on Cas2 that interact with each of the inverted repeat (IR) motifs in the leader (\u003cstrong\u003eFig. 1, 3, Extended Data Fig. 2,3\u003c/strong\u003e). Cas1-2/3 and IHF proteins fold DNA around the leader-repeat junction into a 260° loop that docks the first CRISPR repeat into the Cas1 active site under tension (\u003cstrong\u003eFig. 1, 4, Extended Data Fig. 3, 4\u003c/strong\u003e). The structure suggests that Cas1-mediated strand-transfer releases tension in this DNA loop, which may prevent disintegration of an otherwise isoenergetic strand-transfer reaction, and thereby favour complete integration (\u003cstrong\u003eExtended Data Fig. 4f\u003c/strong\u003e). A similar mechanism has been proposed to favour complete integration in other systems, where both strand transfer events occur simultaneously\u0026nbsp;\u003csup\u003e61,62\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWe show that Cas1-2/3, IHF and the I-F leader facilitate leader-side integration at four different repeat sequences (\u003cstrong\u003eFig. 4b-c\u003c/strong\u003e). These repeats are diverse in sequence identity, length, palindrome and GC-content, but they share a 5’ GT (\u003cstrong\u003eFig. 4b and Extended Data Fig. 8b\u003c/strong\u003e). To determine whether the 5’ GT is a universal feature of CRISPR repeats we analyzed 24,940 CRISPRs. This analysis reveals that CRISPR repeats contain a strongly conserved 5’ GT at the leader-end, and that a 5’ GT is also conserved at the spacer-end of repeats from several CRISPR systems (\u003cstrong\u003eExtended Data Fig. 4g\u003c/strong\u003e). We demonstrate the 5’ GT is critical for leader-side integration and that introducing a 5’ GT at the spacer-end of the repeat increases spacer-side integration. The broad conservation of 5’ GT is consistent with previous reports that type I-A, II-A, and I-E systems require a 5’ G for integration\u0026nbsp;\u003csup\u003e11,63\u003c/sup\u003e. Collectively, these data suggest that Cas1 proteins retain a shared sequence preference for a 5' G or a 5' GT, and lack strict sequence requirements for the central body of the repeat (\u003cstrong\u003eFig. 4b,c\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e64\u003c/sup\u003e.\u0026nbsp;The lack of strict sequence requirements may be advantageous because the CRISPR repeat is at the nexus of foreign DNA integration, processing of the transcribed CRISPR, and loading the mature crRNA into the surveillance complexes (e.g., Cascade, Cas9).\u003c/p\u003e\n\u003cp\u003eGenetic parasites commonly escape CRISPR-based immunity through point mutations\u0026nbsp;\u003csup\u003e65\u003c/sup\u003e. To counter escape mutants, many CRISPR-Cas systems use existing spacer sequences to enhance the acquisition of new spacers from the same foreign genetic element via “primed” acquisition\u0026nbsp;\u003csup\u003e4,30–37\u003c/sup\u003e. In some examples of primed acquisition, the Cas3 nuclease/helicase degrades CRISPR-targeted DNAs into ssDNA fragments enriched in PAM-containing termini\u0026nbsp;\u003csup\u003e66\u003c/sup\u003e. Cas1-2 has been proposed to anneal complementary ssDNA fragments and integrate these into CRISPRs\u0026nbsp;\u003csup\u003e14,30,31,66\u003c/sup\u003e. Single-molecule co-localization and bulk immunoprecipitation suggest that the type I-E Cas1-2 integrase is recruited to a Cas3-Cascade-target DNA complex to facilitate primed acquisition\u0026nbsp;\u003csup\u003e34,67\u003c/sup\u003e. The structure of the type I-F integration complex reveals conformational changes in Cas3 that may enable interactions with DNA-bound Cascade (\u003cstrong\u003eFig. 5a\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e36,68\u003c/sup\u003e. Cascade improves integration efficiency and fidelity \u003cem\u003ein vivo\u0026nbsp;\u003c/em\u003e\u003csup\u003e68,69\u003c/sup\u003e, and the structure suggests a model for the formation of a primed acquisition complex (Cas1-2/3-Cascade-target DNA) that transfers new foreign DNA fragments to the integrase.\u0026nbsp;Additional structures will be necessary to clarify the mechanism(s) of primed adaptation.\u003c/p\u003e\n\u003cp\u003eA comparison of the I-E and I-F integration complexes with a structure of the λ-phage excision complex, reveals structural similarities and differences. In I-E systems, the IHF protein folds the leader DNA to present an upstream motif to a lobe of Cas1. In contrast, the I-F structure highlights extensive cooperation between IHF and Cas1-2/3 in bending the leader into an energetically strained conformation that may increase the specificity of CRISPR recognition (\u003cstrong\u003eFig. 6a,b\u003c/strong\u003e). This cooperation includes Cas1-2/3 kinking the IR DNA motifs presented as parrallel DNA pillars by IHF, and Cas1-2/3 constricting the IHF\u003csub\u003eproximal\u003c/sub\u003e 180° bend ito a 260° (\u003cstrong\u003eFig. 6b\u003c/strong\u003e). Further, the sequestration of the IR-binding surface of Cas2 suggests a unique structural mechanism that prevents Cas1-2/3 interactions with the leader until foreign DNA binding induces a rotation of Cas3 (\u003cstrong\u003eFig. 4e,\u003c/strong\u003e \u003cstrong\u003eSupplementary Video 1 and 3\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e54\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eMany CRISPR leaders contain multiple IHF binding sites and subtype-specific motifs that are reminiscent of motifs found at the ends of transposons, phages, and plasmids, which facilitate integration, excision, or recombination complex assembly (\u003cstrong\u003eExtended Data Fig. 8a\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e3–7,22,40–53\u003c/sup\u003e. For example, the left and right ends of the λ-phage genome contain three IHF binding sites, five copies of a second DNA motif (P1-2,1’-3’), and three copies of third DNA motif (X1,1.5,2), that recruit λ-phage (Int, Xis) and non-λ-phage (IHF, Fis) proteins (\u003cstrong\u003eFig. 6c\u003c/strong\u003e)\u003csup\u003e7\u003c/sup\u003e. These proteins use DNA as a flexible scaffold that organizes enzyme active sites and DNA substrates to facilitate either the integration, or excision, of λ-phage DNA from the bacterial genome (\u003cstrong\u003eFig. 6c\u003c/strong\u003e). Similarly, the I-F CRISPR leader motifs recruit Cas (Cas1-2/3) and non-Cas (IHF) proteins, that bend DNA into a flexible scaffold to organize enzyme active sites and DNA substrates in space, facilitating the integration of foreign DNA into the bacterialgenome (\u003cstrong\u003eFig. 6b\u003c/strong\u003e). Diverse systems thus use DNA as a flexible scaffold to regulate the isoenergetic mobilization of DNA (\u003cstrong\u003eFig. 6a-c\u003c/strong\u003e).”\u003c/p\u003e\n\u003cp\u003eThe CRISPR adaptive immune system has been proposed to have originated through the domestication of a transposon that used a Cas1 homolog to catalyze integration (Casposon)\u0026nbsp;\u003csup\u003e1\u003c/sup\u003e. The pathway for the domestication of proteins from different operons to form the Cas components of CRISPR-Cas system has been previously proposed\u0026nbsp;\u003csup\u003e1,70\u003c/sup\u003e. However, the domestication or evolution of DNA elements that gave rise to the CRISPR repeat and leader has remained unclear\u0026nbsp;\u003csup\u003e1,70\u003c/sup\u003e. The similarities in motif architectures between CRISPR leaders and transposon DNA ends suggests that the Capsoson DNA end may have been domesticated to form the primordial CRISPR “leader” (\u003cstrong\u003eFig. 6d\u003c/strong\u003e). Similarly, the target DNA site that the ancient Casposase integrated at may have been co-opted to form the first CRISPR “repeat” (\u003cstrong\u003eFig. 6d\u003c/strong\u003e). Originally, the transposon end directed Casposase to integrate Casposon (foreign) DNA into the target site. Now, the domesticated leader directs Cas1-2 to integrate foreign DNA into the domesticated repeat (\u003cstrong\u003eFig. 6d\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eDiverse DNA mobilizing enzymes across the tree of life co-opt DNA folding to regulate DNA mobilization\u0026nbsp;\u003csup\u003e3–7\u003c/sup\u003e. In sum, these data provide a mechanistic understanding for the role of DNA as a flexible scaffold that controls DNA mobilization. These insights are critical to developing applications of DNA-mobilizing enzymes in gene therapy, genetic engineering, and chronological DNA recordings\u0026nbsp;\u003csup\u003e6,53,71–75\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eNucleic acid preparation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFour single-stranded DNAs (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e) were synthesized (IDT) and resuspended in 1x TE buffer (10 mM Tris-HCl pH 8, 1 mM EDTA) before being used to assemble the structure of the type I-F integration complex, the assembly is detailed in a below section. The splayed foreign DNAs used in integration assays (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e) were synthesized and resuspended in 1x TE buffer before use. To make \u003csup\u003e32\u003c/sup\u003eP-labeled CRISPR integration substrates, the sequences consisting of the leader and CRISPR arrays were first synthesized and cloned into pUC57 (Genscript). These plasmids have been made available on Addgene (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e). These plasmids were transformed into chemically competent \u003cem\u003eE. coli\u0026nbsp;\u003c/em\u003eDH5α cells and the transformed cells were plated onto LB agar plates containing 100 µg/mL Ampicillin. These cells were cultured in LB media and plasmids were purified using ZymoPURE II Plasmid Midiprep kit (Zymo Research). Each plasmid was then digested with EcoRI-HF and BamHI-HF (NEB) restriction enzymes, and the 294-383 bp inserts of interest were separated from the vector backbone by agarose gel electrophoresis. The gel segments containing the DNA inserts of interest were excised and DNA was purified using a Zymoclean Gel DNA Recovery kit (Zymo Research) (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e). The 5' ends of the CRISPR leader and array fragments were dephosphorylated using Quick Calf Intestinal alkaline Phosphatase (NEB), and the DNAs were purified away from protein using a DNA Clean \u0026amp; Concentrator kit (Zymo Research). Both 5' ends of 1 pmole of the CRISPR leader and array fragments were then labelled with \u003csup\u003e32\u003c/sup\u003eP, by incubation with 4 pmoles of [γ-\u003csup\u003e32\u003c/sup\u003eP]ATP (PerkinElmer) by polynucleotide kinase (NEB) in 1x PNK buffer at 37°C for 45 minutes. PNK was heat-denatured by incubation at 65°C for 20 minutes. Spin column purification (G-25, GE Healthcare) was used to remove unincorporated radioactive nucleotides and to buffer exchange DNAs into 1x TE buffer.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCas1 \u0026amp; Cas2/3 Mutagenesis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe plasmid used to express Cas1 and Cas2/3 (Addgene #89240) was PCR amplified with mutagenic primer pairs using Q5 polymerase (NEB) (\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e). The parental template plasmid was digested with DpnI (NEB) and the PCR amplified products were purified using the DNA Clean and Concentrate kit (Zymo). The purified DNA was 5’ phosphorylated with T4 Polynucleotide Kinase (NEB) and ligated with homemade T4 DNA ligase. The ligation reactions were transformed into DH5α cells. The Cas2 K11D, R12E, R55E, N56D mutant was prepared using the Cas2 K11D, R12E mutant as the template with the PCR primers for Cas2 R55E, N56D in the mutagenic PCR reaction. Plasmids for expressing the mutant Cas1-2/3 complexes, Cas1H25A-2/3, Cas1E184A-2/3, Cas1-2K11D,R12E/3, Cas1-2R55E,N56D/3, and Cas1-2K11D,R12E,R55E,N56D/3 have been deposited at Addgene (#200213, #200214, #200216, #200217, #200218).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProtein purification\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eP. aeruginosa\u003c/em\u003e IHF heterodimer was purified as previously described\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e. Briefly, 6xHis-tagged IHFα and StrepII-tagged IHFβ were co-expressed in \u003cem\u003eE. coli\u0026nbsp;\u003c/em\u003eBL21(DE3). Cell pellets were lysed by sonication in IHF lysis buffer (25 mM HEPES-NaOH pH 7.5, 500 mM NaCl, 10 mM Imidazole, 1 mM TCEP, 5% Glycerol), supplemented with 0.3x Halt Protease Inhibitor Cocktail (ThermoFisher), at 4°C. Lysate was clarified by two rounds of centrifugation at 12,000 rpm for 15 minutes, at 4°C. His-tagged IHF was captured on HisTrap HP resin (Cytiva), and eluted with 500 mM Imidazole. Affinity tags were cleaved using PreScision protease, and the PreScision protease and remaining 6x-His-IHFα were removed by affinity chromatography using HisTrap HP resin (Cytiva). Untagged IHF heterodimer was then further purified on Heparin Sepharose (Cytiva) and eluted with a linear gradient to a buffer containing 2 M NaCl. Fractions containing IHF heterodimer were concentrated and further purified by SEC (\u003cu\u003eS\u003c/u\u003eize \u003cu\u003eE\u003c/u\u003exclusion \u003cu\u003eC\u003c/u\u003ehromatography) on a Superdex 75 column (Cytiva) equilibrated in IHF buffer (25 mM HEPES-NaOH pH 7.5, 200 mM NaCl, 5 % Glycerol) (\u003cstrong\u003eExtended Data Fig. 1a\u003c/strong\u003e). IHF overexpression plasmids have been previously described and deposited at Addgene (#149384, #149385)\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eP. aeruginosa\u003c/em\u003e Cas1-2/3 heterohexamer complexes (wildtype and mutants) were purified as previously described\u0026nbsp;\u003csup\u003e22\u003c/sup\u003e. Briefly, StrepII-tagged Cas1-Cas2/3 was overexpressed in \u003cem\u003eE. coli\u0026nbsp;\u003c/em\u003eBL21(DE3). Cell pellets were lysed via sonication in Cas1-2/3 lysis buffer (50 mM HEPES pH 7.5, 500 mM KCl, 10% Glycerol, 1 mM DTT), supplemented with 0.3x Halt Protease Inhibitor Cocktail (ThermoFisher), at 4°C. Lysate was clarifed as above. StrepII-tagged Cas1-Cas2/3 complexes were affinity purified on StrepTrap HP resin (GE Healthcare) and eluted with Cas1-2/3 Lysis Buffer containing 3 mM desthiobiotin (Sigma-Aldrich). Eluate was concentrated at 4°C (Corning Spin-X concentrators), before purification over a Superdex 200 size-exclusion column (Cytiva) equilibrated in 10 mM HEPES pH 7.5, 500 mM Potassium Glutamate, and 10% Glycerol (\u003cstrong\u003eExtended Data Fig. 1c and 5e\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eIn vitro\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003cstrong\u003eintegration assays\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEnd-point integration reactions were performed in triplicate using 300 nM of splayed foreign DNA fragments, containing or lacking a PAM (IDT), 200 nM of Cas1-2/3, 300 nM of IHF heterodimer and roughly 1 nM of a given \u003csup\u003e32\u003c/sup\u003eP-labelled CRISPR variant fragment in Integration Buffer (20 mM HEPES pH 7.5, 150 mM potassium glutamate, 5 mM MnCl\u003csub\u003e2\u003c/sub\u003e, 1 mM TCEP, 1% Glycerol) (\u003cstrong\u003eSupplementary Table S1)\u003c/strong\u003e. Reactions were assembled on ice and then incubated at 35°C for 20 minutes.Timecourse integration reactions were performed in triplicate using 300 nM of splayed, or fully complementary, foreign DNA fragments lacking a PAM (IDT), 200 nM of Cas1-2/3, 300 nM of IHF heterodimer and roughly 1 nM of a given \u003csup\u003e32\u003c/sup\u003eP-labelled CRISPR variant fragment in modified Integration Buffer (20 mM HEPES pH 7.5, 150 mM potassium glutamate, 5 mM MnCl\u003csub\u003e2\u003c/sub\u003e, 10 mM TCEP, 1% Glycerol) (\u003cstrong\u003eSupplementary Table S1)\u003c/strong\u003e. We noticed that IHF and Cas1-2/3 protein stocks exhibited a propensity to precipitate when diluted into pre-chilled buffer to form working dilution stocks. Therefore all working dilutions of IHF and Cas1-2/3 prepared for timecourse assays were made by first diluting the proteins into room-temperature buffer, mixed, and then chilled on ice. Reactions were assembled on ice and then incubated at 20 minutes. Timepoints were taken at 0, 1, 2, 4 and 8 minutes. Reactions were stopped by the addition of phenol. The aqueous (nucleic acid-containing) layer was mixed 1:1 with 2x formamide loading buffer (95% formamide, 20 mM EDTA, 0.05% bromophenol blue, 0.05% xylene cyanol) and then denatured at 95°C for 5 minutes, before resolving the \u003csup\u003e32\u003c/sup\u003eP-labeled CRISPR substrates and integration products on a 7% (w/v) (29:1 mono:bis) polyacrylamide Urea gel in 1x TBE (100 mM Tris-Borate pH 8.3, 2 mM EDTA). Gels were dried and quantified using a Typhoon phosphorimager (GE Healthcare). The intensity of full-length CRISPR variant, leader-side integration fragments and spacer-side integration fragments were quantified with Multi Gauge v3 (Fujifilm). These readings were then used to calculate leader- and spacer-side integration events as percentages of all events. Images of all gels that resolved integration reactions are shown \u003cstrong\u003eExtended Data Fig. 6,7 and 9\u003c/strong\u003e), and additional control gels show that Cas1-2/3 is required for integration, and show how the custom \u003csup\u003e32\u003c/sup\u003eP-labelled ladder was generated by restriction enzyme digestion of \u003csup\u003e32\u003c/sup\u003eP-labelled CRISPR variant DNAs (\u003cstrong\u003eExtended Data Fig. 8\u003c/strong\u003e). Timecourse integration data was fit to a plateau followed by one phase association (GraphPad Prism).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAssembly and purification of I-F integration complex\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of four single-stranded DNAs (ssDNAs) synthesized to mimic a half-site integration intermediate were annealed in a step-wise manner. 2 nanomoles of ssDNAs mostly corresponding to the sense and anti-sense strands of the CRISPR leader (\"strand_1\" and \"strand_2\" were denatured at 100°C and then slow annealed using a PCR program that cooled the samples to 25°C in 6°C over an hour, in 100 µL of hybridization buffer (20 mM Tris-HCl pH 7.5, 100 mM monopotassium glutamate, 5 mM EDTA, 1 mM TCEP). 2 nanomoles of ssDNAs mostly corresponding to the corresponding to the sense and anti-sense of the strands of the foreign DNA (\"strand_3\" and \"strand_4\") were slow annealed using the same protocol (\u003cstrong\u003eExtended Data Fig. 1a and\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eSupplementary Table S1\u003c/strong\u003e). The two sets of annealed DNAs (tube 1: \"Strand_1\" and \"Strand_2\"; tube 2: \"Strand_3\" and \"Strand_4\") were mixed together, heated to 80°C and then slow annealed using a PCR program that cooled the samples to 25°C in 6°C over an hour, to anneal the complementary sense and anti-sense regions of the CRISPR repeat included in Strand_2 and Strand_3 together. Next, 6 nanomoles of IHF heterodimer in 50 µL of hybridization buffer was warmed to 25°C and then mixed and incubated with the annealed DNAs at 25°C for 10 minutes. Next, 3 nanomoles of Cas1-2/3 in 250 µL of hybridization buffer was warmed to 25°C and mixed with the prepared DNA and IHF mixture, and incubated at 25°C for 10 minutes. The total concentration of monopotassium glutamate in the mixture at this stage was ~200 mM, due to carryover from the stored protein stocks. This sample was centrifuged at 22,000 g, 4°C for 20 minutes to remove precipitates. The type I-F CRISPR integration complex was then purified on a Superdex 200 10/300 column (Cytiva) equilibrated in SEC buffer (20 mM Tris-HCl pH 7.5, 200 mM monopotassium glutamate, 5 mM EDTA, 1 mM TCEP, 2 % Glycerol). 0.5 ml fractions were individually concentrated and stored. The sixth SEC fraction contained all DNAs and proteins of interest and was further analyzed by cryo-EM (\u003cstrong\u003eExtended Data Fig. 1d-f\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCryo-EM sample preparation and data acquisition\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePurified integration complex was diluted to a concentration of 1 µM in SEC buffer lacking glycerol (20 mM Tris-HCl pH 7.5, 200 mM monopotassium glutamate, 5 mM EDTA, 1 mM TCEP), such that the final glycerol concentration was 0.2% within 1 hour of freezing. Sample was applied to Quantifoil R2/2 Cu 200 mesh grids that were glow discharged using 15 mA for 15 seconds with a 10 second hold (easiGlow, Pelco). 4 µl of diluted integration complex was applied to the grids, and then the grids were blotted for 5-6 seconds using Vitrobot\u003csup\u003eTM\u003c/sup\u003e Filter paper (Electron Microscopy Sciences) with a blot force of 6, at 100% humidity, 8°C, followed by plunge freezing into liquid ethane using a Vitrobot (Mk. IV, ThermoFisher Scientific). A preliminary dataset of 230 movies was collected on Montana State University's Talos Arctica transmission electron microscope (ThermoFisher Scientific), with a field emission gun operating at an accelerating voltage of 200 kV using parallel illumination conditions\u0026nbsp;\u003csup\u003e76\u003c/sup\u003e. Movies were acquired using a Gatan K3 direct electron detector, operated in electron counting mode applying a total electron exposure of 50 e-/Å\u003csup\u003e2\u003c/sup\u003e over 50 frames (3.995 s exposure, 0.08 s frame time). The SerialEM data collection software was used to collect micrographs at 36,000-fold nominal magnification (1.152 Å/pixel at the specimen level) with a nominal defocus set to 0.5 µm - 2.0 µm\u0026nbsp;\u003csup\u003e77\u003c/sup\u003e. Stage movement was used to target the center of four 2.0 µm holes for focusing, and image shift was used to acquire high magnification images in the center of each of the holes. An preliminary reconstruction was determined from a curated set of 160 images that had CTF fits less than 9 Å and a full-frame motion less than 40 pixels. Briefly, a round of blob picking (150-270 Å) followed by 2D classification was used to identify 2D classes used as templates for template picking in cryoSPARC\u0026nbsp;\u003csup\u003e78\u003c/sup\u003e. Template picking identified 33,403 initial particles from the above 160 images. 2D classification of these 33,403 particles into 50 classes was used to identify 5 classes with strong structural features containing 4,002 particles. Non-uniform refinement of these 4,002 particles resulted in a ~14.7Å reconstruction that appeared to contain a complete integration complex, therefore new grids were prepared as described above and shipped to National Center for CryoEM Access and Training (NCCAT) and the Simons Electron Microscopy Center located at the New York Structural Biology Center (NYSBC) for additional data collection (\u003cstrong\u003eTable 1\u003c/strong\u003e). At NCCAT, grids were imaged using a 300 kV Titan Krios G3i (Thermo Fisher Scientific) equipped with a GIF BioQuantum and K3 camera (Gatan). 10,740 images were recorded with Leginon\u0026nbsp;\u003csup\u003e79\u003c/sup\u003e (Suloway et al., 2005) with a calibrated pixel size of 0.5335 Å/px (micrograph dimension of 11520 x 8184 px) over a nominal defocus range of −0.7 μm to −2.1 μm and 20 eV slit. Movies were recorded in “super-resolution mode” (native K3 camera binning 1) with subframes of 50 ms over a 2.5 s exposure (50 frames) to give a total exposure of\u0026nbsp;∼69 e-/Å\u003csup\u003e2\u003c/sup\u003e (\u003cstrong\u003eTable 1\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCryo-EM image processing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePatch motion correction and patch CTF correction were performed in cryoSPARC\u0026nbsp;\u003csup\u003e78\u003c/sup\u003e. 3,792 of 10,740 total images (\u003cem\u003eCTF \u0026lt; 8Å, Full-frame motion \u0026lt; 30Å\u003c/em\u003e) were processed first to build an initial template. Blob picking was used to pick particles with diameters ranging from 120-280 Å. These ~1.8 million particles were extracted, Fourier-binned 2x2, and then subjected to 2D classification (\u003cem\u003eCustom parameters: Initial classification uncertainty factor = 3; \u0026nbsp;Number of online-EM iterations = 30; Batchsize per class = 200\u003c/em\u003e) (\u003cstrong\u003eExtended Data Fig. 1g\u003c/strong\u003e). Particles from 82 of the 200 2D classes were selected for an initial round of \u003cem\u003eab initio\u003c/em\u003e reconstruction and heterogeneous refinement. Particles from 1 of 5 of these classes were selected for a second round of \u003cem\u003eab initio\u003c/em\u003e reconstruction and heterogeneous refinement. Particles from 1 of 3 of these classes (174,000 particles) were selected for Non-uniform refinement (\u003cem\u003eCustom parameters: Optimize per-particle defocus = true; Optimize per-group CTF params = true\u003c/em\u003e) to create an initial reconstruction with a resolution of ~3.7Å, that was used to calculate templates\u0026nbsp;\u003csup\u003e80\u003c/sup\u003e. These templates were used to choose particles from 9,858 of 10,740 total images (CTF \u0026lt; 8Å cutoff). These ~5.85 million particles were classified into a total of 6 classes by Heterogeneous refinement, that were seeded with 1 good volumes and 5 junk volumes taken from the above Heterogeneous refinement analyses. The ~1.31 million particles in the single selected class, were passed through a round of 2D classification (\u003cem\u003eCustom parameters: Batchsize per class = 200\u003c/em\u003e) (\u003cstrong\u003eExtended Data Fig. 1h\u003c/strong\u003e). ~1.29 million particles from 49 of the 50 2D classes were selected for a round of \u003cem\u003eab initio\u003c/em\u003e reconstruction followed by heterogeneous refinement into two classes. ~1.1 million particles from one of these two classes were subjected to 3D classification into 4 classes (\u003cem\u003eCustom parameters: Batchsize per class = 20,000; Initialization mode = PCA; Target resolution = 2Å; Particles per reconstruction = 500; Class similarity = 0.3\u003c/em\u003e) , followed by separate Non-uniform refinements of particles from each of these four classes (\u003cem\u003eCustom parameters: Optimize per-particle defocus = true; Optimize per-group CTF params = true\u003c/em\u003e)\u003csup\u003e80\u003c/sup\u003e. The final set of 366,794 particles were re-extracted and re-centered, Fourier-binned 2x2, and subjected to Non-uniform refinement to generate a final reconstruction refined to a global resolution of 3.48 Å based on the 0.143 FSC cutoff (\u003cstrong\u003eExtended Data Fig. 1h-k and Table 1\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e81\u003c/sup\u003e. The 3D FSC was calculated using webserver (3dfsc.salk.edu)\u003csup\u003e82\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel building and validation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe map was sharpened from two half-maps using the local anisotropic sharpening job in Phenix\u0026nbsp;\u003csup\u003e83\u003c/sup\u003e. The published structure of the \u003cem\u003eP. aeruginosa\u003c/em\u003e Cas1 homodimer was used as starting models\u003csup\u003e56\u003c/sup\u003e, because Colabfold consistently failed to predict the alternative fold that one Cas1 subunit adopts to form the assymetric homodimer interface\u0026nbsp;\u003csup\u003e84\u003c/sup\u003e, even when provided template structures. Whereas the Colabfold-predicted models for the \u003cem\u003eP. aeruginosa\u003c/em\u003e IHF heterodimer, and the \u003cem\u003eP. aeruginosa\u003c/em\u003e Cas2/3 subunit were used as starting models. The conformation of the DNA sequences within the \u003cem\u003eE. coli\u003c/em\u003e IHF-DNA co-crystal structure (PDB: 1IHF)\u0026nbsp;\u003csup\u003e38\u003c/sup\u003e was used as starting models for DNA segments within the IHF\u003csub\u003edistal\u003c/sub\u003e and IHF\u003csub\u003eproximal\u003c/sub\u003e DNA bends. For all other double stranded DNA segments, B-form DNA was used as a starting model. Single-stranded DNA segments were built in \u003cem\u003ede novo\u003c/em\u003e. Protein and DNA segments were individually rigid-body fitted into the EM density map. The relative orientation of the Cas2, Cas2/3 linker and Cas3 domains were corrected by real-space refinement into the EM density map in WinCoot\u0026nbsp;\u003csup\u003e85\u003c/sup\u003e. The ReadySet job in Phenix was used to generate hydrogens on all proteins and nucleic acids and prepare the model for further refinement. Then protein and DNA segments were real-space refined in WinCoot\u0026nbsp;\u003csup\u003e85\u003c/sup\u003e, restrained to ideal geometry, secondary structure and German McClure distance restraints generated in ProSMART from the input models\u0026nbsp;\u003csup\u003e86\u003c/sup\u003e. The models were iteratively real-space refined in WinCoot and in Phenix using Ramachandran and secondary structure restraints\u0026nbsp;\u003csup\u003e83,85\u003c/sup\u003e. The starting model was used as a reference model, and harmonic restraints on the starting coordinates were enabled. MolProbity\u0026nbsp;\u003csup\u003e87\u003c/sup\u003e and the PDB validation service server (https://validate-rcsb-1.wwpdb.org/) were used to identify problem regions subsequently corrected in WinCoot\u0026nbsp;\u003csup\u003e85\u003c/sup\u003e. For regions of the reconstruction where side chains are not visible (resolution \u0026gt;4.0Å) the atomic model was truncated to the peptide backbone. For regions of the reconstruction where the backbone was ambigous the sections of the peptide or DNA model were removed. Contacts and hydrogen bonds between residues were identified by ChimeraX v1.4 using the “contacts” and “hbonds” commands respectively, with default parameters\u0026nbsp;\u003csup\u003e88,89\u003c/sup\u003e. The DNAproDB webserver (\u003ca href=\"https://dnaprodb.usc.edu/\"\u003ehttps://dnaprodb.usc.edu/\u003c/a\u003e) was further used to analyze DNA-protein contacts (\u003cstrong\u003eExtended Data Fig. d,e\u003c/strong\u003e)\u0026nbsp;\u003csup\u003e90\u003c/sup\u003e. Structure-guided mutagenesis was used to further validate key Cas1-2/3-DNA contacts in the above biochemical assays.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCas1, Cas2/3 and repeat conservation analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo build a list of type I-F Cas1 sequences, CRISPRDetect v2.4 with default parameters was used to identify CRISPR arrays within a total\u0026nbsp;of 18,225 bacterial and 376 archaeal complete genomes accessed from the NCBI Assembly database on June 10\u003csup\u003eth\u003c/sup\u003e of 2019 as previously described\u0026nbsp;\u003csup\u003e22,91\u003c/sup\u003e. 15,274 high-confidence CRISPR arrays were classified with a CRISPR subtype by CRISPRDetect v2.4 (by matching to a list of repeats with known subtype annotations), and by genetic proximity to subtype-specific \u003cem\u003ecas\u0026nbsp;\u003c/em\u003egenes (within 20,000 bp). To identify \u003cem\u003ecas\u003c/em\u003e genes, the 20,000 bp flanking the CRISPR were submitted to PRODIGAL v2.6.3 (default parameters) to predict all potential \u003cu\u003eo\u003c/u\u003epen \u003cu\u003er\u003c/u\u003eeading \u003cu\u003ef\u003c/u\u003erames (ORFs)\u0026nbsp;\u003csup\u003e92\u003c/sup\u003e. This ORF database was then used as input to search for \u003cem\u003ecas\u003c/em\u003e gene clusters with MacsyFinder v1.0.5\u0026nbsp;\u003csup\u003e93\u003c/sup\u003e. The following parameters were used: “\u003cem\u003emacsyfinder --sequence-db \u0026lt;peptide_database\u0026gt; --db-type gembase -d \u0026lt;CRISPR_subtype_definitions\u0026gt; -p \u0026lt;HMM_profiles\u0026gt; -w 50 -vv all\u003c/em\u003e”. HMM profiles and classification definitions used in MacsyFinder were acquired from the local version of CRISPRCasFinder v4.2.20\u0026nbsp;\u003csup\u003e94\u003c/sup\u003e. Next, the first repeat and 200 nucleotides upstream of CRISPR arrays (leader) which were classified as Type I-F (1,683 arrays) were collected. A non-redundant list of I-F CRISPR leaders (536\u0026nbsp;leaders) was generated using CD-HIT v4.8.1 with a 95% identity cutoff\u0026nbsp;\u003csup\u003e95\u003c/sup\u003e. A local copy of FIMO was used to find significant matches to the position weight matrix representing I-F IHF binging site as previously described\u0026nbsp;\u003csup\u003e22,96\u003c/sup\u003e. I-F CRISPR arrays that possess more than one IHF site (IHF\u003csub\u003eproximal\u003c/sub\u003e and/or IHF\u003csub\u003edistal\u003c/sub\u003e) in the leader sequences were extracted for downstream analyses. Cas1 homologs were identified within the 20,000 base-pair flanking regions of extracted 444 I-F CRISPR arrays by using PRODIGAL and MacsyFinder with the same parameters described above. 371 Cas1 homologs associated with Type I-F CRISPRs and possessing at least one IHF site in the leader sequences were identified. A non-redundant list of Cas1 sequences was generated CD-HIT v4.8.1 with a 95% identity cutoff, resulting in 222 sequences\u0026nbsp;\u003csup\u003e95\u003c/sup\u003e. Sequences smaller than 200 residues and larger than 500 residues were removed, and the remaining 205 sequences were further curated with MaxAlign, which selected a list of 144 unique type I-F Cas1 sequences\u0026nbsp;\u003csup\u003e97\u003c/sup\u003e. The \u003cem\u003eP. aeruginosa\u0026nbsp;\u003c/em\u003ePA14Cas1 sequence was then added to a final list of 145 type I-F Cas1 sequences.\u0026nbsp;To build a list of type I-F Cas2/3 sequences, the \u003cem\u003eP. aeruginosa\u003c/em\u003e PA14 Cas2/3 sequence was used as an input for HHMER for a search for homologs using 3 iterations, an E-value cutoff of 0.0001, against the UNIREF-90 database\u0026nbsp;\u003csup\u003e98,99\u003c/sup\u003e. A list of 500 representative sequences was further curated with MaxAlign, to generate a final list of 458 unique Cas2/3 sequences. Type I-F Cas1 and Cas2/3 sequences were aligned using the MAFFT webserver with the E-INS-I iterative refinment methods to result in alignments with the highest number of gap-free sites\u0026nbsp;\u003csup\u003e100\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eTo build an updated list of CRISPR repeat sequences, CRISPRDetect v3.0 with default parameters was used to identify CRISPR arrays within a total of 25,502 bacterial and 398 archaeal complete genomes and chromosomes accessed from the NCBI\u0026nbsp;RefSeq Assembly database (accessed on June 10\u003csup\u003eth\u003c/sup\u003e, 2021)\u003csup\u003e91\u003c/sup\u003e. This search identified CRISPR loci within 58,864 genomic and plasmid sequences, resulting in 24,940 high-confidence CRISPR loci predictions (array quality score \u0026gt;3). Similar to above, CRISPRDetect annotated the subtype of 14,446 of these CRISPR loci, based on the sequence similarity of the repeats in these loci to known CRISPR repeats. The subtypes of the remaining 10,494 CRISPR loci were determined by their proximity to subtype-specific \u003cem\u003ecas\u003c/em\u003e genes as described above. 5,321 of the 10,494 unclassified CRISPR loci were assigned a subtype using this protocol, such that 5,173 CRISPR loci remained unclassified. The consensus repeat for each of the 24,940 CRISPR loci, as reported by CRISPRDetect, were used for downstream analyses. To ensure the repeats were arranged in the correct orientation, the 24,940 repeats were grouped by subtype, and each group was individually aligned by MAFFT using the “\u003cem\u003e--adjustdirection\u003c/em\u003e” parameter. Sequence logos of the first and last three bps of CRISPR repeats were made using Weblogo v3.7.1 for CRISPR subtypes and across all subtypes\u003csup\u003e101,102\u003c/sup\u003e (\u003cstrong\u003eExtended Data Fig. 4g\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eReporting summary\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFurther information on research design is available in the Nature Research Reporting Summary linked to this paper.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data that support the findings of this study are available from the corresponding author Blake Wiedenheft upon request. Cryo-EM maps were deposited in the Electron Microscopy Data Bank under accession number EMD-29280. The atomic model of the type I-F integration complex was deposited in the PDB under accession number 8FLJ. Plasmids generated in this study are available from Addgene.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCode will be made available upon request and without restriction.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThanks to members of the B.W. laboratory for feedback and discussions. We thank Dr. Mariusz Matyszewski and Dr. Jeliazko Jeliazkov for helpful discussions. Thanks to Coltran Hophan-Nichols for computational support. A.S-F. is a postdoctoral fellow of the Life Science Research Foundation that is supported by the Simons Foundation. A.S-F. is supported by the Postdoctoral Enrichment Program Award from the Burroughs Wellcome Fund. This work was supported by National Institutes of Health, United States grant 1K99GM147842 (A.S-F.). L.T., A.B.G. is supported by Montana State University’s Undergraduate Scholars Program, and by the NIH NIGMS IDeA program (P20GM103474). This work was performed using the cryo-EM Facility at Montana State University (NSF 1828765 and the M.J. Murdock Charitable Trust). Microscopy was also performed at the National Center for CryoEM Access and Training (NCCAT) and the Simons Electron Microscopy Center located at the New York Structural Biology Center, supported by the NIH Common Fund Transformative High Resolution Cryo-Electron Microscopy program (U24 GM129539), and by grants from the Simons Foundation (SF349247) and NY State Assembly. Research in the Wiedenheft lab is supported by the NIH (R35GM134867), the M.J. Murdock Charitable Trust, a young investigator award from Amgen, and the Montana State University Agricultural Experimental Station (USDA NIFA), and a sponsored research agreement from VIRIS Detection Systems. Molecular graphics and analyses performed with UCSF ChimeraX, developed by the Resource for Biocomputing, Visualization, and Informatics at the University of California, San Francisco, with support from National Institutes of Health R01-GM129325 and the Office of Cyber Infrastructure and Computational Biology, National Institute of Allergy and Infectious Diseases. Funders had no role in designing, performing, interpreting, or submitting the work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA.S.-F.: Conceptualization, Data Curation, Formal Analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualization, Writing – original draft. W.S.H., T.W., M.B.: Data curation, Investigation, Methodology, Visualization, Writing – review \u0026amp; editing. A.B.G., R.A.W.: Investigation, Methodology. L.T.: Visualization. C.C.G.:\u0026nbsp;Software, Resources, Writing – review \u0026amp; editing. K.N. and E.E.: Investigation, Resources. G.C.L.: Methodology, Supervision, Visualization, Writing – review \u0026amp; editing. B.W.: Funding acquisition, Project administration, Resources, Supervision, Visualization, Writing – review \u0026amp; editing.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eB.W. is the founder of SurGene and VIRIS Detection Systems. B.W. and A.S.-F. are inventors on patent applications related to CRISPR-Cas systems and applications thereof.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eKoonin, E. V. \u0026amp; Krupovic, M. Evolution of adaptive immunity from transposable elements combined with innate immune systems. \u003cem\u003eNat. Rev. Genet.\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 184\u0026ndash;192 (2015).\u003c/li\u003e\n\u003cli\u003eMcCLINTOCK, B. The origin and behavior of mutable loci in maize. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 344\u0026ndash;355 (1950).\u003c/li\u003e\n\u003cli\u003eNu\u0026ntilde;ez, J. K., Bai, L., Harrington, L. B., Hinder, T. L. \u0026amp; Doudna, J. A. CRISPR Immunological Memory Requires a Host Factor for Specificity. \u003cem\u003eMol. Cell\u003c/em\u003e \u003cstrong\u003e62\u003c/strong\u003e, 824\u0026ndash;833 (2016).\u003c/li\u003e\n\u003cli\u003eFagerlund, R. D. \u003cem\u003eet al.\u003c/em\u003e Spacer capture and integration by a type I-F Cas1\u0026ndash;Cas2-3 CRISPR adaptation complex. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e \u003cstrong\u003e114\u003c/strong\u003e, 201618421 (2017).\u003c/li\u003e\n\u003cli\u003eWright, A. V. \u003cem\u003eet al.\u003c/em\u003e Structures of the CRISPR genome integration complex. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e357\u003c/strong\u003e, 1113\u0026ndash;1118 (2017).\u003c/li\u003e\n\u003cli\u003eHickman, A. B. \u0026amp; Dyda, F. Mechanisms of DNA transposition. \u003cem\u003eMob. DNA III\u003c/em\u003e 529\u0026ndash;553 (2015) doi:10.1128/9781555819217.ch25.\u003c/li\u003e\n\u003cli\u003eLaxmikanthan, G. \u003cem\u003eet al.\u003c/em\u003e Structure of a holliday junction complex reveals mechanisms governing a highly regulated DNA transaction. \u003cem\u003eeLife\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 1\u0026ndash;23 (2016).\u003c/li\u003e\n\u003cli\u003eLee, H. \u0026amp; Sashital, D. G. Creating memories: molecular mechanisms of CRISPR adaptation. \u003cem\u003eTrends Biochem. Sci.\u003c/em\u003e 1\u0026ndash;13 (2022) doi:10.1016/j.tibs.2022.02.004.\u003c/li\u003e\n\u003cli\u003eWang, J. \u003cem\u003eet al.\u003c/em\u003e Structural and Mechanistic Basis of PAM-Dependent Spacer Acquisition in CRISPR-Cas Systems. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e163\u003c/strong\u003e, 840\u0026ndash;853 (2015).\u003c/li\u003e\n\u003cli\u003eNu\u0026ntilde;ez, J. K., Harrington, L. B., Kranzusch, P. J., Engelman, A. N. \u0026amp; Doudna, J. A. Foreign DNA capture during CRISPR-Cas adaptive immunity. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e527\u003c/strong\u003e, 535\u0026ndash;538 (2015).\u003c/li\u003e\n\u003cli\u003eXiao, Y., Ng, S., Nam, K. H. \u0026amp; Ke, A. How type II CRISPR\u0026ndash;Cas establish immunity through Cas1\u0026ndash;Cas2-mediated spacer integration. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e550\u003c/strong\u003e, 137\u0026ndash;141 (2017).\u003c/li\u003e\n\u003cli\u003eJackson, S. A. \u003cem\u003eet al.\u003c/em\u003e CRISPR-Cas: Adapting to change. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e356\u003c/strong\u003e, eaal5056 (2017).\u003c/li\u003e\n\u003cli\u003eMojica, F. J. M., D\u0026iacute;ez-Villase\u0026ntilde;or, C., Garc\u0026iacute;a-Mart\u0026iacute;nez, J. \u0026amp; Almendros, C. Short motif sequences determine the targets of the prokaryotic CRISPR defence system. \u003cem\u003eMicrobiology\u003c/em\u003e \u003cstrong\u003e155\u003c/strong\u003e, 733\u0026ndash;740 (2009).\u003c/li\u003e\n\u003cli\u003eKim, S. \u003cem\u003eet al.\u003c/em\u003e Selective loading and processing of prespacers for precise CRISPR adaptation. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e579\u003c/strong\u003e, 141\u0026ndash;145 (2020).\u003c/li\u003e\n\u003cli\u003eHu, C. \u003cem\u003eet al.\u003c/em\u003e Mechanism for Cas4-assisted directional spacer acquisition in CRISPR\u0026ndash;Cas. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e598\u003c/strong\u003e, 515\u0026ndash;520 (2021).\u003c/li\u003e\n\u003cli\u003eRamachandran, A., Summerville, L., Learn, B. A., DeBell, L. \u0026amp; Bailey, S. Processing and integration of functionally oriented prespacers in the Escherichia coli CRISPR system depends on bacterial host exonucleases. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e \u003cstrong\u003e295\u003c/strong\u003e, 3403\u0026ndash;3414 (2020).\u003c/li\u003e\n\u003cli\u003eLiao, C. \u003cem\u003eet al.\u003c/em\u003e Spacer prioritization in CRISPR\u0026ndash;Cas9 immunity is enabled by the leader RNA. \u003cem\u003eNat. Microbiol.\u003c/em\u003e (2022) doi:10.1038/s41564-022-01074-3.\u003c/li\u003e\n\u003cli\u003eMcGinn, J. \u0026amp; Marraffini, L. A. CRISPR-Cas Systems Optimize Their Immune Response by Specifying the Site of Spacer Integration. \u003cem\u003eMol. Cell\u003c/em\u003e \u003cstrong\u003e64\u003c/strong\u003e, 616\u0026ndash;623 (2016).\u003c/li\u003e\n\u003cli\u003eWang, R., Li, M., Gong, L., Hu, S. \u0026amp; Xiang, H. DNA motifs determining the accuracy of repeat duplication during CRISPR adaptation in Haloarcula hispanica. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, 4266\u0026ndash;4277 (2016).\u003c/li\u003e\n\u003cli\u003eGoren, M. G. \u003cem\u003eet al.\u003c/em\u003e Repeat Size Determination by Two Molecular Rulers in the Type I-E CRISPR Array. \u003cem\u003eCell Rep.\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 2811\u0026ndash;2818 (2016).\u003c/li\u003e\n\u003cli\u003eLinheiro, R. S. \u0026amp; Bergman, C. M. Testing the palindromic target site model for DNA transposon insertion using the Drosophila melanogaster P-element. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 6199\u0026ndash;6208 (2008).\u003c/li\u003e\n\u003cli\u003eSantiago-Frangos, A., Buyukyoruk, M., Wiegand, T., Krishna, P. \u0026amp; Wiedenheft, B. Distribution and phasing of sequence motifs that facilitate CRISPR adaptation. \u003cem\u003eCurr. Biol.\u003c/em\u003e 1\u0026ndash;10 (2021) doi:10.1016/j.cub.2021.05.068.\u003c/li\u003e\n\u003cli\u003eKieper, S. N., Almendros, C. \u0026amp; Brouns, S. J. J. Conserved motifs in the CRISPR leader sequence control spacer acquisition levels in Type I-D CRISPR-Cas systems. \u003cem\u003eFEMS Microbiol. Lett.\u003c/em\u003e \u003cstrong\u003e366\u003c/strong\u003e, 2016\u0026ndash;2020 (2019).\u003c/li\u003e\n\u003cli\u003eRollie, C., Graham, S., Rouillon, C. \u0026amp; White, M. F. Prespacer processing and specific integration in a Type I-A CRISPR system. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e46\u003c/strong\u003e, 1007\u0026ndash;1020 (2018).\u003c/li\u003e\n\u003cli\u003eYosef, I., Goren, M. G. \u0026amp; Qimron, U. Proteins and DNA elements essential for the CRISPR adaptation process in Escherichia coli. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 5569\u0026ndash;5576 (2012).\u003c/li\u003e\n\u003cli\u003eWei, Y., Chesne, M. T., Terns, R. M. \u0026amp; Terns, M. P. Sequences spanning the leader-repeat junction mediate CRISPR adaptation to phage in Streptococcus thermophilus. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e43\u003c/strong\u003e, 1749\u0026ndash;1758 (2015).\u003c/li\u003e\n\u003cli\u003eWright, A. V. \u0026amp; Doudna, J. A. Protecting genome integrity during CRISPR immune adaptation. \u003cem\u003eNat. Struct. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 876\u0026ndash;883 (2016).\u003c/li\u003e\n\u003cli\u003eWestra, E. R. \u003cem\u003eet al.\u003c/em\u003e Parasite Exposure Drives Selective Evolution of Constitutive versus Inducible Defense. \u003cem\u003eCurr. Biol.\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 1043\u0026ndash;1049 (2015).\u003c/li\u003e\n\u003cli\u003eMakarova, K. S. \u003cem\u003eet al.\u003c/em\u003e Evolutionary classification of CRISPR\u0026ndash;Cas systems: a burst of class 2 and derived variants. \u003cem\u003eNat. Rev. Microbiol.\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 67\u0026ndash;83 (2020).\u003c/li\u003e\n\u003cli\u003eRichter, C. \u003cem\u003eet al.\u003c/em\u003e Priming in the Type I-F CRISPR-Cas system triggers strand-independent spacer acquisition, bi-directionally from the primed protospacer. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e42\u003c/strong\u003e, 8516\u0026ndash;8526 (2014).\u003c/li\u003e\n\u003cli\u003eDatsenko, K. A. \u003cem\u003eet al.\u003c/em\u003e Molecular memory of prior infections activates the CRISPR/Cas adaptive bacterial immunity system. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 945 (2012).\u003c/li\u003e\n\u003cli\u003eXiao, Y. \u003cem\u003eet al.\u003c/em\u003e Structure basis for RNA-guided DNA degradation by Cascade and Cas3. \u003cstrong\u003e0839\u003c/strong\u003e, 1\u0026ndash;12 (2018).\u003c/li\u003e\n\u003cli\u003eNicholson, T. J. \u003cem\u003eet al.\u003c/em\u003e Bioinformatic evidence of widespread priming in type I and II CRISPR-Cas systems. \u003cem\u003eRNA Biol.\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 566\u0026ndash;576 (2019).\u003c/li\u003e\n\u003cli\u003eBrown, M. W. \u003cem\u003eet al.\u003c/em\u003e Assembly and translocation of a CRISPR-Cas primed acquisition complex. \u003cem\u003ebioRxiv\u003c/em\u003e \u003cstrong\u003e41\u003c/strong\u003e, 1\u0026ndash;11 (2017).\u003c/li\u003e\n\u003cli\u003eLi, M., Wang, R., Zhao, D. \u0026amp; Xiang, H. Adaptation of the Haloarcula hispanica CRISPR-Cas system to a purified virus strictly requires a priming process. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e42\u003c/strong\u003e, 2483\u0026ndash;2492 (2014).\u003c/li\u003e\n\u003cli\u003eSemenova, E. \u003cem\u003eet al.\u003c/em\u003e Highly efficient primed spacer acquisition from targets destroyed by the \u003cem\u003eEscherichia coli\u003c/em\u003e type I-E CRISPR-Cas interfering complex. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e \u003cstrong\u003e113\u003c/strong\u003e, 7626\u0026ndash;7631 (2016).\u003c/li\u003e\n\u003cli\u003eFineran, P. C. \u003cem\u003eet al.\u003c/em\u003e Degenerate target sites mediate rapid primed CRISPR adaptation. \u003cem\u003eProc. Natl. Acad. Sci. U. S. A.\u003c/em\u003e \u003cstrong\u003e111\u003c/strong\u003e, (2014).\u003c/li\u003e\n\u003cli\u003eRice, P. A., Yang, S., Mizuuchi, K. \u0026amp; Nash, H. A. Crystal Structure of an IHF-DNA Complex: A Protein-Induced DNA U-Turn. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e87\u003c/strong\u003e, 1295\u0026ndash;1306 (1996).\u003c/li\u003e\n\u003cli\u003eRohs, R. \u003cem\u003eet al.\u003c/em\u003e Origins of specificity in protein-DNA recognition. \u003cem\u003eAnnu. Rev. Biochem.\u003c/em\u003e \u003cstrong\u003e79\u003c/strong\u003e, 233\u0026ndash;269 (2010).\u003c/li\u003e\n\u003cli\u003eZayed, H. The DNA-bending protein HMGB1 is a cellular cofactor of Sleeping Beauty transposition. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 2313\u0026ndash;2322 (2003).\u003c/li\u003e\n\u003cli\u003eLittle, A. J., Corbett, E., Ortega, F. \u0026amp; Schatz, D. G. Cooperative recruitment of HMGB1 during V(D)J recombination through interactions with RAG1 and DNA. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e41\u003c/strong\u003e, 3289\u0026ndash;3301 (2013).\u003c/li\u003e\n\u003cli\u003eNash, H. A. \u0026amp; Robertson, C. A. Purification and properties of the Escherichia coli protein factor required for lambda integrative recombination. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e \u003cstrong\u003e256\u003c/strong\u003e, 9246\u0026ndash;9253 (1981).\u003c/li\u003e\n\u003cli\u003eLavoie, B. D. \u0026amp; Chaconas, G. Site-specific HU binding in the Mu transpososome: conversion of a sequence-independent DNA-binding protein into a chemical nuclease. \u003cem\u003eGenes Dev.\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, 2510\u0026ndash;2519 (1993).\u003c/li\u003e\n\u003cli\u003eChalmers, R., Guhathakurta, A., Benjamin, H. \u0026amp; Kleckner, N. IHF Modulation of Tn10 Transposition: Sensory Transduction of Supercoiling Status via a Proposed Protein/DNA Molecular Spring. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e93\u003c/strong\u003e, 897\u0026ndash;908 (1998).\u003c/li\u003e\n\u003cli\u003eHaniford, D. B. Transpososome Dynamics and Regulation in Tn10 Transposition. \u003cem\u003eCrit. Rev. Biochem. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e41\u003c/strong\u003e, 407\u0026ndash;424 (2006).\u003c/li\u003e\n\u003cli\u003eWhitfield, C. R., Wardle, S. J. \u0026amp; Haniford, D. B. The global bacterial regulator H-NS promotes transpososome formation and transposition in the Tn5 system. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 309\u0026ndash;321 (2009).\u003c/li\u003e\n\u003cli\u003eLiu, D., Haniford, D. B. \u0026amp; Chalmers, R. M. H-NS mediates the dissociation of a refractory protein-DNA complex during Tn10/IS10 transposition. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 6660\u0026ndash;6668 (2011).\u003c/li\u003e\n\u003cli\u003evan Gent, D. C., Hiom, K., Paull, T. T. \u0026amp; Gellert, M. Stimulation of V(D)J cleavage by high mobility group proteins. \u003cem\u003eEMBO J.\u003c/em\u003e \u003cstrong\u003e16\u003c/strong\u003e, 2665\u0026ndash;2670 (1997).\u003c/li\u003e\n\u003cli\u003eRowland, S.-J., Stark, W. M. \u0026amp; Boocock, M. R. Sin recombinase from Staphylococcus aureus: synaptic complex architecture and transposon targeting: Sin recombinase. \u003cem\u003eMol. Microbiol.\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, 607\u0026ndash;619 (2002).\u003c/li\u003e\n\u003cli\u003eAlonso, J. C., Weise, F. \u0026amp; Rojo, F. The Bacillus subtilis Histone-like Protein Hbsu Is Required for DNA Resolution and DNA Inversion Mediated by the \u0026beta; Recombinase of Plasmid pSM19035. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e \u003cstrong\u003e270\u003c/strong\u003e, 2938\u0026ndash;2945 (1995).\u003c/li\u003e\n\u003cli\u003ePetit, M.-A., Ehrlich, D. \u0026amp; Janni\u0026egrave;re, L. pAM\u0026beta;1 resolvase has an atypical recombination site and requires a histone-like protein HU. \u003cem\u003eMol. Microbiol.\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 271\u0026ndash;282 (1995).\u003c/li\u003e\n\u003cli\u003eRojo, F. \u0026amp; Alonso, J. C. The \u0026beta; recombinase of plasmid pSM19035 binds to two adjacent sites, making different contacts at each of them. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 3181\u0026ndash;3188 (1995).\u003c/li\u003e\n\u003cli\u003eWalker, M. W. G., Klompe, S. E., Zhang, D. J. \u0026amp; Sternberg, S. H. \u003cem\u003eTransposon mutagenesis libraries reveal novel molecular requirements during CRISPR RNA-guided DNA integration\u003c/em\u003e. http://biorxiv.org/lookup/doi/10.1101/2023.01.19.524723 (2023) doi:10.1101/2023.01.19.524723.\u003c/li\u003e\n\u003cli\u003eRollins, M. F. \u003cem\u003eet al.\u003c/em\u003e Cas1 and the Csy complex are opposing regulators of Cas2/3 nuclease activity. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e \u003cstrong\u003e114\u003c/strong\u003e, 201616395 (2017).\u003c/li\u003e\n\u003cli\u003eWang, X. \u003cem\u003eet al.\u003c/em\u003e Structural basis of Cas3 inhibition by the bacteriophage protein AcrF3. \u003cem\u003eNat. Struct. Mol. Biol.\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 868\u0026ndash;870 (2016).\u003c/li\u003e\n\u003cli\u003eWiedenheft, B. \u003cem\u003eet al.\u003c/em\u003e Structural Basis for DNase Activity of a Conserved Protein Implicated in CRISPR-Mediated Genome Defense. \u003cem\u003eStructure\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 904\u0026ndash;912 (2009).\u003c/li\u003e\n\u003cli\u003eKunin, V., Sorek, R. \u0026amp; Hugenholtz, P. Evolutionary conservation of sequence and secondary structures in CRISPR repeats. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, R61 (2007).\u003c/li\u003e\n\u003cli\u003eNethery, M. A. \u003cem\u003eet al.\u003c/em\u003e CRISPRclassify: Repeat-Based Classification of CRISPR Loci. \u003cem\u003eCRISPR J.\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 558\u0026ndash;574 (2021).\u003c/li\u003e\n\u003cli\u003eDhingra, Y., Suresh, S. K., Juneja, P. \u0026amp; Sashital, D. G. PAM binding ensures orientational integration during Cas4-Cas1-Cas2-mediated CRISPR adaptation. \u003cem\u003eMol. Cell\u003c/em\u003e \u003cstrong\u003e82\u003c/strong\u003e, 4353-4367.e6 (2022).\u003c/li\u003e\n\u003cli\u003eAli Azam, T., Iwata, A., Nishimura, A., Ueda, S. \u0026amp; Ishihama, A. Growth Phase-Dependent Variation in Protein Composition of the \u003cem\u003eEscherichia coli\u003c/em\u003e Nucleoid. \u003cem\u003eJ. Bacteriol.\u003c/em\u003e \u003cstrong\u003e181\u003c/strong\u003e, 6361\u0026ndash;6370 (1999).\u003c/li\u003e\n\u003cli\u003eMonta\u0026ntilde;o, S. P., Pigli, Y. Z. \u0026amp; Rice, P. A. The Mu transpososome structure sheds light on DDE recombinase evolution. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e491\u003c/strong\u003e, 413\u0026ndash;417 (2012).\u003c/li\u003e\n\u003cli\u003eMaertens, G. N., Hare, S. \u0026amp; Cherepanov, P. The mechanism of retroviral integration from X-ray structures of its key intermediates. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e468\u003c/strong\u003e, 326\u0026ndash;329 (2010).\u003c/li\u003e\n\u003cli\u003eRollie, C., Schneider, S., Brinkmann, A. S., Bolt, E. L. \u0026amp; White, M. F. Intrinsic sequence specificity of the Cas1 integrase directs new spacer acquisition. \u003cem\u003eeLife\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 1\u0026ndash;19 (2015).\u003c/li\u003e\n\u003cli\u003eMakarova, K. S., Wolf, Y. I. \u0026amp; Koonin, E. V. Classification and Nomenclature of CRISPR-Cas Systems: Where from Here? \u003cem\u003eCRISPR J.\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 325\u0026ndash;336 (2018).\u003c/li\u003e\n\u003cli\u003eDeveau, H. \u003cem\u003eet al.\u003c/em\u003e Phage response to CRISPR-encoded resistance in Streptococcus thermophilus. \u003cem\u003eJ. Bacteriol.\u003c/em\u003e \u003cstrong\u003e190\u003c/strong\u003e, 1390\u0026ndash;1400 (2008).\u003c/li\u003e\n\u003cli\u003eK\u0026uuml;nne, T. \u003cem\u003eet al.\u003c/em\u003e Cas3-Derived Target DNA Degradation Fragments Fuel Primed CRISPR Adaptation. \u003cem\u003eMol. Cell\u003c/em\u003e \u003cstrong\u003e63\u003c/strong\u003e, 852\u0026ndash;864 (2016).\u003c/li\u003e\n\u003cli\u003eMusharova, O. \u003cem\u003eet al.\u003c/em\u003e Prespacers formed during primed adaptation associate with the Cas1\u0026ndash;Cas2 adaptation complex and the Cas3 interference nuclease\u0026ndash;helicase. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e \u003cstrong\u003e118\u003c/strong\u003e, e2021291118 (2021).\u003c/li\u003e\n\u003cli\u003eWiegand, T. \u003cem\u003eet al.\u003c/em\u003e Reproducible Antigen Recognition by the Type I-F CRISPR-Cas System. \u003cem\u003eCRISPR J.\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, 378\u0026ndash;387 (2020).\u003c/li\u003e\n\u003cli\u003eVorontsova, D. \u003cem\u003eet al.\u003c/em\u003e Foreign DNA acquisition by the I-F CRISPR\u0026ndash;Cas system requires all components of the interference machinery. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e43\u003c/strong\u003e, 10848\u0026ndash;10860 (2015).\u003c/li\u003e\n\u003cli\u003eKoonin, E. V. \u0026amp; Makarova, K. S. Evolutionary plasticity and functional versatility of CRISPR systems. \u003cem\u003ePLOS Biol.\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, e3001481 (2022).\u003c/li\u003e\n\u003cli\u003eCavazzana-Calvo, M. \u003cem\u003eet al.\u003c/em\u003e Gene Therapy of Human Severe Combined Immunodeficiency (SCID)-X1 Disease. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e288\u003c/strong\u003e, 669\u0026ndash;672 (2000).\u003c/li\u003e\n\u003cli\u003eStrecker, J. \u003cem\u003eet al.\u003c/em\u003e RNA-guided DNA insertion with CRISPR-associated transposases. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e365\u003c/strong\u003e, 48\u0026ndash;53 (2019).\u003c/li\u003e\n\u003cli\u003eKlompe, S. E., Vo, P. L. H., Halpin-Healy, T. S. \u0026amp; Sternberg, S. H. Transposon-encoded CRISPR\u0026ndash;Cas systems direct RNA-guided DNA integration. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e571\u003c/strong\u003e, 219\u0026ndash;225 (2019).\u003c/li\u003e\n\u003cli\u003eShipman, S. L., Nivala, J., Macklis, J. D. \u0026amp; Church, G. M. Molecular recordings by directed CRISPR spacer acquisition. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e353\u003c/strong\u003e, aaf1175 (2016).\u003c/li\u003e\n\u003cli\u003eSchmidt, F., Cherepkova, M. Y. \u0026amp; Platt, R. J. Transcriptional recording by CRISPR spacer acquisition from RNA. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e562\u003c/strong\u003e, 380\u0026ndash;385 (2018).\u003c/li\u003e\n\u003cli\u003eHerzik, M. A., Wu, M. \u0026amp; Lander, G. C. High-resolution structure determination of sub-100 kDa complexes using conventional cryo-EM. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 1\u0026ndash;9 (2019).\u003c/li\u003e\n\u003cli\u003eMastronarde, D. N. Automated electron microscope tomography using robust prediction of specimen movements. \u003cem\u003eJ. Struct. Biol.\u003c/em\u003e \u003cstrong\u003e152\u003c/strong\u003e, 36\u0026ndash;51 (2005).\u003c/li\u003e\n\u003cli\u003ePunjani, A., Rubinstein, J. L., Fleet, D. J. \u0026amp; Brubaker, M. A. CryoSPARC: Algorithms for rapid unsupervised cryo-EM structure determination. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 290\u0026ndash;296 (2017).\u003c/li\u003e\n\u003cli\u003eSuloway, C. \u003cem\u003eet al.\u003c/em\u003e Automated molecular microscopy: The new Leginon system. \u003cem\u003eJ. Struct. Biol.\u003c/em\u003e \u003cstrong\u003e151\u003c/strong\u003e, 41\u0026ndash;60 (2005).\u003c/li\u003e\n\u003cli\u003ePunjani, A., Zhang, H. \u0026amp; Fleet, D. J. Non-uniform refinement: adaptive regularization improves single-particle cryo-EM reconstruction. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 1214\u0026ndash;1221 (2020).\u003c/li\u003e\n\u003cli\u003eScheres, S. H. W. \u0026amp; Chen, S. Prevention of overfitting in cryo-EM structure determination. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 853\u0026ndash;854 (2012).\u003c/li\u003e\n\u003cli\u003eTan, Y. Z. \u003cem\u003eet al.\u003c/em\u003e Addressing preferred specimen orientation in single-particle cryo-EM through tilting. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 793\u0026ndash;796 (2017).\u003c/li\u003e\n\u003cli\u003eLiebschner, D. \u003cem\u003eet al.\u003c/em\u003e Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in \u003cem\u003ePhenix\u003c/em\u003e. \u003cem\u003eActa Crystallogr. Sect. Struct. Biol.\u003c/em\u003e \u003cstrong\u003e75\u003c/strong\u003e, 861\u0026ndash;877 (2019).\u003c/li\u003e\n\u003cli\u003eMirdita, M. \u003cem\u003eet al.\u003c/em\u003e ColabFold: making protein folding accessible to all. \u003cem\u003eNat. Methods\u003c/em\u003e \u003cstrong\u003e19\u003c/strong\u003e, 679\u0026ndash;682 (2022).\u003c/li\u003e\n\u003cli\u003eEmsley, P., Lohkamp, B., Scott, W. G. \u0026amp; Cowtan, K. Features and development of Coot. \u003cem\u003eActa Crystallogr. Sect. D\u003c/em\u003e \u003cstrong\u003e66\u003c/strong\u003e, 486\u0026ndash;501 (2010).\u003c/li\u003e\n\u003cli\u003eNicholls, R. A. Conformation-independent comparison of protein structures. (2011).\u003c/li\u003e\n\u003cli\u003eWilliams, C. J. \u003cem\u003eet al.\u003c/em\u003e MolProbity: More and better reference data for improved all-atom structure validation: PROTEIN SCIENCE.ORG. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 293\u0026ndash;315 (2018).\u003c/li\u003e\n\u003cli\u003eGoddard, T. D. \u003cem\u003eet al.\u003c/em\u003e UCSF ChimeraX: Meeting modern challenges in visualization and analysis. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 14\u0026ndash;25 (2018).\u003c/li\u003e\n\u003cli\u003ePettersen, E. F. \u003cem\u003eet al.\u003c/em\u003e UCSF ChimeraX: Structure visualization for researchers, educators, and developers. \u003cem\u003eProtein Sci.\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, 70\u0026ndash;82 (2021).\u003c/li\u003e\n\u003cli\u003eSagendorf, J. M., Markarian, N., Berman, H. M. \u0026amp; Rohs, R. DNAproDB: an expanded database and web-based tool for structural analysis of DNA\u0026ndash;protein complexes. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e gkz889 (2019) doi:10.1093/nar/gkz889.\u003c/li\u003e\n\u003cli\u003eBiswas, A., Staals, R. H. J., Morales, S. E., Fineran, P. C. \u0026amp; Brown, C. M. CRISPRDetect: A flexible algorithm to define CRISPR arrays. \u003cem\u003eBMC Genomics\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 356 (2016).\u003c/li\u003e\n\u003cli\u003eHyatt, D. \u003cem\u003eet al.\u003c/em\u003e Prodigal: prokaryotic gene recognition and translation initiation site identification. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 119 (2010).\u003c/li\u003e\n\u003cli\u003eAbby, S. S., N\u0026eacute;ron, B., M\u0026eacute;nager, H., Touchon, M. \u0026amp; Rocha, E. P. C. MacSyFinder: A program to mine genomes for molecular systems with an application to CRISPR-Cas systems. \u003cem\u003ePLoS ONE\u003c/em\u003e (2014) doi:10.1371/journal.pone.0110726.\u003c/li\u003e\n\u003cli\u003eCouvin, D. \u003cem\u003eet al.\u003c/em\u003e CRISPRCasFinder, an update of CRISRFinder, includes a portable version, enhanced performance and integrates search for Cas proteins. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e46\u003c/strong\u003e, W246\u0026ndash;W251 (2018).\u003c/li\u003e\n\u003cli\u003eLi, W. \u0026amp; Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 1658\u0026ndash;1659 (2006).\u003c/li\u003e\n\u003cli\u003eGrant, C. E., Bailey, T. L. \u0026amp; Noble, W. S. FIMO: scanning for occurrences of a given motif. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e27\u003c/strong\u003e, 1017\u0026ndash;1018 (2011).\u003c/li\u003e\n\u003cli\u003eGouveia-Oliveira, R., Sackett, P. W. \u0026amp; Pedersen, A. G. MaxAlign: maximizing usable data in an alignment. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, 312 (2007).\u003c/li\u003e\n\u003cli\u003eFinn, R. D., Clements, J. \u0026amp; Eddy, S. R. HMMER web server: interactive sequence similarity searching. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, W29\u0026ndash;W37 (2011).\u003c/li\u003e\n\u003cli\u003eSuzek, B. E. \u003cem\u003eet al.\u003c/em\u003e UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 926\u0026ndash;932 (2015).\u003c/li\u003e\n\u003cli\u003eKatoh, K. \u0026amp; Standley, D. M. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. \u003cem\u003eMol. Biol. Evol.\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, 772\u0026ndash;780 (2013).\u003c/li\u003e\n\u003cli\u003eSchneider, T. D. \u0026amp; Stephens, R. M. Sequence logos: A new way to display consensus sequences. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 6097\u0026ndash;6100 (1990).\u003c/li\u003e\n\u003cli\u003eCrooks, G. E., Hon, G., Chandonia, J.-M. \u0026amp; Brenner, S. E. WebLogo: a sequence logo generator. \u003cem\u003eGenome Res.\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 1188\u0026ndash;90 (2004).\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Table","content":"\u003cp\u003e\u003cstrong\u003eTable 1. Cryo-EM data collection, refinement, and validation statistics\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003eI-F Integration Complex\u003c/p\u003e\n \u003cp\u003e(EMD-\u003cstrong\u003e29280\u003c/strong\u003e)\u003c/p\u003e\n \u003cp\u003e(PDB \u003cstrong\u003e8FLJ\u003c/strong\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"100%\" colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eData collection and processing\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eMicroscope \u0026nbsp;\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003eKrios\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eVoltage (keV)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e300\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eDetector\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003eGatan K3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eMagnification (nominal/calibrated) \u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e81,000x/46,860x\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eExposure navigation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003eBeam image-shift\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eData acquisition software\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003eLeginon\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eElectron exposure (e\u003csup\u003e-\u003c/sup\u003e/\u0026Aring;\u003csup\u003e2\u003c/sup\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e69.09\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eExposure rate (e\u003csup\u003e-\u003c/sup\u003e/\u0026Aring;\u003csup\u003e2\u003c/sup\u003e/s)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e27.64\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eFrame length (ms)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e50\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eNumber of frames per micrograph\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e50\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eEnergy filter width (eV)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003ePixel size (\u0026Aring;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e1.067\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eDefocus range (\u0026mu;m)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.7 \u0026ndash; 2.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eStage Tilt (\u0026deg;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eMicrographs collected (no.)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e10,740\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"100%\" colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eReconstruction\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eImage-processing package\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003ecryoSPARC\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eTotal extracted particles (no.)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e5,846,923\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eRefined particles (no.)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e1,103,036\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eFinal particles (no.)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e366,794\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eSymmetry imposed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eC\u003c/em\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eGlobal resolution (\u0026Aring;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;FSC 0.143 (unmasked/masked)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e3.88/3.47\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;FSC 0.5 (unmasked/masked)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e4.22/3.77\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Resolution range (local) (\u0026Aring;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e2.3 \u0026ndash; 9.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e3DFSC sphericity\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.985 out of 1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eModel sharpening \u003cem\u003eB\u003c/em\u003e factor (\u0026Aring;\u003csup\u003e2\u003c/sup\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e-55\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"100%\" colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eRefinement\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eRefinement package\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003ePhenix\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eModel composition\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Non-hydrogen atoms\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e31,794\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Protein residues\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e3,549\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;DNA nucleotides\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e350\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Ligands\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eCC (volume/mask)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.76/0.76\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eR.m.s deviations\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Bond lengths (\u0026Aring;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.005\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; \u0026nbsp;Bond angles (\u0026deg;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.835\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"100%\" colspan=\"2\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eValidation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eRamachandran plot\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; Outliers (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; Allowed (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e4.28\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp; Favored (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e95.72\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eMolProbity score\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e1.54\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003ePoor rotamers (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.26\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eClashscore (all atoms)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e4.67\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eC-beta deviations (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eCaBLAM Outliers (%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e1.70\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"57.729468599033815%\" valign=\"top\"\u003e\n \u003cp\u003eEMRinger score\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"42.270531400966185%\" valign=\"top\"\u003e\n \u003cp\u003e2.55\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"CRISPR-Cas, integration, CRISPR adaptation, Cas1-2/3, IHF, DNA recording devices, transposition, transposase, DNA bending","lastPublishedDoi":"10.21203/rs.3.rs-2982802/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2982802/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBacteria and archaea acquire resistance to viruses and plasmids by integrating fragments of foreign DNA into the first repeat of a CRISPR array. However, the mechanism of site-specific integration remains poorly understood. Here, we determine a 560 kDa integration complex structure that explains how Cas (Cas1-2/3) and non-Cas proteins (IHF) fold 150 base-pairs of host DNA into a U-shaped bend and a loop that protrude from Cas1-2/3 at right angles. The U-shaped bend traps foreign DNA on one face of the Cas1-2/3 integrase, while the loop places the first CRISPR repeat in the Cas1 active site. Both Cas3s rotate 100-degrees to expose DNA binding sites on either side of the Cas2 homodimer, that each bind an inverted repeat motif in the leader. Leader sequence motifs direct Cas1-2/3-mediated integration to diverse repeat sequences that have a 5\u0026rsquo;-GT.\u003c/p\u003e","manuscriptTitle":"Protein-mediated folding of the genome is essential for site-specific integration of foreign DNA into CRISPR loci","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-06-14 20:15:59","doi":"10.21203/rs.3.rs-2982802/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"nature-structural-and-molecular-biology","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"nsmb","sideBox":"Learn more about [Nature Structural \u0026 Molecular Biology](http://www.nature.com/nsmb/)","snPcode":"","submissionUrl":"","title":"Nature Structural \u0026 Molecular Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"53fd9828-9a0e-4048-8070-bb9ba7ae54b2","owner":[],"postedDate":"June 14th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":22322979,"name":"Biological sciences/Structural biology/Electron microscopy/Cryoelectron microscopy"},{"id":22322980,"name":"Biological sciences/Biochemistry/Proteins"}],"tags":[],"updatedAt":"2023-09-29T15:34:11+00:00","versionOfRecord":{"articleIdentity":"rs-2982802","link":"https://doi.org/10.1038/s41594-023-01097-2","journal":{"identity":"nature-structural-and-molecular-biology","isVorOnly":false,"title":"Nature Structural \u0026 Molecular Biology"},"publishedOn":"2023-09-14 04:00:00","publishedOnDateReadable":"September 14th, 2023"},"versionCreatedAt":"2023-06-14 20:15:59","video":"","vorDoi":"10.1038/s41594-023-01097-2","vorDoiUrl":"https://doi.org/10.1038/s41594-023-01097-2","workflowStages":[]},"version":"v1","identity":"rs-2982802","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2982802","identity":"rs-2982802","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.