Detailed mechanisms for unintended large DNA deletions with CRISPR, base editors, and prime editors | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Detailed mechanisms for unintended large DNA deletions with CRISPR, base editors, and prime editors Sangsu Bae, Gue-Ho Hwang, Seok-Hoon Lee, Minsik Oh, Segi Kim, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3835370/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 04 Nov, 2024 Read the published version in Nature Biomedical Engineering → Version 1 posted You are reading this latest preprint version Abstract CRISPR-Cas9 nucleases are versatile tools for genetic engineering cells and function by producing targeted double-strand breaks (DSBs) in the DNA sequence. However, the unintended production of large deletions (> 100 bp) represents a challenge to the effective application of this genome-editing system. We optimized a long-range amplicon sequencing system and developed a k-mer sequence-alignment algorithm to simultaneously detect small DNA alteration events and large DNA deletions. With this workflow, we determined that CRISPR-Cas9 induced large deletions at varying frequencies in cancer cell lines, stem cells, and primary T cells. With CRISPR interference screening, we determined that end resection and the subsequent TMEJ [DNA polymerase theta-mediated end joining] repair process produce most large deletions. Furthermore, base editors and prime editors also generated large deletions despite employing mutated Cas9 “nickases” that produce single-strand breaks. Our findings reveal an important limitation of current genome-editing tools and identify strategies for mitigating unwanted large deletion events. Biological sciences/Biological techniques/Genetic techniques/Gene targeting/CRISPR-Cas9 genome editing Biological sciences/Biological techniques/Genetic engineering Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction The CRISPR-Cas system, derived from a prokaryotic immune response, uses single-guide RNAs (sgRNAs) to recognize target DNA sequences and then introduces DNA double-strand breaks (DSBs) at the target sites 1 – 3 . In mammalian cells, DSBs are repaired through multiple DNA repair pathways: (i) non-homologous end joining (NHEJ), (ii) DNA polymerase theta (Pol θ)-mediated end joining (TMEJ) that operates in a microhomology-dependent manner with as few as 2 to 16 nucleotides of homologous sequence 4 , (iii) single-strand annealing (SSA) that uses homologous repeat sequences to bridge DSB ends, (iv) homologous recombination (HR), and (v) homology-directed repair (HDR) 5 , 6 . Because NHEJ, TMEJ, and SSA are error-prone, they frequently lead to insertion or deletion or both (collectively referred to as indel) mutations during the repair process 7 – 9 . In contrast, the HR and HDR pathways are error-free, thus genome editing that uses single-stranded oligodeoxynucleotides (ssODN) or double-stranded donor plasmids flanking homology arms has higher fidelity by leveraging these repair pathways 10 , 11 . Although most deletions generated by CRISPR nucleases were shorter than 20 base pairs (bp), it was observed that more than 20% of the mutations in mouse embryos induced by CRISPR-Cas9 were unintended deletions longer than 250 bp 12 . Moreover, DSBs induced by CRISPR nucleases can induce chromosomal rearrangements, including chromosomal depletion and translocation 13 – 15 . Because chromosomal translocation involves off-target cleavage sites and the on-target cleavage site, high-fidelity Cas9 nucleases reduce translocation rates 16 . Unfortunately, high-fidelity Cas9 nuclease do not overcome the issue of unintended large deletions that occur at on-target sites. Another variation of the CRISPR-Cas9 system involves using base editors (BEs) and prime editors (PEs), which employ partially inactivated Cas9 (nCas9) that generates single-strand nicks 17 – 19 . Whether these nCas9-based genome editing systems also generate large deletion is not clearly established 20 – 24 . Identifying the repair pathways that produce large deletions following genome-editing tools, with Cas9 or nCas9, will not only provide key insight into the biology of this process but enable the development of strategies to mitigate large deletion events. Determining the frequency of deletion events > 100 bp is challenging using short-range high-throughput sequencing methods. Instead, such events are detected with either a long-range amplicon sequencing method using short-read sequencers, such as Illumina platforms, or a long-read sequencing method, using Pacific Biosciences (PacBio) or Oxford Nanopore Technologies (ONT) platforms 25 . In this study, we modified a long-range amplicon sequencing method and developed a novel k-mer sequence alignment algorithm, and we used this strategy to accurately determine CRISPR-induced large deletions. We observed that large deletions were mostly accompanied by small indels in human cells, including cancer-like cells (HeLa, HEK293T, U2OS, K562), fibroblasts, primary T cells, and human embryo stem cells (H9). We then identified large deletion-associated repair pathways through CRISPR interference (CRISPRi) screening experiments with ONT Nanopore sequencing, establishing that TMEJ is the dominant pathway generating large deletions. We found that BEs generate large deletions through the base excision repair (BER) pathway, especially in the presence of an apurinic/apyrimidinic site (AP site) and that PEs particularly generate large deletions in the presence of additional nicking guide RNAs (ngRNAs). Our findings reveal mechanistic challenges in using genome-editing tools and provide insight for avoiding these issues in research and therapeutic applications. Results Optimizing a long-range amplicon sequencing method to determine small indels and large deletions with a high accuracy Neither the PacBio Single-molecule real-time (SMRT) sequencing nor ONT Nanopore sequencing platforms for long-read sequencing have sufficient sequencing accuracy (PacBio ~ 10% and ONT 2%-15%) 26 to distinguish 1-bp deletion, 1-bp insertion, and substitution events. Furthermore, these methods have a length bias: shorter DNA fragments are read more easily and thereby are overestimated compared to longer DNA fragments ( Supplementary Fig. 1 ) 27 . Thus, currently available long-read sequencing platforms lack sufficient accuracy to identify both large deletions and small indels. Our first task was to establish a precise sequencing method to simultaneously determine small indels and large deletions with a high accuracy. The Illumina platform has the highest sequencing accuracy (> 99.2%) 28 . Therefore, we optimized long-range amplicon sequencing using Illumina sequencers combined with long-range polymerase chain reaction (PCR) and DNA fragmentation (Fig. 1 a ) . The four steps in our sequencing were as follows: (i) extract genomic DNA (gDNA) from CRISPR-Cas9-treated and non-treated cell lines, (ii) amplify ~ 10–15 kb of DNA involving the CRISPR target sites using PCR optimized for both gDNAs, (iii) fragment the amplified products into ~ 300 bp and prepare next generation sequencing (NGS) libraries through end-repair, dA-tailing, adaptor ligation, and PCR enrichment, (iv) acquire the sequencing data with an Illumina Miniseq. To quantify large deletions, small deletions, and small insertions in the bulk cell populations, we used a novel k-mer alignment algorithm ( Supplementary Fig. 2 ). In this protocol, PCR-based DNA amplification of the long DNA region (~ 10–15 kb ) is a critical step. To mitigate the length bias of DNA polymerases, we tested five DNA polymerases: Phusion HF, Q5 that is used in DNA library prep kit for Illumina, Accuprime pfx 29 , KOD -Multi & Epi- DNA polymerase that is a modified KOD DNA polymerase to reduce amplification bias during PCR by addition of the elongation accelerator 30 , 31 , and SUN-PCR blend 32 – 34 . To evaluate the PCR length bias, we prepared nine synthesized DNA fragments in increments of 20 bp from 120 bp to 280 bp with a common forward/reverse primer pair, mixed the fragments in equimolar amounts, and conducted PCR experiments for the mixture with each polymerase ( Supplementary Fig. 3 ). KOD -Multi & Epi- DNA polymerase had the least length bias of the 5 polymerases tested (Fig. 1 b). Therefore, we used this polymerase for the long-range PCR amplification step. Another complicated step in the workflow is sequence alignment of short-read NGS data. Pre-existing tools, such as a burrows-wheeler aligner (BWA)-mem, were developed to rapidly map DNA sequence reads to whole genome sequences. Our workflow requires querying sequences that are limited to the PCR-amplified region to detect indels near the CRISPR-mediated cleavage sites. This is a more limited application that does not require a heuristic approach. Thus, we developed a dedicated k-mer alignment algorithm that is compact and is potentially available as an online tool for analysis of CRISPR-treated samples ( Supplementary Fig. 4 ). To confirm the functionality of our optimized long-range amplicon sequencing workflow, we prepared and applied the workflow to two cell lines: One has two wild-type ST3GAL4 alleles and the other has a 10-bp or a 1,075-bp deletion in each allele (Fig. 1 c). After mixing gDNAs from the two cell lines at ratios of 0%, 20%, 50%, 80%, and 100%, the mixtures were prepared and subjected to the optimized long-range amplicon sequencing. The optimized long-range amplicon sequencing was highly accurate in detecting both the 10-bp and 1,075-bp deletion events (Figs. 1 d and 1 e), showing that our method is an effective tool for measuring large deletions and small indels simultaneously. Determination of CRISPR-mediated small indels and large deletions in various human cell lines Based on the optimized long-range amplicon sequencing, we quantified large deletions and small indels induced by CRISPR-Cas9 nuclease in various human cell lines. We evaluated four cancer or transformed cell lines (HeLa, HEK293T, U2OS, K562), human fibroblasts, and a human embryo stem cell line (H9) (Figs. 2 a- 2 f). For HeLa and HEK293T cells, we analyzed the frequency of CRISPR-Cas9-mediated deletion or insertion events for one target site each in HPRT1 and ST3GAL4 and five target sites in the WISP3 gene. We defined a small insertion and a small deletion event as shorter than 100 bp, and a large deletion event as larger than 100 bp. Across these seven target sites, the average frequency of small indel events were 37.5 ± 13.2% in HeLa cells and 26.4 ± 11.7% in HEK293T cells, whereas the average frequency of large deletion event was 6.4 ± 4.4% in HeLa cells and 4.42 ± 3.6% in HEK293T cells (Figs. 2 a and 2 b). Small indels were the most common CRISPR-Cas9-mediated events in all of the cells with an average frequency of 17.8 ± 7.8% in U2OS cells, 18.4 ± 10.1% in K562 cells, 10.6 ± 8.5% in fibroblast cells, and 49.3 ± 21.7% in H9 cells. CRISPR-Cas9-mediated large deletion events were less common with an average frequency of 3.3 ± 1.8% in U2OS cells, 2.4 ± 1.5% in K562 cells, 1.5 ± 1.7% in fibroblast cells, and 2.6 ± 1.6% in H9 cells (Figs. 2 c- 2 f). The frequency of CRISPR-Cas9-mediated large deletions across all target sites and all six cell lines ranged from 0.2–17.5%. Because CRISPR-Cas9-based genome editing has been widely adopted in engineering T cells for cancer immunotherapy (clinicaltrials.gov: NCT03399448, NCT03081715, NCT03398967, NCT02793856, NCT03044743), we evaluated CRISPR-Cas9-mediated mutation patterns in human primary T cells. We chose sgRNAs targeting PD1 and TRAC1 loci, because these genes are frequently targeted in T cells to enhance T-cell activity 35 . Human T cells were isolated from the peripheral blood mononuclear cells (PBMCs) of two healthy donors (ASAN-04 and ASAN-68) using a magnetic-activated cell sorting (MACS) system and then stimulated by anti-CD3/CD28 beads before introduction of Cas9-sgRNA ribonucleoprotein (RNP) complexes by electroporation (Fig. 2 g). The optimized long-range amplicon sequencing revealed that the average frequencies of small indels were 70.99 ± 11.37% in ASAN04 and 72.03 ± 15.46% in ASAN68, whereas the average frequencies of large deletion events were 15.46 ± 6.68% in ASAN04 and 14.24 ± 6.82% in ASAN68 (Figs. 2 h and 2 i), similar to the previous results 36 , 37 . The average frequency of large deletion events in human primary T cells was higher than the average frequency of these events in the cell lines tested. Using all of the data, we observed that the large deletion frequencies positively correlated with the small deletion frequencies (Pearson corr. = 0.72) and were negatively correlated with small insertion frequencies (Pearson corr. = -0.15), resulting in a low correlation with small indels (Pearson corr. = 0.47) ( Supplementary Fig. 5 ). Determination by CRISPRi screening that end resection and TMEJ processes play positive roles for generating large deletions To examine the underlying mechanism of CRISPR-driven large DNA deletion events, we performed a CRISPR interference (CRISPRi) screen targeting 794 genes that were those referenced in Repair-seq 38 plus genes identified as associated with DNA repair in the Human Protein Atlas website (Fig. 3 a). We introduced a plasmid library with three different sgRNAs per gene and the puromycin resistance gene into an engineered HeLa cell line that stably expresses deactivated Cas9 (dCas9) fused with Krüppel associated box (KRAB) domain to produce gene knockdown 39 . DSBs in the puromycin resistance gene were generated by addition of the Cas9 RNP complex, gDNA was prepared, and ~ 5.6 kbp of DNA involving the sgRNA sites or the DSB (puromycin resistance gene) site were amplified and sequenced with the Nanopore MinION. Because we only needed to assess qualitatively large deletion events, we used Nanopore-seq to read long range region. However, to account for the length-dependent bias in Nanopore-seq ( Supplementary Fig. 1 ), each sequencing read was adjusted according to its total length and then classified as wild-type, insertion, deletion, and large deletions (those > 100 bp). We compared Nanopore-seq data from cells with or without inhibition of the 794 target genes ( Supplementary Fig. 6 and Supplementary Table 1 ). DNA repair pathways can be divided into end protection and end resection 40 (Fig. 3 b). After DSBs occur, the end protection pathway proceeds to NHEJ, which induces blunt-end ligation by Ku-XRCC4-DNA ligase IV complex 40 . In contrast, the end resection pathway proceeds to TMEJ, SSA, or HR, according to the 5’ end resection sequence and the existence of homology sequences in neighboring DNA 40 . Using a previously defined classification scheme, we categorized the genes into those associated with NHEJ ( LIG4, XRCC4, XRCC5 , and XRCC6 ), end resection ( MRE11A, NBN , and RAD50 ), TMEJ ( POLQ ), SSA ( ERCC4 and RAD52 ), and HR ( BRCA1, BRCA2, RAD51AP1, RAD51B, RAD51C , and RAD51D ). Z-scores for large deletions in the puromycin resistance gene region for each of the CRISPRi-targeted genes within each category indicated that fewer large deletions occurred when POLQ of TMEJ or genes associated with end resection were knocked down and large deletions were more common when SSA genes were knocked down (Fig. 3 c). Consistent with these findings, we calculated Z-scores of the difference in large deletion frequencies in cells with and without inhibition of target genes and found a significant reduction in large deletions in cells with either reduced activity of TMEJ (P-value = 9.63x10 − 3 ) or the end resection pathway (P-value = 7.71x10 − 5 ) and a significant increase in cells with reduced activity of SSA (P-value = 2.19x10 − 3 ) (Fig. 3 d). Collectively, the screening results indicated that large deletions are primarily generated by TMEJ following the end resection pathway. Confirmation by knockout that end resection and TMEJ processes generate large deletions To confirm the function of key genes in generating large deletions, we constructed individual knock-out (KO) cell lines using CRISPR-Cas9 or disrupted specific repair pathways pharmacologically. We used the following strategies to impair specific repair pathways in HeLa cells: M4344 to pharmacologically inhibit ATR activity and suppress end resection, Ligase IV ( LIG4 ) KO to suppress NHEJ, POLQ KO to suppress TMEJ, and RAD52 KO to suppress SSA (Fig. 3 b). We analyzed the mutation pattern in each cell line or condition for three CRISPR-Cas9 targeted genes ( HPRT1, ST3GAL4, WISP3 ). Consistent with the CRISPRi screen, we found that POLQ KO (P-value = 1.46x10 − 6 ) or the addition of M4344 (P-value = 8.83x10 − 7 ) significantly reduced large deletion rates, whereas RAD52 KO (P-value = 5.64x10 − 4 ) increased the large deletion rates. In contrast to variable effect of LIG4 knockdown in the CRISPRi screen, the LIG4 KO cell line showed a significant increase in large deletion frequency (P-value = 3.06x10 − 4 ), suggesting that blocking either the end protection process or the subsequent NHEJ can enhance the occurrence of large deletions. Both the CRISPRi and KO experiments indicated that the end resection process and the following TMEJ play major roles in generating large deletions after CRISPR-mediated gene editing. A strategy to reduce large deletions in primary T cells As a potential application in T cell engineering, we examined whether M4434 reduced large deletion frequencies in human primary T cells. After introducing CRISPR-Cas9 by electroporation into human T cells, we exposed the cells to various concentrations of M4434 (0 nM, 1 nM, 5 nM, 10 nM, 25 nM, and 50 nM). Cell viability and number were substantially decreased when the M4344 concentration was over 25 nM, and the frequency of large deletions was reduced to similar amounts at M4344 concentrations from 10 nM to 50 nM ( Supplementary Fig. 7 ). Therefore, we selected 10 nM as the optimal concentration M4434 to assess the effect on large deletion events in human T cells. We analyzed the effect of 10 nM M4344 on the CRISPR-Cas9-induced mutation pattern on TRAC and PD1 in human T cells acquired from two healthy donors (ASAN99 and ASAN107). For this analysis we used the optimized long-range amplicon sequencing method. M4344 decreased the frequency of large deletion events between 35% and 80%, and the reductions were significant: TRAC P-value = 5.41x10 -6 and PD1 P-value = 8.52x10 -3 in ASAN99, TRAC P-value = 1.14x10 -5 and PD1 P-value = 8.03x10 -4 in ASAN99 (Figs. 3 f, 3 g and Supplementary Figs. 7d, 7e ). TMEJ-mediated DNA repair is mediated by microhomology, meaning that repair can occur with as few as 2–16 nucleotides of homologous sequence 4 . We found that the microhomology-mediated deletion frequencies were also decreased ( Supplementary Fig. 8 ), indicating that M4344 prevents the large deletion events generated by TMEJ. Collectively, these data support evaluation of M4344 or drugs that reduce end resection or TMEJ pathways as co-effectors with CRISPR-Cas9-mediated genetic engineering in therapeutic applications to mitigate unintended large deletions. Generation of large DNA deletions by cytosine base editors and adenine base editors In contrast with Cas9 nucleases, BEs employ nCas9 containing the D10A mutation, which is impaired in the ability to generate DNA DSBs. Instead of introducing DSBs at target sites, BEs modify single nucleotides in the DNA sequence. BEs, including cytosine base editor (CBE) and adenine base editor (ABE), introduce infrequent indels 17 , 18 , 41 . Thus, it is possible for BEs to generate large deletions. It was reported that CBE variants without a uracil glycosylase inhibitor (UGI) have relatively high indel frequencies 41 , 42 , thus we hypothesized that BE platforms without UGI induce large deletions with a high frequency. The role of the UGI is to inhibit cellular uracil DNA glycosylase (UNG), the enzyme that excises uracil 43 , 44 . After uracil excision, cellular DNA AP lyase introduces a nick on the non-target strand of the sgRNA, leading to BER 45 . Therefore, we predicted that UGI-lacking BEs generate DSBs and consequently large deletions during the BER process through a nick on each strand— one by nCas9 (D10A) and one by AP lyase. ABE variants also induce cytosine editing in preferred motifs, suggesting the possibility of ABE introducing large deletions during bystander cytosine editing 42 , 46 . We measured large deletion frequencies with six different BE platforms: a canonical CBE containing UGI 47 , a CBE variant without UGI [named CBE (∆UGI)], a CBE variant without UGI and with additional UNG (named CGBE1) 41 , a canonical ABE 17 , an ABE variant with UGI (named ABE-UGI), and an ABE variant containing engineered hypoxanthine excision protein N-methylpurine DNA glycosylase (MPG) (named AYBE) 48 (Fig. 4 a). We introduced each BE platform, or inactivated Cas9 (dCas9) as a control, into HEK293T cells and targeted four endogenous targets ( ABLIM3, FANCF, KLHL29, BRD4 ); then we conducted optimized long-range amplicon sequencing to evaluate large deletion frequencies. Each BE platform generated large deletion mutations; however, the frequency varied both for each platform and each targeted gene (Fig. 4 b). Compared to the canonical CBE, which had a maximum frequency of large deletion events of 0.8%, CBE (∆UGI) and CGBE1 generated higher frequencies of large deletion events with maximum frequencies of 4.8% and 3.5%, respectively. The canonical ABE had a large deletion frequency maximum of 1.42%. In Fig. 4 c, each large deletion frequency was divided by the average large deletion frequency with. Overall, introduction of a UGI slightly reduced the frequency and introduction of the eMPG increased the frequency substantially with a range of 4.41–1.21% and 2.21–7.19%, respectively (Fig. 4 c). For ABE platforms, we observed that target sequences including a TC motif within the editing window, as in ABLIM3 and FANCF , had larger frequencies of both small indels and large deletions compared to target sequences without a TC motif, as in KLHL29 and BRD4 (Figs. 4 b and 4 d), indicating that bystander C-to-U conversion by ABE and the resulting AP site produces DSBs and subsequent large deletion events. Both small indels and large deletion events occurred at lower frequencies at targets with TC motifs with ABE-UGI, consistent with the UGI inhibiting the production of DSBs produced through a mechanism similar to that observed for the CBEs (Fig. 4 d). Compared to targets with a TC motif, those without a TC motif had a lower frequency of either large deletion events or small indels due to AYBE (Figs. 4 b and 4 d). However, the frequency of large deletions due to AYBE at the target sites without TC motif was substantially higher than that of ABE. Thus, inosine excision by eMPG of AYBE likely enables DSB generation and large deletions through a mechanism like that with the uracil excision by UNG in which Cas9 (D10A) nickase nicks one strand and the BER process nicks the other strand (Fig. 4 e). Generation of large deletion events by prime editors In contrast to BEs, PEs employ nCas9 containing H840A mutation and there are two representative PE platforms, PE2 and PE3 19 . PE2 introduces a single nick on the non-target strand of the pegRNA, whereas PE3 introduces a double nick on each strand of the DNA through the addition of a nicking guide RNA (ngRNA) 19 . Empirically, we and other groups observed that PEs are often accompanied with unwanted small indels 19 , 49 . In addition, compared with nCas9 (D10A), the cleavage activity of nCas9 (H840A) is not completely impaired and introduces DSBs with a low efficiency 50 . The potential for DSBs that result in large deletions is particularly high for PE3 21, 51 . We measured large deletion frequencies induced by PE2 or PE3 at substitution targets in HEK3 and HEK4 , deletion targets in FANCF and HEK3 , and insertion targets in RNF2 and HEK3 . As expected, PE2 had a lower frequency of large deletion (maximum 1.2%) than PE3 (maximum 24.3%) (Figs. 5 a- 5 c). Similar to the BE platforms, the frequency of PE-induced large deletion events varied by target (Fig. 5 a). Particularly with PE3, even in the same target, large deletion frequencies varied according to the targeted edit: Compare HEK3 1-10del and HEK3 1CTTins, each with + 90 ngRNA and HEK4 + 2GtoT with the ngRNA + 52 or + 74 (Fig. 5 a). These results suggested that the sequence in proximity to the prime editing site might contribute to the potential for large deletion events. We investigated whether the mismatch repair (MMR) pathway or increasing the lifetime of the pegRNA affects the formation of large deletions. To determine the effect of MMR, we introduced PE3 with pegRNA and a MMR inhibitory factor, hMLH1dn 52 . Inhibition of the MMR pathway did not significantly affect the frequency of large deletion events compared with those induced by PE3 with pegRNA (Fig. 5 d), suggesting that the MMR pathway has a low contribution for the occurrence of PE3-mediated large deletions. Using an established method to increase pegRNA stability, we introduced PE3 with the engineered pegRNA (epegRNA) and found that this resulted in an increase in large deletion frequencies 53 (Fig. 5 d). We also compared the effect of epegRNA combined with two enhanced PE systems, PE4max and PE5max 52 , on large deletion frequencies induced during insertion of 1 TAC into RNF2 or of 1 CTT into HEK3 , which were sites with relatively low frequencies of large deletion in the PE2 and PE3 systems with pegRNA (Fig. 5 a). Compared with the occurrence of large deletion events in cells with dCas9, PE4max with epegRNA did not significantly affect the frequency of large deletion events but PE5max with epegRNA induced a significantly greater frequency of large deletion events (P-value = 0.0071) ( Supplementary Fig. 9 ). Discussion Here, we refined a long-range amplicon sequencing method to enable simultaneous analysis of both large deletions and small indels with high accuracy. To develop this long-range amplicon sequencing method, we determined the PCR polymerase with minimal length bias, optimized the PCR protocol, and developed a dedicated k-mer sequence-alignment algorithm along with an optimized analysis program to eliminate false-positive large deletion reads ( Supplementary Fig. 10 ). With this method, we detected and quantified large deletions and small indels in all human cells that we tested. Many groups analyze CRISPR-mediated gene editing outcomes within a short range (< 500 bp) of the target site, which can miss large deletion events. We experienced mischaracterization of a CRISPR nuclease-mediated knockin cell line ( Supplementary Fig. 11 ). Based on PCR amplification and short-range deep sequencing data, we observed only one pattern for a EXD2 Halo tag knockin cell line, suggesting that the cell line had homologous knockin at both alleles However, with long-range amplicon sequencing, we found that one allele had a large deletion (740 bp). Thus, this is a heterologous knock-in line. A similar phenomenon was observed with wheat genetically engineered for pesticide resistance in which a large deletion altered the epigenetic landscape and changed the growth properties of the plant 54 . Hence, it is necessary to test for the presence of large deletions after gene editing with CRISPR nucleases. Reducing unintended chromosomal changes in therapeutic applications, such as in generation of CAR T cells, is critical. Tsuchida et al. demonstrated that large deletions and chromosomal truncations can be reduced by altering the step of T cell activation in experimental protocols 55 . We determined that TMEJ serves as a primary repair pathway for generating large deletions. Understanding the mechanism responsible for large deletion events enables development of mitigation strategies. Indeed, we showed that inhibiting the end resection pathway and the subsequent TMEJ pharmacologically with M4344 reduced CRISPR-Cas9-induced large deletion frequencies in human primary T cells. BEs and PEs that use Cas9 nickase have the potential to generate DSBs at target sites 21 , 50 , 51 , but some studies report that BEs and PEs do not generate large deletions 20 , 22 . However, our data showed that both BEs and PEs generate large deletions, especially through BER pathway-dependent base editing that creates an AP site for BEs. PE systems that use ngRNA or the stabilized epegRNA showed particularly high frequencies of large deletion. Gene editing tools based on Cas9 or nCas9 exhibited frequencies of large deletion events that varied by target site and for the nCas9-based systems by type of genetic modification. The first-in-kind CRISPR gene editing drug, named Exa-cel or CASGEVY, was approved in UK and USA in 2023. This is remarkably fast for a therapeutic based on CRISPR-Cas nucleases and begins a new era for gene editing therapy. Our findings showed that the present genome-editing tools need additional validation to ensure large deletion events are not present or these tools need additional engineering to prevent the generation of DSBs and large deletions. Whether there are additional shortcomings yet to be discovered for the clinical application of these genome editing tools remains an open question. The workflow that we developed provides a mechanism to evaluate unintended large or small changes to DNA arising from the application of these gene editing tools. Declarations ACKNOWLEDGEMENTS Most analysis of sequencing data was carried out using the computing server at the Genomic Medicine Institute Research Service Center. This research was supported by grants from the National Research Foundation of Korea (NRF) no. 2021R1A2C3012908, no. 2021M3A9H3015389 to S.B. The authors thank Nancy R. Gough (BioSerendipity, LLC) for editorial service. AUTHOR CONTRIBUTIONS S.B. and G.H. conceived this project; G.H., S.-H.L., and H.S.K., performed and analyzed screening experiments; M.O. and G.H. developed bioinformatics algorithms; G.H., S.-H.L., and O.H. performed cell experiments; S.K. and H.-K.J. performed T-cell experiments; C.H.K., S.K., and S.B. supervised this project; G.H., S.-H.L., and S.B. wrote the manuscript with the help of all other authors. Competing interests The authors declare no competing interests. Data availability High-throughput sequencing data have been deposited in the NCBI Sequence Read Archive database (SRA; https://www.ncbi.nlm.nih.gov/sra) under accession number PRJNA1055687. The code of k-mer alignment program is available at https://github.com/ailab-mju/CRISPR-LargeDel. Materials and methods Generation of sgRNA-encoding plasmids Target sequences were designed using Cas-Designer 56 to avoid off-target sequence (up to 2 mistmaches). The list of oligomers for target sequence is in Supplementary Table 1 . pRG2 GG expression vector was digested using Bsa1 restriction enzyme. sgRNA oligos, with overhangs complementary to the digested vector, were ordered from Macrogen (Korea) and Cosmogenetech (Korea). These oligos—comprising both the upper and lower strands—were then annealed to produce double-stranded oligo deoxynucleotides (dsODN). The annealed dsODNs were ligated into the digested expression vector using T4 DNA ligase (Enzynomics) and incubated for 1h at room temperature. The ligation mixture was transformed into DH5a competent cells using the heat-shock method and cultured overnight at 37°C. Individual colonies were selected and grown in LB media for 16 hours at 37°C in a shaking incubator. The plasmids were isolated using Exprep™ Plasmid SV kit (GeneAll). Cell culture and transfection for cancer cell lines and fibroblasts HeLa (ATCC®, CCL-2™), HEK293T (ATCC®, CCL-3216™) cells, and U2OS (ATCC HTB-96) cells were maintained in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS), 100 unit/mL penicillin, and 100 unit/mL streptomycin. K562 (ATCC CCL-243) cells were maintained in Roswell Park Memorial Institute (RPMI) 1640 Medium supplemented with 10% fetal bovine serum (FBS), 100 unit/mL penicillin, and 100 unit/mL streptomycin. Normal fibroblasts [ThermoFisher, Human Dermal Fibroblasts (C0045C)] were maintained in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 20% fetal bovine serum (FBS), 100 unit/mL penicillin, and 100 unit/mL streptomycin. HeLa and HEK293T were transfected with Lipofectamin 2000 (Invitrogen). Before transfection, 1 × 10 5 cells from each well were seeded in 24-well plates. SpCas9 expression plasmids (750 ng) and sgRNA expression plasmids (250 ng) were mixed with 100 µl of Opti-MEM medium and 2 µl of Lipofectamin 2000 and incubated for 20 min at room temperature. The prepared mixture was added to the seeded wells. After 24 hours, the culture media were replaced with fresh media. U2OS and K562 were transfected with Neon transfection system (Invitrogen). cells (2.5 × 10 5 ) were transfected with 750 ng of Cas9 expression plasmids (750 ng) and sgRNA expression plasmids (250 ng) with the following parameters: 1050 V, 30 ms, 2 pulse for U2OS and 1350 V, 10 ms, 4 pulse for K562. Normal fibroblasts were transfected with Amaxa P3 primary cell 4D-nucleofector kit using program DS-137. All cells were analyzed 3 days after transfection. Cell culture and transfection for H9 cells H9 human embryonic stem cells were maintained in Essential 8 (E8) medium (Gibco A1517001) on iMatrix-511 (Matrixome, 892 021). The dissociation of H9 cells into clusters for subculturing was facilitated using ReLeSR (Stemcell Tech., 05873). Subsequently, the cells were transferred and replated in E8 medium supplemented with p160-Rho-associated coiled-coil kinase (ROCK) inhibitor Y-27632. Before electroporation, TrypLE (Gibco, 12604013) was used to generate a suspension of single cells. Cells (1 × 10 5 ) were electroporated with 250 ng of sgRNA-encoding plasmid and 750 ng of Cas9 expression plasmid using a NEON system (ThermoFisher) at 1050 V for 30 ms (two pulses). Cells were then seeded in 48-well plates in E8 supplemented with Y-27632 (10 µM) for 24 hours. After three days of culturing, gDNA was isolated. Generation of Cas9 ribonucleoprtein (RNP) complexes for CRISPRi screening The puromycin targeting sgRNA was synthesized by in vitro transcription using T7 RNA polymerase (NEB) and template oligos, and the sgRNA product was purified using RNeasy Mini Kit (Qiagen). Streptococcus pyogenes Cas9 (SpCas9) was ordered from Enzynomics. To generate Cas9 RNP complex, SpCas9 and sgRNA were mixed in a ratio of 1:3 and incubated at room temperature for 30 min. These Cas9 RNP complexes were added to the CRISPRi-stable HeLa cell line after lentiviral transduction and puromycin selection for CRISPRi screening. Isolation, culture and editing of human primary T cells Whole blood samples from healthy donors were taken under a protocol approved by the committee of Asan Medical Center. Peripheral blood mononuclear cells (PBMCs) were isolated from the whole blood samples using SepMate PBMC isolation tubes (STEMCEL). The PBMCs were further processed to isolate human primary T cells using MACS based PAN-T isolation kits (Miltenyi Biotec). RPMI-1640 (Gibco) supplemented with fetal bovine serum (10%, gibco), GlutaMAX (2 mM, gibco), sodium pyruvate (1 mM, gibco), non-essential amino acids (0.1 mM, gibco), beta-mercaptoethanol (55 µM, gibco), HEPES (10 mM, sigma) and penicillin-streptomycin (1%, gibco) was used to culture human primary T cells with IL-2 (300 IU/ml, BMI KOREA). For gene editing, the human primary T cells were stimulated with Dynabeads human T-Activator CD3/CD28 (Thermo Fischer Scientific) at a cell-to-bead ratio of 1:1 for 48 h. After separating T cells from beads using a magnet, the stimulated T cells were electroporated with Neon transfection system (Thermo Fisher). Briefly, 5 µg of recombinant Cas9 (Enzynomics) and 5 µg of in vitro -transcribed sgRNA was incubated at 37°C for 10 min to form Cas9 RNP complex, immediately before electroporation. Assembled CRISPR RNPs were added to 0.5 million of activated human T cells resuspended in T buffer and electroporated with a Neon electroporation device (1400 V, 10 ms, 3 pulse). Electroporated cells were transferred into culture vessels containing culture medium without antibiotics. One day after electroporation, culture medium was changed into fresh medium containing antibiotics and cells were maintained at a concentration of approximately 1 million cells per ml of medium. Long-range amplicon sequencing The PCR primers were designed using Primer3Plus 57 ( https://www.bioinformatics.nl/cgi-bin/primer3plus/primer3plus.cgi ) or Primer-BLAST 58 ( https://www.ncbi.nlm.nih.gov/tools/primer-blast ). The targeted region (~ 8 to 15 kb) in gDNA was amplified using KOD multi & epi DNA polymerase (TOYOBO) according to the manufacturer’s protocol. For sequences that were difficult to amplify, primers were replaced or gDNA was re-extracted. Amplified products (1 µg) were purified using AMPure XP bead-based reagent (Beckman Coulter) with 0.95X. The purified DNA samples were fragmented to ~ 300 bp with M220 Focused-ultrasonicator (Covaris) according to the manufacturer’s protocol. The fragmented samples were purified with Expin™ PCR SV kit (GeneAll) and prepared as an NGS library with NEBNext® Ultra™ II DNA Library Prep Kit for Illumina® (NEB). The prepared samples were sequenced with MiniSeq High Output Reagent Kit (300-cycles) using MiniSeq to obtain about ~ 400,000 to 500,000 reads. The NGS data FASTQ files from CRISPR-treated and non-treated cells were analyzed using the k-mer alignment program. Developing k-mer alignment program We developed a k -mer based algorithm for detecting CRISPR-induced DNA alterations without requiring a supercomputer and that can run on a personal computer. The k -mer alignment program is efficient to run with limited computational resources. For the cases investigated here, we used a personal computer with (16 GB) memory and (3.4 GHz / 8 cores) CPU. Our software program accepts as input a 10-kbp reference sequence that includes the cleavage site and the paired-end sequencing data in FASTQ format from both CRISPR-treated and non-treated samples. The output quantifies CRISPR-treated large deletions and small indels and provides read alignment results. The alignment program consists of 3 steps: (i) the short read alignment step, (i) short read classification step, and (iii) removal of false-positive large deletions ( Supplementary Fig. 4 ). The first step is the short read alignment task based on a k -mer hash table that is constructed using a reference genome. This hash table is used to obtain positions on the reference genome of ( l-k + 1 ) k -mers for a given read of length l . To determine alignment of given input, our program identifies the longest region of consecutive overlapping k -mers on the reference sequence through the Longest Increasing Subsequence (LIS) algorithm. The second step is the short read classification task. Based on the read alignment results, our program categorizes the read pairs into three distinct classes, (i) all read pairs are skewed to the cleavage site, (ii) all read pairs are mapped without splitting and the cleavage site passes between reads, and (iii) one of the reads is split and passes through the cleavage site. Read pairs in category (i) are considered wildtype. Read pairs in categories (ii) and (iii) are considered as potential CRISPR-derived variant candidates and are compiled into a candidate list. The third step involves the elimination of false-positive large deletions. It is possible for variant candidate read pairs to be present in non-treated samples, primarily due to inherent biases associated with PCR amplification. This phenomenon is not specific to CRISPR-treated samples, but it is a consequence of the characteristics of the reference sequence and affects both CRISPR-treated and non-treated datasets. To reduce the influence of false positives, our program performs a mapping of candidate read pairs to the left and right positions within the reference sequence and identifies clusters of these candidates by implementing k -means clustering. The appropriate value for k (the number of clusters) is determined using the Bayesian information criterion, which enables the automated selection of the optimal k selected based on the characteristics of the dataset. After cluster selection, those clusters that are common to both CRISPR-treated and non-treated datasets are categorized as a result from PCR bias rather than CRISPR-induced variation. These co-occurring clusters are then excluded from the candidate list. Finally, the remaining read pairs in the candidate list are identified as CRISPR-derived variants. Read pairs with a deletion length of 50 base pairs or more within the curated list of candidate variants are determined as a large deletion. Read pairs with insertions or deletions within a range of ± 13 base pairs around the cleavage site are categorized as small indels. To quantify the identified reads, our program enumerates the occurrence of normal mappings, small indels, and large deletions within read pairs passing through the cleavage site (± 13 base pairs). Our program then calculates the relative frequencies at which large deletions, small insertions, and small deletions occur. Let \({N}_{T}\) be the total number of reads passing the cut site (± 13 base pairs), including large deletions. \({N}_{L}\) is the number of large deletions. \({N}_{S}, {N}_{D}\) are the number of reads for small insertions and small deletions. The relative frequency of large deletions calculated by \({N}_{L}\) / \({N}_{T}\) . Small indels are calculated by \({N}_{S}+ {N}_{D}\) / \({N}_{T}\) . Deletion analysis was performed with the developed k -mer alignment program using a local computer version with an added web front-end. CRISPRi library construction Our custom CRISPRi library was constructed using pAX198 (Addgene #173042) that includes pU6-sgRNA-EF1a-Puro-T2A-BFP. Using the Repair-seq dataset and genes associated with DNA repair in the Human Protein Atlas website, we selected 794 genes associated with DNA repair and categorized them into specific repair processes or pathways. For each gene, three CRISPRi gRNA sequences were sourced from the hCRISPRi-v2.1 library developed by the Weissman group ( Supplementary Table 2 ) and 60 non-targeting gRNAs from Repair-seq were added in the CRISPRi library. The oligonucleotide library was procured from GenScript (USA) and subsequently amplified employing the Phusion® High-Fidelity DNA Polymerase (NEB). The amplified product and pAX198 plasmids were digested using FastDigest Bpu1102I and FastDigest BstX1 (ThermoFisher). Desired DNA fragments from the digested pAX198 were isolated through a 1% agarose gel and subsequently purified using the Expin™ Gel SV kit (GeneAll). The digested oligo library was separated and the required DNA sequence was extracted from a 10% PAGE gel. This gel-extracted DNA was then purified using isopropanol precipitation. The digested oligo library and plasmid backbone were ligated using T4 DNA ligase. The ligation mixture was then purified with AMPure XP beads. The ligated product was transformed into MegaX DH10B T1R Electrocomp™ Cells (ThermoFisher) using MicroPulser Electroporator (BioRad). After confirming more than 90,000 colonies, the plasmid library was obtained using NucleoBond Xtra Midi EF kit (Macherey-Nagel). The oligo library was confirmed using nested PCR and Illumina sequencing. Lentivirus preparation Lentivirus was produced in the Lenti-X 23T cell line (Takara Bio). Transfection was carried out with psPAX2 (Addgene #12260), pMD2.G Addgene #12259), and library plasmids using polyethyleneimine (Sigma-Aldrich). The cultured medium was replaced one day post-transfection. Two days later, lentivirus-containing medium was harvested and filtered through a 0.45-µm syringe filter. The lentivirus was concentrated using the Lenti-X concentrator (Takara Bio). The viral titer was determined by performing lentiviral transductions at varying concentrations in a 48-well plate format. After titration, the lentiviral library was aliquoted and stored at -80°C. CRISPRi screening with Nanopore sequencing HeLa CRISPRi cells were generated by lentiviral integration (~ 3 to 5 MOI) using the dCas9-KRAB-blast plasmid (Addgene #89567), followed by single cell isolation. Prior to lentiviral transduction, HeLa CRISPRi cells were cultured at a density of 5 × 10 5 cells in a 100-mm dish. The following day, the sgRNA library lentivirus was added to the cultured HeLa CRISPRi cells in the presence of 8 µg/mL polybrene. After 24 hours, the culture medium was replaced with fresh medium. Another 24 hours later, cell selection was initiated with 2 µg/mL puromycin and continued for 2 to 3 days. Post-selection, the culture medium was replaced with fresh medium and the cells were cultured for 6 to 8 days to allow gene repression to occur. Subsequently, Cas9 RNP complex targeting the transduced puromycin resistance gene was transfected into the sgRNA library transduced CRISPRi stable cell line using Neon transfection system (Fig. 3 a ). Three days post-transfection, gDNA was extracted using the NucleoSpin Blood XL, Maxi kit (Macherey-Nagel). Cell culture was performed whenever the cells had grownto 90% of the cell plate. Half of the extracted gDNA was amplified to generate fragments of ~ 5 to 6 kb using the KOD multi & epi DNA polymerase. These amplified fragments were then purified with AMPure XP beads. The purified samples were sequenced on MinION (Oxford Nanopore) using ligation sequencing kit V14 (Oxford Nanopore) and MinION flow cell R10.4.1 (Oxford Nanopore) according to the manufacturer’s protocol. The sequencing process ran at a speed of 260 bps, and base calls were made on the resulting data using guppy (Oxford Nanopore) with super high accuracy mode. Analysis for CRISPRi screening with Nanopore sequencing Fastq files were aligned to the reference genome using the guppy aligner with default settings. To identify gRNA sequences within the sequencing data, we utilized BWA-mem. A gRNA reference FASTA file was constructed by appending 10 bp from the reference sequence to both ends of the gRNA sequence and the gRNA reference FASTA file was indexed using BWA. The gRNA sequence of each read was obtained and aligned with BWA-mem, applying the parameters “-k10 -A4 -B2 -O2”. The sequencing results were saved as files according to gRNA. To ensure data accuracy, we discarded reads where sequences downstream of both the gRNA and the the Blue fluorescent protein (BFP) did not align. If the deletion was more than 100 bp and the deletion spanned a region within 100 bp of the cleavage site, the deletion was classified as a large deletion mutation. Because Nanopore sequencing has a bias depending on the length of the DNA fragment ( Supplementary Fig. 1 ), the ratio of the length of the DNA fragment to the reference was used instead of the count so that the ratio would decrease as the length of the deletion became longer. Based on the results of 60 non_targeting gRNAs, the Z-Score of large deletion for each gRNA was calculated. Knock-out cell line generation To generate three different gene ( LIG IV , POLQ , RAD52 ) knock-out HeLa cell line, 750 ng of Cas9 expression plasmid and 250 ng of sgRNA targeting upstream exon of each gene were transfected with 2µl of Lipofectamine 2000 reagent (Invitrogen) into HeLa cells. After 72 hours, CRISPR-treated HeLa cells were distributed as a single cell into each well of 96-well plates. The cell lines were cultured for two weeks, and each genotype was confirmed by an Illumina Miniseq instrument. The Miniseq results were analyzed using Cas-Analyzer ( http://www.rgenome.net/cas-analyzer/ ) 59 . For the complete knock-out, the cell line harboring a frameshift mutation in both alleles was selected. M4344 toxicity and effect on CRISPR-induced large deletion events To inhibit the ATR protein in HeLa cells, we used M4344 (Selleckchem, S9639) according to manufacturer’s protocol. HeLa cells (1 × 10 5 ) were seeded in 24-well plates. After 24 hours, the cells were exposed to M4344 (1nM, 5nM, 10nM, 25nM and 50nM) for 1 hour after which 750 ng of Cas9 and 250 ng of sgRNA expression plasmids were transfected with 2 µl of lipofectamine 2000 reagent (Invitrogen) into HeLa cells. After 72 hours from transfection, the cells are detached for genomic DNA extraction. Since the chemical was treated to the cells, the cells were maintained with chemical-containing media. Microhomology-dependent deletion events The alignment information about deletion reads was extracted from alignment results SAM files from k-mer alignment analysis program. Deletion position and the near sequence were calculated using the alignment information. The homology length was calculated by comparing the both sequences in 1bp increments from the start or end position of the deletion sites. If the homology length is 2–16 bp, the reads were categorized as microhomology-dependent. Transfection for base editing and prime editing HEK293T cells (1 x 10 5 cells per well) were cultivated in a 24-well plate for 24 hours. A mixture of 0.5 µl jetOPTIMUS reagent (Polyplus, 101000006), 500ng plasmid DNA (375ng BE expression plasmid and 125 ng sgRNA expression plasmid) or 543 ng plasmid DNA (365 ng PE expression plasmid, 125 ng pegRNA expression plasmid and 43 ng ngRNA expression plasmid) were added to the cells. After 72 hours, gDNA was isolated. References Hsu, P.D., Lander, E.S. & Zhang, F. Development and applications of CRISPR-Cas9 for genome engineering. Cell 157 , 1262-1278 (2014). Doudna, J.A. & Charpentier, E. Genome editing. The new frontier of genome engineering with CRISPR-Cas9. Science 346 , 1258096 (2014). Garneau, J.E. et al. The CRISPR/Cas bacterial immune system cleaves bacteriophage and plasmid DNA. Nature 468 , 67-71 (2010). Sfeir, A. & Symington, L.S. Microhomology-Mediated End Joining: A Back-up Survival Mechanism or Dedicated Pathway? Trends Biochem Sci 40 , 701-714 (2015). Ramsden, D.A., Carvajal-Garcia, J. & Gupta, G.P. Mechanism, cellular functions and cancer roles of polymerase-theta-mediated DNA end joining. Nat Rev Mol Cell Biol 23 , 125-140 (2022). Zhao, B., Rothenberg, E., Ramsden, D.A. & Lieber, M.R. The molecular basis and disease relevance of non-homologous DNA end joining. Nat Rev Mol Cell Biol 21 , 765-781 (2020). Smith, J. et al. Impact of DNA ligase IV on the fidelity of end joining in human cells. Nucleic Acids Res 31 , 2157-2167 (2003). Brambati, A., Barry, R.M. & Sfeir, A. DNA polymerase theta (Poltheta) - an error-prone polymerase necessary for genome stability. Curr Opin Genet Dev 60 , 119-126 (2020). Elliott, B., Richardson, C. & Jasin, M. Chromosomal translocation mechanisms at intronic alu elements in mammalian cells. Mol Cell 17 , 885-894 (2005). Yoshimi, K. et al. ssODN-mediated knock-in with CRISPR-Cas for large genomic regions in zygotes. Nat Commun 7 , 10431 (2016). Zhu, Z., Verma, N., Gonzalez, F., Shi, Z.D. & Huangfu, D. A CRISPR/Cas-Mediated Selection-free Knockin Strategy in Human Embryonic Stem Cells. Stem Cell Reports 4 , 1103-1111 (2015). Kosicki, M., Tomberg, K. & Bradley, A. Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nat Biotechnol 36 , 765-771 (2018). Zuccaro, M.V. et al. Allele-Specific Chromosome Removal after Cas9 Cleavage in Human Embryos. Cell 183 , 1650-1664 e1615 (2020). Liu, M. et al. Global detection of DNA repair outcomes induced by CRISPR-Cas9. Nucleic Acids Res 49 , 8732-8742 (2021). Papathanasiou, S. et al. Whole chromosome loss and genomic instability in mouse embryos after CRISPR-Cas9 genome editing. Nat Commun 12 , 5855 (2021). Turchiano, G. et al. Quantitative evaluation of chromosomal rearrangements in gene-edited human stem cells by CAST-Seq. Cell Stem Cell 28 , 1136-1147 e1135 (2021). Gaudelli, N.M. et al. Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551 , 464-471 (2017). Komor, A.C., Kim, Y.B., Packer, M.S., Zuris, J.A. & Liu, D.R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533 , 420-+ (2016). Anzalone, A.V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576 , 149-157 (2019). Song, Y. et al. Large-Fragment Deletions Induced by Cas9 Cleavage while Not in the BEs System. Mol Ther Nucleic Acids 21 , 523-526 (2020). Owens, D.D.G. et al. Microhomologies are prevalent at Cas9-induced larger deletions. Nucleic Acids Res 47 , 7402-7417 (2019). Peterka, M. et al. Harnessing DSB repair to promote efficient homology-dependent and -independent prime editing. Nat Commun 13 , 1240 (2022). Newby, G.A. et al. Base editing of haematopoietic stem cells rescues sickle cell disease in mice. Nature 595 , 295-302 (2021). Aida, T. et al. Prime editing primarily induces undesired outcomes in mice. bioRxiv , 2020.2008.2006.239723 (2020). Park, S.H. et al. Comprehensive analysis and accurate quantification of unintended large gene modifications induced by CRISPR-Cas9 gene editing. Sci Adv 8 , eabo7676 (2022). Hu, T., Chitnis, N., Monos, D. & Dinh, A. Next-generation sequencing technologies: An overview. Hum Immunol 82 , 801-811 (2021). Yu, S.C.Y. et al. Comparison of Single Molecule, Real-Time Sequencing and Nanopore Sequencing for Analysis of the Size, End-Motif, and Tissue-of-Origin of Long Cell-Free DNA in Plasma. Clin Chem 69 , 168-179 (2023). Cheng, C., Fei, Z. & Xiao, P. Methods to improve the accuracy of next-generation sequencing. Front Bioeng Biotechnol 11 , 982111 (2023). Dabney, J. & Meyer, M. Length and GC-biases during sequencing library amplification: a comparison of various polymerase-buffer systems with ancient and modern DNA sequencing libraries. Biotechniques 52 , 87-94 (2012). Mizuguchi, H., Nakatsuji, M., Fujiwara, S., Takagi, M. & Imanaka, T. Characterization and application to hot start PCR of neutralizing monoclonal antibodies against KOD DNA polymerase. J Biochem 126 , 762-768 (1999). Takagi, M. et al. Characterization of DNA polymerase from Pyrococcus sp. strain KOD1 and its application to PCR. Appl Environ Microbiol 63 , 4504-4510 (1997). Kim, D., Kim, S., Kim, S., Park, J. & Kim, J.S. Genome-wide target specificities of CRISPR-Cas9 nucleases revealed by multiplex Digenome-seq. Genome Res 26 , 406-415 (2016). Jeong, Y.K., Yu, J. & Bae, S. Construction of non-canonical PAM-targeting adenosine base editors by restriction enzyme-free DNA cloning using CRISPR-Cas9. Scientific Reports 9 , 4939 (2019). Yoon, H.H. et al. CRISPR-Cas9 Gene Editing Protects from the A53T-SNCA Overexpression-Induced Pathology of Parkinson's Disease In Vivo. CRISPR J 5 , 95-108 (2022). Stadtmauer, E.A. et al. CRISPR-engineered T cells in patients with refractory cancer. Science 367 (2020). Wen, W. et al. Effective control of large deletions after double-strand breaks by homology-directed repair and dsODN insertion. Genome Biol 22 , 236 (2021). Wu, J. et al. CRISPR/Cas9-induced structural variations expand in T lymphocytes in vivo. Nucleic Acids Res 50 , 11128-11137 (2022). Hussmann, J.A. et al. Mapping the genetic landscape of DNA double-strand break repair. Cell 184 , 5653-5669 e5625 (2021). Gilbert, L.A. et al. CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes. Cell 154 , 442-451 (2013). Chang, H.H.Y., Pannunzio, N.R., Adachi, N. & Lieber, M.R. Non-homologous DNA end joining and alternative pathways to double-strand break repair. Nat Rev Mol Cell Biol 18 , 495-506 (2017). Kurt, I.C. et al. CRISPR C-to-G base editors for inducing targeted DNA transversions in human cells. Nat Biotechnol 39 , 41-46 (2021). Jeong, Y.K. et al. Adenine base editor engineering reduces editing of bystander cytosines. Nat Biotechnol 39 , 1426-1433 (2021). Kunz, C., Saito, Y. & Schar, P. DNA Repair in mammalian cells: Mismatched repair: variations on a theme. Cell Mol Life Sci 66 , 1021-1038 (2009). Mol, C.D. et al. Crystal structure of human uracil-DNA glycosylase in complex with a protein inhibitor: protein mimicry of DNA. Cell 82 , 701-708 (1995). Hegde, M.L., Hazra, T.K. & Mitra, S. Early steps in the DNA base excision/single-strand interruption repair pathway in mammalian cells. Cell Res 18 , 27-47 (2008). Kim, H.S., Jeong, Y.K., Hur, J.K., Kim, J.S. & Bae, S. Adenine base editors catalyze cytosine conversions in human cells. Nat Biotechnol 37 , 1145-1148 (2019). Koblan, L.W. et al. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat Biotechnol 36 , 843-846 (2018). Tong, H. et al. Programmable A-to-Y base editing by fusing an adenine base editor with an N-methylpurine DNA glycosylase. Nat Biotechnol 41 , 1080-1084 (2023). Habib, O., Habib, G., Hwang, G.H. & Bae, S. Comprehensive analysis of prime editing outcomes in human embryonic stem cells. Nucleic Acids Res 50 , 1187-1197 (2022). Lee, J. et al. Prime editing with genuine Cas9 nickases minimizes unwanted indels. Nat Commun 14 , 1786 (2023). Ran, F.A. et al. Double nicking by RNA-guided CRISPR Cas9 for enhanced genome editing specificity. Cell 154 , 1380-1389 (2013). Chen, P.J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184 , 5635-5652 e5629 (2021). Nelson, J.W. et al. Engineered pegRNAs improve prime editing efficiency. Nat Biotechnol 40 , 402-410 (2022). Li, S. et al. Genome-edited powdery mildew resistance in wheat without growth penalties. Nature 602 , 455-460 (2022). Tsuchida, C.A. et al. Mitigation of chromosome loss in clinical CRISPR-Cas9-engineered T cells. bioRxiv , 2023.2003.2022.533709 (2023). Park, J., Bae, S. & Kim, J.S. Cas-Designer: a web-based tool for choice of CRISPR-Cas9 target sites. Bioinformatics 31 , 4014-4016 (2015). Untergasser, A. et al. Primer3Plus, an enhanced web interface to Primer3. Nucleic Acids Res 35 , W71-74 (2007). Ye, J. et al. Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction. BMC Bioinformatics 13 , 134 (2012). Park, J., Lim, K., Kim, J.S. & Bae, S. Cas-analyzer: an online tool for assessing genome editing results using NGS data. Bioinformatics 33 , 286-288 (2017). Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryInformation0105.docx Supplementary Information SupplementaryTable1.xlsx SupplementaryTable2.xlsx Cite Share Download PDF Status: Published Journal Publication published 04 Nov, 2024 Read the published version in Nature Biomedical Engineering → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3835370","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":266368626,"identity":"d965da65-23ef-4956-b4d5-5e1e2fab4137","order_by":0,"name":"Sangsu Bae","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAtUlEQVRIiWNgGAWjYDACCQYG5j8VMN4BIrUw8JwhWQtvGyla5Gc3P3wgOa8uz+AA88MPDGfuEdZicOeYsYHhtsPFBgfYjCUYbhQToUUiwUwicduBxA0HGMwYGD4kEOGwGenffxycUwfUwv6NOC0MN3LMGBsbmIFaeIC23CBCi8GdM8XSDMcOJ848zFMskXCGGIfNbt/4maGmLrHvePvGDx+OEeMwOGAGYpI0jIJRMApGwSjADQCwQjxRnnLEGwAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-3615-8566","institution":"Department of Biomedical Sciences, Seoul National University College of Medicine","correspondingAuthor":true,"prefix":"","firstName":"Sangsu","middleName":"","lastName":"Bae","suffix":""},{"id":266368627,"identity":"8975bf90-f98f-44e4-8786-1b553bfad8b0","order_by":1,"name":"Gue-Ho Hwang","email":"","orcid":"https://orcid.org/0000-0002-4201-0974","institution":"Hanyang University","correspondingAuthor":false,"prefix":"","firstName":"Gue-Ho","middleName":"","lastName":"Hwang","suffix":""},{"id":266368628,"identity":"96c25975-2747-4dc6-831e-e74e09513cba","order_by":2,"name":"Seok-Hoon Lee","email":"","orcid":"","institution":"Seoul National University College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Seok-Hoon","middleName":"","lastName":"Lee","suffix":""},{"id":266368629,"identity":"d62f4649-c3f0-481e-a0b8-affd2e3e081b","order_by":3,"name":"Minsik Oh","email":"","orcid":"","institution":"Myongji University","correspondingAuthor":false,"prefix":"","firstName":"Minsik","middleName":"","lastName":"Oh","suffix":""},{"id":266368630,"identity":"bacd1456-5c9f-4232-be22-8ca92ec6a14e","order_by":4,"name":"Segi Kim","email":"","orcid":"","institution":"Korea Advanced Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Segi","middleName":"","lastName":"Kim","suffix":""},{"id":266368631,"identity":"a4d224b3-fe7c-4715-aadd-151b4e081155","order_by":5,"name":"Omer Habib","email":"","orcid":"","institution":"RedGene Inc","correspondingAuthor":false,"prefix":"","firstName":"Omer","middleName":"","lastName":"Habib","suffix":""},{"id":266368632,"identity":"25022db7-af52-4b45-b59a-ef3dfb12ebe2","order_by":6,"name":"Hyeon-Ki Jang","email":"","orcid":"","institution":"Kangwon National University","correspondingAuthor":false,"prefix":"","firstName":"Hyeon-Ki","middleName":"","lastName":"Jang","suffix":""},{"id":266368633,"identity":"29539b45-7362-4600-bc79-7865862a4de9","order_by":7,"name":"Heon Seok Kim","email":"","orcid":"","institution":"Hanyang University","correspondingAuthor":false,"prefix":"","firstName":"Heon","middleName":"Seok","lastName":"Kim","suffix":""},{"id":266368634,"identity":"56ff4029-42b7-48cb-98d8-39843b3f7b84","order_by":8,"name":"Chan Hyuk Kim","email":"","orcid":"https://orcid.org/0000-0001-9649-1892","institution":"Korea Advanced Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Chan","middleName":"Hyuk","lastName":"Kim","suffix":""},{"id":266368635,"identity":"3b5ab626-8261-4108-8799-f6a394949695","order_by":9,"name":"Sun Kim","email":"","orcid":"","institution":"Seoul National University","correspondingAuthor":false,"prefix":"","firstName":"Sun","middleName":"","lastName":"Kim","suffix":""}],"badges":[],"createdAt":"2024-01-04 21:00:58","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3835370/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3835370/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41551-024-01277-5","type":"published","date":"2024-11-04T05:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":49509891,"identity":"db2ed86a-ccfa-4b02-96a6-a7eae48c4d62","added_by":"auto","created_at":"2024-01-12 05:49:27","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":259012,"visible":true,"origin":"","legend":"\u003cp\u003eScheme of long-range deep sequencing and proof of concept. \u003cstrong\u003ea\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eSchematic of long-range amplicon sequencing. DNA sequences (10 kb) are amplified from genomic DNA isolated from CRISPR-treated and non-treated cells. After fragmentation to ~300 bp, the fragmented DNA is sequenced by Illumina. The resulting long-range amplicon sequencing data are analyzed with our k-mer alignment program and the frequencies of small mutations and large deletions are calculated. \u003cstrong\u003eb\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eComparison of length bias according to type of DNA polymerase and PCR cycle. Nine sizes of DNA fragments were mixed and sequenced with Illumina under various PCR cycle conditions.\u003cstrong\u003e c\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eSchematic of validation experiments for the long-range amplicon sequencing workflow. Wildtype (WT) genomic DNA (gray) and genomic DNA (orange) from a cell line with a 10 bp and 1 kb deletion in \u003cem\u003eST3GAL4\u003c/em\u003e were mixed in a ratio of 10:0, 8:2, 5:5, 2:8, and 0:10. \u003cstrong\u003ed-e, \u003c/strong\u003eComparing the expected mutation frequency and measured mutation frequency for large deletions (\u003cstrong\u003ed\u003c/strong\u003e) and small deletions (\u003cstrong\u003ee\u003c/strong\u003e).\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/2fd413af63497c004b05adfd.png"},{"id":49509892,"identity":"7959d244-1e88-4da5-ae4c-5f2932cbf074","added_by":"auto","created_at":"2024-01-12 05:49:27","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":213568,"visible":true,"origin":"","legend":"\u003cp\u003eMutation frequency in various cell lines and human primary T cells. \u003cstrong\u003ea-f\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eThe frequencies of small deletion, small insertion, and large deletion events in HeLa (\u003cstrong\u003ea)\u003c/strong\u003e, HEK293T (\u003cstrong\u003eb)\u003c/strong\u003e, U2OS (\u003cstrong\u003ec\u003c/strong\u003e)\u003cstrong\u003e,\u003c/strong\u003e K562 (\u003cstrong\u003ed\u003c/strong\u003e), normal fibroblast (\u003cstrong\u003ee\u003c/strong\u003e), stem cells (H9) (\u003cstrong\u003ef\u003c/strong\u003e) after CRISPR-Cas9 gene editing at the indicated genes as determined with long-range amplicon sequencing (n=3).\u003cstrong\u003e g, \u003c/strong\u003eSchematic of isolation and long-rangd deep sequencing for primary human T cells.\u003cstrong\u003e h \u003c/strong\u003eand\u003cstrong\u003e i\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eThe frequencies of small deletion, small insertion, and large deletion events in human T cells from two healthy donors, ASAN04\u003cstrong\u003e \u003c/strong\u003e(\u003cstrong\u003eh\u003c/strong\u003e) and ASAN68 (\u003cstrong\u003ei\u003c/strong\u003e) subjected to CRISPR-Cas9 editing at the indicated genes as determined with long-range amplicon sequencing (n=3). Error bars, mean±SD.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/10cdca2d21e7b5d9fa86e09b.jpeg"},{"id":49509900,"identity":"6a08a4cb-c501-4114-ac68-e5953f817193","added_by":"auto","created_at":"2024-01-12 05:49:30","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":301904,"visible":true,"origin":"","legend":"\u003cp\u003eMutation frequency in knockdown and knockout cell lines. \u003cstrong\u003ea\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eSchematic of CRISPRi screening with Nanopore-seq. CRISPRi stable cells were infected with a lentiviral library. The infected cells were selected with puromycin and transfected with Cas9 RNP complex targeting the puromycin resistance gene using electroporation. The regions with gRNA or repair outcome (puromycin resistance gene) were amplified and sequenced using Nanopore sequencing. \u003cstrong\u003eb\u003c/strong\u003e, Model of DNA double-strand break (DSB) repair pathway. Red colored words are genes and chemical (M4344) used in experiments for KO cell lines or pharmacological inhibition of specific DNA repair processes. \u003cstrong\u003ec\u003c/strong\u003e, Data from CRISPRi screening with Nanopore-seq plotted as Z-scores of the frequency of large deletion events in cells with the indicated gene knocked down. \u003cstrong\u003ed\u003c/strong\u003e, Z-score comparison of the difference in large deletion frequency between negative control (non-targeting) cells and those with the indicated repair pathway compromised. Z-score was calculated using the mean and distribution of non-targeting sample. P-values were calculated using Mann-Whitney tests. \u003cstrong\u003ee\u003c/strong\u003e, Box plot graph to compare large deletion rate in wild-type and the indicated KO cells subjected to CRISPR-Cas9. \u003cstrong\u003ef\u003c/strong\u003e and \u003cstrong\u003eg\u003c/strong\u003e,\u003cstrong\u003e \u003c/strong\u003eThe frequency of large deletion events in human primary T cells, which was analyzed in two healthy donors ASAN99 (\u003cstrong\u003ef\u003c/strong\u003e) and ASAN107 (\u003cstrong\u003eg\u003c/strong\u003e) using long-range amplicon sequencing (n=3; error bar, mean±SD). P-values in \u003cstrong\u003ee\u003c/strong\u003e-g were calculated using student’s \u003cem\u003et\u003c/em\u003e-test (n.s., not significant, * \u003cem\u003eP \u003c/em\u003e\u0026lt; 0.05, ** \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.01, *** \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001, ****\u003cem\u003e P \u003c/em\u003e\u0026lt; 0.0001).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/7d9a2fea0a3b1054a52e7225.png"},{"id":49509894,"identity":"844aab86-5f10-44f0-ac41-264ec10f1c14","added_by":"auto","created_at":"2024-01-12 05:49:27","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":335267,"visible":true,"origin":"","legend":"\u003cp\u003eDNA large deletions are generated by base editors. \u003cstrong\u003ea.\u003c/strong\u003e Schematic diagrams of six different BEs. N and C indicate the amino-terminal and carboxy-terminal ends, respectively. The version of CBE and ABE is AncBE4max and ABEmax, respectively. \u003cstrong\u003eb.\u003c/strong\u003e Editing outcomes of dCas9 (inactivated Cas9) and the indicated BEs targeting \u003cem\u003eABLIM3\u003c/em\u003e, \u003cem\u003eFANCF\u003c/em\u003e, \u003cem\u003eKLHL29\u003c/em\u003e, and \u003cem\u003eBRD4\u003c/em\u003e (\u003cem\u003en\u003c/em\u003e=3, mean ±SEM). \u003cstrong\u003ec.\u003c/strong\u003e Summarized box plot graph to compare the large deletion frequency among the BEs targeting \u003cem\u003eABLIM3\u003c/em\u003e, \u003cem\u003eFANCF\u003c/em\u003e, \u003cem\u003eKLHL29 \u003c/em\u003eand \u003cem\u003eBRD4\u003c/em\u003e. \u003cem\u003eP\u003c/em\u003e = 0.0249 for CBE, \u003cem\u003eP\u003c/em\u003e = 0.0028 for CBE(DUGI), \u003cem\u003eP\u003c/em\u003e = 0.0020 for CGBE1, \u003cem\u003eP\u003c/em\u003e = 0.0005 for ABE, \u003cem\u003eP\u003c/em\u003e = 0.0062 for ABE-UGI, and \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.0001 for AYBE compared to the dCas9 by student’s \u003cem\u003et\u003c/em\u003e-test. \u003cstrong\u003ed.\u003c/strong\u003e Analysis of the large deletion frequency of ABE with or without UGI at targets containing a TC-motif (TC target, \u003cem\u003eABLIM3\u003c/em\u003e and \u003cem\u003eFANCF\u003c/em\u003e) and targets without a TC-motif (Non-TC target, \u003cem\u003eKLHL29\u003c/em\u003e and \u003cem\u003eBRD4\u003c/em\u003e). Statistical analysis was performed by student’s t test. Within the TC target, \u003cem\u003eP\u003c/em\u003e\u0026lt; 0.0001 for ABE, \u003cem\u003eP\u003c/em\u003e = 0.0260 for ABE-UGI and \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.0001 for AYBE compared to the dCas9. \u003cem\u003eP\u003c/em\u003e = 0.0007 for ABE comparing with ABE-UGI at the TC target. Within the non-TC target, \u003cem\u003eP\u003c/em\u003e = 0.0564 for ABE, \u003cem\u003eP\u003c/em\u003e = 0.1264 for ABE-UGI and \u003cem\u003eP\u003c/em\u003e = 0.0010 for AYBE compared to the dCas9. \u003cem\u003eP\u003c/em\u003e= 0.0001 for ABE at TC target compared to the ABE at non-TC target. \u003cstrong\u003ee. \u003c/strong\u003eSchematic of potential cellular mechanisms generating DNA large deletion by BEs. Uracil or inosine excision by endogenous uracil DNA glycosylase (UNG) or engineered protein N-methylpurine DNA glycosylase (eMPG), nicking of the non-target strand by nickase Cas9(D10A), AP site generation by DNA Apurinic or apyrimidinic site lyase (AP lyase), followed by DNA double strand break (DSB) formation. *\u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05, **\u003cem\u003eP\u003c/em\u003e \u0026lt; 0.01, ***\u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001 and ****\u003cem\u003eP\u003c/em\u003e \u0026lt; 0.0001. The description of n.s means non-significant.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/4214cdb500aede5d543c3ae4.png"},{"id":49510105,"identity":"c90fc9ba-4fec-4633-8987-b81b37c1174c","added_by":"auto","created_at":"2024-01-12 05:57:27","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":224447,"visible":true,"origin":"","legend":"\u003cp\u003ePrime editors can generate DNA large deletion especially with the additional nicking gRNA. \u003cstrong\u003ea. \u003c/strong\u003eEditing outcomes of dCas9, PE2, and PE3 with substitution, deletion, or insertion edits. Two targets were tested for each type of editing: \u003cem\u003eHEK4 \u003c/em\u003eand \u003cem\u003eHEK3\u003c/em\u003e for substitution, \u003cem\u003eFANCF\u003c/em\u003e and \u003cem\u003eHEK3\u003c/em\u003e for deletion, \u003cem\u003eRNF2 \u003c/em\u003eand \u003cem\u003eHEK3\u003c/em\u003e for insertion. For PE3 strategy, two ngRNAs were tested individually. (\u003cem\u003en\u003c/em\u003e=3, mean ± SEM). \u003cstrong\u003eb\u003c/strong\u003e and \u003cstrong\u003ec. \u003c/strong\u003eSummarized box plot graph describing the large deletion frequency of PE2 (b) or PE3 (c). \u003cem\u003eP\u003c/em\u003e = 0.0103 for PE2 and \u003cem\u003eP\u003c/em\u003e = 0.0056 for PE3 (Student’s \u003cem\u003et\u003c/em\u003e-test, two-sided). \u003cstrong\u003ed.\u003c/strong\u003e The effect of hMLH1dn or epegRNA on DNA large deletions induced by the PE3 system. Analysis was performed with targets +2GtoC with -38 ngRNA in \u003cem\u003eHEK3\u003c/em\u003e, 5-7del with +48 ngRNA in \u003cem\u003eFANCF\u003c/em\u003e. \u003cem\u003eP\u003c/em\u003e = 0.7507 for PE3 + hMLH1dn and \u003cem\u003eP\u003c/em\u003e = 0.0347 for PE3 + epegRNA compared to the original PE3. *\u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05, **\u003cem\u003eP\u003c/em\u003e \u0026lt; 0.01; n.s means non-significant.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/fd41fa414dbd681b7cd50275.png"},{"id":68240447,"identity":"fd1e96f8-4854-4bcd-8cb8-e8def2ea115c","added_by":"auto","created_at":"2024-11-05 08:05:31","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2190053,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/811e271b-8bba-48d3-8e8d-1aa7d383b9d9.pdf"},{"id":49509896,"identity":"f12f566d-be94-4107-a336-d2ca51671789","added_by":"auto","created_at":"2024-01-12 05:49:28","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":6536919,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Information\u003c/p\u003e","description":"","filename":"SupplementaryInformation0105.docx","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/0453cf1b75e38920e22aa1a2.docx"},{"id":49509895,"identity":"25733b8a-bf18-4520-aab1-164737f457e2","added_by":"auto","created_at":"2024-01-12 05:49:27","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":212887,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/bd73edef862c58fb97794052.xlsx"},{"id":49510363,"identity":"6ebf664c-d0e5-42dd-8dca-bb1e5610f83e","added_by":"auto","created_at":"2024-01-12 06:05:27","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":78049,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cbr\u003e\u003c/p\u003e","description":"","filename":"SupplementaryTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3835370/v1/17920a3a94947440cabe82a9.xlsx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Detailed mechanisms for unintended large DNA deletions with CRISPR, base editors, and prime editors","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe CRISPR-Cas system, derived from a prokaryotic immune response, uses single-guide RNAs (sgRNAs) to recognize target DNA sequences and then introduces DNA double-strand breaks (DSBs) at the target sites\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. In mammalian cells, DSBs are repaired through multiple DNA repair pathways: (i) non-homologous end joining (NHEJ), (ii) DNA polymerase theta (Pol θ)-mediated end joining (TMEJ) that operates in a microhomology-dependent manner with as few as 2 to 16 nucleotides of homologous sequence\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, (iii) single-strand annealing (SSA) that uses homologous repeat sequences to bridge DSB ends, (iv) homologous recombination (HR), and (v) homology-directed repair (HDR)\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Because NHEJ, TMEJ, and SSA are error-prone, they frequently lead to insertion or deletion or both (collectively referred to as indel) mutations during the repair process\u003csup\u003e\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. In contrast, the HR and HDR pathways are error-free, thus genome editing that uses single-stranded oligodeoxynucleotides (ssODN) or double-stranded donor plasmids flanking homology arms has higher fidelity by leveraging these repair pathways\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAlthough most deletions generated by CRISPR nucleases were shorter than 20 base pairs (bp), it was observed that more than 20% of the mutations in mouse embryos induced by CRISPR-Cas9 were unintended deletions longer than 250 bp\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Moreover, DSBs induced by CRISPR nucleases can induce chromosomal rearrangements, including chromosomal depletion and translocation\u003csup\u003e\u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. Because chromosomal translocation involves off-target cleavage sites and the on-target cleavage site, high-fidelity Cas9 nucleases reduce translocation rates\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. Unfortunately, high-fidelity Cas9 nuclease do not overcome the issue of unintended large deletions that occur at on-target sites. Another variation of the CRISPR-Cas9 system involves using base editors (BEs) and prime editors (PEs), which employ partially inactivated Cas9 (nCas9) that generates single-strand nicks\u003csup\u003e\u003cspan additionalcitationids=\"CR18\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Whether these nCas9-based genome editing systems also generate large deletion is not clearly established\u003csup\u003e\u003cspan additionalcitationids=\"CR21 CR22 CR23\" citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. Identifying the repair pathways that produce large deletions following genome-editing tools, with Cas9 or nCas9, will not only provide key insight into the biology of this process but enable the development of strategies to mitigate large deletion events.\u003c/p\u003e \u003cp\u003eDetermining the frequency of deletion events\u0026thinsp;\u0026gt;\u0026thinsp;100 bp is challenging using short-range high-throughput sequencing methods. Instead, such events are detected with either a long-range amplicon sequencing method using short-read sequencers, such as Illumina platforms, or a long-read sequencing method, using Pacific Biosciences (PacBio) or Oxford Nanopore Technologies (ONT) platforms\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. In this study, we modified a long-range amplicon sequencing method and developed a novel k-mer sequence alignment algorithm, and we used this strategy to accurately determine CRISPR-induced large deletions. We observed that large deletions were mostly accompanied by small indels in human cells, including cancer-like cells (HeLa, HEK293T, U2OS, K562), fibroblasts, primary T cells, and human embryo stem cells (H9). We then identified large deletion-associated repair pathways through CRISPR interference (CRISPRi) screening experiments with ONT Nanopore sequencing, establishing that TMEJ is the dominant pathway generating large deletions. We found that BEs generate large deletions through the base excision repair (BER) pathway, especially in the presence of an apurinic/apyrimidinic site (AP site) and that PEs particularly generate large deletions in the presence of additional nicking guide RNAs (ngRNAs). Our findings reveal mechanistic challenges in using genome-editing tools and provide insight for avoiding these issues in research and therapeutic applications.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eOptimizing a long-range amplicon sequencing method to determine small indels and large deletions with a high accuracy\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNeither the PacBio Single-molecule real-time (SMRT) sequencing nor ONT Nanopore sequencing platforms for long-read sequencing have sufficient sequencing accuracy (PacBio\u0026thinsp;~\u0026thinsp;10% and ONT 2%-15%)\u003csup\u003e26\u003c/sup\u003e to distinguish 1-bp deletion, 1-bp insertion, and substitution events. Furthermore, these methods have a length bias: shorter DNA fragments are read more easily and thereby are overestimated compared to longer DNA fragments (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;1\u003c/strong\u003e)\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. Thus, currently available long-read sequencing platforms lack sufficient accuracy to identify both large deletions and small indels. Our first task was to establish a precise sequencing method to simultaneously determine small indels and large deletions with a high accuracy.\u003c/p\u003e\n\u003cp\u003eThe Illumina platform has the highest sequencing accuracy (\u0026gt;\u0026thinsp;99.2%)\u003csup\u003e28\u003c/sup\u003e. Therefore, we optimized long-range amplicon sequencing using Illumina sequencers combined with long-range polymerase chain reaction (PCR) and DNA fragmentation (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ea\u003cstrong\u003e)\u003c/strong\u003e. The four steps in our sequencing were as follows: (i) extract genomic DNA (gDNA) from CRISPR-Cas9-treated and non-treated cell lines, (ii) amplify\u0026thinsp;~\u0026thinsp;10\u0026ndash;15 kb of DNA involving the CRISPR target sites using PCR optimized for both gDNAs, (iii) fragment the amplified products into ~\u0026thinsp;300 bp and prepare next generation sequencing (NGS) libraries through end-repair, dA-tailing, adaptor ligation, and PCR enrichment, (iv) acquire the sequencing data with an Illumina Miniseq.\u0026nbsp;To quantify large deletions, small deletions, and small insertions in the bulk cell populations, we used a novel k-mer alignment algorithm (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;2\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eIn this protocol, PCR-based DNA amplification of the long DNA region (~\u0026thinsp;10\u0026ndash;15 kb ) is a critical step. To mitigate the length bias of DNA polymerases, we tested five DNA polymerases: Phusion HF, Q5 that is used in DNA library prep kit for Illumina, Accuprime pfx\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, KOD -Multi \u0026amp; Epi- DNA polymerase that is a modified KOD DNA polymerase to reduce amplification bias during PCR by addition of the elongation accelerator\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e, and SUN-PCR blend\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. To evaluate the PCR length bias, we prepared nine synthesized DNA fragments in increments of 20 bp from 120 bp to 280 bp with a common forward/reverse primer pair, mixed the fragments in equimolar amounts, and conducted PCR experiments for the mixture with each polymerase (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;3\u003c/strong\u003e). KOD -Multi \u0026amp; Epi- DNA polymerase had the least length bias of the 5 polymerases tested (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eb). Therefore, we used this polymerase for the long-range PCR amplification step. Another complicated step in the workflow is sequence alignment of short-read NGS data. Pre-existing tools, such as a burrows-wheeler aligner (BWA)-mem, were developed to rapidly map DNA sequence reads to whole genome sequences. Our workflow requires querying sequences that are limited to the PCR-amplified region to detect indels near the CRISPR-mediated cleavage sites. This is a more limited application that does not require a heuristic approach. Thus, we developed a dedicated k-mer alignment algorithm that is compact and is potentially available as an online tool for analysis of CRISPR-treated samples (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;4\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eTo confirm the functionality of our optimized long-range amplicon sequencing workflow, we prepared and applied the workflow to two cell lines: One has two wild-type \u003cem\u003eST3GAL4\u003c/em\u003e alleles and the other has a 10-bp or a 1,075-bp deletion in each allele (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ec). After mixing gDNAs from the two cell lines at ratios of 0%, 20%, 50%, 80%, and 100%, the mixtures were prepared and subjected to the optimized long-range amplicon sequencing. The optimized long-range amplicon sequencing was highly accurate in detecting both the 10-bp and 1,075-bp deletion events (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ed and \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003ee), showing that our method is an effective tool for measuring large deletions and small indels simultaneously.\u003c/p\u003e\n\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003eDetermination of CRISPR-mediated small indels and large deletions in various human cell lines\u003c/h2\u003e\n\u003cp\u003eBased on the optimized long-range amplicon sequencing, we quantified large deletions and small indels induced by CRISPR-Cas9 nuclease in various human cell lines. We evaluated four cancer or transformed cell lines (HeLa, HEK293T, U2OS, K562), human fibroblasts, and a human embryo stem cell line (H9) (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ea-\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ef). For HeLa and HEK293T cells, we analyzed the frequency of CRISPR-Cas9-mediated deletion or insertion events for one target site each in \u003cem\u003eHPRT1\u003c/em\u003e and \u003cem\u003eST3GAL4\u003c/em\u003e and five target sites in the \u003cem\u003eWISP3\u003c/em\u003e gene. We defined a small insertion and a small deletion event as shorter than 100 bp, and a large deletion event as larger than 100 bp. Across these seven target sites, the average frequency of small indel events were 37.5\u0026thinsp;\u0026plusmn;\u0026thinsp;13.2% in HeLa cells and 26.4\u0026thinsp;\u0026plusmn;\u0026thinsp;11.7% in HEK293T cells, whereas the average frequency of large deletion event was 6.4\u0026thinsp;\u0026plusmn;\u0026thinsp;4.4% in HeLa cells and 4.42\u0026thinsp;\u0026plusmn;\u0026thinsp;3.6% in HEK293T cells (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ea and \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eb). Small indels were the most common CRISPR-Cas9-mediated events in all of the cells with an average frequency of 17.8\u0026thinsp;\u0026plusmn;\u0026thinsp;7.8% in U2OS cells, 18.4\u0026thinsp;\u0026plusmn;\u0026thinsp;10.1% in K562 cells, 10.6\u0026thinsp;\u0026plusmn;\u0026thinsp;8.5% in fibroblast cells, and 49.3\u0026thinsp;\u0026plusmn;\u0026thinsp;21.7% in H9 cells. CRISPR-Cas9-mediated large deletion events were less common with an average frequency of 3.3\u0026thinsp;\u0026plusmn;\u0026thinsp;1.8% in U2OS cells, 2.4\u0026thinsp;\u0026plusmn;\u0026thinsp;1.5% in K562 cells, 1.5\u0026thinsp;\u0026plusmn;\u0026thinsp;1.7% in fibroblast cells, and 2.6\u0026thinsp;\u0026plusmn;\u0026thinsp;1.6% in H9 cells (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ec-\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ef). The frequency of CRISPR-Cas9-mediated large deletions across all target sites and all six cell lines ranged from 0.2\u0026ndash;17.5%.\u003c/p\u003e\n\u003cp\u003eBecause CRISPR-Cas9-based genome editing has been widely adopted in engineering T cells for cancer immunotherapy (clinicaltrials.gov: NCT03399448, NCT03081715, NCT03398967, NCT02793856, NCT03044743), we evaluated CRISPR-Cas9-mediated mutation patterns in human primary T cells. We chose sgRNAs targeting \u003cem\u003ePD1\u003c/em\u003e and \u003cem\u003eTRAC1\u003c/em\u003e loci, because these genes are frequently targeted in T cells to enhance T-cell activity\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e. Human T cells were isolated from the peripheral blood mononuclear cells (PBMCs) of two healthy donors (ASAN-04 and ASAN-68) using a magnetic-activated cell sorting (MACS) system and then stimulated by anti-CD3/CD28 beads before introduction of Cas9-sgRNA ribonucleoprotein (RNP) complexes by electroporation (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eg). The optimized long-range amplicon sequencing revealed that the average frequencies of small indels were 70.99\u0026thinsp;\u0026plusmn;\u0026thinsp;11.37% in ASAN04 and 72.03\u0026thinsp;\u0026plusmn;\u0026thinsp;15.46% in ASAN68, whereas the average frequencies of large deletion events were 15.46\u0026thinsp;\u0026plusmn;\u0026thinsp;6.68% in ASAN04 and 14.24\u0026thinsp;\u0026plusmn;\u0026thinsp;6.82% in ASAN68 (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eh and \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003ei), similar to the previous results\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. The average frequency of large deletion events in human primary T cells was higher than the average frequency of these events in the cell lines tested.\u003c/p\u003e\n\u003cp\u003eUsing all of the data, we observed that the large deletion frequencies positively correlated with the small deletion frequencies (Pearson corr. = 0.72) and were negatively correlated with small insertion frequencies (Pearson corr. = -0.15), resulting in a low correlation with small indels (Pearson corr. = 0.47) (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;5\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDetermination by CRISPRi screening that end resection and TMEJ processes play positive roles for generating large deletions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo examine the underlying mechanism of CRISPR-driven large DNA deletion events, we performed a CRISPR interference (CRISPRi) screen targeting 794 genes that were those referenced in Repair-seq\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e plus genes identified as associated with DNA repair in the Human Protein Atlas website (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003ea). We introduced a plasmid library with three different sgRNAs per gene and the puromycin resistance gene into an engineered HeLa cell line that stably expresses deactivated Cas9 (dCas9) fused with Kr\u0026uuml;ppel associated box (KRAB) domain to produce gene knockdown\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e. DSBs in the puromycin resistance gene were generated by addition of the Cas9 RNP complex, gDNA was prepared, and ~\u0026thinsp;5.6 kbp of DNA involving the sgRNA sites or the DSB (puromycin resistance gene) site were amplified and sequenced with the Nanopore MinION. Because we only needed to assess qualitatively large deletion events, we used Nanopore-seq to read long range region. However, to account for the length-dependent bias in Nanopore-seq (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;1\u003c/strong\u003e), each sequencing read was adjusted according to its total length and then classified as wild-type, insertion, deletion, and large deletions (those\u0026thinsp;\u0026gt;\u0026thinsp;100 bp). We compared Nanopore-seq data from cells with or without inhibition of the 794 target genes (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;6\u003c/strong\u003e and \u003cstrong\u003eSupplementary Table\u0026nbsp;1\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eDNA repair pathways can be divided into end protection and end resection\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eb). After DSBs occur, the end protection pathway proceeds to NHEJ, which induces blunt-end ligation by Ku-XRCC4-DNA ligase IV complex\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e. In contrast, the end resection pathway proceeds to TMEJ, SSA, or HR, according to the 5\u0026rsquo; end resection sequence and the existence of homology sequences in neighboring DNA\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e. Using a previously defined classification scheme, we categorized the genes into those associated with NHEJ (\u003cem\u003eLIG4, XRCC4, XRCC5\u003c/em\u003e, and \u003cem\u003eXRCC6\u003c/em\u003e), end resection (\u003cem\u003eMRE11A, NBN\u003c/em\u003e, and \u003cem\u003eRAD50\u003c/em\u003e), TMEJ (\u003cem\u003ePOLQ\u003c/em\u003e), SSA (\u003cem\u003eERCC4\u003c/em\u003e and \u003cem\u003eRAD52\u003c/em\u003e), and HR (\u003cem\u003eBRCA1, BRCA2, RAD51AP1, RAD51B, RAD51C\u003c/em\u003e, and \u003cem\u003eRAD51D\u003c/em\u003e). Z-scores for large deletions in the puromycin resistance gene region for each of the CRISPRi-targeted genes within each category indicated that fewer large deletions occurred when \u003cem\u003ePOLQ\u003c/em\u003e of TMEJ or genes associated with end resection were knocked down and large deletions were more common when SSA genes were knocked down (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003ec). Consistent with these findings, we calculated Z-scores of the difference in large deletion frequencies in cells with and without inhibition of target genes and found a significant reduction in large deletions in cells with either reduced activity of TMEJ (P-value\u0026thinsp;=\u0026thinsp;9.63x10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e) or the end resection pathway (P-value\u0026thinsp;=\u0026thinsp;7.71x10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e) and a significant increase in cells with reduced activity of SSA (P-value\u0026thinsp;=\u0026thinsp;2.19x10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003ed). Collectively, the screening results indicated that large deletions are primarily generated by TMEJ following the end resection pathway.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n\u003ch2\u003eConfirmation by knockout that end resection and TMEJ processes generate large deletions\u003c/h2\u003e\n\u003cp\u003eTo confirm the function of key genes in generating large deletions, we constructed individual knock-out (KO) cell lines using CRISPR-Cas9 or disrupted specific repair pathways pharmacologically. We used the following strategies to impair specific repair pathways in HeLa cells: M4344 to pharmacologically inhibit ATR activity and suppress end resection, \u003cem\u003eLigase IV\u003c/em\u003e (\u003cem\u003eLIG4\u003c/em\u003e) KO to suppress NHEJ, \u003cem\u003ePOLQ\u003c/em\u003e KO to suppress TMEJ, and \u003cem\u003eRAD52\u003c/em\u003e KO to suppress SSA (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eb). We analyzed the mutation pattern in each cell line or condition for three CRISPR-Cas9 targeted genes (\u003cem\u003eHPRT1, ST3GAL4, WISP3\u003c/em\u003e). Consistent with the CRISPRi screen, we found that \u003cem\u003ePOLQ\u003c/em\u003e KO (P-value\u0026thinsp;=\u0026thinsp;1.46x10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e) or the addition of M4344 (P-value\u0026thinsp;=\u0026thinsp;8.83x10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e) significantly reduced large deletion rates, whereas \u003cem\u003eRAD52\u003c/em\u003e KO (P-value\u0026thinsp;=\u0026thinsp;5.64x10\u003csup\u003e\u0026minus;\u0026thinsp;4\u003c/sup\u003e) increased the large deletion rates. In contrast to variable effect of \u003cem\u003eLIG4\u003c/em\u003e knockdown in the CRISPRi screen, the \u003cem\u003eLIG4\u003c/em\u003e KO cell line showed a significant increase in large deletion frequency (P-value\u0026thinsp;=\u0026thinsp;3.06x10\u003csup\u003e\u0026minus;\u0026thinsp;4\u003c/sup\u003e), suggesting that blocking either the end protection process or the subsequent NHEJ can enhance the occurrence of large deletions. Both the CRISPRi and KO experiments indicated that the end resection process and the following TMEJ play major roles in generating large deletions after CRISPR-mediated gene editing.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n\u003ch2\u003eA strategy to reduce large deletions in primary T cells\u003c/h2\u003e\n\u003cp\u003eAs a potential application in T cell engineering, we examined whether M4434 reduced large deletion frequencies in human primary T cells. After introducing CRISPR-Cas9 by electroporation into human T cells, we exposed the cells to various concentrations of M4434 (0 nM, 1 nM, 5 nM, 10 nM, 25 nM, and 50 nM). Cell viability and number were substantially decreased when the M4344 concentration was over 25 nM, and the frequency of large deletions was reduced to similar amounts at M4344 concentrations from 10 nM to 50 nM (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;7\u003c/strong\u003e). Therefore, we selected 10 nM as the optimal concentration M4434 to assess the effect on large deletion events in human T cells.\u003c/p\u003e\n\u003cp\u003eWe analyzed the effect of 10 nM M4344 on the CRISPR-Cas9-induced mutation pattern on \u003cem\u003eTRAC\u003c/em\u003e and \u003cem\u003ePD1\u003c/em\u003e in human T cells acquired from two healthy donors (ASAN99 and ASAN107). For this analysis we used the optimized long-range amplicon sequencing method. M4344 decreased the frequency of large deletion events between 35% and 80%, and the reductions were significant: \u003cem\u003eTRAC\u003c/em\u003e P-value\u0026thinsp;=\u0026thinsp;5.41x10\u003csup\u003e-6\u003c/sup\u003e and \u003cem\u003ePD1\u003c/em\u003e P-value\u0026thinsp;=\u0026thinsp;8.52x10\u003csup\u003e-3\u003c/sup\u003e in ASAN99, \u003cem\u003eTRAC\u003c/em\u003e P-value\u0026thinsp;=\u0026thinsp;1.14x10\u003csup\u003e-5\u003c/sup\u003e and \u003cem\u003ePD1\u003c/em\u003e P-value\u0026thinsp;=\u0026thinsp;8.03x10\u003csup\u003e-4\u003c/sup\u003e in ASAN99 (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003ef, \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eg and \u003cstrong\u003eSupplementary Figs.\u0026nbsp;7d, 7e\u003c/strong\u003e). TMEJ-mediated DNA repair is mediated by microhomology, meaning that repair can occur with as few as 2\u0026ndash;16 nucleotides of homologous sequence\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. We found that the microhomology-mediated deletion frequencies were also decreased (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;8\u003c/strong\u003e), indicating that M4344 prevents the large deletion events generated by TMEJ. Collectively, these data support evaluation of M4344 or drugs that reduce end resection or TMEJ pathways as co-effectors with CRISPR-Cas9-mediated genetic engineering in therapeutic applications to mitigate unintended large deletions.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003eGeneration of large DNA deletions by cytosine base editors and adenine base editors\u003c/h2\u003e\n\u003cp\u003eIn contrast with Cas9 nucleases, BEs employ nCas9 containing the D10A mutation, which is impaired in the ability to generate DNA DSBs. Instead of introducing DSBs at target sites, BEs modify single nucleotides in the DNA sequence. BEs, including cytosine base editor (CBE) and adenine base editor (ABE), introduce infrequent indels\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e. Thus, it is possible for BEs to generate large deletions. It was reported that CBE variants without a uracil glycosylase inhibitor (UGI) have relatively high indel frequencies\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e, thus we hypothesized that BE platforms without UGI induce large deletions with a high frequency. The role of the UGI is to inhibit cellular uracil DNA glycosylase (UNG), the enzyme that excises uracil\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e43\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. After uracil excision, cellular DNA AP lyase introduces a nick on the non-target strand of the sgRNA, leading to BER\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. Therefore, we predicted that UGI-lacking BEs generate DSBs and consequently large deletions during the BER process through a nick on each strand\u0026mdash; one by nCas9 (D10A) and one by AP lyase. ABE variants also induce cytosine editing in preferred motifs, suggesting the possibility of ABE introducing large deletions during bystander cytosine editing\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eWe measured large deletion frequencies with six different BE platforms: a canonical CBE containing UGI\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e, a CBE variant without UGI [named CBE (∆UGI)], a CBE variant without UGI and with additional UNG (named CGBE1)\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e, a canonical ABE\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, an ABE variant with UGI (named ABE-UGI), and an ABE variant containing engineered hypoxanthine excision protein N-methylpurine DNA glycosylase (MPG) (named AYBE)\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ea). We introduced each BE platform, or inactivated Cas9 (dCas9) as a control, into HEK293T cells and targeted four endogenous targets (\u003cem\u003eABLIM3, FANCF, KLHL29, BRD4\u003c/em\u003e); then we conducted optimized long-range amplicon sequencing to evaluate large deletion frequencies.\u003c/p\u003e\n\u003cp\u003eEach BE platform generated large deletion mutations; however, the frequency varied both for each platform and each targeted gene (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eb). Compared to the canonical CBE, which had a maximum frequency of large deletion events of 0.8%, CBE (∆UGI) and CGBE1 generated higher frequencies of large deletion events with maximum frequencies of 4.8% and 3.5%, respectively. The canonical ABE had a large deletion frequency maximum of 1.42%. In Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ec, each large deletion frequency was divided by the average large deletion frequency with. Overall, introduction of a UGI slightly reduced the frequency and introduction of the eMPG increased the frequency substantially with a range of 4.41\u0026ndash;1.21% and 2.21\u0026ndash;7.19%, respectively (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ec).\u003c/p\u003e\n\u003cp\u003eFor ABE platforms, we observed that target sequences including a TC motif within the editing window, as in \u003cem\u003eABLIM3\u003c/em\u003e and \u003cem\u003eFANCF\u003c/em\u003e, had larger frequencies of both small indels and large deletions compared to target sequences without a TC motif, as in \u003cem\u003eKLHL29\u003c/em\u003e and \u003cem\u003eBRD4\u003c/em\u003e (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eb and \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ed), indicating that bystander C-to-U conversion by ABE and the resulting AP site produces DSBs and subsequent large deletion events. Both small indels and large deletion events occurred at lower frequencies at targets with TC motifs with ABE-UGI, consistent with the UGI inhibiting the production of DSBs produced through a mechanism similar to that observed for the CBEs (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ed). Compared to targets with a TC motif, those without a TC motif had a lower frequency of either large deletion events or small indels due to AYBE (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eb and \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ed). However, the frequency of large deletions due to AYBE at the target sites without TC motif was substantially higher than that of ABE. Thus, inosine excision by eMPG of AYBE likely enables DSB generation and large deletions through a mechanism like that with the uracil excision by UNG in which Cas9 (D10A) nickase nicks one strand and the BER process nicks the other strand (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003ee).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n\u003ch2\u003eGeneration of large deletion events by prime editors\u003c/h2\u003e\n\u003cp\u003eIn contrast to BEs, PEs employ nCas9 containing H840A mutation and there are two representative PE platforms, PE2 and PE3\u003csup\u003e19\u003c/sup\u003e. PE2 introduces a single nick on the non-target strand of the pegRNA, whereas PE3 introduces a double nick on each strand of the DNA through the addition of a nicking guide RNA (ngRNA)\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Empirically, we and other groups observed that PEs are often accompanied with unwanted small indels\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. In addition, compared with nCas9 (D10A), the cleavage activity of nCas9 (H840A) is not completely impaired and introduces DSBs with a low efficiency\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. The potential for DSBs that result in large deletions is particularly high for PE3\u003csup\u003e21, 51\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eWe measured large deletion frequencies induced by PE2 or PE3 at substitution targets in \u003cem\u003eHEK3\u003c/em\u003e and \u003cem\u003eHEK4\u003c/em\u003e, deletion targets in \u003cem\u003eFANCF\u003c/em\u003e and \u003cem\u003eHEK3\u003c/em\u003e, and insertion targets in \u003cem\u003eRNF2\u003c/em\u003e and \u003cem\u003eHEK3\u003c/em\u003e. As expected, PE2 had a lower frequency of large deletion (maximum 1.2%) than PE3 (maximum 24.3%) (Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ea-\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ec). Similar to the BE platforms, the frequency of PE-induced large deletion events varied by target (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ea). Particularly with PE3, even in the same target, large deletion frequencies varied according to the targeted edit: Compare \u003cem\u003eHEK3\u003c/em\u003e 1-10del and \u003cem\u003eHEK3\u003c/em\u003e 1CTTins, each with +\u0026thinsp;90 ngRNA and \u003cem\u003eHEK4\u003c/em\u003e\u0026thinsp;+\u0026thinsp;2GtoT with the ngRNA\u0026thinsp;+\u0026thinsp;52 or +\u0026thinsp;74 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ea). These results suggested that the sequence in proximity to the prime editing site might contribute to the potential for large deletion events.\u003c/p\u003e\n\u003cp\u003eWe investigated whether the mismatch repair (MMR) pathway or increasing the lifetime of the pegRNA affects the formation of large deletions. To determine the effect of MMR, we introduced PE3 with pegRNA and a MMR inhibitory factor, hMLH1dn\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e. Inhibition of the MMR pathway did not significantly affect the frequency of large deletion events compared with those induced by PE3 with pegRNA (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ed), suggesting that the MMR pathway has a low contribution for the occurrence of PE3-mediated large deletions. Using an established method to increase pegRNA stability, we introduced PE3 with the engineered pegRNA (epegRNA) and found that this resulted in an increase in large deletion frequencies\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ed). We also compared the effect of epegRNA combined with two enhanced PE systems, PE4max and PE5max\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e, on large deletion frequencies induced during insertion of 1 TAC into \u003cem\u003eRNF2\u003c/em\u003e or of 1 CTT into \u003cem\u003eHEK3\u003c/em\u003e, which were sites with relatively low frequencies of large deletion in the PE2 and PE3 systems with pegRNA (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003ea). Compared with the occurrence of large deletion events in cells with dCas9, PE4max with epegRNA did not significantly affect the frequency of large deletion events but PE5max with epegRNA induced a significantly greater frequency of large deletion events (P-value\u0026thinsp;=\u0026thinsp;0.0071) (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;9\u003c/strong\u003e).\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eHere, we refined a long-range amplicon sequencing method to enable simultaneous analysis of both large deletions and small indels with high accuracy. To develop this long-range amplicon sequencing method, we determined the PCR polymerase with minimal length bias, optimized the PCR protocol, and developed a dedicated k-mer sequence-alignment algorithm along with an optimized analysis program to eliminate false-positive large deletion reads (\u003cb\u003eSupplementary Fig.\u0026nbsp;10\u003c/b\u003e). With this method, we detected and quantified large deletions and small indels in all human cells that we tested. Many groups analyze CRISPR-mediated gene editing outcomes within a short range (\u0026lt;\u0026thinsp;500 bp) of the target site, which can miss large deletion events. We experienced mischaracterization of a CRISPR nuclease-mediated knockin cell line (\u003cb\u003eSupplementary Fig.\u0026nbsp;11\u003c/b\u003e). Based on PCR amplification and short-range deep sequencing data, we observed only one pattern for a \u003cem\u003eEXD2\u003c/em\u003e Halo tag knockin cell line, suggesting that the cell line had homologous knockin at both alleles However, with long-range amplicon sequencing, we found that one allele had a large deletion (740 bp). Thus, this is a heterologous knock-in line. A similar phenomenon was observed with wheat genetically engineered for pesticide resistance in which a large deletion altered the epigenetic landscape and changed the growth properties of the plant\u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e. Hence, it is necessary to test for the presence of large deletions after gene editing with CRISPR nucleases.\u003c/p\u003e \u003cp\u003eReducing unintended chromosomal changes in therapeutic applications, such as in generation of CAR T cells, is critical. Tsuchida et al. demonstrated that large deletions and chromosomal truncations can be reduced by altering the step of T cell activation in experimental protocols\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e. We determined that TMEJ serves as a primary repair pathway for generating large deletions. Understanding the mechanism responsible for large deletion events enables development of mitigation strategies. Indeed, we showed that inhibiting the end resection pathway and the subsequent TMEJ pharmacologically with M4344 reduced CRISPR-Cas9-induced large deletion frequencies in human primary T cells.\u003c/p\u003e \u003cp\u003eBEs and PEs that use Cas9 nickase have the potential to generate DSBs at target sites\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e, \u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e, but some studies report that BEs and PEs do not generate large deletions\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. However, our data showed that both BEs and PEs generate large deletions, especially through BER pathway-dependent base editing that creates an AP site for BEs. PE systems that use ngRNA or the stabilized epegRNA showed particularly high frequencies of large deletion.\u003c/p\u003e \u003cp\u003eGene editing tools based on Cas9 or nCas9 exhibited frequencies of large deletion events that varied by target site and for the nCas9-based systems by type of genetic modification. The first-in-kind CRISPR gene editing drug, named Exa-cel or CASGEVY, was approved in UK and USA in 2023. This is remarkably fast for a therapeutic based on CRISPR-Cas nucleases and begins a new era for gene editing therapy. Our findings showed that the present genome-editing tools need additional validation to ensure large deletion events are not present or these tools need additional engineering to prevent the generation of DSBs and large deletions. Whether there are additional shortcomings yet to be discovered for the clinical application of these genome editing tools remains an open question. The workflow that we developed provides a mechanism to evaluate unintended large or small changes to DNA arising from the application of these gene editing tools.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eACKNOWLEDGEMENTS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMost analysis of sequencing data was carried out using the computing server at the Genomic Medicine Institute Research Service Center. This research was supported by grants from the National Research Foundation of Korea (NRF) no. 2021R1A2C3012908, no. 2021M3A9H3015389 to S.B. The authors thank Nancy R. Gough (BioSerendipity, LLC) for editorial service.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAUTHOR CONTRIBUTIONS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eS.B. and G.H. conceived this project; G.H., S.-H.L., and H.S.K., performed and analyzed screening experiments; M.O. and G.H. developed bioinformatics algorithms; G.H., S.-H.L., and O.H. performed cell experiments; S.K. and H.-K.J. performed T-cell experiments; C.H.K., S.K., and S.B. supervised this project; G.H., S.-H.L., and S.B. wrote the manuscript with the help of all other authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHigh-throughput sequencing data have been deposited in the NCBI Sequence Read Archive database (SRA; https://www.ncbi.nlm.nih.gov/sra) under accession number PRJNA1055687. The code of k-mer alignment program is available at https://github.com/ailab-mju/CRISPR-LargeDel.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cp\u003e\u003cstrong\u003eGeneration of sgRNA-encoding plasmids\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTarget sequences were designed using Cas-Designer\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e to avoid off-target sequence (up to 2 mistmaches). The list of oligomers for target sequence is in \u003cstrong\u003eSupplementary Table\u0026nbsp;1\u003c/strong\u003e. pRG2 GG expression vector was digested using \u003cem\u003eBsa1\u003c/em\u003e restriction enzyme. sgRNA oligos, with overhangs complementary to the digested vector, were ordered from Macrogen (Korea) and Cosmogenetech (Korea). These oligos\u0026mdash;comprising both the upper and lower strands\u0026mdash;were then annealed to produce double-stranded oligo deoxynucleotides (dsODN). The annealed dsODNs were ligated into the digested expression vector using T4 DNA ligase (Enzynomics) and incubated for 1h at room temperature. The ligation mixture was transformed into DH5a competent cells using the heat-shock method and cultured overnight at 37\u0026deg;C. Individual colonies were selected and grown in LB media for 16 hours at 37\u0026deg;C in a shaking incubator. The plasmids were isolated using Exprep\u0026trade; Plasmid SV kit (GeneAll).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCell culture and transfection for cancer cell lines and fibroblasts\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHeLa (ATCC\u0026reg;, CCL-2\u0026trade;), HEK293T (ATCC\u0026reg;, CCL-3216\u0026trade;) cells, and U2OS (ATCC HTB-96) cells were maintained in Dulbecco\u0026rsquo;s Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS), 100 unit/mL penicillin, and 100 unit/mL streptomycin. K562 (ATCC CCL-243) cells were maintained in Roswell Park Memorial Institute (RPMI) 1640 Medium supplemented with 10% fetal bovine serum (FBS), 100 unit/mL penicillin, and 100 unit/mL streptomycin. Normal fibroblasts [ThermoFisher, Human Dermal Fibroblasts (C0045C)] were maintained in Dulbecco\u0026rsquo;s Modified Eagle Medium (DMEM) supplemented with 20% fetal bovine serum (FBS), 100 unit/mL penicillin, and 100 unit/mL streptomycin.\u003c/p\u003e\n\u003cp\u003eHeLa and HEK293T were transfected with Lipofectamin 2000 (Invitrogen). Before transfection, 1 \u0026times; 10\u003csup\u003e5\u003c/sup\u003e cells from each well were seeded in 24-well plates. SpCas9 expression plasmids (750 ng) and sgRNA expression plasmids (250 ng) were mixed with 100 \u0026micro;l of Opti-MEM medium and 2 \u0026micro;l of Lipofectamin 2000 and incubated for 20 min at room temperature. The prepared mixture was added to the seeded wells. After 24 hours, the culture media were replaced with fresh media. U2OS and K562 were transfected with Neon transfection system (Invitrogen). cells (2.5 \u0026times; 10\u003csup\u003e5\u003c/sup\u003e) were transfected with 750 ng of Cas9 expression plasmids (750 ng) and sgRNA expression plasmids (250 ng) with the following parameters: 1050 V, 30 ms, 2 pulse for U2OS and 1350 V, 10 ms, 4 pulse for K562. Normal fibroblasts were transfected with Amaxa P3 primary cell 4D-nucleofector kit using program DS-137. All cells were analyzed 3 days after transfection.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCell culture and transfection for H9 cells\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eH9 human embryonic stem cells were maintained in Essential 8 (E8) medium (Gibco A1517001) on iMatrix-511 (Matrixome, 892 021). The dissociation of H9 cells into clusters for subculturing was facilitated using ReLeSR (Stemcell Tech., 05873). Subsequently, the cells were transferred and replated in E8 medium supplemented with p160-Rho-associated coiled-coil kinase (ROCK) inhibitor Y-27632. Before electroporation, TrypLE (Gibco, 12604013) was used to generate a suspension of single cells. Cells (1 \u0026times; 10\u003csup\u003e5\u003c/sup\u003e) were electroporated with 250 ng of sgRNA-encoding plasmid and 750 ng of Cas9 expression plasmid using a NEON system (ThermoFisher) at 1050 V for 30 ms (two pulses). Cells were then seeded in 48-well plates in E8 supplemented with Y-27632 (10 \u0026micro;M) for 24 hours. After three days of culturing, gDNA was isolated.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGeneration of Cas9 ribonucleoprtein (RNP) complexes for CRISPRi screening\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe puromycin targeting sgRNA was synthesized by \u003cem\u003ein vitro\u003c/em\u003e transcription using T7 RNA polymerase (NEB) and template oligos, and the sgRNA product was purified using RNeasy Mini Kit (Qiagen). \u003cem\u003eStreptococcus pyogenes\u003c/em\u003e Cas9 (SpCas9) was ordered from Enzynomics. To generate Cas9 RNP complex, SpCas9 and sgRNA were mixed in a ratio of 1:3 and incubated at room temperature for 30 min. These Cas9 RNP complexes were added to the CRISPRi-stable HeLa cell line after lentiviral transduction and puromycin selection for CRISPRi screening.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIsolation, culture and editing of human primary T cells\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWhole blood samples from healthy donors were taken under a protocol approved by the committee of Asan Medical Center. Peripheral blood mononuclear cells (PBMCs) were isolated from the whole blood samples using SepMate PBMC isolation tubes (STEMCEL). The PBMCs were further processed to isolate human primary T cells using MACS based PAN-T isolation kits (Miltenyi Biotec). RPMI-1640 (Gibco) supplemented with fetal bovine serum (10%, gibco), GlutaMAX (2 mM, gibco), sodium pyruvate (1 mM, gibco), non-essential amino acids (0.1 mM, gibco), beta-mercaptoethanol (55 \u0026micro;M, gibco), HEPES (10 mM, sigma) and penicillin-streptomycin (1%, gibco) was used to culture human primary T cells with IL-2 (300 IU/ml, BMI KOREA).\u003c/p\u003e\n\u003cp\u003eFor gene editing, the human primary T cells were stimulated with Dynabeads human T-Activator CD3/CD28 (Thermo Fischer Scientific) at a cell-to-bead ratio of 1:1 for 48 h. After separating T cells from beads using a magnet, the stimulated T cells were electroporated with Neon transfection system (Thermo Fisher). Briefly, 5 \u0026micro;g of recombinant Cas9 (Enzynomics) and 5 \u0026micro;g of \u003cem\u003ein vitro\u003c/em\u003e-transcribed sgRNA was incubated at 37\u0026deg;C for 10 min to form Cas9 RNP complex, immediately before electroporation. Assembled CRISPR RNPs were added to 0.5\u0026nbsp;million of activated human T cells resuspended in T buffer and electroporated with a Neon electroporation device (1400 V, 10 ms, 3 pulse). Electroporated cells were transferred into culture vessels containing culture medium without antibiotics. One day after electroporation, culture medium was changed into fresh medium containing antibiotics and cells were maintained at a concentration of approximately 1\u0026nbsp;million cells per ml of medium.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLong-range amplicon sequencing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe PCR primers were designed using Primer3Plus\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.bioinformatics.nl/cgi-bin/primer3plus/primer3plus.cgi\u003c/span\u003e\u003c/span\u003e) or Primer-BLAST\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/tools/primer-blast\u003c/span\u003e\u003c/span\u003e). The targeted region (~\u0026thinsp;8 to 15 kb) in gDNA was amplified using KOD multi \u0026amp; epi DNA polymerase (TOYOBO) according to the manufacturer\u0026rsquo;s protocol. For sequences that were difficult to amplify, primers were replaced or gDNA was re-extracted. Amplified products (1 \u0026micro;g) were purified using AMPure XP bead-based reagent (Beckman Coulter) with 0.95X. The purified DNA samples were fragmented to ~\u0026thinsp;300 bp with M220 Focused-ultrasonicator (Covaris) according to the manufacturer\u0026rsquo;s protocol. The fragmented samples were purified with Expin\u0026trade; PCR SV kit (GeneAll) and prepared as an NGS library with NEBNext\u0026reg; Ultra\u0026trade; II DNA Library Prep Kit for Illumina\u0026reg; (NEB). The prepared samples were sequenced with MiniSeq High Output Reagent Kit (300-cycles) using MiniSeq to obtain about\u0026thinsp;~\u0026thinsp;400,000 to 500,000 reads. The NGS data FASTQ files from CRISPR-treated and non-treated cells were analyzed using the k-mer alignment program.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeveloping k-mer alignment program\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe developed a \u003cem\u003ek\u003c/em\u003e-mer based algorithm for detecting CRISPR-induced DNA alterations without requiring a supercomputer and that can run on a personal computer. The \u003cem\u003ek\u003c/em\u003e-mer alignment program is efficient to run with limited computational resources. For the cases investigated here, we used a personal computer with (16 GB) memory and (3.4 GHz / 8 cores) CPU. Our software program accepts as input a 10-kbp reference sequence that includes the cleavage site and the paired-end sequencing data in FASTQ format from both CRISPR-treated and non-treated samples. The output quantifies CRISPR-treated large deletions and small indels and provides read alignment results. The alignment program consists of 3 steps: (i) the short read alignment step, (i) short read classification step, and (iii) removal of false-positive large deletions (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;4\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eThe first step is the short read alignment task based on a \u003cem\u003ek\u003c/em\u003e-mer hash table that is constructed using a reference genome. This hash table is used to obtain positions on the reference genome of (\u003cem\u003el-k\u0026thinsp;+\u0026thinsp;1\u003c/em\u003e) \u003cem\u003ek\u003c/em\u003e-mers for a given read of length \u003cem\u003el\u003c/em\u003e. To determine alignment of given input, our program identifies the longest region of consecutive overlapping \u003cem\u003ek\u003c/em\u003e-mers on the reference sequence through the Longest Increasing Subsequence (LIS) algorithm.\u003c/p\u003e\n\u003cp\u003eThe second step is the short read classification task. Based on the read alignment results, our program categorizes the read pairs into three distinct classes, (i) all read pairs are skewed to the cleavage site, (ii) all read pairs are mapped without splitting and the cleavage site passes between reads, and (iii) one of the reads is split and passes through the cleavage site. Read pairs in category (i) are considered wildtype. Read pairs in categories (ii) and (iii) are considered as potential CRISPR-derived variant candidates and are compiled into a candidate list.\u003c/p\u003e\n\u003cp\u003eThe third step involves the elimination of false-positive large deletions. It is possible for variant candidate read pairs to be present in non-treated samples, primarily due to inherent biases associated with PCR amplification. This phenomenon is not specific to CRISPR-treated samples, but it is a consequence of the characteristics of the reference sequence and affects both CRISPR-treated and non-treated datasets. To reduce the influence of false positives, our program performs a mapping of candidate read pairs to the left and right positions within the reference sequence and identifies clusters of these candidates by implementing \u003cem\u003ek\u003c/em\u003e-means clustering. The appropriate value for \u003cem\u003ek\u003c/em\u003e (the number of clusters) is determined using the Bayesian information criterion, which enables the automated selection of the optimal \u003cem\u003ek\u003c/em\u003e selected based on the characteristics of the dataset. After cluster selection, those clusters that are common to both CRISPR-treated and non-treated datasets are categorized as a result from PCR bias rather than CRISPR-induced variation. These co-occurring clusters are then excluded from the candidate list. Finally, the remaining read pairs in the candidate list are identified as CRISPR-derived variants.\u003c/p\u003e\n\u003cp\u003eRead pairs with a deletion length of 50 base pairs or more within the curated list of candidate variants are determined as a large deletion. Read pairs with insertions or deletions within a range of \u0026plusmn;\u0026thinsp;13 base pairs around the cleavage site are categorized as small indels. To quantify the identified reads, our program enumerates the occurrence of normal mappings, small indels, and large deletions within read pairs passing through the cleavage site (\u0026plusmn;\u0026thinsp;13 base pairs). Our program then calculates the relative frequencies at which large deletions, small insertions, and small deletions occur. Let \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{T}\\)\u003c/span\u003e\u003c/span\u003e be the total number of reads passing the cut site (\u0026plusmn;\u0026thinsp;13 base pairs), including large deletions. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{L}\\)\u003c/span\u003e\u003c/span\u003e is the number of large deletions. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{S}, {N}_{D}\\)\u003c/span\u003e\u003c/span\u003e are the number of reads for small insertions and small deletions. The relative frequency of large deletions calculated by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{L}\\)\u003c/span\u003e\u003c/span\u003e/\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{T}\\)\u003c/span\u003e\u003c/span\u003e. Small indels are calculated by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{S}+ {N}_{D}\\)\u003c/span\u003e\u003c/span\u003e/\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{T}\\)\u003c/span\u003e\u003c/span\u003e. Deletion analysis was performed with the developed \u003cem\u003ek\u003c/em\u003e-mer alignment program using a local computer version with an added web front-end.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCRISPRi library construction\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOur custom CRISPRi library was constructed using pAX198 (Addgene #173042) that includes pU6-sgRNA-EF1a-Puro-T2A-BFP. Using the Repair-seq dataset and genes associated with DNA repair in the Human Protein Atlas website, we selected 794 genes associated with DNA repair and categorized them into specific repair processes or pathways. For each gene, three CRISPRi gRNA sequences were sourced from the hCRISPRi-v2.1 library developed by the Weissman group (\u003cstrong\u003eSupplementary Table\u0026nbsp;2\u003c/strong\u003e) and 60 non-targeting gRNAs from Repair-seq were added in the CRISPRi library. The oligonucleotide library was procured from GenScript (USA) and subsequently amplified employing the Phusion\u0026reg; High-Fidelity DNA Polymerase (NEB). The amplified product and pAX198 plasmids were digested using FastDigest Bpu1102I and FastDigest BstX1 (ThermoFisher). Desired DNA fragments from the digested pAX198 were isolated through a 1% agarose gel and subsequently purified using the Expin\u0026trade; Gel SV kit (GeneAll). The digested oligo library was separated and the required DNA sequence was extracted from a 10% PAGE gel. This gel-extracted DNA was then purified using isopropanol precipitation. The digested oligo library and plasmid backbone were ligated using T4 DNA ligase. The ligation mixture was then purified with AMPure XP beads. The ligated product was transformed into MegaX DH10B T1R Electrocomp\u0026trade; Cells (ThermoFisher) using MicroPulser Electroporator (BioRad). After confirming more than 90,000 colonies, the plasmid library was obtained using NucleoBond Xtra Midi EF kit (Macherey-Nagel). The oligo library was confirmed using nested PCR and Illumina sequencing.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLentivirus preparation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLentivirus was produced in the Lenti-X 23T cell line (Takara Bio). Transfection was carried out with psPAX2 (Addgene #12260), pMD2.G Addgene #12259), and library plasmids using polyethyleneimine (Sigma-Aldrich). The cultured medium was replaced one day post-transfection. Two days later, lentivirus-containing medium was harvested and filtered through a 0.45-\u0026micro;m syringe filter. The lentivirus was concentrated using the Lenti-X concentrator (Takara Bio). The viral titer was determined by performing lentiviral transductions at varying concentrations in a 48-well plate format. After titration, the lentiviral library was aliquoted and stored at -80\u0026deg;C.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCRISPRi screening with Nanopore sequencing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHeLa CRISPRi cells were generated by lentiviral integration (~\u0026thinsp;3 to 5 MOI) using the dCas9-KRAB-blast plasmid (Addgene #89567), followed by single cell isolation. Prior to lentiviral transduction, HeLa CRISPRi cells were cultured at a density of 5 \u0026times; 10\u003csup\u003e5\u003c/sup\u003e cells in a 100-mm dish. The following day, the sgRNA library lentivirus was added to the cultured HeLa CRISPRi cells in the presence of 8 \u0026micro;g/mL polybrene. After 24 hours, the culture medium was replaced with fresh medium. Another 24 hours later, cell selection was initiated with 2 \u0026micro;g/mL puromycin and continued for 2 to 3 days. Post-selection, the culture medium was replaced with fresh medium and the cells were cultured for 6 to 8 days to allow gene repression to occur. Subsequently, Cas9 RNP complex targeting the transduced puromycin resistance gene was transfected into the sgRNA library transduced CRISPRi stable cell line using Neon transfection system (Fig.\u0026nbsp;3\u003cstrong\u003ea\u003c/strong\u003e). Three days post-transfection, gDNA was extracted using the NucleoSpin Blood XL, Maxi kit (Macherey-Nagel). Cell culture was performed whenever the cells had grownto 90% of the cell plate.\u003c/p\u003e\n\u003cp\u003eHalf of the extracted gDNA was amplified to generate fragments of ~\u0026thinsp;5 to 6 kb using the KOD multi \u0026amp; epi DNA polymerase. These amplified fragments were then purified with AMPure XP beads. The purified samples were sequenced on MinION (Oxford Nanopore) using ligation sequencing kit V14 (Oxford Nanopore) and MinION flow cell R10.4.1 (Oxford Nanopore) according to the manufacturer\u0026rsquo;s protocol. The sequencing process ran at a speed of 260 bps, and base calls were made on the resulting data using guppy (Oxford Nanopore) with super high accuracy mode.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAnalysis for CRISPRi screening with Nanopore sequencing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFastq files were aligned to the reference genome using the guppy aligner with default settings. To identify gRNA sequences within the sequencing data, we utilized BWA-mem. A gRNA reference FASTA file was constructed by appending 10 bp from the reference sequence to both ends of the gRNA sequence and the gRNA reference FASTA file was indexed using BWA. The gRNA sequence of each read was obtained and aligned with BWA-mem, applying the parameters \u0026ldquo;-k10 -A4 -B2 -O2\u0026rdquo;. The sequencing results were saved as files according to gRNA. To ensure data accuracy, we discarded reads where sequences downstream of both the gRNA and the the Blue fluorescent protein (BFP) did not align. If the deletion was more than 100 bp and the deletion spanned a region within 100 bp of the cleavage site, the deletion was classified as a large deletion mutation. Because Nanopore sequencing has a bias depending on the length of the DNA fragment (\u003cstrong\u003eSupplementary Fig.\u0026nbsp;1\u003c/strong\u003e), the ratio of the length of the DNA fragment to the reference was used instead of the count so that the ratio would decrease as the length of the deletion became longer. Based on the results of 60 non_targeting gRNAs, the Z-Score of large deletion for each gRNA was calculated.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eKnock-out cell line generation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo generate three different gene (\u003cem\u003eLIG IV\u003c/em\u003e, \u003cem\u003ePOLQ\u003c/em\u003e, \u003cem\u003eRAD52\u003c/em\u003e) knock-out HeLa cell line, 750 ng of Cas9 expression plasmid and 250 ng of sgRNA targeting upstream exon of each gene were transfected with 2\u0026micro;l of Lipofectamine 2000 reagent (Invitrogen) into HeLa cells. After 72 hours, CRISPR-treated HeLa cells were distributed as a single cell into each well of 96-well plates. The cell lines were cultured for two weeks, and each genotype was confirmed by an Illumina Miniseq instrument. The Miniseq results were analyzed using Cas-Analyzer (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.rgenome.net/cas-analyzer/\u003c/span\u003e\u003c/span\u003e)\u003csup\u003e59\u003c/sup\u003e. For the complete knock-out, the cell line harboring a frameshift mutation in both alleles was selected.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eM4344 toxicity and effect on CRISPR-induced large deletion events\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo inhibit the ATR protein in HeLa cells, we used M4344 (Selleckchem, S9639) according to manufacturer\u0026rsquo;s protocol. HeLa cells (1 \u0026times; 10\u003csup\u003e5\u003c/sup\u003e) were seeded in 24-well plates. After 24 hours, the cells were exposed to M4344 (1nM, 5nM, 10nM, 25nM and 50nM) for 1 hour after which 750 ng of Cas9 and 250 ng of sgRNA expression plasmids were transfected with 2 \u0026micro;l of lipofectamine 2000 reagent (Invitrogen) into HeLa cells. After 72 hours from transfection, the cells are detached for genomic DNA extraction. Since the chemical was treated to the cells, the cells were maintained with chemical-containing media.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMicrohomology-dependent deletion events\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe alignment information about deletion reads was extracted from alignment results SAM files from k-mer alignment analysis program. Deletion position and the near sequence were calculated using the alignment information. The homology length was calculated by comparing the both sequences in 1bp increments from the start or end position of the deletion sites. If the homology length is 2\u0026ndash;16 bp, the reads were categorized as microhomology-dependent.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTransfection for base editing and prime editing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHEK293T cells (1 x 10\u003csup\u003e5\u003c/sup\u003e cells per well) were cultivated in a 24-well plate for 24 hours. A mixture of 0.5 \u0026micro;l jetOPTIMUS reagent (Polyplus, 101000006), 500ng plasmid DNA (375ng BE expression plasmid and 125 ng sgRNA expression plasmid) or 543 ng plasmid DNA (365 ng PE expression plasmid, 125 ng pegRNA expression plasmid and 43 ng ngRNA expression plasmid) were added to the cells. After 72 hours, gDNA was isolated.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eHsu, P.D., Lander, E.S. \u0026amp; Zhang, F. Development and applications of CRISPR-Cas9 for genome engineering. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e157\u003c/strong\u003e, 1262-1278 (2014).\u003c/li\u003e\n\u003cli\u003eDoudna, J.A. \u0026amp; Charpentier, E. Genome editing. The new frontier of genome engineering with CRISPR-Cas9. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e346\u003c/strong\u003e, 1258096 (2014).\u003c/li\u003e\n\u003cli\u003eGarneau, J.E. et al. The CRISPR/Cas bacterial immune system cleaves bacteriophage and plasmid DNA. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e468\u003c/strong\u003e, 67-71 (2010).\u003c/li\u003e\n\u003cli\u003eSfeir, A. \u0026amp; Symington, L.S. Microhomology-Mediated End Joining: A Back-up Survival Mechanism or Dedicated Pathway? \u003cem\u003eTrends Biochem Sci\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 701-714 (2015).\u003c/li\u003e\n\u003cli\u003eRamsden, D.A., Carvajal-Garcia, J. \u0026amp; Gupta, G.P. Mechanism, cellular functions and cancer roles of polymerase-theta-mediated DNA end joining. \u003cem\u003eNat Rev Mol Cell Biol\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 125-140 (2022).\u003c/li\u003e\n\u003cli\u003eZhao, B., Rothenberg, E., Ramsden, D.A. \u0026amp; Lieber, M.R. The molecular basis and disease relevance of non-homologous DNA end joining. \u003cem\u003eNat Rev Mol Cell Biol\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 765-781 (2020).\u003c/li\u003e\n\u003cli\u003eSmith, J. et al. Impact of DNA ligase IV on the fidelity of end joining in human cells. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 2157-2167 (2003).\u003c/li\u003e\n\u003cli\u003eBrambati, A., Barry, R.M. \u0026amp; Sfeir, A. DNA polymerase theta (Poltheta) - an error-prone polymerase necessary for genome stability. \u003cem\u003eCurr Opin Genet Dev\u003c/em\u003e \u003cstrong\u003e60\u003c/strong\u003e, 119-126 (2020).\u003c/li\u003e\n\u003cli\u003eElliott, B., Richardson, C. \u0026amp; Jasin, M. Chromosomal translocation mechanisms at intronic alu elements in mammalian cells. \u003cem\u003eMol Cell\u003c/em\u003e \u003cstrong\u003e17\u003c/strong\u003e, 885-894 (2005).\u003c/li\u003e\n\u003cli\u003eYoshimi, K. et al. ssODN-mediated knock-in with CRISPR-Cas for large genomic regions in zygotes. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, 10431 (2016).\u003c/li\u003e\n\u003cli\u003eZhu, Z., Verma, N., Gonzalez, F., Shi, Z.D. \u0026amp; Huangfu, D. A CRISPR/Cas-Mediated Selection-free Knockin Strategy in Human Embryonic Stem Cells. \u003cem\u003eStem Cell Reports\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 1103-1111 (2015).\u003c/li\u003e\n\u003cli\u003eKosicki, M., Tomberg, K. \u0026amp; Bradley, A. Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 765-771 (2018).\u003c/li\u003e\n\u003cli\u003eZuccaro, M.V. et al. Allele-Specific Chromosome Removal after Cas9 Cleavage in Human Embryos. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e183\u003c/strong\u003e, 1650-1664 e1615 (2020).\u003c/li\u003e\n\u003cli\u003eLiu, M. et al. Global detection of DNA repair outcomes induced by CRISPR-Cas9. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, 8732-8742 (2021).\u003c/li\u003e\n\u003cli\u003ePapathanasiou, S. et al. Whole chromosome loss and genomic instability in mouse embryos after CRISPR-Cas9 genome editing. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 5855 (2021).\u003c/li\u003e\n\u003cli\u003eTurchiano, G. et al. Quantitative evaluation of chromosomal rearrangements in gene-edited human stem cells by CAST-Seq. \u003cem\u003eCell Stem Cell\u003c/em\u003e \u003cstrong\u003e28\u003c/strong\u003e, 1136-1147 e1135 (2021).\u003c/li\u003e\n\u003cli\u003eGaudelli, N.M. et al. Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e551\u003c/strong\u003e, 464-471 (2017).\u003c/li\u003e\n\u003cli\u003eKomor, A.C., Kim, Y.B., Packer, M.S., Zuris, J.A. \u0026amp; Liu, D.R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e533\u003c/strong\u003e, 420-+ (2016).\u003c/li\u003e\n\u003cli\u003eAnzalone, A.V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e576\u003c/strong\u003e, 149-157 (2019).\u003c/li\u003e\n\u003cli\u003eSong, Y. et al. Large-Fragment Deletions Induced by Cas9 Cleavage while Not in the BEs System. \u003cem\u003eMol Ther Nucleic Acids\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 523-526 (2020).\u003c/li\u003e\n\u003cli\u003eOwens, D.D.G. et al. Microhomologies are prevalent at Cas9-induced larger deletions. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e47\u003c/strong\u003e, 7402-7417 (2019).\u003c/li\u003e\n\u003cli\u003ePeterka, M. et al. Harnessing DSB repair to promote efficient homology-dependent and -independent prime editing. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 1240 (2022).\u003c/li\u003e\n\u003cli\u003eNewby, G.A. et al. Base editing of haematopoietic stem cells rescues sickle cell disease in mice. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e595\u003c/strong\u003e, 295-302 (2021).\u003c/li\u003e\n\u003cli\u003eAida, T. et al. Prime editing primarily induces undesired outcomes in mice. \u003cem\u003ebioRxiv\u003c/em\u003e, 2020.2008.2006.239723 (2020).\u003c/li\u003e\n\u003cli\u003ePark, S.H. et al. Comprehensive analysis and accurate quantification of unintended large gene modifications induced by CRISPR-Cas9 gene editing. \u003cem\u003eSci Adv\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, eabo7676 (2022).\u003c/li\u003e\n\u003cli\u003eHu, T., Chitnis, N., Monos, D. \u0026amp; Dinh, A. Next-generation sequencing technologies: An overview. \u003cem\u003eHum Immunol\u003c/em\u003e \u003cstrong\u003e82\u003c/strong\u003e, 801-811 (2021).\u003c/li\u003e\n\u003cli\u003eYu, S.C.Y. et al. Comparison of Single Molecule, Real-Time Sequencing and Nanopore Sequencing for Analysis of the Size, End-Motif, and Tissue-of-Origin of Long Cell-Free DNA in Plasma. \u003cem\u003eClin Chem\u003c/em\u003e \u003cstrong\u003e69\u003c/strong\u003e, 168-179 (2023).\u003c/li\u003e\n\u003cli\u003eCheng, C., Fei, Z. \u0026amp; Xiao, P. Methods to improve the accuracy of next-generation sequencing. \u003cem\u003eFront Bioeng Biotechnol\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 982111 (2023).\u003c/li\u003e\n\u003cli\u003eDabney, J. \u0026amp; Meyer, M. Length and GC-biases during sequencing library amplification: a comparison of various polymerase-buffer systems with ancient and modern DNA sequencing libraries. \u003cem\u003eBiotechniques\u003c/em\u003e \u003cstrong\u003e52\u003c/strong\u003e, 87-94 (2012).\u003c/li\u003e\n\u003cli\u003eMizuguchi, H., Nakatsuji, M., Fujiwara, S., Takagi, M. \u0026amp; Imanaka, T. Characterization and application to hot start PCR of neutralizing monoclonal antibodies against KOD DNA polymerase. \u003cem\u003eJ Biochem\u003c/em\u003e \u003cstrong\u003e126\u003c/strong\u003e, 762-768 (1999).\u003c/li\u003e\n\u003cli\u003eTakagi, M. et al. Characterization of DNA polymerase from Pyrococcus sp. strain KOD1 and its application to PCR. \u003cem\u003eAppl Environ Microbiol\u003c/em\u003e \u003cstrong\u003e63\u003c/strong\u003e, 4504-4510 (1997).\u003c/li\u003e\n\u003cli\u003eKim, D., Kim, S., Kim, S., Park, J. \u0026amp; Kim, J.S. Genome-wide target specificities of CRISPR-Cas9 nucleases revealed by multiplex Digenome-seq. \u003cem\u003eGenome Res\u003c/em\u003e \u003cstrong\u003e26\u003c/strong\u003e, 406-415 (2016).\u003c/li\u003e\n\u003cli\u003eJeong, Y.K., Yu, J. \u0026amp; Bae, S. Construction of non-canonical PAM-targeting adenosine base editors by restriction enzyme-free DNA cloning using CRISPR-Cas9. \u003cem\u003eScientific Reports\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 4939 (2019).\u003c/li\u003e\n\u003cli\u003eYoon, H.H. et al. CRISPR-Cas9 Gene Editing Protects from the A53T-SNCA Overexpression-Induced Pathology of Parkinson\u0026apos;s Disease In Vivo. \u003cem\u003eCRISPR J\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 95-108 (2022).\u003c/li\u003e\n\u003cli\u003eStadtmauer, E.A. et al. CRISPR-engineered T cells in patients with refractory cancer. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e367\u003c/strong\u003e (2020).\u003c/li\u003e\n\u003cli\u003eWen, W. et al. Effective control of large deletions after double-strand breaks by homology-directed repair and dsODN insertion. \u003cem\u003eGenome Biol\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 236 (2021).\u003c/li\u003e\n\u003cli\u003eWu, J. et al. CRISPR/Cas9-induced structural variations expand in T lymphocytes in vivo. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e50\u003c/strong\u003e, 11128-11137 (2022).\u003c/li\u003e\n\u003cli\u003eHussmann, J.A. et al. Mapping the genetic landscape of DNA double-strand break repair. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e184\u003c/strong\u003e, 5653-5669 e5625 (2021).\u003c/li\u003e\n\u003cli\u003eGilbert, L.A. et al. CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e154\u003c/strong\u003e, 442-451 (2013).\u003c/li\u003e\n\u003cli\u003eChang, H.H.Y., Pannunzio, N.R., Adachi, N. \u0026amp; Lieber, M.R. Non-homologous DNA end joining and alternative pathways to double-strand break repair. \u003cem\u003eNat Rev Mol Cell Biol\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 495-506 (2017).\u003c/li\u003e\n\u003cli\u003eKurt, I.C. et al. CRISPR C-to-G base editors for inducing targeted DNA transversions in human cells. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 41-46 (2021).\u003c/li\u003e\n\u003cli\u003eJeong, Y.K. et al. Adenine base editor engineering reduces editing of bystander cytosines. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, 1426-1433 (2021).\u003c/li\u003e\n\u003cli\u003eKunz, C., Saito, Y. \u0026amp; Schar, P. DNA Repair in mammalian cells: Mismatched repair: variations on a theme. \u003cem\u003eCell Mol Life Sci\u003c/em\u003e \u003cstrong\u003e66\u003c/strong\u003e, 1021-1038 (2009).\u003c/li\u003e\n\u003cli\u003eMol, C.D. et al. Crystal structure of human uracil-DNA glycosylase in complex with a protein inhibitor: protein mimicry of DNA. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e82\u003c/strong\u003e, 701-708 (1995).\u003c/li\u003e\n\u003cli\u003eHegde, M.L., Hazra, T.K. \u0026amp; Mitra, S. Early steps in the DNA base excision/single-strand interruption repair pathway in mammalian cells. \u003cem\u003eCell Res\u003c/em\u003e \u003cstrong\u003e18\u003c/strong\u003e, 27-47 (2008).\u003c/li\u003e\n\u003cli\u003eKim, H.S., Jeong, Y.K., Hur, J.K., Kim, J.S. \u0026amp; Bae, S. Adenine base editors catalyze cytosine conversions in human cells. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, 1145-1148 (2019).\u003c/li\u003e\n\u003cli\u003eKoblan, L.W. et al. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 843-846 (2018).\u003c/li\u003e\n\u003cli\u003eTong, H. et al. Programmable A-to-Y base editing by fusing an adenine base editor with an N-methylpurine DNA glycosylase. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e41\u003c/strong\u003e, 1080-1084 (2023).\u003c/li\u003e\n\u003cli\u003eHabib, O., Habib, G., Hwang, G.H. \u0026amp; Bae, S. Comprehensive analysis of prime editing outcomes in human embryonic stem cells. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e50\u003c/strong\u003e, 1187-1197 (2022).\u003c/li\u003e\n\u003cli\u003eLee, J. et al. Prime editing with genuine Cas9 nickases minimizes unwanted indels. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 1786 (2023).\u003c/li\u003e\n\u003cli\u003eRan, F.A. et al. Double nicking by RNA-guided CRISPR Cas9 for enhanced genome editing specificity. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e154\u003c/strong\u003e, 1380-1389 (2013).\u003c/li\u003e\n\u003cli\u003eChen, P.J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e184\u003c/strong\u003e, 5635-5652 e5629 (2021).\u003c/li\u003e\n\u003cli\u003eNelson, J.W. et al. Engineered pegRNAs improve prime editing efficiency. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e40\u003c/strong\u003e, 402-410 (2022).\u003c/li\u003e\n\u003cli\u003eLi, S. et al. Genome-edited powdery mildew resistance in wheat without growth penalties. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e602\u003c/strong\u003e, 455-460 (2022).\u003c/li\u003e\n\u003cli\u003eTsuchida, C.A. et al. Mitigation of chromosome loss in clinical CRISPR-Cas9-engineered T cells. \u003cem\u003ebioRxiv\u003c/em\u003e, 2023.2003.2022.533709 (2023).\u003c/li\u003e\n\u003cli\u003ePark, J., Bae, S. \u0026amp; Kim, J.S. Cas-Designer: a web-based tool for choice of CRISPR-Cas9 target sites. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 4014-4016 (2015).\u003c/li\u003e\n\u003cli\u003eUntergasser, A. et al. Primer3Plus, an enhanced web interface to Primer3. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e35\u003c/strong\u003e, W71-74 (2007).\u003c/li\u003e\n\u003cli\u003eYe, J. et al. Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 134 (2012).\u003c/li\u003e\n\u003cli\u003ePark, J., Lim, K., Kim, J.S. \u0026amp; Bae, S. Cas-analyzer: an online tool for assessing genome editing results using NGS data. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cstrong\u003e33\u003c/strong\u003e, 286-288 (2017).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3835370/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3835370/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eCRISPR-Cas9 nucleases are versatile tools for genetic engineering cells and function by producing targeted double-strand breaks (DSBs) in the DNA sequence. However, the unintended production of large deletions (\u0026gt;\u0026thinsp;100 bp) represents a challenge to the effective application of this genome-editing system. We optimized a long-range amplicon sequencing system and developed a k-mer sequence-alignment algorithm to simultaneously detect small DNA alteration events and large DNA deletions. With this workflow, we determined that CRISPR-Cas9 induced large deletions at varying frequencies in cancer cell lines, stem cells, and primary T cells. With CRISPR interference screening, we determined that end resection and the subsequent TMEJ [DNA polymerase theta-mediated end joining] repair process produce most large deletions. Furthermore, base editors and prime editors also generated large deletions despite employing mutated Cas9 \u0026ldquo;nickases\u0026rdquo; that produce single-strand breaks. Our findings reveal an important limitation of current genome-editing tools and identify strategies for mitigating unwanted large deletion events.\u003c/p\u003e","manuscriptTitle":"Detailed mechanisms for unintended large DNA deletions with CRISPR, base editors, and prime editors","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-01-12 05:49:22","doi":"10.21203/rs.3.rs-3835370/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"nature-biomedical-engineering","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"natbiomedeng","sideBox":"Learn more about [Nature Biomedical Engineering](http://www.nature.com/natbiomedeng/)","snPcode":"41551","submissionUrl":"https://mts-natbiomedeng.nature.com/cgi-bin/main.plex","title":"Nature Biomedical Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"c435fdcb-70f8-44ee-923f-262d7c8c4088","owner":[],"postedDate":"January 12th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":28072870,"name":"Biological sciences/Biological techniques/Genetic techniques/Gene targeting/CRISPR-Cas9 genome editing"},{"id":28072871,"name":"Biological sciences/Biological techniques/Genetic engineering"}],"tags":[],"updatedAt":"2024-11-05T08:05:20+00:00","versionOfRecord":{"articleIdentity":"rs-3835370","link":"https://doi.org/10.1038/s41551-024-01277-5","journal":{"identity":"nature-biomedical-engineering","isVorOnly":false,"title":"Nature Biomedical Engineering"},"publishedOn":"2024-11-04 05:00:00","publishedOnDateReadable":"November 4th, 2024"},"versionCreatedAt":"2024-01-12 05:49:22","video":"","vorDoi":"10.1038/s41551-024-01277-5","vorDoiUrl":"https://doi.org/10.1038/s41551-024-01277-5","workflowStages":[]},"version":"v1","identity":"rs-3835370","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3835370","identity":"rs-3835370","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.