A single strand-based library preparation method for unbiased virome characterization

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background The gut virome is an integral component of the gut microbiome, playing a crucial role in maintaining gut health. However, accurately depicting the entire gut virome is challenging due to the inherent diversity of genome types (dsDNA, ssDNA, dsRNA, and ssRNA) and topologies (linear, circular, or fragments), with subsequently biases associated with current sequencing library preparation methods. To overcome these problems and improve reproducibility and comparability across studies, universal or standardized virome sequencing library construction methods are highly needed in the gut virome study. Results We repurposed the ligation-based single-stranded library (SSLR) preparation method for virome studies. We demonstrate that the SSLR method exhibits exceptional efficiency in quantifying viral DNA genomes (both dsDNA and ssDNA) and outperforms existing double-stranded (Nextera) and single-stranded (xGen, MDA + Nextera) library preparation approaches in terms of minimal amplification bias, evenness of coverage, and integrity of assembling viral genomes. The SSLR method can be utilized for the simultaneous library preparation of both DNA and RNA viral genomes. Furthermore, the SSLR method showed its ability to capture highly modified phage genomes, which were often lost using other library preparation approaches. Conclusion We introduce and improve a fast, simple and efficient ligation-based single-stranded DNA library preparation for gut virome study. This method is compatible with Illumina sequencing platforms and only requires ligation reagents within 3 h library preparation, which is similar or even better than the advanced library preparation method (xGen). We hope this method can be further optimized, validated and widely used to make gut virome study more comparable and reproducibly.
Full text 168,966 characters · extracted from preprint-html · click to expand
A single strand-based library preparation method for unbiased virome characterization | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article A single strand-based library preparation method for unbiased virome characterization Xichuan Zhai, Alex Gobbi, Witold Kot, Lukasz Krych, Dennis Sandris Nielsen, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4304844/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 24 Oct, 2024 Read the published version in Microbiome → Version 1 posted 9 You are reading this latest preprint version Abstract Background The gut virome is an integral component of the gut microbiome, playing a crucial role in maintaining gut health. However, accurately depicting the entire gut virome is challenging due to the inherent diversity of genome types (dsDNA, ssDNA, dsRNA, and ssRNA) and topologies (linear, circular, or fragments), with subsequently biases associated with current sequencing library preparation methods. To overcome these problems and improve reproducibility and comparability across studies, universal or standardized virome sequencing library construction methods are highly needed in the gut virome study. Results We repurposed the ligation-based single-stranded library (SSLR) preparation method for virome studies. We demonstrate that the SSLR method exhibits exceptional efficiency in quantifying viral DNA genomes (both dsDNA and ssDNA) and outperforms existing double-stranded (Nextera) and single-stranded (xGen, MDA + Nextera) library preparation approaches in terms of minimal amplification bias, evenness of coverage, and integrity of assembling viral genomes. The SSLR method can be utilized for the simultaneous library preparation of both DNA and RNA viral genomes. Furthermore, the SSLR method showed its ability to capture highly modified phage genomes, which were often lost using other library preparation approaches. Conclusion We introduce and improve a fast, simple and efficient ligation-based single-stranded DNA library preparation for gut virome study. This method is compatible with Illumina sequencing platforms and only requires ligation reagents within 3 h library preparation, which is similar or even better than the advanced library preparation method (xGen). We hope this method can be further optimized, validated and widely used to make gut virome study more comparable and reproducibly. single-stranded library phage mock community ssDNA virome dsDNA virome RNA virome gut virome modification Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction The gut microbiome constitutes a diverse array of microbes, comprising bacteria, archaea, viruses/bacteriophages, and fungi, residing in the human intestine [ 1 ]. Bacteriophages (or phages for short), among them, are believed to be highly abundant constituents of the human microbiome possibly equal in numbers to bacteria. They play important roles in modulating the diversity and abundance of gut bacteria, maintaining a dynamic equilibrium [ 1 , 2 ]. There is growing evidence linking gut virome dysbiosis to human diseases such as inflammatory bowel disease [ 3 – 5 ], severe acute malnutrition [ 6 ], alcoholic liver disease [ 7 ], type 2 diabetes [ 8 ] stunting [ 9 ] and recently also asthma [ 10 ]. Studies have also demonstrated that in fecal microbiota transplantation to treat e.g. recurrent Clostridioides difficile ( rCdiff ) the virome component is important for treatment efficacy and even that sterile filtered feces used for so-called fecal filtrate transplantation is able to cure rCdiff [ 11 – 13 ]. Consequently, there is growing interest in studying the gut virome to understand its role in various diseases and its potential clinical applications. However, this effort is hampered by persistent biases and inconsistencies across studies due to non-standardized and non-optimized pipelines. Every stage of the virome exploration process, from sample collection to bioinformatic analysis, is crucial, with virome characterization of complex microbial communities relying on essential steps such as virus-like particle isolation/purification, viral DNA/RNA purification, and sequencing library preparation and shotgun high-throughput sequencing [ 14 – 17 ]. Furthermore, despite advances during recent years, virus identification is still challenging due to their variability, small genome sizes, and genetic mosaicism. Further, virome studies (distinct from bacteriome studies) present a unique challenge with diverse genome types (dsDNA, ssDNA, dsRNA, and ssRNA) and different topologies (linear, circular, or fragments) [ 18 , 19 ], complicating library preparation and down-stream analysis [ 1 , 3 , 20 ], emphasizing the need for universal or standardized virome sequencing library construction methods for improved reproducibility and comparability across studies. Commonly used library preparation methods, such as randomly amplified and linker-amplified shotgun libraries, have limitations, as these methods are restricted to target only dsDNA viruses and can result in uneven coverage of viral genomes, particularly when the viruses are very different in proportional abundances [ 19 , 21 ]. The transposon-based method, while fast and requiring low-input template, is still restricted to dsDNA [ 22 ]. To overcome these limitations, multiple displacement amplification (MDA) has been utilized, increasing DNA amounts for low biomass virome samples and converting ssDNA into dsDNA for further library preparation. However, MDA distorts the ratios of different viral DNA forms by over-amplifying small circular ssDNA genomes and unevenly amplifying linear genomes [ 23 – 25 ]. More recently, single-stranded based methods for library construction have been applied, such as the xGen ssDNA & Low-Input DNA library preparation kit from IDT (previous name Accel-NGS kit from Swift Biosciences). Though promising, the cost per sample is relatively high and it involves multiple purification steps lowering throughput. Currently, RNA viruses, especially dsRNA viruses, are significantly undersampled due to the instability of RNA and its incompatibility with common DNA sequencing library preparations [ 26 , 27 ]. As a result, DNA and RNA viruses are analyzed separately [ 28 ]. While recent studies have revealed a greater diversity of environmental RNA viruses than previously thought [ 26 , 29 , 30 ], little information is available on human gut RNA viral communities [ 31 ]. The commonly used RNA-Seq technology for investigating RNA viruses requires large amounts of sample material, the preparation process is expensive and time-consuming, and can be further complicated by DNA contamination [ 32 – 34 ]. To address these challenges, we drew inspiration from the Single Reaction Single-stranded LibrarY (SRSLY) method, originally designed for the library preparation of cell-free DNA and oligo sequencing [ 35 ]. We have extended and improved this method for gut virome sequencing (both DNA and RNA) with cost-effectiveness and timesaving in mind. We compared the capability and quantitative accuracy of SRSLY with three different virome library preparation methods (Nextera, MDA amplification and xGen) for Illumina short-read sequencing. We evaluated the types and levels of sequencing bias generated by these protocols using a wide range of DNA/RNA phage mock communities and performed qualitative and quantitative analyses using a diverse mock community with different ratios of 4 different genome types (dsDNA, ssDNA, dsRNA and ssRNA) to further verify the method's applicability for virome studies. Additionally, we assessed reliability, including error rates, composition bias and assembling accuracy of these methods. Finally, we validated the utility and performance of the methods for human fecal virome analysis. Figure 1 | Overview of study workflow. The reproposed library preparation method (SSLR) was indicated with 7 steps (detailed information can be found from Fig. S1 and Supplementary file1 ). Five different libraries were used to prepare and sequence three artificial bacteriophage mocks containing different proportions of the ssDNA phages (phiX174 and M13mp18) mixed with the dsDNA phages. These phage genome abundance values were calculated based on the quantity of dsDNA and ssDNA phages measured Qubit dsDNA (or ssDNA) HS Assay kit. MA: Mock A with a ratio of ~ 90:10 for dsDNA and ssDNA ( Fig. S1 E ); MB: Mock B with a ratio of ~ 50:50 for dsDNA and ssDNA; MC: Mock C with a ratio of ~ 10:90 for dsDNA and ssDNA. MD: Mock D with a ratio of ~ 90:10 for DNA and RNA; ME: Mock E with a ratio of ~ 50:50 for DNA and RNA; MF: Mock F with a ratio of 10:90 for DNA and RNA ( Fig. S1 F ). MG: Mock G contains high modification T4 genome (T4) with equal ratio of all genomes; MH: Mock H contains lower modification T4 genome (T4-c) with equal ratio of all genomes ( Fig. S1 G ). Results SSLR is an effective method for quantifying DNA phage genomes We conducted sequencing library preparation using three DNA mock communities (Mock A, B and C; Fig. 1 , Fig. S1 A and S1E ) containing different ratios of dsDNA to ssDNA phages (10%, 50%, and 90% of dsDNA, Table S1 ) with five different methods, namely, Nextera, MDA_0.5h, MDA_1.5h (0.5h and 1.5h referring to the amplification time), xGen and SSLR. Overall, all the phage genomes were detectable even when the total dsDNA inputs were as low as 0.20 ng in the high ssDNA genomes mock (Mock C, Fig. 2 A). However, the efficiency of quantification varied between methods. Nextera library preparation significantly underestimated ssDNA phages (24 to 35 fold, Fig. 2 A and Table S1 ), where only 7 of 9 genomes could be captured ( Table S2 ), making it unsuitable for ssDNA genomes studies. Although MDA can be very helpful for ssDNA genome studies, we observed a selective amplification bias of ssDNAs (3-fold, Table S1 ) even with short-time amplification (MDA_0.5h). Notably, no significant amplification biases were observed in mocks with a high percentage of ssDNA genomes (≥ 50%, about 1-fold, Table S1 ). In contrast, both SSLR and xGen accurately recovered the percentage of ssDNA genomes when present in high ratios (≥ 50%, Mock B and C), but slightly underestimated ssDNA genomes when present in low ratio inputs (about 10%, Mock A, Fig. S2 A ). Interestingly, we observed quite different quantifications of the T4 genome prepared with different methods (Fig. 2 A, 2 B, 2 D and Table S1 ). Less than 2% of T4 genomes were sequenced in all the MDA-related libraries while 23–34% were reached in the Nextera libraries compared to their corresponding inputs. In contrast, both SSLR and xGen overestimated T4 genomes (1.8 to 3.6 folds, we later found this is due to under-quantification of T4 DNA concentration using Qubit for mock construction, see below), especially in the mock with low T4 genome input (Mock C, Fig. 2 A, 2 D and Table S1 ). We further examined the quantification of viral genomes without considering T4 in the mocks (Fig. 2 C and 2 D) and found that the percentage of each genome was consistently accurate across different mock communities when the libraries were constructed by SSLR or xGen. These methods showed strong linear correlations between expected and observed percentage distributions ( p < 2.2e-16, Fig. 2 D). Genome size and GC-content of each genome had minimal influence on these quantifications ( Fig. S2 B and S2C ), and both methods provided a similar overview of DNA phage communities (Fig. 2 B). While the variations could be minimized for the MDA-related methods without including the T4 genome, the quantifications of ssDNA (especially phiX174) were less consistent than the quantification of dsDNA genomes due to their much higher sequencing depths ( Fig. S3 and Table S1 ). The quantification of phages using the Nextera method appeared to be more consistent when the dsDNA phages were calculated exclusively (Fig. 2 B and 2 C). Figure 2 | Comparison of different library strategies for DNA mock communities . (A) Percentage of phage genomes with T4 generated by different library preparation methods. Different colors indicate different phage genomes. The percentage of phage genomes was calculated by aligning the clean reads to customed database ( https://github.com/XC-Zhai/SSLR/tree/main/Customed_db ) with bowtie2 to assess the relative abundance of each DNA phage (Table 1 ). (B) Principal coordinates analysis (PCoA) plots of bray-Curtis distance matrices. PCoA was used to plot the beta diversity of mock-associated communities using the bray matrix. Different colors indicate different library preparation methods, different shapes indicate different DNA mock communities. The dark-filled shapes display the mock with T4 genome and the non-filled without T4. For each axis, in square brackets, the percentage of variation explained was reported. (C) Percentage of phage genomes without T4. (D) Pearson correlation coefficient (r) and two-tailed p-value between the expected and obtained read distributions (in percentage) in different DNAs phage mock communities with (top panel) or without (bottom panel) phage T4 genome. Five different libraries were used to prepare and sequence three artificial bacteriophage mocks containing different proportions of the ssDNA phage phiX174 and M13mp18 mixed with the dsDNA phage. These phage genome abundance values were calculated based on the quantity of dsDNA and ssDNA phages measured Qubit dsDNA (or ssDNA) HS Assay kit. MA: Mock A with a ratio of ~ 90:10 for dsDNA and ssDNA; MB: Mock B with a ratio of ~ 50:50 for dsDNA and ssDNA; MC: Mock C with a ratio of ~ 10:90 for dsDNA and ssDNA. Table 1 Overview of phage genome characteristics included in the mock communities from the present study Phage Family (Species)* Structure Genomic type Genomic Topology Number of genomic segment(s) Genomic length (bp) GC% C2 **(Ceduovirus c2) Non-enveloped dsDNA Linear 1 22172 36.31% T4 (T4-c) # Straboviridae (Tequatrovirus T4) Non-enveloped dsDNA Linear 1 168903 35.30% Phi29 Salasmaviridae (Salasvirus phi29) Non-enveloped dsDNA Linear 1 19282 40.00% P1 **(Punavirus P1) Non-enveloped dsDNA Linear 2 94800 47.31% T7 Autographiviridae (Teseptimavirus T7) Non-enveloped dsDNA Linear 1 39937 48.40% P35 (Listeria phage P35) Non-enveloped dsDNA Linear 1 35822 40.82% Lambda **(Lambdavirus lambda) Non-enveloped dsDNA Linear 1 48502 49.86% PhiX174 Microviridae (Sinsheimervirus phiX174) Non-enveloped ssDNA (> 85%) Circular 1 5386 44.76% M13mp18 Inoviridae (Inovirus M13) Non-enveloped ssDNA Circular 1 7249 42.27% MS2 Fiersviridae (Emesvirus MS2) Non-enveloped ssRNA Linear 1 3569 52.12% Phi6 Cystoviridae (Cystovirus phi6) Enveloped dsRNA Linear 3 13385 55.82% * Current taxonomy of ICTV217. **- No family levels. # T4-c has the same genome characteristic as T4, but with heavily modified groups. SSLR has less bias and higher accuracy in sequencing DNA phages compared to existing methods We then assessed the efficiency and accuracy of the library prepared using different methods with the presence of the T4 genome. The overall mapping rate to the mock genomes from SSLR method was comparable to xGen but higher than Nextera and MDA-related libraries ( Fig. S4 A ). Furthermore, SSLR had a lower error rate for the entire mock community compared to Nextera and MDA methods, similar to the rate produced by xGen method ( Fig. S4 B and S4C ). When we analyzed individual genomes from SSLR and xGen, and observed mostly even coverage among the genomes, except for sharp spikes near the start and end of each genome ( Fig. S3 ). However, the coverage of ssDNA genomes was significantly increased in MDA-related methods and largely neglected by the Nextera method ( Fig. S3 and Fig. S4 D ). Additionally, the effect of GC content on individual genome profiles (GC bias) varied among different library preparation methods. MDA-related and Nextera libraries showed a similar pattern with coverage increasing as GC content increased. However, coverage biases from SSLR and xGen were not apparently associated to their GC content ( Fig. S5 ). Finally, we evaluated the assembly accuracy and efficiency of the DNA phage genomes in the mock community prepared with different library methods. At a same sequencing depth, the assembly from the SSLR prepared library exhibited perfect quality, including the fully reconstructed phage T4 genome ( Table S3 ). In contrast, the other library methods provided less integrity with large gaps in the reconstructed genomes. Moreover, the hits of assembled contigs to the reference genomes from the SSLR prepared library showed one contig per reference genome (except for phage P1, where SSLR led to 2 contigs for P1 genome, Table S4 ), where the other methods resulted wide range hits from 0 to 22. SSLR can quantify highly modified viral DNA Notably, the T4 genome, in which all cytosines are modified to 5-hydroxymethylcytosine (5-HMC) and further glucosylated (glc-HMC) [ 36 ], showed a distinct coverage pattern when prepared with MDA-related methods. The sequencing depths at all genome positions were quite low regardless of the MDA time ( Fig. S3 ). This finding prompted further investigation by introducing an unmodified C-sites of T4-c genome, which instead have amber mutations in dCTPase and dHMase genes [ 37 ]. Quantification of the T4 genome by Nanodrop showed significantly higher concentration (3.03–3.45 folds) than Qubit measurements of the same samples, while the less modified T4-c genome showed less differences (1.65-fold, Fig. 3 A). MDA amplification of T4 genome resulted in relatively lower amount of product with shorter fragments compared to phage genomes P1, T7 and T4-c (Fig. 3 B and 3 C). Equal amounts of phage genome (11.11% of each, quantified by Nanodrop) was used for preparation of Mock G (with T4) and Mock H (with T4-c), we observed that T4-c still yield lower percentage after sequencing ( Fig.S6 ), but had higher ratios (about 2-fold) than T4 in Nextera, SSLR and xGen method ( Fig. S1 C , Fig. S1 G , Fig. 3 D and Table S5 ). Figure 3 | Quantification and MDA bias for highly modified T4 genome. (A) The concentrations of DNAs used in this study were measured by two different quantification methods (Qubit and Nanodrop). (B) MDA amplification differs in genomes. One and 10 ng of DNA from P1, T4, T4-c and T7 were amplified by 30 min of MDA and purified with Zymol Genomic Purification kit. The y-axis indicates the total yield of amplified MDA products (µg) and measured either by Nanodrop (red) or Qubit (blue). (C) The MDA amplification products (T4, T4-c, P1 and T7) were visualized by Tapestation4200 with genomic ScreenTape. (D) Percentage of phage genomes with T4 or T4-c generated by different library preparation methods. Different colors indicate different phage genomes with equal input (11.11%). MG: Mock G contains high modification T4 genome (T4) with equal ratio of all genomes; MH: Mock H contains lower modification T4 genome (T4-c) with equal ratio of all genomes ( Fig. S1 G ). The percentage of phage genomes was calculated by aligning the clean reads to customed database ( https://github.com/XC-Zhai/SSLR/tree/main/Customed_db ) with bowtie2 to assess the relative abundance of each DNA phage ( Table S5 ). SSLR can simultaneously identify DNA and RNA viruses Next, we applied the SSLR method for simultaneous sequencing of DNA and RNA phage genomes (Table 1 , Fig. 1 , Fig. S1 B and S1F ). We used DMSO or heat treatment to denature the dsRNA to facilitate the reverse transcription (RT) process. We found both treatments including the control (No DMSO) could increase the ratio of ssRNA compared Heat or No Heat, especially for the low ssRNA mock (5% of ssRNA, Mock D) (Fig. 4 A and Fig.S7 ). Although pretreatment with DMSO increased the number of reads of dsRNA compared to non-DMSO treatment in the mocks with less dsRNA (5% ~ 25% of dsRNA, Mock D and E, TableS6 ), it was not as effective as heat treatment in terms of read accuracy and sequence coverage (Fig. 4 B, Fig.S7 and TableS6 ). Figure 4 | Comparison of the efficiency of SSLR on the simultaneous identification of DNA/RNA mock communities. SSLR was used to prepare and sequence 4 artificial virome containing different proportions of the DNA phage and RNA phage. (A) Effect of DMSO and heat treatment on the percentage of 4 phage genomes. The percentage of phage genomes was calculated by aligning the clean reads to customed database ( https://github.com/XC-Zhai/SSLR/tree/main/Customed_db ) with bowtie2 to assess the relative abundance of each DNA phage (TableS6). These phage genome abundance values were calculated based on the quantity of DNA and RNA phages measured by Nanodrop. MD: Mock D with ratio of 90:10 for DNA and RNA; ME: Mock E with ratio of 50:50 for DNA and RNA; MF: Mock F with ratio of 10:90 for DNA and RNA. phi6 has 3 segments, large (6,374 bp), medium (4,063 bp) and small (2,948 bp). (B) Heat treatment does not adversely affect sequencing error rates. The R package ShadowRegression estimates reference-free error rates (inset) based on a transform of the slope of read counts and their ‘shadows’ (main plot line graphs). Shadows (y-axis) are a measure of the variation in read counts across different sequencing runs for the same sample. They are calculated by taking the logarithm of the ratio of read counts in one run to another run. A higher shadow value means a larger difference in read counts between the two runs. Tags (x-axis) are a measure of the abundance of reads for a given nucleotide position in a sample. They are calculated by taking the logarithm of the read count at that position. A higher tag value means a higher number of reads at that position. The figure shows the relationship between shadows and tags for different samples treated with or without Heat/DMSO. The slope of this relationship is used to estimate the sequencing error rate for each sample, which is shown in the inset plots. The figures suggest that DMSO treatment does not affect the sequencing error rate significantly. SSLR provides high quality fecal virome characterization and allows detection of highly modified phage genome Finally, we tested the capability for determining fecal virome composition using the methods described above (Fig. 1 and Fig. S1 D ). Although both MDA based methods yielded more reads that can be assigned to Reference Viral Database (RVDB) (range, 5.77–6.26%, Fig.S8A and TableS7 ) compared to SSLR (0.83–1.20%) and xGen (1.23–1.38%) ( Fig.S8B ), 96% of viral operational taxonomic units (vOTUs) were shared among different methods ( Fig.S8C ). We also observed that although the MDA method resulted in higher alignment rates, it led to lower virome Shannon diversity compared to the other library preparation methods ( Fig.S8D and 8E ). And for both short and long-time MDA, proportions of ssDNA viral families were increased compared to other library preparation methods ( TableS8 ). The SSLR could capture around 2-times less of single genome virome (ssDNA of Microviridae and ssRNA of Virgaviridae ) than xGen method ( TableS8 , Fig. 5 B and Fig.S8F ). Figure 5 | Comparison of different library strategies for fecal virome communities . (A) Principal coordinates analysis (PCoA) plots of Bray-Curtis distance matrices. PCoA was used to plot the beta diversity of viral-associated communities using the bray matrix. Different colors indicate different library preparation methods. For each axis, in square brackets, the percentage of variation explained was reported. (B) Representative taxonomic distribution (relative abundance) of the sequenced virome. The relative distribution is described at the taxonomical level of the family. The taxonomy of contigs was determined by querying the viral contigs against a database containing taxon signature genes for virus orthologous group hosted at www.vogdb.org . The unclassified are the contigs that cannot be assigned to any known viral taxonomy at the family level, the unknown is the contigs that are related to “viral dark matter”. (C) Volcano plot illustrates differential vOTUs were derived from DESeq2 analysis that had more than or equal to two-fold changes in SSLR versus xGen. Each dot represents a vOTU contig and is colored to indicate significance. Grey: not significant (NS), Green – significant by log2 fold change (> 2), Blue: significant by p-value (< 0.01), Red – significant by log2 fold change and p-value. Figure made with EnhancedVolcano implemented in R. (D) Heatmap of the high abundance (RPKM) viral contigs that were highly enriched in SSLR method that does not present in the rest of methods. This figure shows a subset of the data presented in Fig.S8A . Genomic maps of the open reading frames (ORFs) which are predicted by prodigal and then annotated by the blast from the NCBI protein database, the best hits were used for the annotation. Different colors indicate different annotated protein, directional boxes indicate ORFs in the respective orientation (detailed contigs legend can be found from Fig.S9D . Additionally, the different library preparation methods led to differential abundant core species as compared with the Nextera method although the majority of vOTUs are not identified at family level ( Fig.S8F ). Based on the virome abundance (log10(RPKM + 1)), 4 clusters were identified ( Fig.S9A ). Notably, the abundance of cluster 3 from the SSLR method was approximately 10 3 times higher than the other methods (Fig. 5 D and Fig.S9B ). We found that these 34 vOTUs were associated with multiple hits to Sphingomonas phage PAU (previous Myoviridae ) (Fig. 5 C, Fig. 5 D and Fig.S9C ). As it is evident that the cytosines in Sphingomonas phage PAU genome were highly modified [ 38 ], this further confirmed that SSLR method can efficiently include modified viral genome in the sequencing libraries. Discussion The viral community in the human gut is highly complex and likely plays an important role in maintaining gut homeostasis [ 1 , 39 ]. However, whether the true diversity and distribution of gut virome can be accurately determined by traditional nucleic acids library construction and sequencing methods and with what kind and degree of biases are only partly known. In this study, we addressed these problems by introducing a series of phage genome mock communities reflecting different ratios of four types of viral genomes (dsDNA, ssDNA, dsRNA, ssRNA). To ensure accuracy and avoid extraction efficiency variations [ 40 , 41 ], we used phage genome (DNA/RNA quantified by Nanodrop or Qubit) as initial inputs instead of phage particles by plaque assay or SYBR green counts. The single-stranded library preparation method (SSLR) we introduced for gut virome studies proved to be highly accurate, less biased and comparably efficient, with the added benefits of being fast and cost-efficient. SSLR demonstrated excellent efficiency in quantifying both dsDNA and ssDNA genomes, overcoming the major limitation of the Nextera protocol, which only works for dsDNA genomes [ 7 , 42 ]. While MDA can be applied for turning ssDNA into dsDNA before using Nextera, and it is known for its amplification preference for ssDNA genomes [ 23 – 25 ]. In line with the previous studies [ 24 ], we showed that MDA led to 2-3-fold overestimation of ssDNA in low ssDNA mocks (Mock A, around 10% of ssDNA) but not in high ssDNA genome mocks (≥ 50%, Mock B and C). Additionally, the amplification bias appeared to depend on the type of ssDNA genome, as MDA preferentially amplified the alpha3 genome more than PhiX174 and M13mp18, which might be associated to the circular ssDNA of alpha3 being amplified faster than linear ssDNA or the different responses of ssDNA genomes to the heat-denaturing before MDA [ 23 ]. Similar to previous studies we also found that SSLR and xGen slightly underestimated ssDNA in mocks with a low ratio of ssDNA (≤ 10%), this might be due to the ssDNA being more sensitive to shearing during the sonication step compared to dsDNA [ 24 , 43 ]. T4 is an exception among dsDNA genomes in the mocks, as its relative abundance in SSLR and xGen was consistently higher than its input (1.6–3.7 fold) and higher than its abundances in Nextera (1.96–9.59 fold) and MDA (99–586 folds). Interestingly, we observed a significant and consistent difference between the quantification of DNA concentration by Qubit and Nanodrop when measuring T4 genomes. This discrepancy may be attributed to glucosylated 5-hydroxymethyl-cytosine (glu-5HMC) modifications present in the T4 genome [ 36 , 37 ]. These modifications could interfere with or alter the binding modes of the Qubit dye, leading to lower concentration values than those obtained through UV absorbance-based quantification (Nanodrop) [ 44 ]. Considering the high quality of our T4 genome, which was extracted, purified, and recovered from agarose gel (see the method part), the concentration values measured from Nanodrop should be more accurate, as 3 times concentration higher from nanodrop is corresponding to the ~ 3-time higher sequencing output (Fig. 3 A vs Fig. 2 A). Our results also shed light on the reason why the 7-deazaguanine modified Cellulophaga phage phi38:2 (previous Myoviridae) was overestimated in the A-LA method (xGen) and underestimated in MDA [ 24 ], as it was demonstrated recently that this phage has extensive modifications on its genome to protect its DNA from bacterial defense systems [ 45 ]. It was also reported that when measuring potentially highly modified viral DNA using DNA-binding fluorescent dyes, greater caution should be exercised to avoid pitfalls [ 46 ]. In comparison to the T4 genome, the T4-c genome is less modified, and its genome can be sequenced deeper using the Nextera method ( Fig.S6 ). However, in the MDA-related method, the genome yield was not increased for T4-c suggesting the remaining modifications in T4-c genome could still inhibit the MDA process on a similar level. SSLR has an advantage for sequencing highly modified viral genomes (such as T4 with 5-HMC) compared to existing methods, which is further demonstrated by the detection of PAU phage in the fecal virome. We failed to identify PAU-like phage from four gut virome databases (GVD, GPD, MGV and IMG_VR4.1, TableS10 ) [ 47 – 50 ], probably due to the modified phages were not included in the original metagenomic sequencing libraries. This finding suggested that many virome studies may have missed or at least underestimated the existence of viruses with highly modified genomes. Unlike the widely used Nextera library preparation method, which can be affected by genome size and GC-content [ 51 , 52 ], the proposed SSLR protocol exhibits comparable alignment rate, low read errors and provides comprehensive ability for variable genomes with exceptional evenness of coverage and near-complete assembly of phage genomes, as also exemplified for the otherwise “difficult” phage T4 ( TableS11 ). We also observed that SSLR outperforms xGen in several performance metrics (including assembly quality, evenness of genome coverage and error rate, ( Fig. S3 and Table S4 )) probably due to the fact that it does not need synthesizing a second strand for ssDNA viral genomes prior to adaptor ligation, thus retaining the native termini and avoiding possible artifacts or errors presenting during sequencing library preparation [ 35 ]. SSLR also offers the advantage of using cDNA directly, the first-strand products of RT, as ligation templates for sequencing of RNA viruses. This simplifies the RT step and preserves the strand information [ 35 ]. Our DNA and RNA genome mocks tests demonstrated that the RT step did not significantly affect the distribution of DNA viruses, although the ssRNA virus MS2 was overrepresented. This overrepresentation can be attributed to the faster RT conversion rate of ssRNA compared to dsRNA, leading to non-uniformity in read coverage, which is consistent with previous observations [ 53 ]. Furthermore, for denaturing of dsRNA, the heat-shock treatment appears to be a more cost-effective and easier approach compared to DMSO treatment, it not only reduces the denaturation time but also eliminates the necessity to remove DMSO after treatment, thereby conserving more sample [ 53 ]. In conclusion, we introduce an improved, rapid, straightforward, and efficient ligation-based single-stranded DNA library preparation method tailored for virome studies. Materials and Methods Mock communities with phage DNA and RNA genomes Customized bacteriophage mock community samples were prepared, comprising of a mixture of 4 to 9 bacteriophage genomes with different genome sizes (3.5 to 168.9 kbp), genome types (dsDNA, ssDNA, dsRNA and ssRNA), and G + C content (35 to 56%) (Table 1 ). The phages were propagated using the conditions listed ( TableS10 ) and the genomic DNA/RNA extractions were conducted using commercial extraction kits (DNeasy Blood & Tissue Kit for DNA or RNeasy Mini Kit for RNA, both from Qiagen) and purified with GeneJET Genomic DNA Purification Kit (Thermo Scientific, #K0722) for recover the high-quality genome. The quality and integrity of DNA and RNA was checked by agarose gel electrophoresis and Agilent TapeStation 4200 High Sensitivity (HS) RNA ScreenTape. Their concentrations were quantified by dsDNA (or ssDNA) HS Assay Kit or RNA HS assay kit (ThermoFisher Scientific, Waltham, MA, USA) with Qubit4.0, along with Nanodrop1000 (V3.8.1). DNA phages library preparation and sequencing For the DNA phage library preparation, we mixed genomic DNA of 9 phages (7 dsDNAs and 2 ssDNAs) with different ratios (Mock A, Mock B and Mock C) to test 5 different library preparation methods (Fig. 1 , Fig. S1 A, Fig. S1 E and Table S1 ). Nextera XT libraries (Nextera) were prepared according to the manufacturer’s protocol (FC-131-1096) (Illumina). Briefly, 1.0 ng of each DNA library was enzymatically fragmented and tagged by tagmentation. Amplification was performed using Illumina dual index (i7 + i5) adapters and cleaned by AMPure XP bead cleanup (A63880l; Beckman Coulter). Two technical replicates were generated for each biological sample, resulting in 6 Nextera-prepared libraries. The same mixture of phage DNA as mentioned above was amplified by Multiple Displacement Amplification (MDA) using the Genomephi V3 kit (GE Healthcare Life Science, Marlborough, MA, USA) with 10 ng input and incubated at 30°C for 0.5 h (MDA_0.5h) or 1.5 h (MDA_1.5h). Subsequently, the amplified DNA was cleaned using Genomic DNA Clean & Concentrator™ kit (Zymo Research, Irvine, CA, USA). Qubit® 1X dsDNA HS Assay Kit on Qubit Fluorometer (Life Technologies, CA, USA) was used to measure the concentration of the purified DNA. Finally, the library was constructed using the Nextera XT kit (Illumina, San Diego, CA, USA). Two technical replicates were generated for each biological sample, resulting in a total of 12 MDA-prepared libraries. xGen™ ssDNA & Low-Input DNA Library (xGen, previously known as Swift Bioscience Accel-NGS® 1S) was prepared following the manufacturer’s protocol (Catalog # 10009859, Integrated DNA Technologies, IDT). The Mock (15 ng in 20 µl TE buffer) was randomly fragmented by BioRuptor (Diagenode, Liege, Belgium) at default intensity for 15 sec on and 90 sec off, with 7 cycles at 4°C as described by Hoang et al [ 55 ]. Sheared DNAs were firstly subjected to 3’ end tailing and ligation with adapter1, followed by extension and ligation with adapter2. The purified ligation products were indexed with xGen™ Unique Dual Index (UDI) Primers (Catalog # 10005975, IDT) and amplified through a 10-cycle PCR. The amplified products were cleaned up with AMPure XP beads (Beckman Coulter Genomic, CA, USA). For each biological sample, two technical replicates were conducted, resulting in a total of 6 xGen-prepared libraries. For the Single-stranded library preparation (SSLR, detailed in Supplementary file1 ), the same mocks were used. After random fragmentation, the end modified adapters ( Supplementary file2 ) synthesized by IDT (IDT, Leuven, Belgium) were ligated to DNA fragments employing the method from Troll et al [ 35 ]. Subsequently, the ligation mix was purified using MinElute Reaction Kit (Qiagen, Hilden, Germany) and the ligation products were eluted with 23 µl of TE buffer (10 mM, pH8.0). The cleaned ligation products were then amplified and indexed for sequencing with Illumina dual Indexes (i7 + i5) and further purified using AMPure XP beads (Beckman Coulter Genomic, CA, USA). An additional bead cleanup step can be applied to mitigate the formation of byproducts and adapter dimers resulting from self-ligation of the adapters ( Fig.S10 ). Two technical replicates were performed for each biological sample, resulting in a total of 6 SSLR-prepared libraries. A detailed protocol can be found on ( Supplementary file1 ). The DNA concentrations of Nextera-prepared, MDA-prepared, xGen-prepared and SSLR-prepared libraries were measured by Qubit® 1X dsDNA HS Assay Kit on Qubit Fluorometer (Life Technologies, CA, USA) and fragment length distributions were determined using TapeStation 4200 (Agilent, CA, USA). Equal amounts of DNA from each library (except xGen-prepared) were pooled and sequenced using 2×150 bp paired-end settings on an Illumina NextSeq550 platform (Illumina, CA, USA). The xGen-prepared libraries were sequenced separately on an iSeq100 System (Illumina, CA, USA) using 2×150 bp paired-end settings due to their indexes being incompatible with the Illumina index. RNA phages library preparation and sequencing Four types of phage genome (Table 1 , dsDNA, ssDNA, dsRNA and ssRNA) with different ratios were prepared for phage DNA/RNA mock communities (Mock D, Mock E and Mock F). We used three mock communities to determine the applicability of the SSLR method for quantification of RNA and DNA phages simultaneously (Table 1 , Fig. S1 B and Fig. S1 F ). Specifically, 20 µl mock community was treated with DMSO at a final concentration of 50% and incubated at 65°C for 90 min [ 53 ]. The DMSO was subsequently removed using a QIAmp viral RNA mini kit (Qiagen, Hilden, Germany), according to the manufacturer’s instructions. As a control, a DMSO-free treatment was included. Alternatively, the same mock community was denatured at 95°C for 3 min and immediately snap-cooled on ice, maintaining the nucleic acids as single-stranded. Again, a sample without heat treatment was included as a control library. All the DMSO-treated samples, heat-treated samples and their corresponding controls were subjected to reverse transcription in a 20-µl reaction system according to the user guide (SuperScript™ IV VILO™ Master Mix, Invitrogen™) and the reaction mix was purified using AMPure XP beads, and a 20 µl of Tris buffer (10 mM, pH8.0) was used for the final elution. Finally, the library preparations with SSLR were performed according to the above-mentioned steps and sequenced on a NovaSeq6000 platform (Illumina, CA, USA). Sequencing of highly modified phages Two types of DNA phage mock communities (Mock G and Mock H) were prepared with equal input (11.11% for each of the 9 genomes, Fig. S1 C and Fig. S1 G ). Mock G contained a highly modified T4 genome (T4), while Mock H contained a less modified T4-c genome (T4-c). All the libraries were prepared as described above and sequenced with the NovaSeq6000 platform (Illumina, CA, USA). Human fecal virome isolation and sequencing library preparation For environmental sample, a fresh fecal sample was obtained from an anonymous healthy adult (Ethical Committee of the Capital Region of Denmark registration number H-20028549) and thoroughly mixed with SM buffer. Fecal virome isolation and purification were carried out according to our previous method with the following modifications [ 16 ]: the Centriprep 50K was replaced by Centrisart® I centrifugal ultrafiltration unit (MWCO 100 kDa, Sartorius Stedim Biotech GmbH) and the enrichment step was conducted at 2500 × g for 30 min at 4°C. The QIAmp viral RNA mini kit (Qiagen, Hilden, Germany) was used for the extraction of viral DNA/RNA from the concentrated virome solution ( Fig. S1 D ). The library preparations were conducted with the aforementioned methods in triplicate. An equal amount of DNA from each library was pooled and sequenced using 150 bp paired-end settings on an NextSeq550 platform (Illumina, CA, USA). Metavirome sequencing, data pre-processing and data analyses Metavirome reads were quality filtered, trimmed and assembled using a previously published pipeline from GitHub: https://github.com/jcame/virome_analysis‑FOOD . Detailed data analyses for sequencing of the mock communities and fecal virome can be found from Supplementary file3 . For mock community data, plots were generated using the ggplot2 [ 56 ] package in R (v4.2.0). Analyses of viral community α- and β-diversity were performed using packages Phyloseq (v1. 36.0) [ 57 ] and Vegan (v2.6.2) [ 58 ] in R. For α-diversity analyses, indexes of observed taxa and Shannon diversity were calculated with t.test using the package ggsignif (v0.6.3) [ 59 ]. For β-diversity analyses, Bray-Curtis distance metrics were calculated and unconstrained ordination was performed using principal coordinate analysis (PCoA). To identify differentially enriched vira on the summarized family level, DESeq2 (v1.42.0) was adopted [ 60 ]. The results were then visualized in a heatmap using the R package complexheatmap (v2.18.0) [ 61 ]. Estimation of sequencing error rates was conducted according to Wilcox et al [ 53 ]. Coverage biases were visualized using the python script from: https://github.com/padbr/gcbias , with modifications made by adding the Pearson correlation. To visualize the functional genes of annotated viral contigs, DnaFeaturesViewer (v3.1.0) [ 62 ] and R package gggenomes (v 0.9.12.9000) [ 63 ] was used. Abbreviations SSLR single-stranded library Single Reaction Single-stranded LibrarY SRSLY RT Reverse transcription MDA multiple Displacement Amplification PCoA principal coordinates analysis vOTUs viral-operational taxonomic units ORFs open reading frames NCLDV nucleocytoplasmic large DNA viruses LCA Lowest Common Ancestor SGBs species-level genome bins NS not significant gut virome database GVD gut phage database GPD metagenomic gut virus MGV high- Integrated Microbial Genomes /Viruses 4.1 IMG/VR4.1 High Sensitivity HS Unique Dual Index UDI Integrated DNA Technologies IDT Declarations Ethics approval and consent to participate A healthy adult anonymous donor donated fecal samples for the study after being informed orally and in writing and provided written consent (Ethical Committee of the Capital Region of Denmark registration number H-20028549). Consent for publication Not applicable. Funding This work was supported by research grants (23145 and 36242) from VILLUM FONDEN. Availability of data and materials The authors declare that the data supporting the findings of this study are available within the paper and its supplemental information files. All the raw viral metagenome sequences data produced in this study are available through the National Center for Biotechnology Informationʼs Sequence Read Archive under BioProject accession number PRJNA1094595 with submission No. SUB13375064 (DNA mocks), No. SUB14350434 (DNA/RNA mock), No. SUB14350532 (Modification mock), and No. SUB14350824 (fecal virome). Code for metavirome pre-analysis is available from GitHub: https://github.com/jcame/virome_analysis-FOOD. Original python code from (https://github.com/padbr/gcbias) with modifications for fitting the GC bias can be found from GitHub: https://github.com/padbr/gcbias. The read accuracy and mutation analysis can be found on GitHub: https://github.com/awilcox83/dsRNA-sequencing. The public databases for checkV (v1.5) at https://portal.nersc.gov/CheckV, VIBRANT (v1.2.1) at (https://github.com/AnantharamanLab/VIBRANT/tree/master/databases), Virsoter2 at (https://osf.io/u3t4j), VOG217 at (https://fileshare.csb.univie.ac.at/vog/vog217) and Iphop (v1.3.2) at (https://portal.nersc.gov/cfs/m342/iphop/db/) are accessible online. Competing interests The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Authors' Contributions This research was designed and directed by LK, DSN and LD. The experiment was performed by XZ. Data collection was performed by XZ, AG and WK. Data analysis was done by XZ and LD. The manuscript was written and revised by XZ, DSN and LD. All authors read and approved the final manuscript. Acknowledgements We thank the donors and our colleagues at Section for Food Microbiology, Gut Health and Fermentation, Department of Food Science, University of Copenhagen (KU FOOD) for their participation and cooperation. We acknowledge Dr. Samuel Kilcher and Prof. Martin Loessner from Institute of Food, Nutrition, and Health, ETH Zurich for providing phage P35 and its host strain Listeria monocytogenes Mack. We acknowledge Dr. Yuvaraj Bhoobalan-Chitty from Department of Biology, University of Copenhagen (Ole Maaløes Vej 5, 2200 Copenhagen N, Denmark) for providing phage T4-c. XZ was supported by China Scholarship Council (CSC) Grant #201906870027. References Shkoporov, A.N. & Hill, C. Bacteriophages of the Human Gut: The “Known Unknown” of the Microbiome. Cell Host & Microbe 25, 195–209 (2019). Van Belleghem, J.D., Dąbrowska, K., Vaneechoutte, M., Barr, J.J. & Bollyky, P.L. Interactions between Bacteriophage, Bacteria, and the Mammalian Immune System. Viruses 11, 10 (2019). Liang, G., Cobian-Guemes, A.G., Albenberg, L. & Bushman, F. The gut virome in inflammatory bowel diseases. Curr Opin Virol 51, 190–198 (2021). Norman, J.M. et al. Disease-Specific Alterations in the Enteric Virome in Inflammatory Bowel Disease. Cell 160, 447–460 (2015). Fernandes, M.A. et al. Enteric Virome and Bacterial Microbiota in Children With Ulcerative Colitis and Crohn Disease. J Pediatr Gastroenterol Nutr 68, 30–36 (2019). Reyesa, A. et al. Gut DNA viromes of Malawian twins discordant for severe acute malnutrition. P Natl Acad Sci USA 112, 11941–11946 (2015). Duan, Y. et al. Bacteriophage targeting of gut bacterium attenuates alcoholic liver disease. Nature 575, 505–511 (2019). Rasmussen, T.S. et al. Faecal virome transplantation decreases symptoms of type 2 diabetes and obesity in a murine model. Gut 69, 2122–2130 (2020). Mirzaei, M.K. et al. Bacteriophages Isolated from Stunted Children Can Regulate Gut Bacterial Communities in an Age-Specific Manner. Cell Host Microbe 27, 199–212.e5 (2020). Leal Rodríguez, C. et al. The infant gut virome is associated with preschool asthma risk independently of bacteria. Nat Med 30, 138–148 (2024) Ott, S.J. et al. Efficacy of Sterile Fecal Filtrate Transfer for Treating Patients With Clostridium difficile Infection. Gastroenterology 152, 799–811 e7(2017). Kortright, K.E., Chan, B.K., Koff, J.L. & Turner, P.E. Phage Therapy: A Renewed Approach to Combat Antibiotic-Resistant Bacteria. Cell Host Microbe 25, 219–232 (2019). Zuo, T. et al. Bacteriophage transfer during faecal microbiota transplantation in infection is associated with treatment outcome. Gut 67, 634–643 (2018). Kazantseva, J., Malv, E., Kaleda, A., Kallastu, A. & Meikas, A. Optimisation of sample storage and DNA extraction for human gut microbiota studies. BMC Microbiol 21, 158 (2021). Maghini, D.G. et al. Quantifying bias introduced by sample collection in relative and absolute microbiome measurements. Nat Biotechnol 42, 328–338 (2024). Zhai, X. et al. The impact of storage buffer and storage conditions on fecal samples for bacteriophage infectivity and metavirome analyses. Microbiome 11, 193 (2023). Shah, S.A. et al. Expanding known viral diversity in the healthy infant gut. Nat Microbiol 8, 986–998 (2023). Zolfo, M., Vol. Ph.D. University of Trento, Trento; 2020. Unterer, M., Khan Mirzaei, M. & Deng, L. Gut Phage Database: phage mining in the cave of wonders. Signal Transduct Target Ther 6, 193 (2021). Reyes, A., Semenkovich, N.P., Whiteson, K., Rohwer, F. & Gordon, J.I. Going viral: next-generation sequencing applied to phage populations in the human gut. Nat Rev Microbiol 10, 607–617 (2012). Khan Mirzaei, M. et al. Challenges of Studying the Human Virome - Relevant Emerging Technologies. Trends Microbiol 29, 171–181 (2021). Marine, R. et al. Evaluation of a transposase protocol for rapid generation of shotgun high-throughput sequencing libraries from nanogram quantities of DNA. Appl Environ Microbiol 77, 8071–8079 (2011). Kim, K.H. & Bae, J.W. Amplification methods bias metagenomic libraries of uncultured single-stranded and double-stranded DNA viruses. Appl Environ Microbiol 77, 7663–7668 (2011). Roux, S. et al. Towards quantitative viromics for both double-stranded and single-stranded DNA viruses. PeerJ 4, e2777 (2016). Chen, M. et al. Comparison of multiple displacement amplification (MDA) and multiple annealing and looping-based amplification cycles (MALBAC) in single-cell sequencing. PLoS One 9, e114520 (2014). Callanan, J. et al. Expansion of known ssRNA phage genomes: From tens to over a thousand. Sci Adv 6, eaay5981 (2020). Starr, E.P., Nuccio, E.E., Pett-Ridge, J., Banfield, J.F. & Firestone, M.K. Metatranscriptomic reconstruction reveals RNA viruses with the potential to shape carbon cycling in soil. Proc Natl Acad Sci U S A 116, 25900–25908 (2019). Sachsenroder, J. et al. Simultaneous Identification of DNA and RNA Viruses Present in Pig Faeces Using Process-Controlled Deep Sequencing. Plos One 7 (2012). Krishnamurthy, S.R., Janowski, A.B., Zhao, G., Barouch, D. & Wang, D. Hyperexpansion of RNA Bacteriophage Diversity. PLoS Biol 14, e1002409 (2016). Neri, U. et al. Expansion of the global RNA virome reveals diverse clades of bacteriophages. Cell 185, 4023–4037 e4018 (2022). Zhang, T. et al. RNA viral community in human feces: Prevalence of plant pathogenic viruses. Plos Biology 4, 108–118 (2006). Jurasz, H., Pawłowski, T. & Perlejewski, K. Contamination Issue in Viral Metagenomics: Problems, Solutions, and Clinical Perspectives. Front microbiol 12 (2021). Martin, S.A.M., Dehler, C.E. & Krol, E. Transcriptomic responses in the fish intestine. Dev Comp Immunol 64, 103–117 (2016). Ozsolak, F. & Milos, P.M. RNA sequencing: advances, challenges and opportunities. Nat Rev Genet 12, 87–98 (2011). Troll, C.J. et al. A ligation-based single-stranded library preparation method to analyze cell-free DNA and synthetic oligos. BMC Genomics 20, 1023 (2019). Wilson, G.G., Young, K.K.Y., Edlin, G.J. & Konigsberg, W. High-Frequency Generalized Transduction by Bacteriophage-T4. Nature 280, 80–82 (1979). Tao, P., Wu, X.R., Tang, W.C., Zhu, J.G. & Rao, V. Engineering of Bacteriophage T4 Genome Using CRISPR-Cas9. Acs Synth Biol 6, 1952–1961 (2017). Hua, J. in University of Pittsburgh, Vol. Doctoral Dissertation (University of Pittsburgh, University of Pittsburgh; 2016). Guerin, E. & Hill, C. Shining Light on Human Gut Bacteriophages. Front Cell Infect Microbiol 10, 481 (2020). Jakociune, D. & Moodley, A. A Rapid Bacteriophage DNA Extraction Method. Method Protocol 1 (2018). Soleimani-Delfan, A., Bouzari, M. & Wang, R. A rapid competitive method for bacteriophage genomic DNA extraction. J Virol Methods 293 (2021). Teng, Y. et al. Gut bacterial isoamylamine promotes age-related cognitive dysfunction by promoting microglial cell death. Cell Host Microbe 30, 944–960.e8 (2022). Maud, B. et al. An alternative method to multiple displacement amplification for preparing virome DNA in a way adapted for the sequencing of both double-strand and single-strand DNA viruses. bioRxiv , 2022.2012.2012.520144 (2022). Singer, V.L., Jones, L.J., Yue, S.T. & Haugland, R.P. Characterization of PicoGreen reagent and development of a fluorescence-based solution assay for double-stranded DNA quantitation. Anal Biochem 249, 228–238 (1997). Cui, L. et al. Four additional natural 7-deazaguanine derivatives in phages and how to make them. Nucleic Acids Research 51, 9214–9226 (2023). Nakayama, Y., Yamaguchi, H., Einaga, N. & Esumi, M. Pitfalls of DNA Quantification Using DNA-Binding Fluorescent Dyes and Suggested Solutions. Plos One 11 (2016). Gregory, A.C. et al. The Gut Virome Database Reveals Age-Dependent Patterns of Virome Diversity in the Human Gut. Cell Host Microbe 28, 724–740 e728 (2020). Camargo, A.P. et al. IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata. Nucleic Acids Research 51, D733-D743 (2023). Nayfach, S. et al. Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome. Nat Microbiol 6, 960–970 (2021). Camarillo-Guerrero, L.F., Almeida, A., Rangel-Pineros, G., Finn, R.D. & Lawley, T.D. Massive expansion of human gut bacteriophage diversity. Cell 184, 1098-1098-1109.e9 (2021). Sato, M.P. et al. Comparison of the sequencing bias of currently available library preparation kits for Illumina sequencing of bacterial genomes and metagenomes. DNA Res 26, 391–398 (2019). Jones, M.B. et al. Library preparation methodology can influence genomic and functional predictions in human microbiome research. Proc Natl Acad Sci U S A 112, 14024–14029 (2015). Wilcox, A.H., Delwart, E. & Díaz-Muñoz, S.L. Next-generation sequencing of dsRNA is greatly improved by treatment with the inexpensive denaturing reagent DMSO. Microbial genomics 5, e000315 (2019). Robertson, G. et al. De novo assembly and analysis of RNA-seq data. Nat Methods 7, 909–912 (2010). Hoang, M.L. et al. Genome-wide quantification of rare somatic mutations in normal human tissues using massively parallel sequencing. Proc Natl Acad Sci U S A 113, 9846–9851 (2016). Wickham, H. ggplot2: Elegant Graphics for Data Analysis. (Springer-Verlag, New York; 2016). McMurdie, P.J. & Holmes, S. phyloseq: An R Package for Reproducible Interactive Analysis and Graphics of Microbiome Census Data. Plos One 8 (2013). Oksanen, F.J., et al. Vegan: Community Ecology Package. R package Version 2.6-2 (2022). Ahlmann-Eltze, C. & Patil, I. ggsignif: R Package for Displaying Significance Brackets for 'ggplot2'. (PsyArXiv, 2021). Love, M.I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 15 (2014). Gu, Z.G., Eils, R. & Schlesner, M. Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847–2849 (2016). Zulkower, V. & Rosser, S. DNA Features Viewer: a sequence annotation formatting and plotting library for Python. Bioinformatics 36, 4350–4352 (2020). Thomas, H.M., J. A.;Kristina, H. gggenomes: A Grammar of Graphics for Comparative Genomics. R package version 0.9.5.9000 (2022). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 24 Oct, 2024 Read the published version in Microbiome → Version 1 posted Editorial decision: Revision requested 27 Aug, 2024 Reviews received at journal 27 Aug, 2024 Reviews received at journal 27 Aug, 2024 Reviewers agreed at journal 12 Aug, 2024 Reviewers agreed at journal 30 Jul, 2024 Reviewers invited by journal 28 Jul, 2024 Editor assigned by journal 07 May, 2024 Submission checks completed at journal 25 Apr, 2024 First submitted to journal 22 Apr, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4304844","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":295279952,"identity":"a6e1eca5-4942-4809-94ee-d47166d878b2","order_by":0,"name":"Xichuan Zhai","email":"","orcid":"","institution":"University of Copenhagen","correspondingAuthor":false,"prefix":"","firstName":"Xichuan","middleName":"","lastName":"Zhai","suffix":""},{"id":295279953,"identity":"8c2917eb-55cf-4c8c-875d-3efb6ba57d17","order_by":1,"name":"Alex Gobbi","email":"","orcid":"","institution":"University of Copenhagen","correspondingAuthor":false,"prefix":"","firstName":"Alex","middleName":"","lastName":"Gobbi","suffix":""},{"id":295279954,"identity":"ca077e98-96a6-46a9-89c2-2eac9adb5a05","order_by":2,"name":"Witold Kot","email":"","orcid":"","institution":"University of Copenhagen","correspondingAuthor":false,"prefix":"","firstName":"Witold","middleName":"","lastName":"Kot","suffix":""},{"id":295279955,"identity":"85f85ba5-fd30-4258-86aa-b921bfe4505e","order_by":3,"name":"Lukasz Krych","email":"","orcid":"","institution":"University of Copenhagen","correspondingAuthor":false,"prefix":"","firstName":"Lukasz","middleName":"","lastName":"Krych","suffix":""},{"id":295279956,"identity":"dea19d90-4a14-4e65-be79-0739ed8d7c28","order_by":4,"name":"Dennis Sandris Nielsen","email":"","orcid":"","institution":"University of Copenhagen","correspondingAuthor":false,"prefix":"","firstName":"Dennis","middleName":"Sandris","lastName":"Nielsen","suffix":""},{"id":295279958,"identity":"a09f55c7-3d4f-44a3-a86b-b6e136222851","order_by":5,"name":"Ling Deng","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1UlEQVRIiWNgGAWjYBACPmYgwQPE/MwHQIwEIDMNvxY2mBbJtgRitTBAtRgcI1oLO+/hF28q7thtPsZ88MEbhrTEBva0BAIO40uznHPmWfK2Y2zJhnMYchIbeJ4dIKCFx8yYt+1wstn9HjNpHoaKxAaJ9AYitPw7nGzcxkO8FuPHvA2H7QzYwFqADpNII+wwxjnHDidIgP1ikAa07VkCXi38/GeMP7ypOWzP3wYKsYpk2X72NAO8WkAWSQCJRIj7DaAxRQAwfwAS9kQoHAWjYBSMgpEKAH5VP6CH+LZ3AAAAAElFTkSuQmCC","orcid":"","institution":"University of Copenhagen","correspondingAuthor":true,"prefix":"","firstName":"Ling","middleName":"","lastName":"Deng","suffix":""}],"badges":[],"createdAt":"2024-04-22 09:47:42","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4304844/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4304844/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s40168-024-01935-5","type":"published","date":"2024-10-24T15:57:17+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":55731987,"identity":"d6020192-b18b-4999-a5fd-edb8c427f257","added_by":"auto","created_at":"2024-05-02 11:25:11","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":243364,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of study workflow.\u003c/strong\u003e The reproposed library preparation method (SSLR) was indicated with 7 steps (detailed information can be found from \u003cstrong\u003eFig.S1\u003c/strong\u003eand \u003cstrong\u003eSupplementary file1\u003c/strong\u003e). Five different libraries were used to prepare and sequence three artificial bacteriophage mocks containing different proportions of the ssDNA phages (phiX174 and M13mp18) mixed with the dsDNA phages. These phage genome abundance values were calculated based on the quantity of dsDNA and ssDNA phages measured Qubit dsDNA (or ssDNA) HS Assay kit. MA: Mock A with a ratio of ~90:10 for dsDNA and ssDNA (\u003cstrong\u003eFig.S1E\u003c/strong\u003e); MB: Mock B with a ratio of ~50:50 for dsDNA and ssDNA; MC: Mock C with a ratio of ~10:90 for dsDNA and ssDNA. MD: Mock D with a ratio of ~90:10 for DNA and RNA; ME: Mock E with a ratio of ~50:50 for DNA and RNA; MF: Mock F with a ratio of 10:90 for DNA and RNA (\u003cstrong\u003eFig.S1F\u003c/strong\u003e). MG: Mock G contains high modification T4 genome (T4) with equal ratio of all genomes; MH: Mock H contains lower modification T4 genome (T4-c) with equal ratio of all genomes (\u003cstrong\u003eFig.S1G\u003c/strong\u003e).\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-4304844/v1/4c3df9954e28a0d1bf6a6152.png"},{"id":55733124,"identity":"32168bc3-757d-4b7c-a176-b99b724a5ab7","added_by":"auto","created_at":"2024-05-02 11:41:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":626027,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparison of different library strategies for DNA mock communities\u003c/strong\u003e. (A) Percentage of phage genomes with T4 generated by different library preparation methods. Different colors indicate different phage genomes. The percentage of phage genomes was calculated by aligning the clean reads to customed database (https://github.com/XC-Zhai/SSLR/tree/main/Customed_db) with bowtie2 to assess the relative abundance of each DNA phage (\u003cstrong\u003eTable1\u003c/strong\u003e). (B) Principal coordinates analysis (PCoA) plots of bray-Curtis distance matrices. PCoA was used to plot the beta diversity of mock-associated communities using the bray matrix. Different colors indicate different library preparation methods, different shapes indicate different DNA mock communities. The dark-filled shapes display the mock with T4 genome and the non-filled without T4. For each axis, in square brackets, the percentage of variation explained was reported. (C) Percentage of phage genomes without T4. (D) Pearson correlation coefficient (r) and two-tailed p-value between the expected and obtained read distributions (in percentage) in different DNAs phage mock communities with (top panel) or without (bottom panel) phage T4 genome. Five different libraries were used to prepare and sequence three artificial bacteriophage mocks containing different proportions of the ssDNA phage phiX174 and M13mp18 mixed with the dsDNA phage. These phage genome abundance values were calculated based on the quantity of dsDNA and ssDNA phages measured Qubit dsDNA (or ssDNA) HS Assay kit. MA: Mock A with a ratio of ~90:10 for dsDNA and ssDNA; MB: Mock B with a ratio of ~50:50 for dsDNA and ssDNA; MC: Mock C with a ratio of ~10:90 for dsDNA and ssDNA.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-4304844/v1/242fb8b1a63c9b81e40a56fd.png"},{"id":55732818,"identity":"3fdcbcd0-0d3f-4667-ac7a-07d3db3c511d","added_by":"auto","created_at":"2024-05-02 11:33:11","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":302635,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eQuantification and MDA bias for highly modified T4 genome.\u003c/strong\u003e (A) The concentrations of DNAs used in this study were measured by two different quantification methods (Qubit and Nanodrop). (B) MDA amplification differs in genomes. One and 10 ng of DNA from P1, T4, T4-c and T7 were amplified by 30 min of MDA and purified with Zymol Genomic Purification kit. The y-axis indicates the total yield of amplified MDA products (μg) and measured either by Nanodrop (red) or Qubit (blue). (C) The MDA amplification products (T4, T4-c, P1 and T7) were visualized by Tapestation4200 with genomic ScreenTape. (D) Percentage of phage genomes with T4 or T4-c generated by different library preparation methods. Different colors indicate different phage genomes with equal input (11.11%). MG: Mock G contains high modification T4 genome (T4) with equal ratio of all genomes; MH: Mock H contains lower modification T4 genome (T4-c) with equal ratio of all genomes (\u003cstrong\u003eFig.S1G\u003c/strong\u003e). The percentage of phage genomes was calculated by aligning the clean reads to customed database (https://github.com/XC-Zhai/SSLR/tree/main/Customed_db) with bowtie2 to assess the relative abundance of each DNA phage (\u003cstrong\u003eTableS5\u003c/strong\u003e).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-4304844/v1/2d6dc19ce8454318595af7c6.png"},{"id":55731991,"identity":"f804484b-f395-4eeb-b693-f4e9071c79c4","added_by":"auto","created_at":"2024-05-02 11:25:11","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":113245,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparison of the efficiency of SSLR on the simultaneous identification of DNA/RNA mock communities.\u003c/strong\u003e SSLR was used to prepare and sequence 4 artificial virome containing different proportions of the DNA phage and RNA phage. (A) Effect of DMSO and heat treatment on the percentage of 4 phage genomes. The percentage of phage genomes was calculated by aligning the clean reads to customed database (https://github.com/XC-Zhai/SSLR/tree/main/Customed_db) with bowtie2 to assess the relative abundance of each DNA phage (TableS6). These phage genome abundance values were calculated based on the quantity of DNA and RNA phages measured by Nanodrop. MD: Mock D with ratio of 90:10 for DNA and RNA; ME: Mock E with ratio of 50:50 for DNA and RNA; MF: Mock F with ratio of 10:90 for DNA and RNA. phi6 has 3 segments, large (6,374 bp), medium (4,063 bp) and small (2,948 bp). (B) Heat treatment does not adversely affect sequencing error rates. The R package ShadowRegression estimates reference-free error rates (inset) based on a transform of the slope of read counts and their ‘shadows’ (main plot line graphs). Shadows (y-axis) are a measure of the variation in read counts across different sequencing runs for the same sample. They are calculated by taking the logarithm of the ratio of read counts in one run to another run. A higher shadow value means a larger difference in read counts between the two runs. Tags (x-axis) are a measure of the abundance of reads for a given nucleotide position in a sample. They are calculated by taking the logarithm of the read count at that position. A higher tag value means a higher number of reads at that position. The figure shows the relationship between shadows and tags for different samples treated with or without Heat/DMSO. The slope of this relationship is used to estimate the sequencing error rate for each sample, which is shown in the inset plots. The figures suggest that DMSO treatment does not affect the sequencing error rate significantly.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-4304844/v1/e4a61779904a928fd80e8007.png"},{"id":55731993,"identity":"7256c6d2-d889-46fb-8f41-9d0878cd58df","added_by":"auto","created_at":"2024-05-02 11:25:11","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":413366,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparison of different library strategies for fecal virome communities\u003c/strong\u003e. (A) Principal coordinates analysis (PCoA) plots of Bray-Curtis distance matrices. PCoA was used to plot the beta diversity of viral-associated communities using the bray matrix. Different colors indicate different library preparation methods. For each axis, in square brackets, the percentage of variation explained was reported. (B) Representative taxonomic distribution (relative abundance) of the sequenced virome. The relative distribution is described at the taxonomical level of the family. The taxonomy of contigs was determined by querying the viral contigs against a database containing taxon signature genes for virus orthologous group hosted at www.vogdb.org. The unclassified are the contigs that cannot be assigned to any known viral taxonomy at the family level, the unknown is the contigs that are related to “viral dark matter”. (C) Volcano plot illustrates differential vOTUs were derived from DESeq2 analysis that had more than or equal to two-fold changes in SSLR versus xGen. Each dot represents a vOTU contig and is colored to indicate significance. Grey: not significant (NS), Green – significant by log2 fold change (\u0026gt;2), Blue: significant by p-value (\u0026lt;0.01), Red – significant by log2 fold change and p-value. Figure made with EnhancedVolcano implemented in R. (D) Heatmap of the high abundance (RPKM) viral contigs that were highly enriched in SSLR method that does not present in the rest of methods. This figure shows a subset of the data presented in \u003cstrong\u003eFig.S8A\u003c/strong\u003e. Genomic maps of the open reading frames (ORFs) which are predicted by prodigal and then annotated by the blast from the NCBI protein database, the best hits were used for the annotation. Different colors indicate different annotated protein, directional boxes indicate ORFs in the respective orientation (detailed contigs legend can be found from \u003cstrong\u003eFig.S9D\u003c/strong\u003e.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-4304844/v1/1ee15cd96770969a2f5a74de.png"},{"id":67681761,"identity":"a3842e76-a96c-4d5b-91dc-5f07beb788af","added_by":"auto","created_at":"2024-10-28 16:10:04","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2542619,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4304844/v1/b5c55347-d968-43b0-9dcc-15e83046aba9.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A single strand-based library preparation method for unbiased virome characterization","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe gut microbiome constitutes a diverse array of microbes, comprising bacteria, archaea, viruses/bacteriophages, and fungi, residing in the human intestine [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Bacteriophages (or phages for short), among them, are believed to be highly abundant constituents of the human microbiome possibly equal in numbers to bacteria. They play important roles in modulating the diversity and abundance of gut bacteria, maintaining a dynamic equilibrium [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. There is growing evidence linking gut virome dysbiosis to human diseases such as inflammatory bowel disease [\u003cspan additionalcitationids=\"CR4\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], severe acute malnutrition [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], alcoholic liver disease [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], type 2 diabetes [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] stunting [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] and recently also asthma [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Studies have also demonstrated that in fecal microbiota transplantation to treat e.g. recurrent \u003cem\u003eClostridioides difficile\u003c/em\u003e (\u003cem\u003erCdiff\u003c/em\u003e) the virome component is important for treatment efficacy and even that sterile filtered feces used for so-called fecal filtrate transplantation is able to cure \u003cem\u003erCdiff\u003c/em\u003e [\u003cspan additionalcitationids=\"CR12\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Consequently, there is growing interest in studying the gut virome to understand its role in various diseases and its potential clinical applications.\u003c/p\u003e \u003cp\u003eHowever, this effort is hampered by persistent biases and inconsistencies across studies due to non-standardized and non-optimized pipelines. Every stage of the virome exploration process, from sample collection to bioinformatic analysis, is crucial, with virome characterization of complex microbial communities relying on essential steps such as virus-like particle isolation/purification, viral DNA/RNA purification, and sequencing library preparation and shotgun high-throughput sequencing [\u003cspan additionalcitationids=\"CR15 CR16\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Furthermore, despite advances during recent years, virus identification is still challenging due to their variability, small genome sizes, and genetic mosaicism. Further, virome studies (distinct from bacteriome studies) present a unique challenge with diverse genome types (dsDNA, ssDNA, dsRNA, and ssRNA) and different topologies (linear, circular, or fragments) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], complicating library preparation and down-stream analysis [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], emphasizing the need for universal or standardized virome sequencing library construction methods for improved reproducibility and comparability across studies.\u003c/p\u003e \u003cp\u003eCommonly used library preparation methods, such as randomly amplified and linker-amplified shotgun libraries, have limitations, as these methods are restricted to target only dsDNA viruses and can result in uneven coverage of viral genomes, particularly when the viruses are very different in proportional abundances [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. The transposon-based method, while fast and requiring low-input template, is still restricted to dsDNA [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. To overcome these limitations, multiple displacement amplification (MDA) has been utilized, increasing DNA amounts for low biomass virome samples and converting ssDNA into dsDNA for further library preparation. However, MDA distorts the ratios of different viral DNA forms by over-amplifying small circular ssDNA genomes and unevenly amplifying linear genomes [\u003cspan additionalcitationids=\"CR24\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. More recently, single-stranded based methods for library construction have been applied, such as the xGen ssDNA \u0026amp; Low-Input DNA library preparation kit from IDT (previous name Accel-NGS kit from Swift Biosciences). Though promising, the cost per sample is relatively high and it involves multiple purification steps lowering throughput.\u003c/p\u003e \u003cp\u003eCurrently, RNA viruses, especially dsRNA viruses, are significantly undersampled due to the instability of RNA and its incompatibility with common DNA sequencing library preparations [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. As a result, DNA and RNA viruses are analyzed separately [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. While recent studies have revealed a greater diversity of environmental RNA viruses than previously thought [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], little information is available on human gut RNA viral communities [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. The commonly used RNA-Seq technology for investigating RNA viruses requires large amounts of sample material, the preparation process is expensive and time-consuming, and can be further complicated by DNA contamination [\u003cspan additionalcitationids=\"CR33\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo address these challenges, we drew inspiration from the Single Reaction Single-stranded LibrarY (SRSLY) method, originally designed for the library preparation of cell-free DNA and oligo sequencing [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. We have extended and improved this method for gut virome sequencing (both DNA and RNA) with cost-effectiveness and timesaving in mind. We compared the capability and quantitative accuracy of SRSLY with three different virome library preparation methods (Nextera, MDA amplification and xGen) for Illumina short-read sequencing. We evaluated the types and levels of sequencing bias generated by these protocols using a wide range of DNA/RNA phage mock communities and performed qualitative and quantitative analyses using a diverse mock community with different ratios of 4 different genome types (dsDNA, ssDNA, dsRNA and ssRNA) to further verify the method's applicability for virome studies. Additionally, we assessed reliability, including error rates, composition bias and assembling accuracy of these methods. Finally, we validated the utility and performance of the methods for human fecal virome analysis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e \u003cb\u003e| Overview of study workflow.\u003c/b\u003e The reproposed library preparation method (SSLR) was indicated with 7 steps (detailed information can be found from \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e and \u003cb\u003eSupplementary file1\u003c/b\u003e). Five different libraries were used to prepare and sequence three artificial bacteriophage mocks containing different proportions of the ssDNA phages (phiX174 and M13mp18) mixed with the dsDNA phages. These phage genome abundance values were calculated based on the quantity of dsDNA and ssDNA phages measured Qubit dsDNA (or ssDNA) HS Assay kit. MA: Mock A with a ratio of ~\u0026thinsp;90:10 for dsDNA and ssDNA (\u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eE\u003c/b\u003e); MB: Mock B with a ratio of ~\u0026thinsp;50:50 for dsDNA and ssDNA; MC: Mock C with a ratio of ~\u0026thinsp;10:90 for dsDNA and ssDNA. MD: Mock D with a ratio of ~\u0026thinsp;90:10 for DNA and RNA; ME: Mock E with a ratio of ~\u0026thinsp;50:50 for DNA and RNA; MF: Mock F with a ratio of 10:90 for DNA and RNA (\u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eF\u003c/b\u003e). MG: Mock G contains high modification T4 genome (T4) with equal ratio of all genomes; MH: Mock H contains lower modification T4 genome (T4-c) with equal ratio of all genomes (\u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eG\u003c/b\u003e).\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eSSLR is an effective method for quantifying DNA phage genomes\u003c/h2\u003e \u003cp\u003eWe conducted sequencing library preparation using three DNA mock communities (Mock A, B and C; Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eA\u003c/b\u003e and \u003cb\u003eS1E\u003c/b\u003e) containing different ratios of dsDNA to ssDNA phages (10%, 50%, and 90% of dsDNA, \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e) with five different methods, namely, Nextera, MDA_0.5h, MDA_1.5h (0.5h and 1.5h referring to the amplification time), xGen and SSLR. Overall, all the phage genomes were detectable even when the total dsDNA inputs were as low as 0.20 ng in the high ssDNA genomes mock (Mock C, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). However, the efficiency of quantification varied between methods. Nextera library preparation significantly underestimated ssDNA phages (24 to 35 fold, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA and \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e), where only 7 of 9 genomes could be captured (\u003cb\u003eTable\u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e), making it unsuitable for ssDNA genomes studies. Although MDA can be very helpful for ssDNA genome studies, we observed a selective amplification bias of ssDNAs (3-fold, \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e) even with short-time amplification (MDA_0.5h). Notably, no significant amplification biases were observed in mocks with a high percentage of ssDNA genomes (\u0026ge;\u0026thinsp;50%, about 1-fold, \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e). In contrast, both SSLR and xGen accurately recovered the percentage of ssDNA genomes when present in high ratios (\u0026ge;\u0026thinsp;50%, Mock B and C), but slightly underestimated ssDNA genomes when present in low ratio inputs (about 10%, Mock A, \u003cb\u003eFig.\u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003eA\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eInterestingly, we observed quite different quantifications of the T4 genome prepared with different methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD and \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e). Less than 2% of T4 genomes were sequenced in all the MDA-related libraries while 23\u0026ndash;34% were reached in the Nextera libraries compared to their corresponding inputs. In contrast, both SSLR and xGen overestimated T4 genomes (1.8 to 3.6 folds, we later found this is due to under-quantification of T4 DNA concentration using Qubit for mock construction, see below), especially in the mock with low T4 genome input (Mock C, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA, \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD and \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWe further examined the quantification of viral genomes without considering T4 in the mocks (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD) and found that the percentage of each genome was consistently accurate across different mock communities when the libraries were constructed by SSLR or xGen. These methods showed strong linear correlations between expected and observed percentage distributions (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;2.2e-16, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). Genome size and GC-content of each genome had minimal influence on these quantifications (\u003cb\u003eFig.\u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003eB\u003c/b\u003e and \u003cb\u003eS2C\u003c/b\u003e), and both methods provided a similar overview of DNA phage communities (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB). While the variations could be minimized for the MDA-related methods without including the T4 genome, the quantifications of ssDNA (especially phiX174) were less consistent than the quantification of dsDNA genomes due to their much higher sequencing depths (\u003cb\u003eFig.\u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e\u003c/b\u003e and \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e). The quantification of phages using the Nextera method appeared to be more consistent when the dsDNA phages were calculated exclusively (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e \u003cb\u003e| Comparison of different library strategies for DNA mock communities\u003c/b\u003e. (A) Percentage of phage genomes with T4 generated by different library preparation methods. Different colors indicate different phage genomes. The percentage of phage genomes was calculated by aligning the clean reads to customed database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/XC-Zhai/SSLR/tree/main/Customed_db\u003c/span\u003e\u003cspan address=\"https://github.com/XC-Zhai/SSLR/tree/main/Customed_db\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) with bowtie2 to assess the relative abundance of each DNA phage (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). (B) Principal coordinates analysis (PCoA) plots of bray-Curtis distance matrices. PCoA was used to plot the beta diversity of mock-associated communities using the bray matrix. Different colors indicate different library preparation methods, different shapes indicate different DNA mock communities. The dark-filled shapes display the mock with T4 genome and the non-filled without T4. For each axis, in square brackets, the percentage of variation explained was reported. (C) Percentage of phage genomes without T4. (D) Pearson correlation coefficient (r) and two-tailed p-value between the expected and obtained read distributions (in percentage) in different DNAs phage mock communities with (top panel) or without (bottom panel) phage T4 genome. Five different libraries were used to prepare and sequence three artificial bacteriophage mocks containing different proportions of the ssDNA phage phiX174 and M13mp18 mixed with the dsDNA phage. These phage genome abundance values were calculated based on the quantity of dsDNA and ssDNA phages measured Qubit dsDNA (or ssDNA) HS Assay kit. MA: Mock A with a ratio of ~\u0026thinsp;90:10 for dsDNA and ssDNA; MB: Mock B with a ratio of ~\u0026thinsp;50:50 for dsDNA and ssDNA; MC: Mock C with a ratio of ~\u0026thinsp;10:90 for dsDNA and ssDNA.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOverview of phage genome characteristics included in the mock communities from the present study\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePhage\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFamily (Species)*\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eStructure\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGenomic type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eGenomic Topology\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eNumber of genomic segment(s)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eGenomic length (bp)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eGC%\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eC2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e**(Ceduovirus c2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e22172\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e36.31%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eT4 (T4-c)\u003c/b\u003e \u003csup\u003e\u003cb\u003e#\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStraboviridae (Tequatrovirus T4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e168903\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e35.30%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePhi29\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSalasmaviridae (Salasvirus phi29)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e19282\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e40.00%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eP1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e**(Punavirus P1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e94800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e47.31%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eT7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAutographiviridae (Teseptimavirus T7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e39937\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e48.40%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eP35\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(Listeria phage P35)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e35822\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e40.82%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eLambda\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e**(Lambdavirus lambda)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edsDNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e48502\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e49.86%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePhiX174\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMicroviridae (Sinsheimervirus phiX174)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003essDNA (\u0026gt;\u0026thinsp;85%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eCircular\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e5386\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e44.76%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eM13mp18\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eInoviridae (Inovirus M13)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003essDNA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eCircular\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e7249\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e42.27%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMS2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFiersviridae (Emesvirus MS2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-enveloped\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003essRNA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e3569\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e52.12%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePhi6\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCystoviridae (Cystovirus phi6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eEnveloped\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003edsRNA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLinear\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e13385\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e55.82%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003e* Current taxonomy of ICTV217.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003e**- No family levels.\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"8\"\u003e# T4-c has the same genome characteristic as T4, but with heavily modified groups.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eSSLR has less bias and higher accuracy in sequencing DNA phages compared to existing methods\u003c/h2\u003e \u003cp\u003eWe then assessed the efficiency and accuracy of the library prepared using different methods with the presence of the T4 genome. The overall mapping rate to the mock genomes from SSLR method was comparable to xGen but higher than Nextera and MDA-related libraries (\u003cb\u003eFig.\u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003eA\u003c/b\u003e). Furthermore, SSLR had a lower error rate for the entire mock community compared to Nextera and MDA methods, similar to the rate produced by xGen method (\u003cb\u003eFig.\u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003eB\u003c/b\u003e and \u003cb\u003eS4C\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWhen we analyzed individual genomes from SSLR and xGen, and observed mostly even coverage among the genomes, except for sharp spikes near the start and end of each genome (\u003cb\u003eFig.\u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e\u003c/b\u003e). However, the coverage of ssDNA genomes was significantly increased in MDA-related methods and largely neglected by the Nextera method (\u003cb\u003eFig.\u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e\u003c/b\u003e and \u003cb\u003eFig.\u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003eD\u003c/b\u003e). Additionally, the effect of GC content on individual genome profiles (GC bias) varied among different library preparation methods. MDA-related and Nextera libraries showed a similar pattern with coverage increasing as GC content increased. However, coverage biases from SSLR and xGen were not apparently associated to their GC content (\u003cb\u003eFig.\u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eFinally, we evaluated the assembly accuracy and efficiency of the DNA phage genomes in the mock community prepared with different library methods. At a same sequencing depth, the assembly from the SSLR prepared library exhibited perfect quality, including the fully reconstructed phage T4 genome (\u003cb\u003eTable\u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e\u003c/b\u003e). In contrast, the other library methods provided less integrity with large gaps in the reconstructed genomes. Moreover, the hits of assembled contigs to the reference genomes from the SSLR prepared library showed one contig per reference genome (except for phage P1, where SSLR led to 2 contigs for P1 genome, \u003cb\u003eTable\u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e\u003c/b\u003e), where the other methods resulted wide range hits from 0 to 22.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eSSLR can quantify highly modified viral DNA\u003c/h2\u003e \u003cp\u003eNotably, the T4 genome, in which all cytosines are modified to 5-hydroxymethylcytosine (5-HMC) and further glucosylated (glc-HMC) [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], showed a distinct coverage pattern when prepared with MDA-related methods. The sequencing depths at all genome positions were quite low regardless of the MDA time (\u003cb\u003eFig.\u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e\u003c/b\u003e). This finding prompted further investigation by introducing an unmodified C-sites of T4-c genome, which instead have amber mutations in dCTPase and dHMase genes [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. Quantification of the T4 genome by Nanodrop showed significantly higher concentration (3.03\u0026ndash;3.45 folds) than Qubit measurements of the same samples, while the less modified T4-c genome showed less differences (1.65-fold, Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). MDA amplification of T4 genome resulted in relatively lower amount of product with shorter fragments compared to phage genomes P1, T7 and T4-c (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB and \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). Equal amounts of phage genome (11.11% of each, quantified by Nanodrop) was used for preparation of Mock G (with T4) and Mock H (with T4-c), we observed that T4-c still yield lower percentage after sequencing (\u003cb\u003eFig.S6\u003c/b\u003e), but had higher ratios (about 2-fold) than T4 in Nextera, SSLR and xGen method (\u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eC\u003c/b\u003e, \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eG\u003c/b\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD and \u003cb\u003eTable\u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e \u003cb\u003e| Quantification and MDA bias for highly modified T4 genome.\u003c/b\u003e (A) The concentrations of DNAs used in this study were measured by two different quantification methods (Qubit and Nanodrop). (B) MDA amplification differs in genomes. One and 10 ng of DNA from P1, T4, T4-c and T7 were amplified by 30 min of MDA and purified with Zymol Genomic Purification kit. The y-axis indicates the total yield of amplified MDA products (\u0026micro;g) and measured either by Nanodrop (red) or Qubit (blue). (C) The MDA amplification products (T4, T4-c, P1 and T7) were visualized by Tapestation4200 with genomic ScreenTape. (D) Percentage of phage genomes with T4 or T4-c generated by different library preparation methods. Different colors indicate different phage genomes with equal input (11.11%). MG: Mock G contains high modification T4 genome (T4) with equal ratio of all genomes; MH: Mock H contains lower modification T4 genome (T4-c) with equal ratio of all genomes (\u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eG\u003c/b\u003e). The percentage of phage genomes was calculated by aligning the clean reads to customed database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/XC-Zhai/SSLR/tree/main/Customed_db\u003c/span\u003e\u003cspan address=\"https://github.com/XC-Zhai/SSLR/tree/main/Customed_db\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) with bowtie2 to assess the relative abundance of each DNA phage (\u003cb\u003eTable\u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eSSLR can simultaneously identify DNA and RNA viruses\u003c/h2\u003e \u003cp\u003eNext, we applied the SSLR method for simultaneous sequencing of DNA and RNA phage genomes (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eB\u003c/b\u003e and \u003cb\u003eS1F\u003c/b\u003e). We used DMSO or heat treatment to denature the dsRNA to facilitate the reverse transcription (RT) process. We found both treatments including the control (No DMSO) could increase the ratio of ssRNA compared Heat or No Heat, especially for the low ssRNA mock (5% of ssRNA, Mock D) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA and \u003cb\u003eFig.S7\u003c/b\u003e). Although pretreatment with DMSO increased the number of reads of dsRNA compared to non-DMSO treatment in the mocks with less dsRNA (5% ~ 25% of dsRNA, Mock D and E, \u003cb\u003eTableS6\u003c/b\u003e), it was not as effective as heat treatment in terms of read accuracy and sequence coverage (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB, \u003cb\u003eFig.S7\u003c/b\u003e and \u003cb\u003eTableS6\u003c/b\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e \u003cb\u003e| Comparison of the efficiency of SSLR on the simultaneous identification of DNA/RNA mock communities.\u003c/b\u003e SSLR was used to prepare and sequence 4 artificial virome containing different proportions of the DNA phage and RNA phage. (A) Effect of DMSO and heat treatment on the percentage of 4 phage genomes. The percentage of phage genomes was calculated by aligning the clean reads to customed database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/XC-Zhai/SSLR/tree/main/Customed_db\u003c/span\u003e\u003cspan address=\"https://github.com/XC-Zhai/SSLR/tree/main/Customed_db\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) with bowtie2 to assess the relative abundance of each DNA phage (TableS6). These phage genome abundance values were calculated based on the quantity of DNA and RNA phages measured by Nanodrop. MD: Mock D with ratio of 90:10 for DNA and RNA; ME: Mock E with ratio of 50:50 for DNA and RNA; MF: Mock F with ratio of 10:90 for DNA and RNA. phi6 has 3 segments, large (6,374 bp), medium (4,063 bp) and small (2,948 bp). (B) Heat treatment does not adversely affect sequencing error rates. The R package ShadowRegression estimates reference-free error rates (inset) based on a transform of the slope of read counts and their \u0026lsquo;shadows\u0026rsquo; (main plot line graphs). Shadows (y-axis) are a measure of the variation in read counts across different sequencing runs for the same sample. They are calculated by taking the logarithm of the ratio of read counts in one run to another run. A higher shadow value means a larger difference in read counts between the two runs. Tags (x-axis) are a measure of the abundance of reads for a given nucleotide position in a sample. They are calculated by taking the logarithm of the read count at that position. A higher tag value means a higher number of reads at that position. The figure shows the relationship between shadows and tags for different samples treated with or without Heat/DMSO. The slope of this relationship is used to estimate the sequencing error rate for each sample, which is shown in the inset plots. The figures suggest that DMSO treatment does not affect the sequencing error rate significantly.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eSSLR provides high quality fecal virome characterization and allows detection of highly modified phage genome\u003c/h2\u003e \u003cp\u003eFinally, we tested the capability for determining fecal virome composition using the methods described above (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eD\u003c/b\u003e). Although both MDA based methods yielded more reads that can be assigned to Reference Viral Database (RVDB) (range, 5.77\u0026ndash;6.26%, \u003cb\u003eFig.S8A\u003c/b\u003e and \u003cb\u003eTableS7\u003c/b\u003e) compared to SSLR (0.83\u0026ndash;1.20%) and xGen (1.23\u0026ndash;1.38%) (\u003cb\u003eFig.S8B\u003c/b\u003e), 96% of viral operational taxonomic units (vOTUs) were shared among different methods (\u003cb\u003eFig.S8C\u003c/b\u003e). We also observed that although the MDA method resulted in higher alignment rates, it led to lower virome Shannon diversity compared to the other library preparation methods (\u003cb\u003eFig.S8D\u003c/b\u003e and \u003cb\u003e8E\u003c/b\u003e). And for both short and long-time MDA, proportions of ssDNA viral families were increased compared to other library preparation methods (\u003cb\u003eTableS8\u003c/b\u003e). The SSLR could capture around 2-times less of single genome virome (ssDNA of \u003cem\u003eMicroviridae\u003c/em\u003e and ssRNA of \u003cem\u003eVirgaviridae\u003c/em\u003e) than xGen method (\u003cb\u003eTableS8\u003c/b\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB and \u003cb\u003eFig.S8F\u003c/b\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e \u003cb\u003e| Comparison of different library strategies for fecal virome communities\u003c/b\u003e. (A) Principal coordinates analysis (PCoA) plots of Bray-Curtis distance matrices. PCoA was used to plot the beta diversity of viral-associated communities using the bray matrix. Different colors indicate different library preparation methods. For each axis, in square brackets, the percentage of variation explained was reported. (B) Representative taxonomic distribution (relative abundance) of the sequenced virome. The relative distribution is described at the taxonomical level of the family. The taxonomy of contigs was determined by querying the viral contigs against a database containing taxon signature genes for virus orthologous group hosted at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003ca href=\"https://github.com/XC-Zhai/SSLR/tree/main/Customed_db\" target=\"_blank\"\u003ewww.vogdb.org\u003c/a\u003e\u003c/span\u003e\u003cspan address=\"http://www.vogdb.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. The unclassified are the contigs that cannot be assigned to any known viral taxonomy at the family level, the unknown is the contigs that are related to \u0026ldquo;viral dark matter\u0026rdquo;. (C) Volcano plot illustrates differential vOTUs were derived from DESeq2 analysis that had more than or equal to two-fold changes in SSLR versus xGen. Each dot represents a vOTU contig and is colored to indicate significance. Grey: not significant (NS), Green \u0026ndash; significant by log2 fold change (\u0026gt;\u0026thinsp;2), Blue: significant by p-value (\u0026lt;\u0026thinsp;0.01), Red \u0026ndash; significant by log2 fold change and p-value. Figure made with EnhancedVolcano implemented in R. (D) Heatmap of the high abundance (RPKM) viral contigs that were highly enriched in SSLR method that does not present in the rest of methods. This figure shows a subset of the data presented in \u003cb\u003eFig.S8A\u003c/b\u003e. Genomic maps of the open reading frames (ORFs) which are predicted by prodigal and then annotated by the blast from the NCBI protein database, the best hits were used for the annotation. Different colors indicate different annotated protein, directional boxes indicate ORFs in the respective orientation (detailed contigs legend can be found from \u003cb\u003eFig.S9D\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eAdditionally, the different library preparation methods led to differential abundant core species as compared with the Nextera method although the majority of vOTUs are not identified at family level (\u003cb\u003eFig.S8F\u003c/b\u003e). Based on the virome abundance (log10(RPKM\u0026thinsp;+\u0026thinsp;1)), 4 clusters were identified (\u003cb\u003eFig.S9A\u003c/b\u003e). Notably, the abundance of cluster 3 from the SSLR method was approximately 10\u003csup\u003e3\u003c/sup\u003e times higher than the other methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD and \u003cb\u003eFig.S9B\u003c/b\u003e). We found that these 34 vOTUs were associated with multiple hits to Sphingomonas phage PAU (previous \u003cem\u003eMyoviridae\u003c/em\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eC, Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD and \u003cb\u003eFig.S9C\u003c/b\u003e). As it is evident that the cytosines in Sphingomonas phage PAU genome were highly modified [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e], this further confirmed that SSLR method can efficiently include modified viral genome in the sequencing libraries.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe viral community in the human gut is highly complex and likely plays an important role in maintaining gut homeostasis [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. However, whether the true diversity and distribution of gut virome can be accurately determined by traditional nucleic acids library construction and sequencing methods and with what kind and degree of biases are only partly known. In this study, we addressed these problems by introducing a series of phage genome mock communities reflecting different ratios of four types of viral genomes (dsDNA, ssDNA, dsRNA, ssRNA). To ensure accuracy and avoid extraction efficiency variations [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e], we used phage genome (DNA/RNA quantified by Nanodrop or Qubit) as initial inputs instead of phage particles by plaque assay or SYBR green counts. The single-stranded library preparation method (SSLR) we introduced for gut virome studies proved to be highly accurate, less biased and comparably efficient, with the added benefits of being fast and cost-efficient.\u003c/p\u003e \u003cp\u003eSSLR demonstrated excellent efficiency in quantifying both dsDNA and ssDNA genomes, overcoming the major limitation of the Nextera protocol, which only works for dsDNA genomes [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. While MDA can be applied for turning ssDNA into dsDNA before using Nextera, and it is known for its amplification preference for ssDNA genomes [\u003cspan additionalcitationids=\"CR24\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. In line with the previous studies [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], we showed that MDA led to 2-3-fold overestimation of ssDNA in low ssDNA mocks (Mock A, around 10% of ssDNA) but not in high ssDNA genome mocks (\u0026ge;\u0026thinsp;50%, Mock B and C). Additionally, the amplification bias appeared to depend on the type of ssDNA genome, as MDA preferentially amplified the alpha3 genome more than PhiX174 and M13mp18, which might be associated to the circular ssDNA of alpha3 being amplified faster than linear ssDNA or the different responses of ssDNA genomes to the heat-denaturing before MDA [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Similar to previous studies we also found that SSLR and xGen slightly underestimated ssDNA in mocks with a low ratio of ssDNA (\u0026le;\u0026thinsp;10%), this might be due to the ssDNA being more sensitive to shearing during the sonication step compared to dsDNA [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eT4 is an exception among dsDNA genomes in the mocks, as its relative abundance in SSLR and xGen was consistently higher than its input (1.6\u0026ndash;3.7 fold) and higher than its abundances in Nextera (1.96\u0026ndash;9.59 fold) and MDA (99\u0026ndash;586 folds). Interestingly, we observed a significant and consistent difference between the quantification of DNA concentration by Qubit and Nanodrop when measuring T4 genomes. This discrepancy may be attributed to glucosylated 5-hydroxymethyl-cytosine (glu-5HMC) modifications present in the T4 genome [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. These modifications could interfere with or alter the binding modes of the Qubit dye, leading to lower concentration values than those obtained through UV absorbance-based quantification (Nanodrop) [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. Considering the high quality of our T4 genome, which was extracted, purified, and recovered from agarose gel (see the method part), the concentration values measured from Nanodrop should be more accurate, as 3 times concentration higher from nanodrop is corresponding to the ~\u0026thinsp;3-time higher sequencing output (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA vs Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). Our results also shed light on the reason why the 7-deazaguanine modified Cellulophaga phage phi38:2 (previous Myoviridae) was overestimated in the A-LA method (xGen) and underestimated in MDA [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], as it was demonstrated recently that this phage has extensive modifications on its genome to protect its DNA from bacterial defense systems [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. It was also reported that when measuring potentially highly modified viral DNA using DNA-binding fluorescent dyes, greater caution should be exercised to avoid pitfalls [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. In comparison to the T4 genome, the T4-c genome is less modified, and its genome can be sequenced deeper using the Nextera method (\u003cb\u003eFig.S6\u003c/b\u003e). However, in the MDA-related method, the genome yield was not increased for T4-c suggesting the remaining modifications in T4-c genome could still inhibit the MDA process on a similar level.\u003c/p\u003e \u003cp\u003eSSLR has an advantage for sequencing highly modified viral genomes (such as T4 with 5-HMC) compared to existing methods, which is further demonstrated by the detection of PAU phage in the fecal virome. We failed to identify PAU-like phage from four gut virome databases (GVD, GPD, MGV and IMG_VR4.1, \u003cb\u003eTableS10\u003c/b\u003e) [\u003cspan additionalcitationids=\"CR48 CR49\" citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e], probably due to the modified phages were not included in the original metagenomic sequencing libraries. This finding suggested that many virome studies may have missed or at least underestimated the existence of viruses with highly modified genomes.\u003c/p\u003e \u003cp\u003eUnlike the widely used Nextera library preparation method, which can be affected by genome size and GC-content [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e, \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e], the proposed SSLR protocol exhibits comparable alignment rate, low read errors and provides comprehensive ability for variable genomes with exceptional evenness of coverage and near-complete assembly of phage genomes, as also exemplified for the otherwise \u0026ldquo;difficult\u0026rdquo; phage T4 (\u003cb\u003eTableS11\u003c/b\u003e). We also observed that SSLR outperforms xGen in several performance metrics (including assembly quality, evenness of genome coverage and error rate, (\u003cb\u003eFig.\u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e\u003c/b\u003e and \u003cb\u003eTable\u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e\u003c/b\u003e)) probably due to the fact that it does not need synthesizing a second strand for ssDNA viral genomes prior to adaptor ligation, thus retaining the native termini and avoiding possible artifacts or errors presenting during sequencing library preparation [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSSLR also offers the advantage of using cDNA directly, the first-strand products of RT, as ligation templates for sequencing of RNA viruses. This simplifies the RT step and preserves the strand information [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. Our DNA and RNA genome mocks tests demonstrated that the RT step did not significantly affect the distribution of DNA viruses, although the ssRNA virus MS2 was overrepresented. This overrepresentation can be attributed to the faster RT conversion rate of ssRNA compared to dsRNA, leading to non-uniformity in read coverage, which is consistent with previous observations [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. Furthermore, for denaturing of dsRNA, the heat-shock treatment appears to be a more cost-effective and easier approach compared to DMSO treatment, it not only reduces the denaturation time but also eliminates the necessity to remove DMSO after treatment, thereby conserving more sample [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn conclusion, we introduce an improved, rapid, straightforward, and efficient ligation-based single-stranded DNA library preparation method tailored for virome studies.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eMock communities with phage DNA and RNA genomes\u003c/h2\u003e \u003cp\u003eCustomized bacteriophage mock community samples were prepared, comprising of a mixture of 4 to 9 bacteriophage genomes with different genome sizes (3.5 to 168.9 kbp), genome types (dsDNA, ssDNA, dsRNA and ssRNA), and G\u0026thinsp;+\u0026thinsp;C content (35 to 56%) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The phages were propagated using the conditions listed (\u003cb\u003eTableS10\u003c/b\u003e) and the genomic DNA/RNA extractions were conducted using commercial extraction kits (DNeasy Blood \u0026amp; Tissue Kit for DNA or RNeasy Mini Kit for RNA, both from Qiagen) and purified with GeneJET Genomic DNA Purification Kit (Thermo Scientific, #K0722) for recover the high-quality genome. The quality and integrity of DNA and RNA was checked by agarose gel electrophoresis and Agilent TapeStation 4200 High Sensitivity (HS) RNA ScreenTape. Their concentrations were quantified by dsDNA (or ssDNA) HS Assay Kit or RNA HS assay kit (ThermoFisher Scientific, Waltham, MA, USA) with Qubit4.0, along with Nanodrop1000 (V3.8.1).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eDNA phages library preparation and sequencing\u003c/h2\u003e \u003cp\u003eFor the DNA phage library preparation, we mixed genomic DNA of 9 phages (7 dsDNAs and 2 ssDNAs) with different ratios (Mock A, Mock B and Mock C) to test 5 different library preparation methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eA, Fig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eE\u003c/b\u003e and \u003cb\u003eTable\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eNextera XT libraries (Nextera) were prepared according to the manufacturer\u0026rsquo;s protocol (FC-131-1096) (Illumina). Briefly, 1.0 ng of each DNA library was enzymatically fragmented and tagged by tagmentation. Amplification was performed using Illumina dual index (i7\u0026thinsp;+\u0026thinsp;i5) adapters and cleaned by AMPure XP bead cleanup (A63880l; Beckman Coulter). Two technical replicates were generated for each biological sample, resulting in 6 Nextera-prepared libraries.\u003c/p\u003e \u003cp\u003eThe same mixture of phage DNA as mentioned above was amplified by Multiple Displacement Amplification (MDA) using the Genomephi V3 kit (GE Healthcare Life Science, Marlborough, MA, USA) with 10 ng input and incubated at 30\u0026deg;C for 0.5 h (MDA_0.5h) or 1.5 h (MDA_1.5h). Subsequently, the amplified DNA was cleaned using Genomic DNA Clean \u0026amp; Concentrator\u0026trade; kit (Zymo Research, Irvine, CA, USA). Qubit\u0026reg; 1X dsDNA HS Assay Kit on Qubit Fluorometer (Life Technologies, CA, USA) was used to measure the concentration of the purified DNA. Finally, the library was constructed using the Nextera XT kit (Illumina, San Diego, CA, USA). Two technical replicates were generated for each biological sample, resulting in a total of 12 MDA-prepared libraries.\u003c/p\u003e \u003cp\u003exGen\u0026trade; ssDNA \u0026amp; Low-Input DNA Library (xGen, previously known as Swift Bioscience Accel-NGS\u0026reg; 1S) was prepared following the manufacturer\u0026rsquo;s protocol (Catalog # 10009859, Integrated DNA Technologies, IDT). The Mock (15 ng in 20 \u0026micro;l TE buffer) was randomly fragmented by BioRuptor (Diagenode, Liege, Belgium) at default intensity for 15 sec on and 90 sec off, with 7 cycles at 4\u0026deg;C as described by Hoang et al [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e]. Sheared DNAs were firstly subjected to 3\u0026rsquo; end tailing and ligation with adapter1, followed by extension and ligation with adapter2. The purified ligation products were indexed with xGen\u0026trade; Unique Dual Index (UDI) Primers (Catalog # 10005975, IDT) and amplified through a 10-cycle PCR. The amplified products were cleaned up with AMPure XP beads (Beckman Coulter Genomic, CA, USA). For each biological sample, two technical replicates were conducted, resulting in a total of 6 xGen-prepared libraries.\u003c/p\u003e \u003cp\u003eFor the Single-stranded library preparation (SSLR, detailed in \u003cb\u003eSupplementary file1\u003c/b\u003e), the same mocks were used. After random fragmentation, the end modified adapters (\u003cb\u003eSupplementary file2\u003c/b\u003e) synthesized by IDT (IDT, Leuven, Belgium) were ligated to DNA fragments employing the method from Troll et al [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. Subsequently, the ligation mix was purified using MinElute Reaction Kit (Qiagen, Hilden, Germany) and the ligation products were eluted with 23 \u0026micro;l of TE buffer (10 mM, pH8.0). The cleaned ligation products were then amplified and indexed for sequencing with Illumina dual Indexes (i7\u0026thinsp;+\u0026thinsp;i5) and further purified using AMPure XP beads (Beckman Coulter Genomic, CA, USA). An additional bead cleanup step can be applied to mitigate the formation of byproducts and adapter dimers resulting from self-ligation of the adapters (\u003cb\u003eFig.S10\u003c/b\u003e). Two technical replicates were performed for each biological sample, resulting in a total of 6 SSLR-prepared libraries. A detailed protocol can be found on (\u003cb\u003eSupplementary file1\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eThe DNA concentrations of Nextera-prepared, MDA-prepared, xGen-prepared and SSLR-prepared libraries were measured by Qubit\u0026reg; 1X dsDNA HS Assay Kit on Qubit Fluorometer (Life Technologies, CA, USA) and fragment length distributions were determined using TapeStation 4200 (Agilent, CA, USA). Equal amounts of DNA from each library (except xGen-prepared) were pooled and sequenced using 2\u0026times;150 bp paired-end settings on an Illumina NextSeq550 platform (Illumina, CA, USA). The xGen-prepared libraries were sequenced separately on an iSeq100 System (Illumina, CA, USA) using 2\u0026times;150 bp paired-end settings due to their indexes being incompatible with the Illumina index.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eRNA phages library preparation and sequencing\u003c/h2\u003e \u003cp\u003eFour types of phage genome (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, dsDNA, ssDNA, dsRNA and ssRNA) with different ratios were prepared for phage DNA/RNA mock communities (Mock D, Mock E and Mock F). We used three mock communities to determine the applicability of the SSLR method for quantification of RNA and DNA phages simultaneously (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eB\u003c/b\u003e and \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eF\u003c/b\u003e). Specifically, 20 \u0026micro;l mock community was treated with DMSO at a final concentration of 50% and incubated at 65\u0026deg;C for 90 min [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. The DMSO was subsequently removed using a QIAmp viral RNA mini kit (Qiagen, Hilden, Germany), according to the manufacturer\u0026rsquo;s instructions. As a control, a DMSO-free treatment was included. Alternatively, the same mock community was denatured at 95\u0026deg;C for 3 min and immediately snap-cooled on ice, maintaining the nucleic acids as single-stranded. Again, a sample without heat treatment was included as a control library.\u003c/p\u003e \u003cp\u003eAll the DMSO-treated samples, heat-treated samples and their corresponding controls were subjected to reverse transcription in a 20-\u0026micro;l reaction system according to the user guide (SuperScript\u0026trade; IV VILO\u0026trade; Master Mix, Invitrogen\u0026trade;) and the reaction mix was purified using AMPure XP beads, and a 20 \u0026micro;l of Tris buffer (10 mM, pH8.0) was used for the final elution. Finally, the library preparations with SSLR were performed according to the above-mentioned steps and sequenced on a NovaSeq6000 platform (Illumina, CA, USA).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eSequencing of highly modified phages\u003c/h2\u003e \u003cp\u003eTwo types of DNA phage mock communities (Mock G and Mock H) were prepared with equal input (11.11% for each of the 9 genomes, \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eC\u003c/b\u003e and \u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eG\u003c/b\u003e). Mock G contained a highly modified T4 genome (T4), while Mock H contained a less modified T4-c genome (T4-c). All the libraries were prepared as described above and sequenced with the NovaSeq6000 platform (Illumina, CA, USA).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eHuman fecal virome isolation and sequencing library preparation\u003c/h2\u003e \u003cp\u003eFor environmental sample, a fresh fecal sample was obtained from an anonymous healthy adult (Ethical Committee of the Capital Region of Denmark registration number H-20028549) and thoroughly mixed with SM buffer. Fecal virome isolation and purification were carried out according to our previous method with the following modifications [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]: the Centriprep 50K was replaced by Centrisart\u0026reg; I centrifugal ultrafiltration unit (MWCO 100 kDa, Sartorius Stedim Biotech GmbH) and the enrichment step was conducted at 2500 \u0026times; g for 30 min at 4\u0026deg;C. The QIAmp viral RNA mini kit (Qiagen, Hilden, Germany) was used for the extraction of viral DNA/RNA from the concentrated virome solution (\u003cb\u003eFig.\u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eD\u003c/b\u003e). The library preparations were conducted with the aforementioned methods in triplicate. An equal amount of DNA from each library was pooled and sequenced using 150 bp paired-end settings on an NextSeq550 platform (Illumina, CA, USA).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eMetavirome sequencing, data pre-processing and data analyses\u003c/h2\u003e \u003cp\u003eMetavirome reads were quality filtered, trimmed and assembled using a previously published pipeline from GitHub: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/jcame/virome_analysis‑FOOD\u003c/span\u003e\u003cspan address=\"https://github.com/jcame/virome_analysis‑FOOD\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Detailed data analyses for sequencing of the mock communities and fecal virome can be found from \u003cb\u003eSupplementary file3\u003c/b\u003e. For mock community data, plots were generated using the ggplot2 [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e] package in R (v4.2.0). Analyses of viral community α- and β-diversity were performed using packages Phyloseq (v1. 36.0) [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e] and Vegan (v2.6.2) [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e] in R. For α-diversity analyses, indexes of observed taxa and Shannon diversity were calculated with t.test using the package ggsignif (v0.6.3) [\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]. For β-diversity analyses, Bray-Curtis distance metrics were calculated and unconstrained ordination was performed using principal coordinate analysis (PCoA). To identify differentially enriched vira on the summarized family level, DESeq2 (v1.42.0) was adopted [\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e]. The results were then visualized in a heatmap using the R package complexheatmap (v2.18.0) [\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e]. Estimation of sequencing error rates was conducted according to Wilcox et al [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. Coverage biases were visualized using the python script from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/padbr/gcbias\u003c/span\u003e\u003cspan address=\"https://github.com/padbr/gcbias\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, with modifications made by adding the Pearson correlation. To visualize the functional genes of annotated viral contigs, DnaFeaturesViewer (v3.1.0) [\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e] and R package gggenomes (v 0.9.12.9000) [\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e] was used.\u003c/p\u003e \u003c/div\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSSLR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esingle-stranded library\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSingle Reaction Single-stranded LibrarY\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSRSLY\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRT\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eReverse transcription\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMDA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003emultiple Displacement Amplification\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePCoA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eprincipal coordinates analysis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003evOTUs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eviral-operational taxonomic units\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eORFs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eopen reading frames\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNCLDV\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003enucleocytoplasmic large DNA viruses\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLCA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLowest Common Ancestor\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSGBs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003especies-level genome bins\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003enot significant\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003egut virome database\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGVD\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003egut phage database\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGPD\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003emetagenomic gut virus\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMGV\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ehigh- Integrated Microbial Genomes /Viruses 4.1\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIMG/VR4.1\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHigh Sensitivity\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHS\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eUnique Dual Index\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eUDI\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eIntegrated DNA Technologies\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIDT\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA healthy adult anonymous donor donated fecal samples for the study after being informed orally and in writing and provided written consent (Ethical Committee of the Capital Region of Denmark registration number H-20028549).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by research grants (23145 and 36242) from VILLUM FONDEN.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that the data supporting the findings of this study are available within the paper and its supplemental information files. All the raw viral metagenome sequences data produced in this study are available through the National Center for Biotechnology Informationʼs Sequence Read Archive under BioProject accession number PRJNA1094595 with submission No. SUB13375064 (DNA mocks), No. SUB14350434 (DNA/RNA mock), No. SUB14350532 (Modification mock), and No. SUB14350824 (fecal virome).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCode for metavirome pre-analysis is available from GitHub: https://github.com/jcame/virome_analysis-FOOD. Original python code from (https://github.com/padbr/gcbias) with modifications for fitting the GC bias can be found from GitHub: https://github.com/padbr/gcbias. The read accuracy and mutation analysis can be found on GitHub: https://github.com/awilcox83/dsRNA-sequencing.\u003c/p\u003e\n\u003cp\u003eThe public databases for checkV (v1.5) at https://portal.nersc.gov/CheckV, VIBRANT (v1.2.1) at (https://github.com/AnantharamanLab/VIBRANT/tree/master/databases), Virsoter2 at (https://osf.io/u3t4j), VOG217 at (https://fileshare.csb.univie.ac.at/vog/vog217) and Iphop (v1.3.2) at (https://portal.nersc.gov/cfs/m342/iphop/db/) are accessible online.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research was designed and directed by LK, DSN and LD. The experiment was performed by XZ. Data collection was performed by XZ, AG and WK. Data analysis was done by XZ and LD. The manuscript was written and revised by XZ, DSN and LD. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank the donors and our colleagues at Section for Food Microbiology, Gut Health and Fermentation, Department of Food Science, University of Copenhagen (KU FOOD) for their participation and cooperation. We acknowledge Dr. Samuel Kilcher and Prof. Martin Loessner from Institute of Food, Nutrition, and Health, ETH Zurich for providing phage P35 and its host strain \u003cem\u003eListeria monocytogenes\u003c/em\u003e Mack. We acknowledge Dr. Yuvaraj Bhoobalan-Chitty from Department of Biology, University of Copenhagen (Ole Maal\u0026oslash;es Vej 5, 2200 Copenhagen N, Denmark) for providing phage T4-c. XZ was supported by China Scholarship Council (CSC) Grant #201906870027.\u003cstrong\u003e\u003c/strong\u003e\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eShkoporov, A.N. \u0026amp; Hill, C. Bacteriophages of the Human Gut: The \u0026ldquo;Known Unknown\u0026rdquo; of the Microbiome. Cell Host \u0026amp; Microbe 25, 195\u0026ndash;209 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Belleghem, J.D., Dąbrowska, K., Vaneechoutte, M., Barr, J.J. \u0026amp; Bollyky, P.L. Interactions between Bacteriophage, Bacteria, and the Mammalian Immune System. Viruses 11, 10 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang, G., Cobian-Guemes, A.G., Albenberg, L. \u0026amp; Bushman, F. The gut virome in inflammatory bowel diseases. Curr Opin Virol 51, 190\u0026ndash;198 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNorman, J.M. et al. Disease-Specific Alterations in the Enteric Virome in Inflammatory Bowel Disease. Cell 160, 447\u0026ndash;460 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFernandes, M.A. et al. Enteric Virome and Bacterial Microbiota in Children With Ulcerative Colitis and Crohn Disease. J Pediatr Gastroenterol Nutr 68, 30\u0026ndash;36 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReyesa, A. et al. Gut DNA viromes of Malawian twins discordant for severe acute malnutrition. P Natl Acad Sci USA 112, 11941\u0026ndash;11946 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuan, Y. et al. Bacteriophage targeting of gut bacterium attenuates alcoholic liver disease. Nature 575, 505\u0026ndash;511 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRasmussen, T.S. et al. Faecal virome transplantation decreases symptoms of type 2 diabetes and obesity in a murine model. Gut 69, 2122\u0026ndash;2130 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMirzaei, M.K. et al. Bacteriophages Isolated from Stunted Children Can Regulate Gut Bacterial Communities in an Age-Specific Manner. Cell Host Microbe 27, 199\u0026ndash;212.e5 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeal Rodr\u0026iacute;guez, C. et al. The infant gut virome is associated with preschool asthma risk independently of bacteria. Nat Med 30, 138\u0026ndash;148 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOtt, S.J. et al. Efficacy of Sterile Fecal Filtrate Transfer for Treating Patients With Clostridium difficile Infection. Gastroenterology 152, 799\u0026ndash;811 e7(2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKortright, K.E., Chan, B.K., Koff, J.L. \u0026amp; Turner, P.E. Phage Therapy: A Renewed Approach to Combat Antibiotic-Resistant Bacteria. Cell Host Microbe 25, 219\u0026ndash;232 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZuo, T. et al. Bacteriophage transfer during faecal microbiota transplantation in infection is associated with treatment outcome. Gut 67, 634\u0026ndash;643 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKazantseva, J., Malv, E., Kaleda, A., Kallastu, A. \u0026amp; Meikas, A. Optimisation of sample storage and DNA extraction for human gut microbiota studies. BMC Microbiol 21, 158 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaghini, D.G. et al. Quantifying bias introduced by sample collection in relative and absolute microbiome measurements. Nat Biotechnol 42, 328\u0026ndash;338 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhai, X. et al. The impact of storage buffer and storage conditions on fecal samples for bacteriophage infectivity and metavirome analyses. Microbiome 11, 193 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShah, S.A. et al. Expanding known viral diversity in the healthy infant gut. Nat Microbiol 8, 986\u0026ndash;998 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZolfo, M., Vol. Ph.D. University of Trento, Trento; 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUnterer, M., Khan Mirzaei, M. \u0026amp; Deng, L. Gut Phage Database: phage mining in the cave of wonders. Signal Transduct Target Ther 6, 193 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReyes, A., Semenkovich, N.P., Whiteson, K., Rohwer, F. \u0026amp; Gordon, J.I. Going viral: next-generation sequencing applied to phage populations in the human gut. Nat Rev Microbiol 10, 607\u0026ndash;617 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan Mirzaei, M. et al. Challenges of Studying the Human Virome - Relevant Emerging Technologies. Trends Microbiol 29, 171\u0026ndash;181 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarine, R. et al. Evaluation of a transposase protocol for rapid generation of shotgun high-throughput sequencing libraries from nanogram quantities of DNA. Appl Environ Microbiol 77, 8071\u0026ndash;8079 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, K.H. \u0026amp; Bae, J.W. Amplification methods bias metagenomic libraries of uncultured single-stranded and double-stranded DNA viruses. Appl Environ Microbiol 77, 7663\u0026ndash;7668 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoux, S. et al. Towards quantitative viromics for both double-stranded and single-stranded DNA viruses. PeerJ 4, e2777 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, M. et al. Comparison of multiple displacement amplification (MDA) and multiple annealing and looping-based amplification cycles (MALBAC) in single-cell sequencing. PLoS One 9, e114520 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCallanan, J. et al. Expansion of known ssRNA phage genomes: From tens to over a thousand. Sci Adv 6, eaay5981 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStarr, E.P., Nuccio, E.E., Pett-Ridge, J., Banfield, J.F. \u0026amp; Firestone, M.K. Metatranscriptomic reconstruction reveals RNA viruses with the potential to shape carbon cycling in soil. Proc Natl Acad Sci U S A 116, 25900\u0026ndash;25908 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSachsenroder, J. et al. Simultaneous Identification of DNA and RNA Viruses Present in Pig Faeces Using Process-Controlled Deep Sequencing. Plos One 7 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrishnamurthy, S.R., Janowski, A.B., Zhao, G., Barouch, D. \u0026amp; Wang, D. Hyperexpansion of RNA Bacteriophage Diversity. PLoS Biol 14, e1002409 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNeri, U. et al. Expansion of the global RNA virome reveals diverse clades of bacteriophages. Cell 185, 4023\u0026ndash;4037 e4018 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, T. et al. RNA viral community in human feces: Prevalence of plant pathogenic viruses. Plos Biology 4, 108\u0026ndash;118 (2006).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJurasz, H., Pawłowski, T. \u0026amp; Perlejewski, K. Contamination Issue in Viral Metagenomics: Problems, Solutions, and Clinical Perspectives. Front microbiol 12 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMartin, S.A.M., Dehler, C.E. \u0026amp; Krol, E. Transcriptomic responses in the fish intestine. Dev Comp Immunol 64, 103\u0026ndash;117 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOzsolak, F. \u0026amp; Milos, P.M. RNA sequencing: advances, challenges and opportunities. Nat Rev Genet 12, 87\u0026ndash;98 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTroll, C.J. et al. A ligation-based single-stranded library preparation method to analyze cell-free DNA and synthetic oligos. BMC Genomics 20, 1023 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilson, G.G., Young, K.K.Y., Edlin, G.J. \u0026amp; Konigsberg, W. High-Frequency Generalized Transduction by Bacteriophage-T4. Nature 280, 80\u0026ndash;82 (1979).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTao, P., Wu, X.R., Tang, W.C., Zhu, J.G. \u0026amp; Rao, V. Engineering of Bacteriophage T4 Genome Using CRISPR-Cas9. Acs Synth Biol 6, 1952\u0026ndash;1961 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHua, J. in University of Pittsburgh, Vol. Doctoral Dissertation (University of Pittsburgh, University of Pittsburgh; 2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuerin, E. \u0026amp; Hill, C. Shining Light on Human Gut Bacteriophages. Front Cell Infect Microbiol 10, 481 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJakociune, D. \u0026amp; Moodley, A. A Rapid Bacteriophage DNA Extraction Method. Method Protocol 1 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSoleimani-Delfan, A., Bouzari, M. \u0026amp; Wang, R. A rapid competitive method for bacteriophage genomic DNA extraction. J Virol Methods 293 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTeng, Y. et al. Gut bacterial isoamylamine promotes age-related cognitive dysfunction by promoting microglial cell death. Cell Host Microbe 30, 944\u0026ndash;960.e8 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaud, B. et al. An alternative method to multiple displacement amplification for preparing virome DNA in a way adapted for the sequencing of both double-strand and single-strand DNA viruses. \u003cem\u003ebioRxiv\u003c/em\u003e, 2022.2012.2012.520144 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinger, V.L., Jones, L.J., Yue, S.T. \u0026amp; Haugland, R.P. Characterization of PicoGreen reagent and development of a fluorescence-based solution assay for double-stranded DNA quantitation. Anal Biochem 249, 228\u0026ndash;238 (1997).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCui, L. et al. Four additional natural 7-deazaguanine derivatives in phages and how to make them. Nucleic Acids Research 51, 9214\u0026ndash;9226 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNakayama, Y., Yamaguchi, H., Einaga, N. \u0026amp; Esumi, M. Pitfalls of DNA Quantification Using DNA-Binding Fluorescent Dyes and Suggested Solutions. Plos One 11 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGregory, A.C. et al. The Gut Virome Database Reveals Age-Dependent Patterns of Virome Diversity in the Human Gut. \u003cem\u003eCell Host Microbe\u003c/em\u003e 28, 724\u0026ndash;740 e728 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamargo, A.P. et al. IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata. Nucleic Acids Research 51, D733-D743 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNayfach, S. et al. Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome. Nat Microbiol 6, 960\u0026ndash;970 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamarillo-Guerrero, L.F., Almeida, A., Rangel-Pineros, G., Finn, R.D. \u0026amp; Lawley, T.D. Massive expansion of human gut bacteriophage diversity. \u003cem\u003eCell\u003c/em\u003e 184, 1098-1098-1109.e9 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSato, M.P. et al. Comparison of the sequencing bias of currently available library preparation kits for Illumina sequencing of bacterial genomes and metagenomes. DNA Res 26, 391\u0026ndash;398 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJones, M.B. et al. Library preparation methodology can influence genomic and functional predictions in human microbiome research. Proc Natl Acad Sci U S A 112, 14024\u0026ndash;14029 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilcox, A.H., Delwart, E. \u0026amp; D\u0026iacute;az-Mu\u0026ntilde;oz, S.L. Next-generation sequencing of dsRNA is greatly improved by treatment with the inexpensive denaturing reagent DMSO. Microbial genomics 5, e000315 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRobertson, G. et al. De novo assembly and analysis of RNA-seq data. Nat Methods 7, 909\u0026ndash;912 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoang, M.L. et al. Genome-wide quantification of rare somatic mutations in normal human tissues using massively parallel sequencing. Proc Natl Acad Sci U S A 113, 9846\u0026ndash;9851 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWickham, H. ggplot2: Elegant Graphics for Data Analysis. (Springer-Verlag, New York; 2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcMurdie, P.J. \u0026amp; Holmes, S. phyloseq: An R Package for Reproducible Interactive Analysis and Graphics of Microbiome Census Data. Plos One 8 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOksanen, F.J., et al. Vegan: Community Ecology Package. \u003cem\u003eR package Version 2.6-2\u003c/em\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhlmann-Eltze, C. \u0026amp; Patil, I. ggsignif: R Package for Displaying Significance Brackets for 'ggplot2'. (PsyArXiv, 2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLove, M.I., Huber, W. \u0026amp; Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 15 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGu, Z.G., Eils, R. \u0026amp; Schlesner, M. Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847\u0026ndash;2849 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZulkower, V. \u0026amp; Rosser, S. DNA Features Viewer: a sequence annotation formatting and plotting library for Python. Bioinformatics 36, 4350\u0026ndash;4352 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomas, H.M., J. A.;Kristina, H. gggenomes: A Grammar of Graphics for Comparative Genomics. \u003cem\u003eR package version 0.9.5.9000\u003c/em\u003e (2022).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"microbiome","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mbio","sideBox":"Learn more about [Microbiome](http://microbiomejournal.biomedcentral.com/)","snPcode":"40168","submissionUrl":"https://submission.nature.com/new-submission/40168/3","title":"Microbiome","twitterHandle":"@MicrobiomeJ","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"single-stranded library, phage mock community, ssDNA virome, dsDNA virome, RNA virome, gut virome, modification","lastPublishedDoi":"10.21203/rs.3.rs-4304844/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4304844/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe gut virome is an integral component of the gut microbiome, playing a crucial role in maintaining gut health. However, accurately depicting the entire gut virome is challenging due to the inherent diversity of genome types (dsDNA, ssDNA, dsRNA, and ssRNA) and topologies (linear, circular, or fragments), with subsequently biases associated with current sequencing library preparation methods. To overcome these problems and improve reproducibility and comparability across studies, universal or standardized virome sequencing library construction methods are highly needed in the gut virome study.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eWe repurposed the ligation-based single-stranded library (SSLR) preparation method for virome studies. We demonstrate that the SSLR method exhibits exceptional efficiency in quantifying viral DNA genomes (both dsDNA and ssDNA) and outperforms existing double-stranded (Nextera) and single-stranded (xGen, MDA\u0026thinsp;+\u0026thinsp;Nextera) library preparation approaches in terms of minimal amplification bias, evenness of coverage, and integrity of assembling viral genomes. The SSLR method can be utilized for the simultaneous library preparation of both DNA and RNA viral genomes. Furthermore, the SSLR method showed its ability to capture highly modified phage genomes, which were often lost using other library preparation approaches.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eWe introduce and improve a fast, simple and efficient ligation-based single-stranded DNA library preparation for gut virome study. This method is compatible with Illumina sequencing platforms and only requires ligation reagents within 3 h library preparation, which is similar or even better than the advanced library preparation method (xGen). We hope this method can be further optimized, validated and widely used to make gut virome study more comparable and reproducibly.\u003c/p\u003e","manuscriptTitle":"A single strand-based library preparation method for unbiased virome characterization","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-05-02 11:25:06","doi":"10.21203/rs.3.rs-4304844/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-08-28T01:06:23+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-08-27T21:01:48+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-08-27T19:22:34+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"338605160840207638577580691991404716159","date":"2024-08-12T14:12:03+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"41337314203248794666273351107743608495","date":"2024-07-30T15:06:21+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-07-28T12:03:25+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-05-07T09:07:07+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-04-25T06:59:08+00:00","index":"","fulltext":""},{"type":"submitted","content":"Microbiome","date":"2024-04-22T09:31:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"microbiome","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mbio","sideBox":"Learn more about [Microbiome](http://microbiomejournal.biomedcentral.com/)","snPcode":"40168","submissionUrl":"https://submission.nature.com/new-submission/40168/3","title":"Microbiome","twitterHandle":"@MicrobiomeJ","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d01510a4-c50e-430a-ba8e-cdc6b9819071","owner":[],"postedDate":"May 2nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-10-28T16:00:24+00:00","versionOfRecord":{"articleIdentity":"rs-4304844","link":"https://doi.org/10.1186/s40168-024-01935-5","journal":{"identity":"microbiome","isVorOnly":false,"title":"Microbiome"},"publishedOn":"2024-10-24 15:57:17","publishedOnDateReadable":"October 24th, 2024"},"versionCreatedAt":"2024-05-02 11:25:06","video":{"identity":"f44751e7c0eeb1d6eaa42441143ffb4b"},"vorDoi":"10.1186/s40168-024-01935-5","vorDoiUrl":"https://doi.org/10.1186/s40168-024-01935-5","workflowStages":[]},"version":"v1","identity":"rs-4304844","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4304844","identity":"rs-4304844","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00