Genomic diversity of phages infecting the globally widespread genus Sulfurimonas | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Genomic diversity of phages infecting the globally widespread genus Sulfurimonas Ruolin Cheng, Xiaofeng Li, Chuan-Xi Zhang, Zongze Shao This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4432365/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Nov, 2024 Read the published version in Communications Biology → Version 1 posted You are reading this latest preprint version Abstract The bacterial genus Sulfurimonas is globally widespread and occupies a key ecological niche in different habitats. However, phages infecting Sulfurimonas have never been isolated and characterized. Here we systematically investigated the genetic diversity, taxonomy and interaction patterns of Sulfurimonas -associated phages based on sequenced microbial genomes and metagenome datasets. High-confidence phage contigs related to Sulfurimonas were identified from various ecosystems, clustered into 61 viral operational taxonomic units across 3 viral realms. Most Sulfurimonas -associated phages were tailed viruses of Caudoviricetes ; these were assigned to 19 genus-level viral clusters, the majority of which were distantly related to previously known viruses. Phages encoding double jelly-roll major capsid proteins represented another group of double-stranded DNA phage with diverse gene compositions. Inovirus-like single-stranded DNA phages were primarily identified as integrated prophages or extrachromosomal viral genomes, suggesting chronic infections in hosts. Historical and current phage-host interactions were revealed, implying the viral impact on host evolution. Additionally, phages encoding auxiliary metabolic genes might benefit the infected bacteria by compensating or augmenting host metabolisms. This study highlights the remarkable diversity and novelty of Sulfurimonas -associated phages with highly divergent tailless lineages, providing basis for further investigation of phage-host interactions within this genus. Biological sciences/Microbiology/Virology/Metagenomics Biological sciences/Microbiology/Environmental microbiology/Water microbiology Sulfurimonas prophage uncultured viral genomes genomic and phylogenetic analysis novel phage groups Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Bacteria of the genus Sulfurimonas within class Campylobacterota (formerly Epsilonproteobacteria ) 1 are one of the most widespread and dominant groups in deep-sea hydrothermal vents. They can grow chemolithoautotrophically with various electron donors, electron acceptors, and inorganic carbon sources 2 , 3 , playing important roles in the biogeochemical cycles of hydrothermal ecosystems. Besides, sequences of Sulfurimonas have been identified in the pelagic redox cline, coastal sediments, and many different types of terrestrial habitats worldwide 3 . So far, 13 Sulfurimonas species have been isolated and characterized 2 , 4 – 8 , while an increasing number of metagenome-assembled genomes (MAGs) and single-cell amplified genomes (SAGs) have been deposited in the public database. For example, an uncultivated Sulfurimonas species was described recently, which was globally abundant and active in deep-sea oxygen-saturated hydrothermal plumes 9 . As the most abundant and diverse biological entities on earth, bacteriophages (or phages for short) play significant roles in shaping the structure of microbial communities and contribute to the global carbon and nitrogen cycling 10 . During infection, phages can reprogram host metabolism through the expression of viral-encoded auxiliary metabolic genes (AMGs), and phage-mediated horizontal gene transfer (HGT) can also influence host fitness, driving microbial evolution and diversification 11 . Typically, phages follow one of two different lifestyles: the lytic or the lysogenic cycle. While lytic phages kill their hosts through cell lysis, temperate phages can remain latent and replicate as prophages. It is estimated that prophage elements may account for 10 to 20% of the bacterial genomic DNA, and comprise a large proportion of strain-specific differences within species 12 , 13 . Over the last decade, our knowledge of the viral diversity has greatly expanded. With the development of high-throughput sequencing and bioinformatic tools, a large number of novel viral populations and virus-host linkages have been revealed by culture-independent surveys 14 – 18 . These uncultivated virus genomes (UViGs) now already represent the vast majority of the taxonomic diversity in public viral genome databases 19 . In this context, new virus species and higher taxa recovered from sequencing data alone are now accepted by the International Committee on Taxonomy of Viruses (ICTV) 20 . Although the diversity of the viral world has been increasingly unveiled, phages infecting Sulfurimonas have never been isolated and characterized to date. Here, we present a comprehensive study of the phylogenetic diversity, genomic features, biogeographic distribution, and phage-host interaction patterns of the Sulfurimonas -associated phages. The results shed new light on the coevolution of this ecologically important genus and their phages, and provided a robust foundation to further characterize the biology and ecological roles of these novel viral lineages. Materials and Methods Acquisition of Sulfurimonas -associated phage genomes Genomic assemblies of all Sulfurimonas strains were downloaded from the NCBI genome database (June, 2023), and sequences from cultured isolates, MAGs and SAGs retrieved from previous studies 17 were added into the collection. Medium- and high-quality (completeness ≥ 50% and contamination ≤ 10%) genomes evaluated by CheckM v1.1.3 21 were dereplicated at 99.99% average nucleotide identity (ANI) using dRep v2.3.2 22 . The final dataset included 219 strain-level genomes, representing isolates (n = 21), MAGs (n = 188) and SAGs (n = 10) recovered from various environments (Supplementary Table 1). Prophage-like sequences in Sulfurimonas were predicted using a combination of three popular phage finding tools, including Phaster 23 , Virsorter v1.0.6 24 and VIBRANT v1.2.0 25 . Putative viral contigs that were identified as higher confidence predictions (i.e., s “intact” or “questionable” in Phaster, “category 1/2/4/5” in Virsorter, and “complete”, “high-” or “medium-quality” in VIBRANT) by at least two methods were retained for further analysis. Since the current prophage prediction tools are not efficient in identifying non-tailed phages 14 , 24 , more specific strategies were also applied to detect other viral groups. For phages encoding double jelly-roll major capsid proteins (DJR MCPs), we screened the Sulfurimonas genomes using the hmmsearch tool in HMMER v3.1b2 26 , with a collection of DJR MCP profiles representing different clades within Varidnaviria 27 , 28 . In addition, Inovirus_detector 29 was used to search the genome assemblies for inovirus-like phages within Monodnaviria . Phages probably infecting Sulfurimonas strains were also retrieved from the IMG/VR v.4 dataset 19 . Uncultured viral genomes (UViGs) of high-confidence were downloaded for analysis, and Sulfurimonas -associated phages were identified through four computational host prediction methods 14 , 18 . (i) Nucleotide sequence homology. UViGs were searched against genomic sequences of Sulfurimonas strains using BLASTn 30 with the following thresholds: 70% minimum nucleotide identity over 75% of the contig length, and 0.001 maximum e-value. (ii) CRISPR spacer match. Clustered regularly interspaced short palindromic repeats (CRISPR) arrays and cas genes were identified from all Sulfurimonas genomes using CRISPRCasFinder 31 . CRISPR spacers were queried for exact matches (100% identity over 100% spacer length) against the IMG/VR contigs using the BLASTn-short mode. (iii) tRNAs similarity. ARAGORN v1.2.38 32 was used with the “−t” option to detect tRNAs from the Sulfurimonas genomes and IMG/VR contigs, and the identified tRNA sequences were compared using BLASTn to select perfect hits (100% coverage and 100% identity). (iv) k-mer frequencies. WIsH v1.0 33 was run with the default parameters to identify connections between IMG/VR contigs and all reference genomes from the Genome Taxonomy Database (GTDB) 34 , with p < 0.05 being considered as a match. Based on the methodology described above, prophages and UViGs associated with Sulfurimonas were identified. Only contigs ≥ 10 kb were retained for tailed phages from the realm Duplodnaviria . Contigs ≥ 5 kb were retained for nontailed phages within Varidnaviria and Monodnaviria , because the genome size of these groups were much smaller. These contigs were then clustered at 95% shared nucleotide identity and 80% coverage to generate viral operational taxonomic unit (vOTUs) 35 . Completeness and contamination of the vOTUs was estimated using CheckV v1.0.1 36 . Taxonomic assignment and network analysis Taxonomy of the IMG/VR UViGs were retrieved from the metadata, and prophages that clustered with them in vOTUs were assigned to the same taxon. For the remaining prophages, two classification methods were used in combination: (i) marker-based taxonomic assignment by geNomad 37 ; (ii) last common ancestor (LCA) algorithm-based assignment by CAT v5.0.3 38 . Prodigal v2.6.3 39 was used for ORF prediction from the vOTUs, and the resulting protein sequences were used as input for vConTACT2 v0.11.3 40 . Viral RefSeq version 211 was selected as the reference database, and DIAMOND v0.9.21 41 was chosen for all-to-all comparison of the protein sequences. The similarity score between viral contigs was calculated based on the number of shared protein clusters, and viral contigs were clustered through ClusterONE. Lastly, the genome-content based network was visualized in Cytoscape v3.7.2 42 . Host-phage interactions were constructed based on the information of current infections (existence of prophages) and historical infections (inferred from CRISPR spacer match). The interaction network was visualized by Gephi v0.10.1 43 using the Fruchterman Reingold layout. Phylogenetic and comparative genomic analysis A proteomic tree of the tailed Sulfurimonas phages and all prokaryotic dsDNA viruses in the Virus-Host DB (RefSeq release 219) was generated using ViPtree v3.7 44 . For visualization purposes, only queries and related genomes (S G > 0.02) were selected to construct the final tree. For inovirus-like sequences, prokaryotic ssDNA viruses in the Virus-Host DB and previously reported uncultured inoviruses 29 were selected as references. Genome alignment of inoviruses was also performed and visualized by the ViPtree server. The amino acid identity (AAI) between the pI-like proteins were calculated using CompareM v0.0.32( https://github.com/donovan-h-parks/CompareM ). For phylogenetic analysis of DJR MCPs, previously reported MCPs were used as references 27 , 28 . For CysH genes, related sequences were retrieved from the NCBI’s non-redundant protein database using BLASTP 30 . The amino acid sequences were aligned using MUSCLE v3.8.31 45 and the multiple alignments were trimmed with TrimAl v1.2 46 . A maximum-likelihood (ML) tree was inferred using IQ-TREE v2.0 47 with the best substitution model selected by ModelFinder 48 . Support for nodes in the ML tree was evaluated with 1000 ultrafast bootstrap replicates 49 . The constructed tree was then visualized using FigTree v1.4.4 ( http://tree.bio.ed.ac.uk/software/fgtree/ ). Structural modeling was performed using ESMfold 50 and the structures were visualized in ChimeraX v1.5 51 . Identification of auxiliary metabolic genes (AMGs) Viral AMGs were identified and annotated based on VIBRANT v1.2.0 25 and DRAM-v v1.3.5 52 as previously described. (i) VIBRANT pipeline. The vOTU sequences were ran through VIBRANT using default parameters. (ii) DRAM-v pipeline. VirSorter2 (--prep-for-dramv) was run first to produce the affi-contigs, and the resulting files were used as input for DRAM-v annotation. Putative AMGs were assigned an auxiliary score based on the category of flanking genes, and only AMGs with auxiliary scores < 4 were retained. Finally, the AMGs predicted by the two pipelines were combined. Genome maps for contigs encoding AMGs were drawn using the R package gggenes v0.5.1 ( https://cran.r-project.org/web/packages/gggenes ). Results and discussion Overview of prophages and associated UViGs in Sulfurimonas To determine the prevalence of prophage elements within Sulfurimonas genomes, we screened genomic assemblies of 226 Sulfurimonas strains (Supplementary Table 1) using a combination of several different methods (Phaster, Virsorter, VIBRANT, Inovirus_detector and hmmsearch). As a result, 43 putative prophage sequences were identified from 14.2% (32/226) of the Sulfurimonas genomes examined here (Supplementary Table 2). In addition, we mined the IMG/VR v.4 dataset 19 for UViGs that were associated with Sulfurimonas . This dataset composed of > 15 million virus genomes and genome fragments obtained from (meta)genomes and metatranscriptomes, representing the largest collection of UViGs currently available. To be conservative, only the high-confidence viruses (~ 5 million sequences) were used in the analysis. Based on sequence homology, k-mer frequencies, tRNA sequences and CRISPR spacer similarity, a total of 40 UViGs were predicted to infect the genus Sulfurimonas (Supplementary Table 2). The Sulfurimonas -associated prophages and UViGs are widely distributed across global oceans and continents. Most of these phages were recovered from marine ecosystems, including deep-sea hydrothermal vents, pelagic, and coastal environments, while a relatively small part of the phages were of terrestrial origin (Fig. 1 , Supplementary Table 2). All the sequences were combined and clustered at 95% ANI and 80% coverage, resulting in 61 viral operational taxonomic units (vOTUs, Supplementary Table 3). These vOTUs were assigned to three different viral realms, i.e., Duplodnaviria (viruses with dsDNA genomes; 40 vOTUs), Monodnaviria (viruses with ssDNA genomes; 16 vOTUs) and Varidnaviria (a portmanteau of various DNA viruses; 5 vOTUs). According to their CheckV completeness, the viral genomes or genome fragments were categorized into four different groups: high-quality (completeness ≥ 90%; 14 vOTUs), medium-quality (completeness between 50% and 90%; 10 vOTUs), low-quality (completeness 120% or no completeness estimate, 4 vOTUs). The incompleteness of prophages and UViGs was probably due to the fragmented nature of high-throughput sequencing data, as most of them were derived from draft genomes, MAGs, SAGs or metagenome assemblies. Notably, several vOTUs within the Monodnaviria and Varidnaviria showed very low completeness or had no completeness estimate, even for sequences that contained direct terminal repeats (DTR) and were approximately the expected genome length (Supplementary Table 3). These can also be complete genomes that are distantly related to all reference sequences, as viral lineages other than Caudoviricetes are underrepresented in the CheckV database. Despite the widespread occurrence of the genus Sulfurimonas , we knew next to nothing about phages infecting this group thus far. To better reflect the diversity of Sulfurimonas -associated phages, all of the vOTUs described above were retained for subsequent analysis regardless of their completeness. These vOTUs represent species-level phages that are associated with Sulfurimonas and will greatly expand our understanding of the taxonomic, genomic and ecological characteristics of Sulfurimonas phages. Remarkably, most of the vOTUs contained a single sequence, indicating that the genetic diversity of Sulfurimonas phages remained to be fully explored. Taxonomic scope of tailed Sulfurimonas phages As mentioned above, approximately two-thirds of the Sulfurimonas -associated prophages and UViGs were assigned to the viral class Caudoviricetes (Supplementary Table 3) using geNomad’s markers 37 . However, only one out of the 40 vOTUs was classified at the established family level (vOTU_23 as Schitoviridae ), suggesting that these Sulfurimonas phages were of significant phylogenetic novelty. To illustrate the phylogenetic relationship among the tailed Sulfurimonas phages and other known prokaryotic dsDNA viruses, a viral proteomic tree based on genome-wide sequence similarities was built using the ViPtree server 44 . Then a subset of 859 related phage genomes from the Virus-Host DB were selected for final visualization. The resulting tree showed that the 40 vOTUs were distributed in 8 clades, many of which were far from other isolated phages (Fig. 2 A). The largest group (clade B) included 22 vOTUs, followed by clade C and clade D that comprising 7 and 6 vOTUs, respectively. The remaining clades (clade A, E-H) contained only one vOTU, which were unrelated to each other. We also constructed a gene-sharing network to further evaluate the taxonomic position of these vOTUs using vConTACT2 (Fig. 2 B). This tool group viral contigs into viral clusters (VCs) that were ~ 96% concordant with ICTV prokaryotic viral genera 40 . Clustering with genomes from the ProkaryoticViralRefSeq211 database (n = 4,534) in vConTACT2 suggested that 27 vOTUs could be assigned to 6 genus-level clusters, while 9 vOTUs were designated as outliers (which might be related to the VCs they were connected to, but not at the genus level) and 4 vOTUs as singletons (which had few or no gene similarities against other genomes and were not shown in the network). As illustrated in Fig. 2 B, several VCs and outliers were connected and formed 3 larger clusters representing subfamily or family-level relationships. These clusters roughly corresponded to the B, C and D clades defined by ViPtree, except for the affiliation of vOTU_32 and vOTU_37 (Supplementary Table 3). The vOTU_23 was related to viral genomes from the family Schitoviridae but was not clustered with them at the genus level, which was consistent with both the geNomad and ViPtree classification. No previously described phages were clustered with these vOTUs, even when all the NCBI phage genomes (n = 25,903, as of August 2023) were used as references (data not shown). Overall, the tailed phages of Sulfurimonas were classified into 19 genus-level VCs from 8 families, based on two different viral taxonomy approaches. The results generated by ViPtree and vConTACT2 showed a high degree of agreement and both methods performed well for long (≥ 10 kbp) viral contigs in previous studies 53 . All of the genus-level VCs did not include any known tailed phages and might represent novel genera. To date, 29 phages isolated from the phylum Campylobacterota have been deposited in the Virus-Host DB (October, 2023), most of which were associated with the pathogenic genera Campylobacter and Helicobacter. Only one temperate phage was induced from the deep-sea vent Campylobacterota , Nitratiruptor sp. SB155-2 54 . As shown in the Fig. 2 A, phages within the clade C and clade D were clustered with the Campylobacter phage CJIE4 and Nitratiruptor phage NrS-1, respectively. The rest of Sulfurimonas phages, however, were distantly related to any phage isolates infecting Campylobacterota . Putative phages encoding double jelly-roll major capsid proteins (DJR MCPs) Bacteriophages with dsDNA genomes were included in two viral realms, Duplodnaviria and Varidnaviria . Apart from the diverse tailed bacteriophages belonging to Duplodnaviria , a total of 5 DJR MCP-encoding sequences were identified in Sulfurimonas MAGs (Fig. 3 , Supplementary Table 2), which might represent non-tailed dsDNA phages of the realm Varidnaviria. The size of these vOTUs ranged from 7,425 to 10,747 bp, with the percentage of G + C content from 33.6–37.8%. One of the predicted phages located in a long contig with flanking bacterial genes, indicative of a provirus. Others were viral contigs containing no host genes, including one with DTRs and might be a complete genome. The genomic context of the DJR MCP-encoding sequences was quite variable. Only two genes, the MCP and the upstream packaging ATPase were conserved among all these phage elements (Fig. 3 A). Transcription regulators containing a helix-turn-helix (HTH) domain were detected in 4 out of the 5 sequences but was missing in the vOTU_60, probably due to the incompleteness of this contig. The vOTU_57 and vOTU_58 encoded an N-acetyltransferase (NAT) which was conserved in the family Autolykiviridae and Corticoviridae 28 . In comparison, vOTU_60 and vOTU_61 encoded a beta-sandwich jelly-roll fold protein (B_sand), which was present in most members of the STIV group 27 . Several genes were predicted to be involved in cell lysis, including those encoding holin and cell wall hydrolase. Besides, the majority of the predicted phage genes were of unknown function. A maximum likelihood phylogenetic tree was constructed for these DJR MCPs and their homologous protein sequences using IQ-TREE2 (Fig. 3 B). At coarse grain, the MCPs formed two distinct clades, which previously defined as PM2 group and STIV group 27 . The MCPs of vOTU_57 and vOTU_58 were closely related to those of phage isolates from Autolykiviridae and belonged to the PM2 group. This group also consisted of the cultured Corticoviridae , representing the largest group of known DJR MCPs. The vOTU_59 also fell into the PM2 group, but was more distantly related to phages of these two families. The MCPs of vOTU_60 and vOTU_61 were clustered with the STIV group MCPs. The described members of this group included two Sulfolobus turreted icosahedral viruses (STIV1 and STIV2) from family Turriviridae and dozens of highly diverged sequences encoded by archaeal and bacterial (pro)viruses. Comparison of ESMFold-predicted structures (Fig. 3 B) also showed that the MCPs of vOTU_57, vOTU_58 and vOTU_59 were most similar to the PM2 group MCP while those of vOTU_60 and vOTU_61 resembled the STIV group MCP. Tailless phages with dsDNA genomes were thought to be abundant in global surface oceans 55 , 56 . However, the vast majority of cultivated phages are tailed viruses of the realm Duplodnaviria while the non-tailed phages within Varidnaviria are far less investigated. Metagenomic studies had revealed the diversity and prevalence of this group, based on the presence of the hallmark gene encoding DJR MCPs 27 , 28 . In the current study, we used a sensitive profile-based method to screen the Sulfurimonas genomes and discovered several highly diverse DJR MCP-encoding sequences. Despite the conserved tertiary structures, sequence similarities between some MCPs were low (amino acid identity below 30% or not detectable). Thus, we could not rule out the possibility that more divergent MCPs were missed in our analysis. Further explorations are required to unravel the role of this phage group infecting Sulfurimonas. Diverse ssDNA phages infecting Sulfurimonas Bacteriophages from the Inoviridae family are characterized by circular, single-stranded DNA genomes encapsidated in filamentous virions 57 . Among the reported inovirus genomes, a gene encoding the mophogenesis protein (pI) was the only conserved marker gene 29 . Due to their unique and diverse gene content, most of the current phage prediction tools were not able to identify inoviruses from genomic or metagenomics sequences. Thus, we used a machine learning approach 29 based on marker gene and genome features to detect putative inoviruses in Sulfurimonas genomes. As a result, a total of 16 inovirus-like sequences were identified from 15 strains. In addition, 5 UViGs from the IMG/VR v.4 dataset that were predicted to infect Sulfurimonas were classified as inoviruses (Supplementary Table 2). These prophages and UViG sequences were clustered into 16 vOTUs (Supplementary Table 3), including 2 circular contigs, 9 prophages with canonical att sites (direct repeats of ≥ 10bp in a tRNA or next to an integrase) and 5 partial genomes (Fig. 4 A). Genome alignments of the inovirus-like sequences (Fig. 4 A) indicated that they were highly diverse at the amino acid level, yet displayed genomic synteny to some extent. Most of the homologous proteins shared low sequence similarity with AAI values below 50%. An exception was found between the vOTU_44 and vOTU_50, whose hosts were different strains of the same species. These two sequences showed high similarity (AAI > 95%), but the vOTU_44 existed as a circular contig while the vOTU_50 integrated into the host’s chromosome. The complete genomes encoded 13–18 genes, which were organized in functional modules involved in virion structure, assembly/secretion, regulation, DNA replication and integration. Some host strains, such as Sulfurimonas sp. SWIR-19 and Sulfurimonas sp. S012_79_esom, contained two inovirus-like sequences that were distinct from one another. However, the vOTU_54 in Sulfurimonas sp. SWIR-19 and vOTU_52 in Sulfurimonas sp. S012_79_esom lacked the replication genes and might be defective prophages. To determine the evolutionary relationship and taxon classification of these putative inoviruses, a viral proteomic tree was constructed using ViPtree 44 with all related ssDNA phages as references. As shown in Fig. 4 B, the Sulfurimonas -associated inoviruses were clustering with members of the proposed family Amplinoviridae 29 . This proposed family comprised large genomes associated with hosts of Deltaproteobacteria and Campylobacterota , but did not include any viral isolate. Pairwise comparison of the AAI percentage of marker genes (i.e. pI-like proteins) suggested that these 16 vOTUs belonged to 10 genus-level clusters (Supplementary Fig. 1), based on the proposed threshold for delineating Inoviridae genera (50% AAI) 29 . Filamentous phages are extremely diverse and widespread, infecting a broad diversity of bacterial hosts 29 . In the current study, highly divergent inovirus-like sequences were identified in different Sulfurimonas strains as integrated prophages or extrachromosomal viral genomes. Like the other known inoviruses, these phages might establish a chronic infection within the hosts and exert significant effect on their growth, adaptability and evolution 58 , 59 . Bacteriophages of the family Microviridae are another group of ssDNA viruses that have also been frequently observed in viromes 60 . However, we did not find any evidence that microviruses were associated with Sulfurimonas hosts. Interaction between Sulfurimonas phages and their hosts Based on the virus-host linkages revealed via prophage prediction and CRISPR spacer match, we further investigated the potential interactions between Sulfurimonas associated phages and their past/current hosts. CRISPR arrays with the highest evidence-level (annotated by CRISPRCasFinder) were detected in 65 Sulfurimonas strains. Of these, 7 strains contained spacers that target prophages or UViGs sequences identified in this study (Fig. 5 , Supplementary Table 4). CRISPR spacers from 5 strains were associated with the identified inoviruses, and sometimes the viral sequences were targeted by more than one spacer (Supplementary Table 4). These strains did not contain any inovirus-like sequences, indicating that they acquired immunity against the corresponding phages from the historical infection. By contrast, closely related strains or species which lacked the CRISPR spacers might still carry these prophages. For example, Sulfurimonas sp. NW15 contained several spacers matching the vOTU_44 and vOTU_50, which were not detected in this strain but were present in the genomes of Sulfurimonas sp. NW367 and Sulfurimonas sp. NW9, respectively (Fig. 5 ). Infections of dsDNA phages were also recorded by CRISPR-Cas system, but the frequency was relatively low given that dsDNA phages were predominant. Moreover, it seemed that they had narrow host ranges compared with ssDNA phages. While many inovirus-like phages interacted with different strains and species, most of the dsDNA phages infected only one Sulfurimonas strain. This was consistent with the previous findings that non-tailed phages had a broader host range than the tailed group 56 , 61 , 62 . One exception was the vOTU_4, a provirus identified in the genome of Sulfurimonas sp. RIFOXYD12_FULL_33_39. Spacers targeting this phage was detected in Sulfurospirillum sp. SCADC, indicative of its ability to infect hosts of different genera. These broad-host range phages might facilitate horizontal gene transfer 63 within the habitat and play a role in the evolution of the microbial community. In some host strains, multiple phage contigs were detected, suggesting co-infection of two or more phages in a single cell (Fig. 5 ). The frequency of co-infection might be overestimated for tailed phages if the large genomes were mis-assembled into several different contigs 64 . In the case of Sulfurimonas sp. OB8 and Sulfurimonas sp. NORP112, the multiple sequences from Caudoviricetes represented at least two different phages as two distinct copies of the marker gene terL were identified. Co-infections of phages from different viral realms were also observed, especially for inoviruses with persistent infection cycles. The presence of multiple phages favored recombination and genetic exchange between phages, which is important for generation of phage diversity 65 and for escape from CRISPR immunity 66 . Putative AMGs in Sulfurimonas -associated phages AMGs encoded by phages may participate in microbial metabolic pathways during infection. To better understand the impact of phages on host metabolisms, the vOTUs were examined for the presence of AMGs. Using VIBRANT and DRAM-v pipelines, a total of 11 putative AMGs were identified in the genomes of Sulfurimonas -associated phages (Fig. 6 A, Supplementary Table 5). All of these AMGs were surrounded by phage genes, suggesting a true viral origin for them. Based on annotations against the KEGG Orthologs (KO) database, Pfam, and CAZy databases, these AMGs were classified into 4 functional categories, i.e., energy metabolism, amino acid metabolism, nucleotide metabolism and carbohydrate metabolism. The most common AMG within these phage contigs was the gene encoding DNA (cytosine-5)-methyltransferase 1 (DNMT1 or dcm ), which was identified from 5 different vOTUs (vOTU_5, vOTU_27, vOTU_30, vOTU_36 and vOTU_37). This enzyme catalyzes a site-specific methyl transfer reaction, playing a role in cysteine and methionine metabolisms. However, viral dcm was frequently found in diverse hosts and environments, and was thought to perform central functions for virus itself rather than its host 25 . AMGs that encode deoxyuridine 5'-triphosphate nucleotidohydrolase (dUTPase) were present in 3 Sulfurimonas -associated vOTUs (vOTU_8, vOTU_19 and vOTU_32). The dUTPase is a housekeeping enzyme which catalyzes the hydrolysis of dUTP to dUMP, affecting the rate of DNA metabolism 67 . AMGs related to carbohydrate metabolism were detected in 2 vOTUs, including the genes encoding glycoside hydrolases (GH) of the family GH19 (vOTU_38) and GH28 (vOTU_24). Expression of these phage GHs may help the host bacteria with the breakdown of complex carbohydrates. One of the AMGs, the gene encoding phosphoadenosine phosphosulfate reductase (PAPS reductase or CysH), was identified in a complete prophage genome (vOTU_1). CysH is a key enzyme involved in assimilatory sulfate reduction, catalyzing the reduction of PAPS to sulfite. Viral-encoded cysH genes have been found in various marine habitats, including oxygen-deficient water columns 68 , deep-sea hydrothermal vents 17 , cold seeps 15 , seamounts 69 and hadal trenches 70 . The putative CysH of Sulfurimonas prophage contains the conserved domain and structural configuration of CysH enzymes that assimilate sulfates for methionine and cysteine biosynthesis in microorganisms (Fig. 6 C). However, phylogenetic analysis of cysH genes showed that the phage CysH is distantly related to its homologs in Sulfurimonas , indicating the complex evolutionary history of this AMG (Fig. 6 B). Interestingly, while many Sulfurimonas species encode the cysH genes, the host of this prophage ( Sulfurimonas sp. 4561 − 380_metabat1_scaf2bin.038_VB, a MAG recovered from hydrothermal vent metagenome) lacks one. Thus, we suspect that the phage CysH may have potential compensatory effects on host metabolisms. Conclusions Sulfurimonas plays important roles in chemoautotrophic processes and sulfur cycles in various habitats, and infection of phages will inevitably affect their population dynamics and metabolic capacities. In this study, we have revealed a high level of taxonomic and genomic diversity in Sulfurimonas -associated phages for the first time, and illustrated their interactions with different host species or strains. To overcome the limitations of the regular viral prediction tools in identifying non-tailed phages, we applied a combination of several strategies and found that these “cryptic” phage elements were more ubiquitous than supposed. Although smaller than tailed phages, their divergent genome architectures and broader host ranges suggested that they might significantly contribute to the genetic diversity and evolution of this bacterial genus. As this is an in silico analysis, further experiments will be necessary to verify the phage activity and host-phage interactions. Nonetheless, our findings greatly expanded the current understanding of phages infecting Sulfurimonas , and provided good basis for future investigations to explore the ecological and evolutionary role of these phages. Declarations Data availability The genomic sequences used for analysis are publicly available in the NCBI genome repository at https://www.ncbi.nlm.nih.gov/genome. The sequences of the vOTUs generated from the current study have been deposited in the National Omics Data Encyclopedia (NODE) database at https://www.biosino.org/, accession number OEP005305. Acknowledgments This work was funded by Natural Science Foundation of China (No. 42376125; No. 42006088); Natural Science Foundation of Fujian Province of China (No. 2023J011383); the China Ocean Mineral Resources R&D Association (COMRA) program (No. DY-XZ-04; No. DY135-B2-01), and National Key Research and Development Program of China (No. 2018YFC0310701). Author contributions Conceptualization: RC and XL; Methodology: XL; Investigation: XL; Data curation: RC and ZS; Writing-original draft preparation: XL; Writing-review and editing: RC, CZ and ZS; Supervision: CZ and RC; Funding acquisition: RC and ZS. All authors have read and approved the final version of the manuscript. Competing interests The authors declare no financial or non-financial competing interests. References Waite, D.W., et al.: Addendum: Comparative Genomic Analysis of the Class Epsilonproteobacteria and Proposed Reclassification to Epsilonbacteraeota (phyl. nov). Front. Microbiol. 9 , 772 (2018). Erratum Wang, S., et al.: Characterization of Sulfurimonas hydrogeniphila sp. nov., a Novel Bacterium Predominant in Deep-Sea Hydrothermal Vents and Comparative Genomic Analyses of the Genus Sulfurimonas. Front. Microbiol. 12 (2021) Han, Y., Perner, M.: The globally widespread genus Sulfurimonas : versatile energy metabolisms and adaptations to redox clines. Front. Microbiol. 6 (2015) Hu, Q., Wang, S., Lai, Q., Shao, Z., Jiang, L.: Sulfurimonas indica sp. nov., a hydrogen- and sulfur-oxidizing chemolithoautotroph isolated from a hydrothermal sulfide chimney in the Northwest Indian Ocean. Int. J. Syst. Evol. MicroBiol. 71 (2021) Wang, S., et al.: Sulfurimonas sediminis sp. nov., a novel hydrogen- and sulfur-oxidizing chemolithoautotroph isolated from a hydrothermal vent at the Longqi system, southwestern Indian ocean. Antonie van Leeuwenhoek. 114 , 813–822 (2021) Wang, S., Jiang, L., Liu, X., Yang, S., Shao, Z.: Sulfurimonas xiamenensis sp. nov. and Sulfurimonas lithotrophica sp. nov., hydrogen- and sulfur-oxidizing chemolithoautotrophs within the Epsilonproteobacteria isolated from coastal sediments, and an emended description of the genus Sulfurimonas . International Journal of Systematic and Evolutionary Microbiology 70, 2657–2663 (2020) Takai, K., et al.: Sulfurimonas paralvinellae sp. nov., a novel mesophilic, hydrogen- and sulfur-oxidizing chemolithoautotroph within the Epsilonproteobacteria isolated from a deep-sea hydrothermal vent polychaete nest, reclassification of Thiomicrospira denitrificans as Sulfurimonas denitrificans comb. nov. and emended description of the genus Sulfurimonas . International journal of systematic and evolutionary microbiology 56, 1725–1733 (2006) Inagaki, F., Takai, K., Kobayashi, H., Nealson, K.H., Horikoshi, K.: Sulfurimonas autotrophica gen. nov., sp. nov., a novel sulfur-oxidizing ε-proteobacterium isolated from hydrothermal sediments in the Mid-Okinawa Trough. Int. J. Syst. Evol. MicroBiol. 53 , 1801–1805 (2003) Molari, M., et al.: A hydrogenotrophic Sulfurimonas is globally abundant in deep-sea oxygen-saturated hydrothermal plumes. Nat. Microbiol. 8 , 651–665 (2023) Suttle, C.A.: Marine viruses - major players in the global ecosystem. Nat. Rev. Microbiol. 5 , 801–812 (2007) Rohwer, F., Prangishvili, D., Lindell, D.: Roles of viruses in the environment. Environ. Microbiol. 11 , 2771–2774 (2009) Canchaya, C., Proux, C., Fournous, G., Bruttin, A., Brüssow, H.: Prophage genomics. Microbiol. Mol. Biol. Rev. 67 , (2003). 238 – 76, table of contents Casjens, S.: Prophages and bacterial genomics: what have we learned so far? Mol. Microbiol. 49 , 277–300 (2003) Paez-Espino, D., et al.: Uncovering Earth’s virome. Nature. 536 , 425–430 (2016) Li, Z., et al.: Deep sea sediments associated with cold seeps are a subsurface reservoir of viral diversity. ISME J. (2021) Jian, H., et al.: Diversity and distribution of viruses inhabiting the deepest ocean on Earth. ISME J. 15 , 3094–3110 (2021) Cheng, R., et al.: Virus diversity and interactions with hosts in deep-sea hydrothermal vents. Microbiome. 10 , 235 (2022) Emerson, J.B., et al.: Host-linked soil viral ecology along a permafrost thaw gradient. Nat. Microbiol. 3 , 870–880 (2018) Camargo, A.P., et al.: IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata. Nucleic Acids Res. 51 , D733–D743 (2022) Simmonds, P., et al.: Consensus statement: Virus taxonomy in the age of metagenomics. Nat. Rev. Microbiol. 15 , 161–168 (2017) Parks, D.H., Imelfort, M., Skennerton, C.T., Hugenholtz, P., Tyson, G.W.: CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 25 , 1043–1055 (2015) Olm, M.R., Brown, C.T., Brooks, B., Banfield, J.F.: dRep: a tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication. ISME J. 11 , 2864–2868 (2017) Arndt, D., Marcu, A., Liang, Y., Wishart, D.S., PHAST: PHASTER and PHASTEST: Tools for finding prophage in bacterial genomes. Brief. Bioinform. 20 , 1560–1567 (2019) Roux, S., Enault, F., Hurwitz, B.L., Sullivan, M.B.: VirSorter: mining viral signal from microbial genomic data. PeerJ 3 , e985 (2015) Kieft, K., Zhou, Z., Anantharaman, K.: VIBRANT: automated recovery, annotation and curation of microbial viruses, and evaluation of viral community function from genomic sequences. Microbiome. 8 , 90 (2020) Eddy, S.R., Accelerated Profile, H.M.M., Searches: PLoS Comput. Biol. 7 , e1002195 (2011) Yutin, N., Bäckström, D., Ettema, T.J.G., Krupovic, M., Koonin, E.V.: Vast diversity of prokaryotic virus genomes encoding double jelly-roll major capsid proteins uncovered by genomic and metagenomic sequence analysis. Virol. J. 15 , 67 (2018) Yutin, N., et al.: Varidnaviruses in the Human Gut: A Major Expansion of the Order Vinavirales. Viruses 14 (2022) Roux, S., et al.: Cryptic inoviruses revealed as pervasive in bacteria and archaea across Earth's biomes. Nat. Microbiol. 4 , 1895–1906 (2019) Camacho, C., et al.: BLAST+: architecture and applications. BMC Bioinform. 10 , 421 (2009) Couvin, D., et al.: CRISPRCasFinder, an update of CRISRFinder, includes a portable version, enhanced performance and integrates search for Cas proteins. Nucleic Acids Res. 46 , W246–w251 (2018) Laslett, D., Canback, B.: ARAGORN, a program to detect tRNA genes and tmRNA genes in nucleotide sequences. Nucleic Acids Res. 32 , 11–16 (2004) Galiez, C., Siebert, M., Enault, F., Vincent, J., Söding, J.: WIsH: who is the host? Predicting prokaryotic hosts from metagenomic phage contigs. Bioinformatics. 33 , 3113–3114 (2017) Chaumeil, P.A., Mussig, A.J., Hugenholtz, P., Parks, D.H.: GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics (2019) Bolduc, B., Roux, S.: Clustering viral genomes in iVirus. (2017) Nayfach, S., et al.: CheckV assesses the quality and completeness of metagenome-assembled viral genomes. Nat. Biotechnol. 39 , 578–585 (2021) Camargo, A.P., et al.: Identification of mobile genetic elements with geNomad. Nat. Biotechnol. (2023) von Meijenfeldt, F.A.B., Arkhipova, K., Cambuy, D.D., Coutinho, F.H., Dutilh, B.E.: Robust taxonomic classification of uncharted microbial sequences and bins with CAT and BAT. Genome Biol. 20 , 217 (2019) Hyatt, D., et al.: Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinform. 11 , 119 (2010) Jang, B.: Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat. Biotechnol. 37 , 632–639 (2019) Buchfink, B., Xie, C., Huson, D.H.: Fast and sensitive protein alignment using DIAMOND. Nat. Methods. 12 , 59–60 (2015) Shannon, P., et al.: Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 13 , 2498–2504 (2003) Bastian, M., Heymann, S., Jacomy, M.: Gephi: An Open Source Software for Exploring and Manipulating Networks , (2009) Nishimura, Y., et al.: ViPTree: the viral proteomic tree server. Bioinformatics. 33 , 2379–2380 (2017) Robert, C.: & Edgar. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic acids research (2004) Capella-Gutiérrez, S., Silla-Martínez, J.M., Gabaldón, T.: trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 25 , 1972–1973 (2009) Minh, B.Q., et al.: IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol. Biol. Evol. 37 , 1530–1534 (2020) Kalyaanamoorthy, S., Minh, B.Q., Wong, T.K.F., Haeseler, A.V., Jermiin, L.S.: ModelFinder: fast model selection for accurate phylogenetic estimates. Nat. Methods 14 (2017) Hoang, D.T., Chernomor, O., von Haeseler, A., Minh, B.Q., Vinh, L.S.: UFBoot2: Improving the Ultrafast Bootstrap Approximation. Mol. Biol. Evol. 35 , 518–522 (2017) Lin, Z., et al.: Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 379 , 1123–1130 (2023) Pettersen, E.F., et al.: UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci. 30 , 70–82 (2021) Shaffer, M., et al.: DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Res. 48 , 8883–8900 (2020) Pratama, A.A., et al.: Expanding standards in viromics: in silico evaluation of dsDNA viral genome identification, classification, and auxiliary metabolic gene curation. PeerJ 9 , e11447 (2021) Yoshida-Takashima, Y., Takaki, Y., Shimamura, S., Nunoura, T., Takai, K.: Genome sequence of a novel deep-sea vent epsilonproteobacterial phage provides new insight into the co-evolution of Epsilonproteobacteria and their phages. Extremophiles: life under extreme conditions. 17 , 405–419 (2013) Brum, J.R., Schenck, R.O., Sullivan, M.B.: Global morphological analysis of marine viruses shows minimal regional variation and dominance of non-tailed viruses. Isme j. 7 , 1738–1751 (2013) Kauffman, K.M., et al.: A major lineage of non-tailed dsDNA viruses as unrecognized killers of marine bacteria. Nature. 554 , 118–122 (2018) Rakonjac, J., Bennett, N.J., Spagnuolo, J., Gagic, D., Russel, M.: Filamentous bacteriophage: biology, phage display and nanotechnology applications. Curr. Issues Mol. Biol. 13 , 51–76 (2011) Yu, Z.C., et al.: Filamentous phages prevalent in Pseudoalteromonas spp. confer properties advantageous to host survival in Arctic sea ice. Isme j 9, 871 – 81 (2015) Ilyina, T.S.: Filamentous bacteriophages and their role in the virulence and evolution of pathogenic bacteria. Mol. Genet. Microbiol. Virol. 30 , 1–9 (2015) Yoshida, M., et al.: Quantitative Viral Community DNA Analysis Reveals the Dominance of Single-Stranded DNA Viruses in Offshore Upper Bathyal Sediment from Tohoku, Japan. Front. Microbiol. 9 (2018) Ignacio-Espinoza, J.C., Fuhrman, J.A.: A non-tailed twist in the viral tale. Nature. 554 , 38–39 (2018) Kauffman, K.M., et al.: Resolving the structure of phage-bacteria interactions in the context of natural diversity. Nat. Commun. 13 , 372 (2022) Tzipilevich, E., Habusha, M., Ben-Yehuda, S.: Acquisition of Phage Sensitivity by Bacteria through Exchange of Phage Receptors. Cell. 168 , 186–199e12 (2017) Yi, Y., et al.: A systematic analysis of marine lysogens and proviruses. Nat. Commun. 14 , 6013 (2023) Bellas, C.M., Schroeder, D.C., Edwards, A., Barker, G., Anesio, A.M.: Flexible genes establish widespread bacteriophage pan-genomes in cryoconite hole ecosystems. Nat. Commun. 11 , 4403 (2020) Paez-Espino, D., et al.: CRISPR immunity drives rapid phage genome evolution in Streptococcus thermophilus. mBio 6 (2015) Huang, X., Jiao, N., Zhang, R.: The genomic content and context of auxiliary metabolic genes in roseophages. Environ. Microbiol. 23 , 3743–3757 (2021) Mara, P., et al.: Viral elements and their potential influence on microbial processes along the permanently stratified Cariaco Basin redoxcline. ISME J. 14 , 3079–3092 (2020) Yu, M., et al.: Diversity and potential host-interactions of viruses inhabiting deep-sea seamount sediments. Nat. Commun. 15 , 3228 (2024) Zhao, J., et al.: Novel Viral Communities Potentially Assisting in Carbon, Nitrogen, and Sulfur Metabolism in the Upper Slope Sediments of Mariana Trench. mSystems 7, e0135821 (2022) Additional Declarations (Not answered) Supplementary Files FigureS1.tif Supplementary Figure 1. Heatmap showing the amino acid identity (AAI) value between pI proteins. TableS1.pdf Supplementary Table 1. Genomic features of Sulfurimonas strains used in this study. TableS2.pdf Supplementary Table 2. Characteristics of prophages recovered from Sulfurimonas genomes and Sulfurimonas- associated UViG retrieved from the IMG/VR v.4 dataset. TableS3.pdf Supplementary Table 3. Characteristics of Sulfurimonas- associated vOTUs. TableS4.pdf Supplementary Table 4. CRISPR spacer matches between phages and Sulfurimonas genomes. TableS5.pdf Supplementary Table 5. Annotation of putative auxiliary metabolic genes. Cite Share Download PDF Status: Published Journal Publication published 02 Nov, 2024 Read the published version in Communications Biology → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4432365","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":312334315,"identity":"4769acd6-10b6-4fd7-aa32-4d43ac39d494","order_by":0,"name":"Ruolin Cheng","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIie3PIQvCQBTA8TcOtnLu6o3JPsPJYEU/jCJsZeCiYBSW1Cz4RQyGk8EsC0aDYSvLK4qCwRNFNHhqE7w/B/de+HEcgEr1i9HbbSLQcmDfEB0BYvw7Ig7lnwgyG6bFbpE4ulFLB1XUAmKv1nCcSx7ZpoFbLxNXR6a/4cwHaxJG2jh7TRgNPZvypBMj7AmSAMtwG2mxjPT2dxJ9SELdqm4ErsTgUkI3vmcDD8RfsEsz5mNrhGE5lhAy7ZbWgTcdQrJG1T+1HIKNIj9KyCWEHxYxY8blAEA7PO9G/k6oVCrVf3UGg3hIHeSbGXIAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-8568-2252","institution":"Third Institute of Oceanography,Ministry of Natural Resources","correspondingAuthor":true,"prefix":"","firstName":"Ruolin","middleName":"","lastName":"Cheng","suffix":""},{"id":312334316,"identity":"ed6596ae-b9be-4b0d-a0c5-b511dfce5760","order_by":1,"name":"Xiaofeng Li","email":"","orcid":"https://orcid.org/0000-0003-0573-1702","institution":"Ningbo University","correspondingAuthor":false,"prefix":"","firstName":"Xiaofeng","middleName":"","lastName":"Li","suffix":""},{"id":312334317,"identity":"36977e13-83be-48f5-9f94-d72d81ced393","order_by":2,"name":"Chuan-Xi Zhang","email":"","orcid":"https://orcid.org/0000-0002-7784-1188","institution":"Ningbo University","correspondingAuthor":false,"prefix":"","firstName":"Chuan-Xi","middleName":"","lastName":"Zhang","suffix":""},{"id":312334318,"identity":"a5f96a88-1d5f-4cd2-ac60-15506a050f02","order_by":3,"name":"Zongze Shao","email":"","orcid":"","institution":"Third Institute of Oceanography","correspondingAuthor":false,"prefix":"","firstName":"Zongze","middleName":"","lastName":"Shao","suffix":""}],"badges":[],"createdAt":"2024-05-16 16:30:34","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4432365/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4432365/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s42003-024-07079-4","type":"published","date":"2024-11-02T04:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":60416987,"identity":"a390aaa7-1698-4a05-8017-437fa5c2e2b4","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":381409,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGeographical distribution of the predicted\u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eSulfurimonas\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e phages.\u003c/strong\u003e Each location was represented by a symbol proportional to the number of phage contigs. The phages belonging to different viral realms were presented by different shapes, and each ecosystem type is indicated by a unique color.\u003c/p\u003e","description":"","filename":"Figure1401.png","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/f387d3753e4ba47f7b11d160.png"},{"id":60416988,"identity":"8a703b78-88e4-498d-b115-8ae0294c9a76","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":4954197,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePhylogenetic and network analysis of tailed \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eSulfurimonas \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003ephages. \u003c/strong\u003e(A) The genome-wide proteomic tree of tailed \u003cem\u003eSulfurimonas\u003c/em\u003ephages and other related known dsDNA phages constructed using ViPTree. The colored inner and outer rings represent the virus family and host groups, respectively. Phages identified in this study are indicated with red asterisks, and the corresponding branches are colored red. (B) Gene-sharing network analysis of tailed \u003cem\u003eSulfurimonas \u003c/em\u003ephages and reference viral genomes. Phages identified in this study are indicated by triangles.\u003c/p\u003e","description":"","filename":"Figure1402.png","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/56eed021eb97e48c15f3be5c.png"},{"id":60417644,"identity":"239ecfbe-ba00-420d-ac46-07e2624e3c5b","added_by":"auto","created_at":"2024-07-16 14:09:48","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":958719,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenomic characterization and phylogenetic analysis of DJR MCP-encoding phages. \u003c/strong\u003e(A) Genome arrangement of putative phages encoding DJR MCP. ORFs are shown by block arrows drawn to scale. Gene abbreviations: MCP, major capsid protein; HTH, helix-turn-helix domain-containing protein; B_sand, beta-sandwich jelly-roll fold protein; NAT, GNAT family acetyltransferase; Rep, replication protein. (B) Maximum likelihood tree of the MCP proteins. Branch color indicates different DJR MCP groups. DJR MCPs in \u003cem\u003eSulfurimonas\u003c/em\u003e genomes and in cultured \u003cem\u003eTurriviridae, Corticoviridae \u003c/em\u003eand\u003cem\u003e Autolykiviridae\u003c/em\u003e are labeled with green, purple, red and blue circles, respectively.\u003c/p\u003e","description":"","filename":"Figure1403.png","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/36351d43ee926ef7ff41372f.png"},{"id":60417646,"identity":"8b6faf4f-4de9-4188-ae19-599483f875da","added_by":"auto","created_at":"2024-07-16 14:09:48","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":4687960,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenome alignment and phylogenetic analysis of inovirus-like phages. \u003c/strong\u003e(A) Alignment and comparison of putative inoviruses identified in \u003cem\u003eSulfurimonas.\u003c/em\u003eORFs are shown by block arrows drawn to scale and color-coded according to their putative functions. The color of the shading connecting homologous genes indicates amino acid identities between them. (B) The genome-wide proteomic tree of putative inoviruses identified in \u003cem\u003eSulfurimonas\u003c/em\u003eand other related ssDNA phages constructed using ViPTree. The colored left and rignt lines represent the virus family and host groups, respectively. Phages identified in this study are labeled with colored asterisks.\u003c/p\u003e","description":"","filename":"Figure1404.png","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/8f499e5f7eccc0d7b9ec0079.png"},{"id":60418520,"identity":"a9ca94cb-ba64-4a01-8eec-b7cdbee09ca7","added_by":"auto","created_at":"2024-07-16 14:17:48","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":904776,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eInteractions between phages and \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eSulfurimonas \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003estrains\u003c/strong\u003e. Phages and hosts are shown by different node types, and each ecosystem type is indicated by a unique color. Current infections and historical infections are indicated as red and green arrows, respectively.\u003c/p\u003e","description":"","filename":"Figure1405.png","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/8fd3b8ceb09d77754635561a.png"},{"id":60416995,"identity":"5472bf22-9a91-4318-8bdc-cdc528a0a543","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1280284,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenomic context, phylogeny and predicted protein structure of AMGs encoded by \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eSulfurimonas\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003ephages.\u003c/strong\u003e (A) Genomic maps of representative viral contigs containing AMGs, with the AMGs of interest in red, viral hallmark genes in green, viral-like genes in yellow, and uncharacterized genes in grey. Gene abbreviations: CysH, phosphoadenosine phosphosulfate reductase; GH, glycoside hydrolases; dUTPase, deoxyuridine 5'-triphosphate nucleotidohydrolase; DNMT1, DNA (cytosine-5)-methyltransferase 1. (B) Maximum likelihood tree based on the CysH amino acid sequences. Leaves were colored according to the affiliated taxonomic groups. Support for nodes was evaluated with 1000 ultrafast bootstrap replicates, and bootstrap scores \u0026gt;70% are indicated by black dots. GenBank accession number is given for each sequence. (C) Predicted tertiary structure of the phage CysH proteins.\u003c/p\u003e","description":"","filename":"Figure1406.png","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/a4e2f6f9be08b01a27570025.png"},{"id":68105316,"identity":"2a1a1e79-0438-479a-a169-7804a2d4df04","added_by":"auto","created_at":"2024-11-03 08:07:30","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":14266395,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/1c77f936-f44f-40dc-9cef-77fc6f884445.pdf"},{"id":60416999,"identity":"5b6602eb-9ad7-455e-87f2-4e404dee8454","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"tif","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":5069476,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Figure 1. Heatmap showing the amino acid identity (AAI) value between pI proteins.\u003c/p\u003e","description":"","filename":"FigureS1.tif","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/f1722c6c03210782817087e0.tif"},{"id":60416989,"identity":"4887d3ad-ff87-4a2e-ac60-f793565d31b3","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":77323,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Table 1. Genomic features of \u003cem\u003eSulfurimonas \u003c/em\u003estrains used in this study.\u003c/p\u003e","description":"","filename":"TableS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/75b2f498ba15e3c45339ca11.pdf"},{"id":60416993,"identity":"5a1ecd97-84c1-4eb3-96f5-becccaa03614","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":90592,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Table 2. Characteristics of prophages recovered from \u003cem\u003eSulfurimonas \u003c/em\u003egenomes and \u003cem\u003eSulfurimonas-\u003c/em\u003eassociated UViG retrieved from the IMG/VR v.4 dataset.\u003c/p\u003e","description":"","filename":"TableS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/481c791bb385049ad834d4bf.pdf"},{"id":60419327,"identity":"87ccd58a-1fa0-4e44-8276-291f722ad705","added_by":"auto","created_at":"2024-07-16 14:25:48","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":76616,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Table 3. Characteristics of \u003cem\u003eSulfurimonas-\u003c/em\u003eassociated vOTUs.\u003c/p\u003e","description":"","filename":"TableS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/cc66e38f8501971055f468e8.pdf"},{"id":60416997,"identity":"fdb6e170-f6d0-455b-ad43-f16cbb6af965","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":49286,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Table 4. CRISPR spacer matches between phages and\u003cem\u003eSulfurimonas\u003c/em\u003e genomes.\u003c/p\u003e","description":"","filename":"TableS4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/e7e0dd7f382f1ffdab6793d3.pdf"},{"id":60416998,"identity":"82c3d120-eb15-4822-8aea-0194189b414e","added_by":"auto","created_at":"2024-07-16 14:01:48","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":116998,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Table 5. Annotation of putative auxiliary metabolic genes.\u003c/p\u003e","description":"","filename":"TableS5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4432365/v1/19a2913f9c0129c7d4ef8e82.pdf"}],"financialInterests":"(Not answered)","formattedTitle":"Genomic diversity of phages infecting the globally widespread genus Sulfurimonas","fulltext":[{"header":"Introduction","content":"\u003cp\u003eBacteria of the genus \u003cem\u003eSulfurimonas\u003c/em\u003e within class \u003cem\u003eCampylobacterota\u003c/em\u003e (formerly \u003cem\u003eEpsilonproteobacteria\u003c/em\u003e) \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e are one of the most widespread and dominant groups in deep-sea hydrothermal vents. They can grow chemolithoautotrophically with various electron donors, electron acceptors, and inorganic carbon sources\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, playing important roles in the biogeochemical cycles of hydrothermal ecosystems. Besides, sequences of \u003cem\u003eSulfurimonas\u003c/em\u003e have been identified in the pelagic redox cline, coastal sediments, and many different types of terrestrial habitats worldwide\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. So far, 13 \u003cem\u003eSulfurimonas\u003c/em\u003e species have been isolated and characterized \u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan additionalcitationids=\"CR5 CR6 CR7\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e, while an increasing number of metagenome-assembled genomes (MAGs) and single-cell amplified genomes (SAGs) have been deposited in the public database. For example, an uncultivated \u003cem\u003eSulfurimonas\u003c/em\u003e species was described recently, which was globally abundant and active in deep-sea oxygen-saturated hydrothermal plumes\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAs the most abundant and diverse biological entities on earth, bacteriophages (or phages for short) play significant roles in shaping the structure of microbial communities and contribute to the global carbon and nitrogen cycling\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. During infection, phages can reprogram host metabolism through the expression of viral-encoded auxiliary metabolic genes (AMGs), and phage-mediated horizontal gene transfer (HGT) can also influence host fitness, driving microbial evolution and diversification\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. Typically, phages follow one of two different lifestyles: the lytic or the lysogenic cycle. While lytic phages kill their hosts through cell lysis, temperate phages can remain latent and replicate as prophages. It is estimated that prophage elements may account for 10 to 20% of the bacterial genomic DNA, and comprise a large proportion of strain-specific differences within species\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eOver the last decade, our knowledge of the viral diversity has greatly expanded. With the development of high-throughput sequencing and bioinformatic tools, a large number of novel viral populations and virus-host linkages have been revealed by culture-independent surveys\u003csup\u003e\u003cspan additionalcitationids=\"CR15 CR16 CR17\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. These uncultivated virus genomes (UViGs) now already represent the vast majority of the taxonomic diversity in public viral genome databases\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. In this context, new virus species and higher taxa recovered from sequencing data alone are now accepted by the International Committee on Taxonomy of Viruses (ICTV)\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAlthough the diversity of the viral world has been increasingly unveiled, phages infecting \u003cem\u003eSulfurimonas\u003c/em\u003e have never been isolated and characterized to date. Here, we present a comprehensive study of the phylogenetic diversity, genomic features, biogeographic distribution, and phage-host interaction patterns of the \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages. The results shed new light on the coevolution of this ecologically important genus and their phages, and provided a robust foundation to further characterize the biology and ecological roles of these novel viral lineages.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003e \u003cb\u003eAcquisition of\u003c/b\u003e \u003cb\u003eSulfurimonas\u003c/b\u003e\u003cb\u003e-associated phage genomes\u003c/b\u003e\u003c/p\u003e \u003cp\u003eGenomic assemblies of all \u003cem\u003eSulfurimonas\u003c/em\u003e strains were downloaded from the NCBI genome database (June, 2023), and sequences from cultured isolates, MAGs and SAGs retrieved from previous studies\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e were added into the collection. Medium- and high-quality (completeness\u0026thinsp;\u0026ge;\u0026thinsp;50% and contamination\u0026thinsp;\u0026le;\u0026thinsp;10%) genomes evaluated by CheckM v1.1.3\u003csup\u003e21\u003c/sup\u003ewere dereplicated at 99.99% average nucleotide identity (ANI) using dRep v2.3.2\u003csup\u003e22\u003c/sup\u003e. The final dataset included 219 strain-level genomes, representing isolates (n\u0026thinsp;=\u0026thinsp;21), MAGs (n\u0026thinsp;=\u0026thinsp;188) and SAGs (n\u0026thinsp;=\u0026thinsp;10) recovered from various environments (Supplementary Table\u0026nbsp;1).\u003c/p\u003e \u003cp\u003eProphage-like sequences in \u003cem\u003eSulfurimonas\u003c/em\u003e were predicted using a combination of three popular phage finding tools, including Phaster\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, Virsorter v1.0.6\u003csup\u003e24\u003c/sup\u003e and VIBRANT v1.2.0\u003csup\u003e25\u003c/sup\u003e. Putative viral contigs that were identified as higher confidence predictions (i.e., s \u0026ldquo;intact\u0026rdquo; or \u0026ldquo;questionable\u0026rdquo; in Phaster, \u0026ldquo;category 1/2/4/5\u0026rdquo; in Virsorter, and \u0026ldquo;complete\u0026rdquo;, \u0026ldquo;high-\u0026rdquo; or \u0026ldquo;medium-quality\u0026rdquo; in VIBRANT) by at least two methods were retained for further analysis. Since the current prophage prediction tools are not efficient in identifying non-tailed phages\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e, more specific strategies were also applied to detect other viral groups. For phages encoding double jelly-roll major capsid proteins (DJR MCPs), we screened the \u003cem\u003eSulfurimonas\u003c/em\u003e genomes using the hmmsearch tool in HMMER v3.1b2\u003csup\u003e26\u003c/sup\u003e, with a collection of DJR MCP profiles representing different clades within \u003cem\u003eVaridnaviria\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. In addition, Inovirus_detector\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e was used to search the genome assemblies for inovirus-like phages within \u003cem\u003eMonodnaviria\u003c/em\u003e.\u003c/p\u003e \u003cp\u003ePhages probably infecting \u003cem\u003eSulfurimonas\u003c/em\u003e strains were also retrieved from the IMG/VR v.4 dataset\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Uncultured viral genomes (UViGs) of high-confidence were downloaded for analysis, and \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages were identified through four computational host prediction methods\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. (i) Nucleotide sequence homology. UViGs were searched against genomic sequences of \u003cem\u003eSulfurimonas\u003c/em\u003e strains using BLASTn\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e with the following thresholds: 70% minimum nucleotide identity over 75% of the contig length, and 0.001 maximum e-value. (ii) CRISPR spacer match. Clustered regularly interspaced short palindromic repeats (CRISPR) arrays and \u003cem\u003ecas\u003c/em\u003e genes were identified from all \u003cem\u003eSulfurimonas\u003c/em\u003e genomes using CRISPRCasFinder\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. CRISPR spacers were queried for exact matches (100% identity over 100% spacer length) against the IMG/VR contigs using the BLASTn-short mode. (iii) tRNAs similarity. ARAGORN v1.2.38 \u003csup\u003e32\u003c/sup\u003e was used with the \u0026ldquo;\u0026minus;t\u0026rdquo; option to detect tRNAs from the \u003cem\u003eSulfurimonas\u003c/em\u003e genomes and IMG/VR contigs, and the identified tRNA sequences were compared using BLASTn to select perfect hits (100% coverage and 100% identity). (iv) k-mer frequencies. WIsH v1.0\u003csup\u003e33\u003c/sup\u003e was run with the default parameters to identify connections between IMG/VR contigs and all reference genomes from the Genome Taxonomy Database (GTDB)\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e, with p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 being considered as a match.\u003c/p\u003e \u003cp\u003eBased on the methodology described above, prophages and UViGs associated with \u003cem\u003eSulfurimonas\u003c/em\u003e were identified. Only contigs\u0026thinsp;\u0026ge;\u0026thinsp;10 kb were retained for tailed phages from the realm \u003cem\u003eDuplodnaviria\u003c/em\u003e. Contigs\u0026thinsp;\u0026ge;\u0026thinsp;5 kb were retained for nontailed phages within \u003cem\u003eVaridnaviria\u003c/em\u003e and \u003cem\u003eMonodnaviria\u003c/em\u003e, because the genome size of these groups were much smaller. These contigs were then clustered at 95% shared nucleotide identity and 80% coverage to generate viral operational taxonomic unit (vOTUs)\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e. Completeness and contamination of the vOTUs was estimated using CheckV v1.0.1 \u003csup\u003e36\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eTaxonomic assignment and network analysis\u003c/h2\u003e \u003cp\u003eTaxonomy of the IMG/VR UViGs were retrieved from the metadata, and prophages that clustered with them in vOTUs were assigned to the same taxon. For the remaining prophages, two classification methods were used in combination: (i) marker-based taxonomic assignment by geNomad\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e; (ii) last common ancestor (LCA) algorithm-based assignment by CAT v5.0.3 \u003csup\u003e38\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eProdigal v2.6.3 \u003csup\u003e39\u003c/sup\u003e was used for ORF prediction from the vOTUs, and the resulting protein sequences were used as input for vConTACT2 v0.11.3 \u003csup\u003e40\u003c/sup\u003e. Viral RefSeq version 211 was selected as the reference database, and DIAMOND v0.9.21\u003csup\u003e41\u003c/sup\u003ewas chosen for all-to-all comparison of the protein sequences. The similarity score between viral contigs was calculated based on the number of shared protein clusters, and viral contigs were clustered through ClusterONE. Lastly, the genome-content based network was visualized in Cytoscape v3.7.2\u003csup\u003e42\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eHost-phage interactions were constructed based on the information of current infections (existence of prophages) and historical infections (inferred from CRISPR spacer match). The interaction network was visualized by Gephi v0.10.1\u003csup\u003e43\u003c/sup\u003e using the Fruchterman Reingold layout.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003ePhylogenetic and comparative genomic analysis\u003c/h2\u003e \u003cp\u003eA proteomic tree of the tailed \u003cem\u003eSulfurimonas\u003c/em\u003e phages and all prokaryotic dsDNA viruses in the Virus-Host DB (RefSeq release 219) was generated using ViPtree v3.7\u003csup\u003e44\u003c/sup\u003e. For visualization purposes, only queries and related genomes (S\u003csub\u003eG\u003c/sub\u003e \u0026gt; 0.02) were selected to construct the final tree. For inovirus-like sequences, prokaryotic ssDNA viruses in the Virus-Host DB and previously reported uncultured inoviruses\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e were selected as references. Genome alignment of inoviruses was also performed and visualized by the ViPtree server. The amino acid identity (AAI) between the pI-like proteins were calculated using CompareM v0.0.32(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/donovan-h-parks/CompareM\u003c/span\u003e\u003cspan address=\"https://github.com/donovan-h-parks/CompareM\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFor phylogenetic analysis of DJR MCPs, previously reported MCPs were used as references \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. For CysH genes, related sequences were retrieved from the NCBI\u0026rsquo;s non-redundant protein database using BLASTP\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. The amino acid sequences were aligned using MUSCLE v3.8.31 \u003csup\u003e45\u003c/sup\u003e and the multiple alignments were trimmed with TrimAl v1.2 \u003csup\u003e46\u003c/sup\u003e. A maximum-likelihood (ML) tree was inferred using IQ-TREE v2.0 \u003csup\u003e47\u003c/sup\u003e with the best substitution model selected by ModelFinder\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. Support for nodes in the ML tree was evaluated with 1000 ultrafast bootstrap replicates\u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. The constructed tree was then visualized using FigTree v1.4.4 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://tree.bio.ed.ac.uk/software/fgtree/\u003c/span\u003e\u003cspan address=\"http://tree.bio.ed.ac.uk/software/fgtree/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Structural modeling was performed using ESMfold\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e and the structures were visualized in ChimeraX v1.5\u003csup\u003e51\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of auxiliary metabolic genes (AMGs)\u003c/h2\u003e \u003cp\u003eViral AMGs were identified and annotated based on VIBRANT v1.2.0\u003csup\u003e25\u003c/sup\u003e and DRAM-v v1.3.5\u003csup\u003e52\u003c/sup\u003e as previously described. (i) VIBRANT pipeline. The vOTU sequences were ran through VIBRANT using default parameters. (ii) DRAM-v pipeline. VirSorter2 (--prep-for-dramv) was run first to produce the affi-contigs, and the resulting files were used as input for DRAM-v annotation. Putative AMGs were assigned an auxiliary score based on the category of flanking genes, and only AMGs with auxiliary scores\u0026thinsp;\u0026lt;\u0026thinsp;4 were retained. Finally, the AMGs predicted by the two pipelines were combined. Genome maps for contigs encoding AMGs were drawn using the R package gggenes v0.5.1 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cran.r-project.org/web/packages/gggenes\u003c/span\u003e\u003cspan address=\"https://cran.r-project.org/web/packages/gggenes\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e"},{"header":"Results and discussion","content":"\u003cp\u003e \u003cb\u003eOverview of prophages and associated UViGs in\u003c/b\u003e \u003cb\u003eSulfurimonas\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo determine the prevalence of prophage elements within \u003cem\u003eSulfurimonas\u003c/em\u003e genomes, we screened genomic assemblies of 226 \u003cem\u003eSulfurimonas\u003c/em\u003e strains (Supplementary Table\u0026nbsp;1) using a combination of several different methods (Phaster, Virsorter, VIBRANT, Inovirus_detector and hmmsearch). As a result, 43 putative prophage sequences were identified from 14.2% (32/226) of the \u003cem\u003eSulfurimonas\u003c/em\u003e genomes examined here (Supplementary Table\u0026nbsp;2).\u003c/p\u003e \u003cp\u003eIn addition, we mined the IMG/VR v.4 dataset \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e for UViGs that were associated with \u003cem\u003eSulfurimonas\u003c/em\u003e. This dataset composed of \u0026gt;\u0026thinsp;15\u0026nbsp;million virus genomes and genome fragments obtained from (meta)genomes and metatranscriptomes, representing the largest collection of UViGs currently available. To be conservative, only the high-confidence viruses (~\u0026thinsp;5\u0026nbsp;million sequences) were used in the analysis. Based on sequence homology, k-mer frequencies, tRNA sequences and CRISPR spacer similarity, a total of 40 UViGs were predicted to infect the genus \u003cem\u003eSulfurimonas\u003c/em\u003e (Supplementary Table\u0026nbsp;2).\u003c/p\u003e \u003cp\u003eThe \u003cem\u003eSulfurimonas\u003c/em\u003e-associated prophages and UViGs are widely distributed across global oceans and continents. Most of these phages were recovered from marine ecosystems, including deep-sea hydrothermal vents, pelagic, and coastal environments, while a relatively small part of the phages were of terrestrial origin (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Supplementary Table\u0026nbsp;2). All the sequences were combined and clustered at 95% ANI and 80% coverage, resulting in 61 viral operational taxonomic units (vOTUs, Supplementary Table\u0026nbsp;3). These vOTUs were assigned to three different viral realms, i.e., \u003cem\u003eDuplodnaviria\u003c/em\u003e (viruses with dsDNA genomes; 40 vOTUs), \u003cem\u003eMonodnaviria\u003c/em\u003e (viruses with ssDNA genomes; 16 vOTUs) and \u003cem\u003eVaridnaviria\u003c/em\u003e (a portmanteau of various DNA viruses; 5 vOTUs).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAccording to their CheckV completeness, the viral genomes or genome fragments were categorized into four different groups: high-quality (completeness\u0026thinsp;\u0026ge;\u0026thinsp;90%; 14 vOTUs), medium-quality (completeness between 50% and 90%; 10 vOTUs), low-quality (completeness\u0026thinsp;\u0026lt;\u0026thinsp;50%; 33 vOTUs), and not-determined (completeness\u0026thinsp;\u0026gt;\u0026thinsp;120% or no completeness estimate, 4 vOTUs). The incompleteness of prophages and UViGs was probably due to the fragmented nature of high-throughput sequencing data, as most of them were derived from draft genomes, MAGs, SAGs or metagenome assemblies. Notably, several vOTUs within the \u003cem\u003eMonodnaviria\u003c/em\u003e and \u003cem\u003eVaridnaviria\u003c/em\u003e showed very low completeness or had no completeness estimate, even for sequences that contained direct terminal repeats (DTR) and were approximately the expected genome length (Supplementary Table\u0026nbsp;3). These can also be complete genomes that are distantly related to all reference sequences, as viral lineages other than \u003cem\u003eCaudoviricetes\u003c/em\u003e are underrepresented in the CheckV database.\u003c/p\u003e \u003cp\u003eDespite the widespread occurrence of the genus \u003cem\u003eSulfurimonas\u003c/em\u003e, we knew next to nothing about phages infecting this group thus far. To better reflect the diversity of \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages, all of the vOTUs described above were retained for subsequent analysis regardless of their completeness. These vOTUs represent species-level phages that are associated with \u003cem\u003eSulfurimonas\u003c/em\u003e and will greatly expand our understanding of the taxonomic, genomic and ecological characteristics of \u003cem\u003eSulfurimonas\u003c/em\u003e phages. Remarkably, most of the vOTUs contained a single sequence, indicating that the genetic diversity of \u003cem\u003eSulfurimonas\u003c/em\u003e phages remained to be fully explored.\u003c/p\u003e \u003cp\u003e \u003cb\u003eTaxonomic scope of tailed\u003c/b\u003e \u003cb\u003eSulfurimonas\u003c/b\u003e \u003cb\u003ephages\u003c/b\u003e\u003c/p\u003e \u003cp\u003eAs mentioned above, approximately two-thirds of the \u003cem\u003eSulfurimonas\u003c/em\u003e-associated prophages and UViGs were assigned to the viral class \u003cem\u003eCaudoviricetes\u003c/em\u003e (Supplementary Table\u0026nbsp;3) using geNomad\u0026rsquo;s markers\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. However, only one out of the 40 vOTUs was classified at the established family level (vOTU_23 as \u003cem\u003eSchitoviridae\u003c/em\u003e), suggesting that these \u003cem\u003eSulfurimonas\u003c/em\u003e phages were of significant phylogenetic novelty.\u003c/p\u003e \u003cp\u003eTo illustrate the phylogenetic relationship among the tailed \u003cem\u003eSulfurimonas\u003c/em\u003e phages and other known prokaryotic dsDNA viruses, a viral proteomic tree based on genome-wide sequence similarities was built using the ViPtree server\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. Then a subset of 859 related phage genomes from the Virus-Host DB were selected for final visualization. The resulting tree showed that the 40 vOTUs were distributed in 8 clades, many of which were far from other isolated phages (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). The largest group (clade B) included 22 vOTUs, followed by clade C and clade D that comprising 7 and 6 vOTUs, respectively. The remaining clades (clade A, E-H) contained only one vOTU, which were unrelated to each other.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe also constructed a gene-sharing network to further evaluate the taxonomic position of these vOTUs using vConTACT2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB). This tool group viral contigs into viral clusters (VCs) that were ~\u0026thinsp;96% concordant with ICTV prokaryotic viral genera\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e. Clustering with genomes from the ProkaryoticViralRefSeq211 database (n\u0026thinsp;=\u0026thinsp;4,534) in vConTACT2 suggested that 27 vOTUs could be assigned to 6 genus-level clusters, while 9 vOTUs were designated as outliers (which might be related to the VCs they were connected to, but not at the genus level) and 4 vOTUs as singletons (which had few or no gene similarities against other genomes and were not shown in the network). As illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB, several VCs and outliers were connected and formed 3 larger clusters representing subfamily or family-level relationships. These clusters roughly corresponded to the B, C and D clades defined by ViPtree, except for the affiliation of vOTU_32 and vOTU_37 (Supplementary Table\u0026nbsp;3). The vOTU_23 was related to viral genomes from the family \u003cem\u003eSchitoviridae\u003c/em\u003e but was not clustered with them at the genus level, which was consistent with both the geNomad and ViPtree classification. No previously described phages were clustered with these vOTUs, even when all the NCBI phage genomes (n\u0026thinsp;=\u0026thinsp;25,903, as of August 2023) were used as references (data not shown).\u003c/p\u003e \u003cp\u003eOverall, the tailed phages of \u003cem\u003eSulfurimonas\u003c/em\u003e were classified into 19 genus-level VCs from 8 families, based on two different viral taxonomy approaches. The results generated by ViPtree and vConTACT2 showed a high degree of agreement and both methods performed well for long (\u0026ge;\u0026thinsp;10 kbp) viral contigs in previous studies\u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e. All of the genus-level VCs did not include any known tailed phages and might represent novel genera. To date, 29 phages isolated from the phylum \u003cem\u003eCampylobacterota\u003c/em\u003e have been deposited in the Virus-Host DB (October, 2023), most of which were associated with the pathogenic genera \u003cem\u003eCampylobacter\u003c/em\u003e and \u003cem\u003eHelicobacter.\u003c/em\u003e Only one temperate phage was induced from the deep-sea vent \u003cem\u003eCampylobacterota\u003c/em\u003e, \u003cem\u003eNitratiruptor\u003c/em\u003e sp. SB155-2 \u003csup\u003e54\u003c/sup\u003e. As shown in the Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA, phages within the clade C and clade D were clustered with the \u003cem\u003eCampylobacter\u003c/em\u003e phage CJIE4 and \u003cem\u003eNitratiruptor\u003c/em\u003e phage NrS-1, respectively. The rest of \u003cem\u003eSulfurimonas\u003c/em\u003e phages, however, were distantly related to any phage isolates infecting \u003cem\u003eCampylobacterota\u003c/em\u003e.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003ePutative phages encoding double jelly-roll major capsid proteins (DJR MCPs)\u003c/h2\u003e \u003cp\u003eBacteriophages with dsDNA genomes were included in two viral realms, \u003cem\u003eDuplodnaviria\u003c/em\u003e and \u003cem\u003eVaridnaviria\u003c/em\u003e. Apart from the diverse tailed bacteriophages belonging to \u003cem\u003eDuplodnaviria\u003c/em\u003e, a total of 5 DJR MCP-encoding sequences were identified in \u003cem\u003eSulfurimonas\u003c/em\u003e MAGs (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, Supplementary Table\u0026nbsp;2), which might represent non-tailed dsDNA phages of the realm \u003cem\u003eVaridnaviria.\u003c/em\u003e The size of these vOTUs ranged from 7,425 to 10,747 bp, with the percentage of G\u0026thinsp;+\u0026thinsp;C content from 33.6\u0026ndash;37.8%. One of the predicted phages located in a long contig with flanking bacterial genes, indicative of a provirus. Others were viral contigs containing no host genes, including one with DTRs and might be a complete genome.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe genomic context of the DJR MCP-encoding sequences was quite variable. Only two genes, the MCP and the upstream packaging ATPase were conserved among all these phage elements (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). Transcription regulators containing a helix-turn-helix (HTH) domain were detected in 4 out of the 5 sequences but was missing in the vOTU_60, probably due to the incompleteness of this contig. The vOTU_57 and vOTU_58 encoded an N-acetyltransferase (NAT) which was conserved in the family \u003cem\u003eAutolykiviridae\u003c/em\u003e and \u003cem\u003eCorticoviridae\u003c/em\u003e\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. In comparison, vOTU_60 and vOTU_61 encoded a beta-sandwich jelly-roll fold protein (B_sand), which was present in most members of the STIV group\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. Several genes were predicted to be involved in cell lysis, including those encoding holin and cell wall hydrolase. Besides, the majority of the predicted phage genes were of unknown function.\u003c/p\u003e \u003cp\u003eA maximum likelihood phylogenetic tree was constructed for these DJR MCPs and their homologous protein sequences using IQ-TREE2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). At coarse grain, the MCPs formed two distinct clades, which previously defined as PM2 group and STIV group\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. The MCPs of vOTU_57 and vOTU_58 were closely related to those of phage isolates from \u003cem\u003eAutolykiviridae\u003c/em\u003e and belonged to the PM2 group. This group also consisted of the cultured \u003cem\u003eCorticoviridae\u003c/em\u003e, representing the largest group of known DJR MCPs. The vOTU_59 also fell into the PM2 group, but was more distantly related to phages of these two families. The MCPs of vOTU_60 and vOTU_61 were clustered with the STIV group MCPs. The described members of this group included two \u003cem\u003eSulfolobus turreted\u003c/em\u003e icosahedral viruses (STIV1 and STIV2) from family \u003cem\u003eTurriviridae\u003c/em\u003e and dozens of highly diverged sequences encoded by archaeal and bacterial (pro)viruses. Comparison of ESMFold-predicted structures (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB) also showed that the MCPs of vOTU_57, vOTU_58 and vOTU_59 were most similar to the PM2 group MCP while those of vOTU_60 and vOTU_61 resembled the STIV group MCP.\u003c/p\u003e \u003cp\u003eTailless phages with dsDNA genomes were thought to be abundant in global surface oceans\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e,\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e. However, the vast majority of cultivated phages are tailed viruses of the realm \u003cem\u003eDuplodnaviria\u003c/em\u003e while the non-tailed phages within \u003cem\u003eVaridnaviria\u003c/em\u003e are far less investigated. Metagenomic studies had revealed the diversity and prevalence of this group, based on the presence of the hallmark gene encoding DJR MCPs\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. In the current study, we used a sensitive profile-based method to screen the \u003cem\u003eSulfurimonas\u003c/em\u003e genomes and discovered several highly diverse DJR MCP-encoding sequences. Despite the conserved tertiary structures, sequence similarities between some MCPs were low (amino acid identity below 30% or not detectable). Thus, we could not rule out the possibility that more divergent MCPs were missed in our analysis. Further explorations are required to unravel the role of this phage group infecting \u003cem\u003eSulfurimonas.\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eDiverse ssDNA phages infecting\u003c/b\u003e \u003cb\u003eSulfurimonas\u003c/b\u003e\u003c/p\u003e \u003cp\u003eBacteriophages from the \u003cem\u003eInoviridae\u003c/em\u003e family are characterized by circular, single-stranded DNA genomes encapsidated in filamentous virions\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e. Among the reported inovirus genomes, a gene encoding the mophogenesis protein (pI) was the only conserved marker gene\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. Due to their unique and diverse gene content, most of the current phage prediction tools were not able to identify inoviruses from genomic or metagenomics sequences. Thus, we used a machine learning approach \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e based on marker gene and genome features to detect putative inoviruses in \u003cem\u003eSulfurimonas\u003c/em\u003e genomes. As a result, a total of 16 inovirus-like sequences were identified from 15 strains. In addition, 5 UViGs from the IMG/VR v.4 dataset that were predicted to infect \u003cem\u003eSulfurimonas\u003c/em\u003e were classified as inoviruses (Supplementary Table\u0026nbsp;2). These prophages and UViG sequences were clustered into 16 vOTUs (Supplementary Table\u0026nbsp;3), including 2 circular contigs, 9 prophages with canonical att sites (direct repeats of \u0026ge;\u0026thinsp;10bp in a tRNA or next to an integrase) and 5 partial genomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGenome alignments of the inovirus-like sequences (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA) indicated that they were highly diverse at the amino acid level, yet displayed genomic synteny to some extent. Most of the homologous proteins shared low sequence similarity with AAI values below 50%. An exception was found between the vOTU_44 and vOTU_50, whose hosts were different strains of the same species. These two sequences showed high similarity (AAI\u0026thinsp;\u0026gt;\u0026thinsp;95%), but the vOTU_44 existed as a circular contig while the vOTU_50 integrated into the host\u0026rsquo;s chromosome. The complete genomes encoded 13\u0026ndash;18 genes, which were organized in functional modules involved in virion structure, assembly/secretion, regulation, DNA replication and integration. Some host strains, such as \u003cem\u003eSulfurimonas\u003c/em\u003e sp. SWIR-19 and \u003cem\u003eSulfurimonas\u003c/em\u003e sp. S012_79_esom, contained two inovirus-like sequences that were distinct from one another. However, the vOTU_54 in \u003cem\u003eSulfurimonas\u003c/em\u003e sp. SWIR-19 and vOTU_52 in \u003cem\u003eSulfurimonas\u003c/em\u003e sp. S012_79_esom lacked the replication genes and might be defective prophages.\u003c/p\u003e \u003cp\u003eTo determine the evolutionary relationship and taxon classification of these putative inoviruses, a viral proteomic tree was constructed using ViPtree\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e with all related ssDNA phages as references. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB, the \u003cem\u003eSulfurimonas\u003c/em\u003e-associated inoviruses were clustering with members of the proposed family \u003cem\u003eAmplinoviridae\u003c/em\u003e\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. This proposed family comprised large genomes associated with hosts of \u003cem\u003eDeltaproteobacteria\u003c/em\u003e and \u003cem\u003eCampylobacterota\u003c/em\u003e, but did not include any viral isolate. Pairwise comparison of the AAI percentage of marker genes (i.e. pI-like proteins) suggested that these 16 vOTUs belonged to 10 genus-level clusters (Supplementary Fig.\u0026nbsp;1), based on the proposed threshold for delineating \u003cem\u003eInoviridae\u003c/em\u003e genera (50% AAI)\u003csup\u003e29\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eFilamentous phages are extremely diverse and widespread, infecting a broad diversity of bacterial hosts\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. In the current study, highly divergent inovirus-like sequences were identified in different \u003cem\u003eSulfurimonas\u003c/em\u003e strains as integrated prophages or extrachromosomal viral genomes. Like the other known inoviruses, these phages might establish a chronic infection within the hosts and exert significant effect on their growth, adaptability and evolution \u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e. Bacteriophages of the family \u003cem\u003eMicroviridae\u003c/em\u003e are another group of ssDNA viruses that have also been frequently observed in viromes\u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e. However, we did not find any evidence that microviruses were associated with \u003cem\u003eSulfurimonas\u003c/em\u003e hosts.\u003c/p\u003e \u003cp\u003e \u003cb\u003eInteraction between\u003c/b\u003e \u003cb\u003eSulfurimonas\u003c/b\u003e \u003cb\u003ephages and their hosts\u003c/b\u003e\u003c/p\u003e \u003cp\u003eBased on the virus-host linkages revealed via prophage prediction and CRISPR spacer match, we further investigated the potential interactions between \u003cem\u003eSulfurimonas\u003c/em\u003e associated phages and their past/current hosts. CRISPR arrays with the highest evidence-level (annotated by CRISPRCasFinder) were detected in 65 \u003cem\u003eSulfurimonas\u003c/em\u003e strains. Of these, 7 strains contained spacers that target prophages or UViGs sequences identified in this study (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, Supplementary Table\u0026nbsp;4).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eCRISPR spacers from 5 strains were associated with the identified inoviruses, and sometimes the viral sequences were targeted by more than one spacer (Supplementary Table\u0026nbsp;4). These strains did not contain any inovirus-like sequences, indicating that they acquired immunity against the corresponding phages from the historical infection. By contrast, closely related strains or species which lacked the CRISPR spacers might still carry these prophages. For example, \u003cem\u003eSulfurimonas\u003c/em\u003e sp. NW15 contained several spacers matching the vOTU_44 and vOTU_50, which were not detected in this strain but were present in the genomes of \u003cem\u003eSulfurimonas\u003c/em\u003e sp. NW367 and \u003cem\u003eSulfurimonas\u003c/em\u003e sp. NW9, respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eInfections of dsDNA phages were also recorded by CRISPR-Cas system, but the frequency was relatively low given that dsDNA phages were predominant. Moreover, it seemed that they had narrow host ranges compared with ssDNA phages. While many inovirus-like phages interacted with different strains and species, most of the dsDNA phages infected only one \u003cem\u003eSulfurimonas\u003c/em\u003e strain. This was consistent with the previous findings that non-tailed phages had a broader host range than the tailed group\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e,\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e. One exception was the vOTU_4, a provirus identified in the genome of \u003cem\u003eSulfurimonas\u003c/em\u003e sp. RIFOXYD12_FULL_33_39. Spacers targeting this phage was detected in \u003cem\u003eSulfurospirillum\u003c/em\u003e sp. SCADC, indicative of its ability to infect hosts of different genera. These broad-host range phages might facilitate horizontal gene transfer \u003csup\u003e\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e within the habitat and play a role in the evolution of the microbial community.\u003c/p\u003e \u003cp\u003eIn some host strains, multiple phage contigs were detected, suggesting co-infection of two or more phages in a single cell (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). The frequency of co-infection might be overestimated for tailed phages if the large genomes were mis-assembled into several different contigs \u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e. In the case of \u003cem\u003eSulfurimonas\u003c/em\u003e sp. OB8 and \u003cem\u003eSulfurimonas\u003c/em\u003e sp. NORP112, the multiple sequences from \u003cem\u003eCaudoviricetes\u003c/em\u003e represented at least two different phages as two distinct copies of the marker gene \u003cem\u003eterL\u003c/em\u003e were identified. Co-infections of phages from different viral realms were also observed, especially for inoviruses with persistent infection cycles. The presence of multiple phages favored recombination and genetic exchange between phages, which is important for generation of phage diversity\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e and for escape from CRISPR immunity\u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003e \u003cb\u003ePutative AMGs in\u003c/b\u003e \u003cb\u003eSulfurimonas\u003c/b\u003e\u003cb\u003e-associated phages\u003c/b\u003e\u003c/p\u003e \u003cp\u003eAMGs encoded by phages may participate in microbial metabolic pathways during infection. To better understand the impact of phages on host metabolisms, the vOTUs were examined for the presence of AMGs. Using VIBRANT and DRAM-v pipelines, a total of 11 putative AMGs were identified in the genomes of \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, Supplementary Table\u0026nbsp;5). All of these AMGs were surrounded by phage genes, suggesting a true viral origin for them. Based on annotations against the KEGG Orthologs (KO) database, Pfam, and CAZy databases, these AMGs were classified into 4 functional categories, i.e., energy metabolism, amino acid metabolism, nucleotide metabolism and carbohydrate metabolism.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe most common AMG within these phage contigs was the gene encoding DNA (cytosine-5)-methyltransferase 1 (DNMT1 or \u003cem\u003edcm\u003c/em\u003e), which was identified from 5 different vOTUs (vOTU_5, vOTU_27, vOTU_30, vOTU_36 and vOTU_37). This enzyme catalyzes a site-specific methyl transfer reaction, playing a role in cysteine and methionine metabolisms. However, viral \u003cem\u003edcm\u003c/em\u003e was frequently found in diverse hosts and environments, and was thought to perform central functions for virus itself rather than its host\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. AMGs that encode deoxyuridine 5'-triphosphate nucleotidohydrolase (dUTPase) were present in 3 \u003cem\u003eSulfurimonas\u003c/em\u003e-associated vOTUs (vOTU_8, vOTU_19 and vOTU_32). The dUTPase is a housekeeping enzyme which catalyzes the hydrolysis of dUTP to dUMP, affecting the rate of DNA metabolism\u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u003c/sup\u003e. AMGs related to carbohydrate metabolism were detected in 2 vOTUs, including the genes encoding glycoside hydrolases (GH) of the family GH19 (vOTU_38) and GH28 (vOTU_24). Expression of these phage GHs may help the host bacteria with the breakdown of complex carbohydrates.\u003c/p\u003e \u003cp\u003eOne of the AMGs, the gene encoding phosphoadenosine phosphosulfate reductase (PAPS reductase or CysH), was identified in a complete prophage genome (vOTU_1). CysH is a key enzyme involved in assimilatory sulfate reduction, catalyzing the reduction of PAPS to sulfite. Viral-encoded \u003cem\u003ecysH\u003c/em\u003e genes have been found in various marine habitats, including oxygen-deficient water columns\u003csup\u003e\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e\u003c/sup\u003e, deep-sea hydrothermal vents\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e, cold seeps \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e, seamounts\u003csup\u003e\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e\u003c/sup\u003e and hadal trenches\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e. The putative CysH of \u003cem\u003eSulfurimonas\u003c/em\u003e prophage contains the conserved domain and structural configuration of CysH enzymes that assimilate sulfates for methionine and cysteine biosynthesis in microorganisms (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eC). However, phylogenetic analysis of \u003cem\u003ecysH\u003c/em\u003e genes showed that the phage CysH is distantly related to its homologs in \u003cem\u003eSulfurimonas\u003c/em\u003e, indicating the complex evolutionary history of this AMG (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eB). Interestingly, while many \u003cem\u003eSulfurimonas\u003c/em\u003e species encode the \u003cem\u003ecysH\u003c/em\u003e genes, the host of this prophage (\u003cem\u003eSulfurimonas\u003c/em\u003e sp. 4561\u0026thinsp;\u0026minus;\u0026thinsp;380_metabat1_scaf2bin.038_VB, a MAG recovered from hydrothermal vent metagenome) lacks one. Thus, we suspect that the phage CysH may have potential compensatory effects on host metabolisms.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003e \u003cem\u003eSulfurimonas\u003c/em\u003e plays important roles in chemoautotrophic processes and sulfur cycles in various habitats, and infection of phages will inevitably affect their population dynamics and metabolic capacities. In this study, we have revealed a high level of taxonomic and genomic diversity in \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages for the first time, and illustrated their interactions with different host species or strains. To overcome the limitations of the regular viral prediction tools in identifying non-tailed phages, we applied a combination of several strategies and found that these \u0026ldquo;cryptic\u0026rdquo; phage elements were more ubiquitous than supposed. Although smaller than tailed phages, their divergent genome architectures and broader host ranges suggested that they might significantly contribute to the genetic diversity and evolution of this bacterial genus. As this is an \u003cem\u003ein silico\u003c/em\u003e analysis, further experiments will be necessary to verify the phage activity and host-phage interactions. Nonetheless, our findings greatly expanded the current understanding of phages infecting \u003cem\u003eSulfurimonas\u003c/em\u003e, and provided good basis for future investigations to explore the ecological and evolutionary role of these phages.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe genomic sequences used for analysis are publicly available in the NCBI genome repository at https://www.ncbi.nlm.nih.gov/genome. The sequences of the vOTUs generated from the current study have been deposited in the National Omics Data Encyclopedia (NODE) database at https://www.biosino.org/, accession number OEP005305.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was funded by Natural Science Foundation of China (No. 42376125; No. 42006088); Natural Science Foundation of Fujian Province of China (No. 2023J011383); the China Ocean Mineral Resources R\u0026amp;D Association (COMRA) program (No. DY-XZ-04; No. DY135-B2-01), and National Key Research and Development Program of China (No. 2018YFC0310701).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization: RC and XL; Methodology: XL; Investigation: XL; Data curation: RC and ZS; Writing-original draft preparation: XL; Writing-review and editing: RC, CZ and ZS; Supervision: CZ and RC; Funding acquisition: RC and ZS. All authors have read and approved the final version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no financial or non-financial competing interests.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWaite, D.W., et al.: Addendum: Comparative Genomic Analysis of the Class \u003cem\u003eEpsilonproteobacteria\u003c/em\u003e and Proposed Reclassification to Epsilonbacteraeota (phyl. nov). Front. Microbiol. \u003cb\u003e9\u003c/b\u003e, 772 (2018). Erratum\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, S., et al.: Characterization of \u003cem\u003eSulfurimonas hydrogeniphila\u003c/em\u003e sp. nov., a Novel Bacterium Predominant in Deep-Sea Hydrothermal Vents and Comparative Genomic Analyses of the Genus Sulfurimonas. Front. Microbiol. \u003cb\u003e12\u003c/b\u003e (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan, Y., Perner, M.: The globally widespread genus \u003cem\u003eSulfurimonas\u003c/em\u003e: versatile energy metabolisms and adaptations to redox clines. Front. Microbiol. \u003cb\u003e6\u003c/b\u003e (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu, Q., Wang, S., Lai, Q., Shao, Z., Jiang, L.: \u003cem\u003eSulfurimonas indica\u003c/em\u003e sp. nov., a hydrogen- and sulfur-oxidizing chemolithoautotroph isolated from a hydrothermal sulfide chimney in the Northwest Indian Ocean. Int. J. Syst. Evol. MicroBiol. \u003cb\u003e71\u003c/b\u003e (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, S., et al.: \u003cem\u003eSulfurimonas sediminis\u003c/em\u003e sp. nov., a novel hydrogen- and sulfur-oxidizing chemolithoautotroph isolated from a hydrothermal vent at the Longqi system, southwestern Indian ocean. Antonie van Leeuwenhoek. \u003cb\u003e114\u003c/b\u003e, 813\u0026ndash;822 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, S., Jiang, L., Liu, X., Yang, S., Shao, Z.: \u003cem\u003eSulfurimonas xiamenensis\u003c/em\u003e sp. nov. and \u003cem\u003eSulfurimonas lithotrophica\u003c/em\u003e sp. nov., hydrogen- and sulfur-oxidizing chemolithoautotrophs within the \u003cem\u003eEpsilonproteobacteria\u003c/em\u003e isolated from coastal sediments, and an emended description of the genus \u003cem\u003eSulfurimonas\u003c/em\u003e. \u003cem\u003eInternational Journal of Systematic and Evolutionary Microbiology\u003c/em\u003e 70, 2657\u0026ndash;2663 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakai, K., et al.: Sulfurimonas paralvinellae sp. nov., a novel mesophilic, hydrogen- and sulfur-oxidizing chemolithoautotroph within the \u003cem\u003eEpsilonproteobacteria\u003c/em\u003e isolated from a deep-sea hydrothermal vent polychaete nest, reclassification of \u003cem\u003eThiomicrospira denitrificans\u003c/em\u003e as \u003cem\u003eSulfurimonas denitrificans\u003c/em\u003e comb. nov. and emended description of the genus \u003cem\u003eSulfurimonas\u003c/em\u003e. \u003cem\u003eInternational journal of systematic and evolutionary microbiology\u003c/em\u003e 56, 1725\u0026ndash;1733 (2006)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInagaki, F., Takai, K., Kobayashi, H., Nealson, K.H., Horikoshi, K.: \u003cem\u003eSulfurimonas autotrophica\u003c/em\u003e gen. nov., sp. nov., a novel sulfur-oxidizing ε-proteobacterium isolated from hydrothermal sediments in the Mid-Okinawa Trough. Int. J. Syst. Evol. MicroBiol. \u003cb\u003e53\u003c/b\u003e, 1801\u0026ndash;1805 (2003)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMolari, M., et al.: A hydrogenotrophic \u003cem\u003eSulfurimonas\u003c/em\u003e is globally abundant in deep-sea oxygen-saturated hydrothermal plumes. Nat. Microbiol. \u003cb\u003e8\u003c/b\u003e, 651\u0026ndash;665 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuttle, C.A.: Marine viruses - major players in the global ecosystem. Nat. Rev. Microbiol. \u003cb\u003e5\u003c/b\u003e, 801\u0026ndash;812 (2007)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRohwer, F., Prangishvili, D., Lindell, D.: Roles of viruses in the environment. Environ. Microbiol. \u003cb\u003e11\u003c/b\u003e, 2771\u0026ndash;2774 (2009)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCanchaya, C., Proux, C., Fournous, G., Bruttin, A., Br\u0026uuml;ssow, H.: Prophage genomics. Microbiol. Mol. Biol. Rev. \u003cb\u003e67\u003c/b\u003e, (2003). 238\u0026thinsp;\u0026ndash;\u0026thinsp;76, table of contents\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCasjens, S.: Prophages and bacterial genomics: what have we learned so far? Mol. Microbiol. \u003cb\u003e49\u003c/b\u003e, 277\u0026ndash;300 (2003)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaez-Espino, D., et al.: Uncovering Earth\u0026rsquo;s virome. Nature. \u003cb\u003e536\u003c/b\u003e, 425\u0026ndash;430 (2016)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, Z., et al.: Deep sea sediments associated with cold seeps are a subsurface reservoir of viral diversity. ISME J. (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJian, H., et al.: Diversity and distribution of viruses inhabiting the deepest ocean on Earth. ISME J. \u003cb\u003e15\u003c/b\u003e, 3094\u0026ndash;3110 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng, R., et al.: Virus diversity and interactions with hosts in deep-sea hydrothermal vents. Microbiome. \u003cb\u003e10\u003c/b\u003e, 235 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEmerson, J.B., et al.: Host-linked soil viral ecology along a permafrost thaw gradient. Nat. Microbiol. \u003cb\u003e3\u003c/b\u003e, 870\u0026ndash;880 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamargo, A.P., et al.: IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata. Nucleic Acids Res. \u003cb\u003e51\u003c/b\u003e, D733\u0026ndash;D743 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSimmonds, P., et al.: Consensus statement: Virus taxonomy in the age of metagenomics. Nat. Rev. Microbiol. \u003cb\u003e15\u003c/b\u003e, 161\u0026ndash;168 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eParks, D.H., Imelfort, M., Skennerton, C.T., Hugenholtz, P., Tyson, G.W.: CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. \u003cb\u003e25\u003c/b\u003e, 1043\u0026ndash;1055 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOlm, M.R., Brown, C.T., Brooks, B., Banfield, J.F.: dRep: a tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication. ISME J. \u003cb\u003e11\u003c/b\u003e, 2864\u0026ndash;2868 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArndt, D., Marcu, A., Liang, Y., Wishart, D.S., PHAST: PHASTER and PHASTEST: Tools for finding prophage in bacterial genomes. Brief. Bioinform. \u003cb\u003e20\u003c/b\u003e, 1560\u0026ndash;1567 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoux, S., Enault, F., Hurwitz, B.L., Sullivan, M.B.: VirSorter: mining viral signal from microbial genomic data. PeerJ \u003cb\u003e3\u003c/b\u003e, e985 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKieft, K., Zhou, Z., Anantharaman, K.: VIBRANT: automated recovery, annotation and curation of microbial viruses, and evaluation of viral community function from genomic sequences. Microbiome. \u003cb\u003e8\u003c/b\u003e, 90 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEddy, S.R., Accelerated Profile, H.M.M., Searches: PLoS Comput. Biol. \u003cb\u003e7\u003c/b\u003e, e1002195 (2011)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYutin, N., B\u0026auml;ckstr\u0026ouml;m, D., Ettema, T.J.G., Krupovic, M., Koonin, E.V.: Vast diversity of prokaryotic virus genomes encoding double jelly-roll major capsid proteins uncovered by genomic and metagenomic sequence analysis. Virol. J. \u003cb\u003e15\u003c/b\u003e, 67 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYutin, N., et al.: Varidnaviruses in the Human Gut: A Major Expansion of the Order Vinavirales. \u003cem\u003eViruses\u003c/em\u003e 14 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoux, S., et al.: Cryptic inoviruses revealed as pervasive in bacteria and archaea across Earth's biomes. Nat. Microbiol. \u003cb\u003e4\u003c/b\u003e, 1895\u0026ndash;1906 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamacho, C., et al.: BLAST+: architecture and applications. BMC Bioinform. \u003cb\u003e10\u003c/b\u003e, 421 (2009)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCouvin, D., et al.: CRISPRCasFinder, an update of CRISRFinder, includes a portable version, enhanced performance and integrates search for Cas proteins. Nucleic Acids Res. \u003cb\u003e46\u003c/b\u003e, W246\u0026ndash;w251 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaslett, D., Canback, B.: ARAGORN, a program to detect tRNA genes and tmRNA genes in nucleotide sequences. Nucleic Acids Res. \u003cb\u003e32\u003c/b\u003e, 11\u0026ndash;16 (2004)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGaliez, C., Siebert, M., Enault, F., Vincent, J., S\u0026ouml;ding, J.: WIsH: who is the host? Predicting prokaryotic hosts from metagenomic phage contigs. Bioinformatics. \u003cb\u003e33\u003c/b\u003e, 3113\u0026ndash;3114 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChaumeil, P.A., Mussig, A.J., Hugenholtz, P., Parks, D.H.: GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBolduc, B., Roux, S.: Clustering viral genomes in iVirus. (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNayfach, S., et al.: CheckV assesses the quality and completeness of metagenome-assembled viral genomes. Nat. Biotechnol. \u003cb\u003e39\u003c/b\u003e, 578\u0026ndash;585 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamargo, A.P., et al.: Identification of mobile genetic elements with geNomad. Nat. Biotechnol. (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evon Meijenfeldt, F.A.B., Arkhipova, K., Cambuy, D.D., Coutinho, F.H., Dutilh, B.E.: Robust taxonomic classification of uncharted microbial sequences and bins with CAT and BAT. Genome Biol. \u003cb\u003e20\u003c/b\u003e, 217 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHyatt, D., et al.: Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinform. \u003cb\u003e11\u003c/b\u003e, 119 (2010)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJang, B.: Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat. Biotechnol. \u003cb\u003e37\u003c/b\u003e, 632\u0026ndash;639 (2019)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuchfink, B., Xie, C., Huson, D.H.: Fast and sensitive protein alignment using DIAMOND. Nat. Methods. \u003cb\u003e12\u003c/b\u003e, 59\u0026ndash;60 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShannon, P., et al.: Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. \u003cb\u003e13\u003c/b\u003e, 2498\u0026ndash;2504 (2003)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBastian, M., Heymann, S., Jacomy, M.: \u003cem\u003eGephi: An Open Source Software for Exploring and Manipulating Networks\u003c/em\u003e, (2009)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNishimura, Y., et al.: ViPTree: the viral proteomic tree server. Bioinformatics. \u003cb\u003e33\u003c/b\u003e, 2379\u0026ndash;2380 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRobert, C.: \u0026amp; Edgar. MUSCLE: multiple sequence alignment with high accuracy and high throughput. \u003cem\u003eNucleic acids research\u003c/em\u003e (2004)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCapella-Guti\u0026eacute;rrez, S., Silla-Mart\u0026iacute;nez, J.M., Gabald\u0026oacute;n, T.: trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. \u003cb\u003e25\u003c/b\u003e, 1972\u0026ndash;1973 (2009)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMinh, B.Q., et al.: IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol. Biol. Evol. \u003cb\u003e37\u003c/b\u003e, 1530\u0026ndash;1534 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKalyaanamoorthy, S., Minh, B.Q., Wong, T.K.F., Haeseler, A.V., Jermiin, L.S.: ModelFinder: fast model selection for accurate phylogenetic estimates. Nat. Methods \u003cb\u003e14\u003c/b\u003e (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoang, D.T., Chernomor, O., von Haeseler, A., Minh, B.Q., Vinh, L.S.: UFBoot2: Improving the Ultrafast Bootstrap Approximation. Mol. Biol. Evol. \u003cb\u003e35\u003c/b\u003e, 518\u0026ndash;522 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, Z., et al.: Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. \u003cb\u003e379\u003c/b\u003e, 1123\u0026ndash;1130 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePettersen, E.F., et al.: UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci. \u003cb\u003e30\u003c/b\u003e, 70\u0026ndash;82 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShaffer, M., et al.: DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Res. \u003cb\u003e48\u003c/b\u003e, 8883\u0026ndash;8900 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePratama, A.A., et al.: Expanding standards in viromics: in silico evaluation of dsDNA viral genome identification, classification, and auxiliary metabolic gene curation. PeerJ \u003cb\u003e9\u003c/b\u003e, e11447 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYoshida-Takashima, Y., Takaki, Y., Shimamura, S., Nunoura, T., Takai, K.: Genome sequence of a novel deep-sea vent epsilonproteobacterial phage provides new insight into the co-evolution of Epsilonproteobacteria and their phages. Extremophiles: life under extreme conditions. \u003cb\u003e17\u003c/b\u003e, 405\u0026ndash;419 (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrum, J.R., Schenck, R.O., Sullivan, M.B.: Global morphological analysis of marine viruses shows minimal regional variation and dominance of non-tailed viruses. Isme j. \u003cb\u003e7\u003c/b\u003e, 1738\u0026ndash;1751 (2013)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKauffman, K.M., et al.: A major lineage of non-tailed dsDNA viruses as unrecognized killers of marine bacteria. Nature. \u003cb\u003e554\u003c/b\u003e, 118\u0026ndash;122 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRakonjac, J., Bennett, N.J., Spagnuolo, J., Gagic, D., Russel, M.: Filamentous bacteriophage: biology, phage display and nanotechnology applications. Curr. Issues Mol. Biol. \u003cb\u003e13\u003c/b\u003e, 51\u0026ndash;76 (2011)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu, Z.C., et al.: Filamentous phages prevalent in Pseudoalteromonas spp. confer properties advantageous to host survival in Arctic sea ice. \u003cem\u003eIsme j\u003c/em\u003e 9, 871\u0026thinsp;\u0026ndash;\u0026thinsp;81 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIlyina, T.S.: Filamentous bacteriophages and their role in the virulence and evolution of pathogenic bacteria. Mol. Genet. Microbiol. Virol. \u003cb\u003e30\u003c/b\u003e, 1\u0026ndash;9 (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYoshida, M., et al.: Quantitative Viral Community DNA Analysis Reveals the Dominance of Single-Stranded DNA Viruses in Offshore Upper Bathyal Sediment from Tohoku, Japan. Front. Microbiol. \u003cb\u003e9\u003c/b\u003e (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIgnacio-Espinoza, J.C., Fuhrman, J.A.: A non-tailed twist in the viral tale. Nature. \u003cb\u003e554\u003c/b\u003e, 38\u0026ndash;39 (2018)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKauffman, K.M., et al.: Resolving the structure of phage-bacteria interactions in the context of natural diversity. Nat. Commun. \u003cb\u003e13\u003c/b\u003e, 372 (2022)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTzipilevich, E., Habusha, M., Ben-Yehuda, S.: Acquisition of Phage Sensitivity by Bacteria through Exchange of Phage Receptors. Cell. \u003cb\u003e168\u003c/b\u003e, 186\u0026ndash;199e12 (2017)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYi, Y., et al.: A systematic analysis of marine lysogens and proviruses. Nat. Commun. \u003cb\u003e14\u003c/b\u003e, 6013 (2023)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBellas, C.M., Schroeder, D.C., Edwards, A., Barker, G., Anesio, A.M.: Flexible genes establish widespread bacteriophage pan-genomes in cryoconite hole ecosystems. Nat. Commun. \u003cb\u003e11\u003c/b\u003e, 4403 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaez-Espino, D., et al.: CRISPR immunity drives rapid phage genome evolution in Streptococcus thermophilus. mBio \u003cb\u003e6\u003c/b\u003e (2015)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, X., Jiao, N., Zhang, R.: The genomic content and context of auxiliary metabolic genes in roseophages. Environ. Microbiol. \u003cb\u003e23\u003c/b\u003e, 3743\u0026ndash;3757 (2021)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMara, P., et al.: Viral elements and their potential influence on microbial processes along the permanently stratified Cariaco Basin redoxcline. ISME J. \u003cb\u003e14\u003c/b\u003e, 3079\u0026ndash;3092 (2020)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu, M., et al.: Diversity and potential host-interactions of viruses inhabiting deep-sea seamount sediments. Nat. Commun. \u003cb\u003e15\u003c/b\u003e, 3228 (2024)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, J., et al.: Novel Viral Communities Potentially Assisting in Carbon, Nitrogen, and Sulfur Metabolism in the Upper Slope Sediments of Mariana Trench. \u003cem\u003emSystems\u003c/em\u003e 7, e0135821 (2022)\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Sulfurimonas, prophage, uncultured viral genomes, genomic and phylogenetic analysis, novel phage groups","lastPublishedDoi":"10.21203/rs.3.rs-4432365/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4432365/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe bacterial genus \u003cem\u003eSulfurimonas\u003c/em\u003e is globally widespread and occupies a key ecological niche in different habitats. However, phages infecting \u003cem\u003eSulfurimonas\u003c/em\u003e have never been isolated and characterized. Here we systematically investigated the genetic diversity, taxonomy and interaction patterns of \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages based on sequenced microbial genomes and metagenome datasets. High-confidence phage contigs related to \u003cem\u003eSulfurimonas\u003c/em\u003e were identified from various ecosystems, clustered into 61 viral operational taxonomic units across 3 viral realms. Most \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages were tailed viruses of \u003cem\u003eCaudoviricetes\u003c/em\u003e; these were assigned to 19 genus-level viral clusters, the majority of which were distantly related to previously known viruses. Phages encoding double jelly-roll major capsid proteins represented another group of double-stranded DNA phage with diverse gene compositions. Inovirus-like single-stranded DNA phages were primarily identified as integrated prophages or extrachromosomal viral genomes, suggesting chronic infections in hosts. Historical and current phage-host interactions were revealed, implying the viral impact on host evolution. Additionally, phages encoding auxiliary metabolic genes might benefit the infected bacteria by compensating or augmenting host metabolisms. This study highlights the remarkable diversity and novelty of \u003cem\u003eSulfurimonas\u003c/em\u003e-associated phages with highly divergent tailless lineages, providing basis for further investigation of phage-host interactions within this genus.\u003c/p\u003e","manuscriptTitle":"Genomic diversity of phages infecting the globally widespread genus Sulfurimonas","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-07-16 14:01:43","doi":"10.21203/rs.3.rs-4432365/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"communications-biology","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commsbio","sideBox":"Learn more about [Communications Biology](http://www.nature.com/commsbio/)","snPcode":"","submissionUrl":"","title":"Communications Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"940de0ff-6fad-4488-948f-6b3ab03f5ca2","owner":[],"postedDate":"July 16th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":33012747,"name":"Biological sciences/Microbiology/Virology/Metagenomics"},{"id":33012748,"name":"Biological sciences/Microbiology/Environmental microbiology/Water microbiology"}],"tags":[],"updatedAt":"2024-11-03T08:07:18+00:00","versionOfRecord":{"articleIdentity":"rs-4432365","link":"https://doi.org/10.1038/s42003-024-07079-4","journal":{"identity":"communications-biology","isVorOnly":false,"title":"Communications Biology"},"publishedOn":"2024-11-02 04:00:00","publishedOnDateReadable":"November 2nd, 2024"},"versionCreatedAt":"2024-07-16 14:01:43","video":"","vorDoi":"10.1038/s42003-024-07079-4","vorDoiUrl":"https://doi.org/10.1038/s42003-024-07079-4","workflowStages":[]},"version":"v1","identity":"rs-4432365","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4432365","identity":"rs-4432365","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.