Microbial Genome-Resolved Metaproteomic Analyses Frame Intertwined Carbon and Nitrogen Cycles in River Hyporheic Sediments | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Microbial Genome-Resolved Metaproteomic Analyses Frame Intertwined Carbon and Nitrogen Cycles in River Hyporheic Sediments Josué A. Rodríguez-Ramos, Mikayla A. Borton, Bridget B. McGivern, and 11 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-746574/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Rivers serve as a nexus for nutrient transfer between terrestrial and marine ecosystems and as such, have a significant impact on global carbon and nitrogen cycles. In river ecosystems, the sediments found within the hyporheic zone are microbial hotspots that can account for a significant portion of ecosystem respiration and have profound impacts on system biogeochemistry. Despite this, studies using genome-resolved analyses linking microbial and viral communities to nitrogen and carbon biogeochemistry are limited. Results: Here, we characterized the microbial and viral communities of Columbia River hyporheic zone sediments to reveal the metabolisms that actively cycle carbon and nitrogen. Using genome-resolved metagenomics, we created the Hyporheic Uncultured Microbial and Viral (HUM-V) database, containing a dereplicated database of 55 microbial Metagenome-Assembled Genomes (MAGs), representing 12 distinct phyla. We also sampled 111 viral Metagenome Assembled Genomes (vMAGs) from 26 distinct and novel genera. The HUM-V recruited metaproteomes from these same samples, providing the first inventory of microbial gene expression in hyporheic zone sediments. Combining this data with metabolite data, we generated a conceptual model where heterotrophic and autotrophic metabolisms co-occur to drive an integrated carbon and nitrogen cycle, revealing microbial sources and sinks for carbon dioxide and ammonium in these sediments. We uncovered the metabolic handoffs underpinning these processes including mutualistic nitrification by Thermoproteota (formerly Thaumarchaeota) and Nitrospirota, as well as identified possible cooperative and cheating behavior impacting nitrogen mineralization. Finally, by linking vMAGs to microbial genome hosts, we reveal possible viral controls on microbial nitrification and organic carbon degradation. Conclusions: Our multi-omics analyses provide new mechanistic insight into coupled carbon-nitrogen cycling in the hyporheic zone. This is a key step in developing predictive hydrobiogeochemical models that account for microbial cross-feeding and viral influences over potential and expressed microbial metabolisms. Furthermore, the publicly available HUM-V genome resource can be queried and expanded by researchers working in other ecosystems to assess the transferability of our results to other parts of the globe. Applied & Industrial Microbiology General Microbiology microbiome viruses metagenomics river sediment denitrification nitrification carboxydotrophs Thaumarchaeota peptidases Binatia Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Background The hyporheic zone (HZ) acts as a transitional space between river and groundwater compartments in the river corridor where the bidirectional supply of nutrients and organic carbon stimulate microbial activity [ 1 – 3 ]. Characterized as the permanently saturated interface between the river surface channel and underlying sediments, the HZ is considered a biogeochemical hotspot for microbial cycling of carbon and nitrogen [ 1 , 3 ]. These zones have been reported to support microbial heterotrophic respiration, denitrification, and nitrification, as well as the consumption and production of greenhouse gases (such as nitrous oxide and carbon dioxide) [ 4 – 7 ]. In addition to harboring diverse energy metabolisms, these sediments may act as an important sink of carbon and nitrogen in terms of microbial subsurface biomass [ 4 , 8 ]. Overall estimates of ecosystem respiration have revealed that the HZ accounts for 40 to 90% of total river respiration [ 9 , 10 ], highlighting that a substantial amount of respiration is associated with hyporheic microbial activities. Despite the importance of microbial metabolism to river corridor biogeochemistry, the ability to partition metabolic handoffs between organisms, the linked use of carbon and nitrogen by individual organisms, and the mineralization of detritus have yet to be holistically interrogated. Metagenomic studies in river sediments have not fully inventoried carbon and nitrogen cycling metabolisms, instead focusing on specific aspects of the nitrogen cycle (e.g., genes in denitrification [ 11 ]). Moreover, most of these studies were not genome resolved, hindering the assignment of biogeochemical processes to specific microorganisms, and culture independent genomic reconstructions from river sediments are limited to a handful of studies [ 12 , 13 ], all of which focused exclusively on nitrification. Thus, little is known about uncultured microbial communities in river sediments, with the enzymes, interconnected chemical reactions, and microbial metabolic lifestyles mediating carbon and nitrogen transformations in river sediments not currently resolvable from existing HZ microbiome datasets. Here we address this knowledge gap, creating a genome-resolved inventory of the microbial and viral members in HZ sediments collected from the Columbia River in Washington State, USA (Fig. 1 ). This resource was used to recruit metaproteomic data, providing a first of its kind, comprehensive inventory of the active microbial organisms and their enzymatic machinery in river systems. We contextualized these biological findings, using chemical scaffolding provided from metabolomics and geochemistry. Reconstructing the expressed metabolic capabilities of numerous lineages enabled us to resolve microbial contributions to biogeochemistry in these sediments. Our proteome enabled road map outlined the metabolic circuitry coupling carbon and nitrogen biogeochemistry in these HZ sediments, providing a framework to develop hydrobiogeochemical models informed by biochemical mechanisms and ecological interactions. Methods Experimental Design To investigate the microbial processes involved in biogeochemical cycling in HZ sediments, we leveraged previously collected geochemical, metagenomic, metaproteomic, and Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and combined it with new nuclear magnetic resonance (NMR) metabolomic characterization and additional metagenomic sequencing. Samples were collected from the hyporheic zone of the Columbia River (46°22’15.80″N, 119°16’31.52″W) in March 2015 as previously described [ 14 ]. Briefly, liquid nitrogen frozen sediment profiles (0-60cm) were collected along two transects separated by approximately 170 meters (Fig. 1 a). At each transect, three sediment cores up to 60 cm in depth were collected at 5-meter intervals perpendicular to the river flow. All cores were collected during conditions in which the sediments were fully saturated. Each core was sectioned into 10 cm segments from 0-60-centimeter depths and stored at -80ºC. Analyses were carried out at 10-centimeter increments, with the exception of one core that was pooled from 0–30 centimeters to have sufficient input masses. Collectively, this robust, paired multi-omic dataset was made up of previously reported metagenomes (n = 33, 3-4Gbp), metaproteomes (n = 33), FTICR-MS metabolomes (n = 33), and geochemical characterizations (n = 33), as well as new metagenomes (n = 10, 10-25Gbp) and NMR metabolomes (n = 17) (Fig. 1 b). Genome recovery from these samples was improved by increasing the metagenomic sequencing depth per sample (from an average of 3.8 to 25.3 Gbp for selected samples) and employing a hybrid of co-assembly and single assembly methods (Fig. 1 c). DNA extraction and sequencing As described previously [ 14 ], deoxyribonucleic acid (DNA) was extracted from the sediments using the MoBio PowerSoil kit (MoBio Laboratories, Inc., Carlsbad, CA) following manufacturer's instructions, with the addition of a 2-hour proteinase-K incubation at 55°C prior to bead-beating to facilitate cell lysis. Purified genomic DNA was sent to the Joint Genome Institute (JGI, n = 33) under Joint Genome Institute / Environmental Molecular Sciences Laboratory (EMSL) proposal 1781 or to the Genomics Shared Resource facility at The Ohio State University (OSU, n = 10), producing 43 metagenomes with average sequencing depth of 4 (JGI) and 25 Gbp (OSU) per sample, totaling 377Gbp. DNA submitted to JGI were prepared for sequencing using an Illumina Library creation kit (KAPA Biosystems), and then solid-phase reversible immobilization size selection. DNA submitted to OSU were prepared for sequencing with a Nextera XT library System followed by solid-phase reversible immobilization size selection. Libraries at both facilities were quantified to ensure input thresholds, and then sequenced using an Illumina HiSeq 2500 platform. Deeper sequencing was performed at OSU to enhance MAG recovery, resulting in an increase of 252 Gbp of additional sequencing for ten samples, increasing the sequencing depth per sample by at least 3-fold (15.37–49.24 Gbp per sample) ( Fig. 1 , Additional File 1) . Additional File 1 details all sequencing information, including National Center for Biotechnology Information (NCBI) accession numbers. Metagenome assembly and binning Raw reads were trimmed for length and quality using Sickle v1.33 ( https://github.com/najoshi/sickle ) and then subsequently assembled using Iterative De Bruijn Graph De Novo Assembler – Uneven Depth (IDBA-UD) 1.1.0 [ 15 ] with an initial kmer of 40 or metagenomic St. Petersburg genome Assembler (metaSPAdes) 3.13.0 [ 16 ] with default parameters. To further increase genomic recovery, for the ten samples that had shallow and deep sequencing, metagenomic reads were coassembled using IDBA-UD 1.1.0 with an initial kmer of 40. All assemblies, including co-assemblies, were then individually binned using Metabat2 [ 17 ] with default parameters to obtain Metagenome Assembled Genomes (MAGs). For each bin, genome completion was estimated based on the presence of core gene sets (highly conserved genes that occur in single copy) for Bacteria (n = 31 genes) and Archaea (n = 104 genes) using Amphora2 [ 18 ]. Bins were discarded for further analysis if completion was 10% to select for only medium to high quality bins [ 19 ]. This resulted in 102 MAGs that were then dereplicated using dRep [ 20 ] with default parameters and resulted in a final set of 55 MAGs (> 99% ANI). To further assess bin quality, we used the Distilled and Refined Annotation of MAGs (DRAM) [ 21 ] annotation pipeline to identify ribosomal ribonucleic acids (rRNAs) and transfer ribonucleic acids (tRNAs). The 102 MAGs detailed here are deposited on NCBI under the BioProject ID PRJNA576070, with genome quality information reported in Additional File 2. Phylogenetic and metabolic analysis of Metagenome Assembled Genomes (MAG) Medium and high-quality MAGs were taxonomically classified using the Genome Taxonomy Database (GTDB) Toolkit v1.3.0 on September 2020 [ 22 ]. Novel taxonomy was identified as the first taxonomic level with no designation using GTDB taxonomy. For example, MAGs whose GTDB taxonomy string was not designated after the family level (e.g., g__) were identified as novel genera. Of the MAGs that had multiple representatives sharing taxonomy strings up to the Family level (Binatia and CSP1-3), we used average nucleotide identity (ANI) to determine whether they belonged to the same genus. MAG scaffolds were annotated using the DRAM pipeline [ 21 ]. The raw annotations for each genome are deposited in the Zenodo repository under doi 10.5281/zenodo.5128772 and can be accessed here: https://doi.org/10.5281/zenodo.5128772 . Additional File 2 shows the metabolic summary of genomes (product DRAM output) and output is also displayed in Additional File 3: Figure S1 . Target metabolic marker genes of interest recovered in bins were used to query the Integrated Microbial Genomes / Microbiomes Expert Review (IMG/M ER) ( https://img.jgi.doe.gov/cgi-bin/mer/main.cgi ) and NCBI ( https://www.ncbi.nlm.nih.gov ) databases using BLASTp, or retrieved from the Protein Family (PFAM) database by protein family. Returned amino acid sequences were compiled with other known genes not retrieved via sequence homology, and de-replicated to make a reference sequence database. Sequences from the metagenomes were then aligned to the reference sequences using Multiple Sequence Comparison by Log-Expectation (MUSCLE) version 3.8.31 [ 23 ] or Multiple Alignment using Fast Fourier Transform (MAFFT) version 7.427 [ 24 ]. Alignments were manually curated to remove end and other gap regions. These alignments were then used to construct phylogenetic trees using FastTree version 2.1.11 [ 25 ] with default settings. An additional phylogenetic analysis was performed on genes annotated as respiratory nitrate reductase ( nar ) and nitrite oxidoreductase ( nxr ) to resolve novel Binatia role in nitrogen cycling. Specifically, sequences from [ 26 ] were downloaded and combined with nar and nxr amino acid sequences from dereplicated bins, aligned using MUSCLE, version 3.8.31, and run through ProtPipeliner, a Python script developed in-house for generation of phylogenetic trees ( https://github.com/TheWrightonLab ). Phylogenetic trees are shown in Additional File 3: Figure S2 and Additional File 4 . For polyphenol and organic polymer degradation, we used functional annotation in addition to predicted secretion to assess functional potential. To determine if the predicted genes encoded a secreted protein, we used pSortb [ 27 ] and SignalP [ 28 ] to predict location; if those methods did not detect a signal peptide, the amino acid sequence was queried to SecretomeP and a SecP score > 0.5 [ 29 ] was used as a threshold to report non-canonical secretion signals. Metabolic information for each MAG discussed in this manuscript are available in Additional File 2 and Additional File 5 . Viral Analyses Metagenomic assemblies (n = 43) were screened for DNA viral sequences using VirSorter v1.0.3 with the ViromeDB database option [ 30 ], retaining viral contigs ranked 1, 2, 4 or 5 with greater than 10kb in genome length as stated by the Minimum Information about an Uncultivated Virus Genome (MIUViG) standards [ 31 ]. To determine an approximate species level taxonomy for viral scaffolds, they were clustered into viral metagenome assembled genomes (vMAGs) at 95% ANI across 85% of the shortest contig using ClusterGenomes 5.1 ( https://github.com/simroux/ClusterGenome ) [ 31 ]. After clustering, vMAGs were manually confirmed to be viral by assessing the total of viral genes with regards to non-viral genes in the genome, where genomes containing more than 18% of non-viral genes were discarded (J flag, DRAM [ 21 ]).This resulted in 111 vMAGs that were deposited on NCBI under the BioProject ID PRJNA576070 Additional File 6. To determine taxonomic affiliation, vMAGs were clustered to viruses belonging to standard viral reference taxonomy databases NCBI Bacterial and Archaeal Viral RefSeq V85 with the International Committee on Taxonomy of Viruses (ICTV) and NCBI Taxonomy using the network-based protein classification software vContact2 v0.9.8 [ 32 , 33 ]. Default methods were used. To determine geographic distribution of viruses in freshwater ecosystems, we also included viruses mined from publicly available freshwater metagenomes in vContact2 analyses: 1) East River, CO (PRJNA579838) 2) A previous study from the Columbia River, WA (PRJNA375338) 3) Prairie Potholes, ND (PRJNA365086) and 4) the Amazon River (PRJNA237344). Viral contigs were annotated with DRAM-v [ 21 ], with annotations for each viral genome reported in Additional File 6 . Genes that were identified by DRAM-v as being possible auxiliary metabolic genes (categories 1–3) were subjected to protein modeling using Protein Homology / AnalogY Recognition Engine (PHYRE2) in order to improve the accuracy of annotation [ 34 ]. To identify likely vMAG hosts, oligonucleotide frequencies between virus (n = 111) and non-dereplicated hosts (n = 102) were analyzed using VirHostMatcher using a threshold of d2* measurements of < 0.25 [ 35 ]. Per genome results of VirHostMatcher predictions are reported in Additional File 6 . The lowest d2* value for each viral contig < 0.25 was used. Viruses and hosts could not be linked by matching clustered regularly interspaced short palindromic repeats (CRISPR) spacers in host genomes to vMAG genomes. Genome relative abundance calculations Relative abundance for each MAG and vMAG was estimated using in-house scripts available at https://github.com/TheWrightonLab [ 36 , 37 ]. Briefly, all metagenomic reads were concatenated for each sample, rarified to 3Gbp, and multi-mapped to 55 unique MAGs via Bowtie2 [ 38 ]. For MAGs, a minimum scaffold coverage of 75% and depth of 3x required for read recruitment at 7 mismatches. For vMAGs, reads were mapped using Bowtie2 [ 38 ] at a maximum mismatch of 15, a minimum contig coverage of 75% and a minimum depth coverage of 2x. Relative abundances for each MAG and vMAG were calculated as their coverage proportion from the sum of the whole coverage of all bins for each set of metagenomic reads. To identify correlations between MAGs and vMAGs to geochemistry and predictive capability, we also mapped the subset of deep sequencing (n = 10) and rarefied to the lowest deep metagenome available (4.8Gbp). Correlations and sparse Partial Least Squares Regression (sPLS) predictions (PLS R package [ 39 ]) were done using mapping data pertaining to only the 10 deeply sequenced metagenomes. Genome relative abundances per sample for MAGs and vMAGs are reported in Additional File 2 and Additional File 6 . Metaproteome generation and peptide mapping Sediment samples were prepared for metaproteome analysis as previously reported in Graham et al. 2018 [ 14 ] and the protocol outlined by Nicora et al [ 40 ]. For protein identification, spectra were searched against two files that included (i) 55 dereplicated MAG and (ii) 111 clustered vMAGs amino acid sequences. Exact sequence duplicates were removed, and 16 commonly observed contaminants (e.g., tryptic fragments, human keratins, and serum albumin precursors) were included. The tandem mass spectrometry (MS/MS) spectra from all liquid chromatography tandem mass spectrometry (LC-MS/MS) datasets were converted to ASCII text (.dta format) using MSConvert ( http://proteowizard.sourceforge.net/tools/msconvert.html ) which more precisely assigns the charge and parent mass values to an MS/MS spectrum. The data files were then interrogated via target-decoy approach [ 41 ] using MSGF+ [ 42 ] using a ± 20 ppm parent mass tolerance, partially tryptic digestion enzyme settings, and a variable posttranslational modification of oxidized Methionine. All MS/MS search results for each dataset were collated into tab separated ASCII text files listing the best scoring identification for each spectrum. Collated search results were further combined into a single result file. These results were imported into a Microsoft SQL Server database. Results were filtered to N1% false detection rate (FDR) using an MSGF + supplied Q-Value that assesses reversed sequence decoy identifications for a given MSGF score across each dataset. Using the protein references as a grouping term, unique peptides belonging to each protein were counted, as were all peptide spectrum matches (PSMs) belonging to all peptides for that protein (i.e., a protein level observation count value). PSM observation counts were reported for each sample that was analyzed. Crosstabulation tables were created to enumerate protein level PSM observations for each sample, allowing low-precision quantitative comparisons to be made. Microbial metaproteomes were converted to normalized spectral abundance frequency (NSAF) values and subsequently divided into unique, non-unique specialized, and non-unique categories, while viral metaproteomes were analyzed using peptide counts only from unique hits due to low recruitment [ 36 ]. Peptide recruitment for each MAG amino acid sequence per sample is reported in Additional File 5 . Hits were divided into 3 categories: (1) uniques (peptide hits to a single protein), (2) non-unique specialized (peptide hits to multiple amino acid sequences that all had same annotation and MAG taxonomy), (3) non-unique (peptide hits to multiple amino acid sequences with different annotation or from MAGs with different taxonomy) [ 43 ]. This designation was necessary as several hits could not be resolved to the MAG level due to functional conservation across closely related genomes in the HUM-V database. Data in Fig. 2 showcases (1) and (2) categories, with the entire dataset shown in Additional File 3: Figure S3 . Including the non-unique specialized hits assigned an additional 14% of the proteome (grey bar, Fig. 2 b) and confirmed we did not underrepresent the gene expression from genomically well-sampled strains (e.g., Nitrospiraceae). Metaproteome hits for MAGs were used for further metabolic analyses if they were detected in at least three samples. Annotations for the entire metaproteomic dataset are shown in Additional File 3: Figure S4 . Geochemical measurements, FTICR-MS characterization of organic matter, and NMR detected metabolites. As previously reported [ 14 ], total nitrogen, total carbon, and total sulfur were determined using Elementar vario EL cube (Elementar Co., Germany), with details in the Supplementary Information ( Additional File 7 ). To characterize organic matter, we used FTICR-MS to analyze sediments as previously reported [ 14 ], with details in the Supplementary Information ( Additional File 3: Figure S5, Additional File 8 ). To identify the metabolites available to microorganisms in this river system, we performed 1 H Nuclear Magnetic Resonance (NMR) spectroscopy on sediment pore water. Sediment samples were mixed with 200, 300, or 600 µL of MilliQ water depending on the sediment mass ( Additional File 7 ) and centrifuged to remove the sediment. Supernatant (180 µL) was then diluted by 10% (vol/vol) with 5 mM 2,2-dimethyl-2-silapentane-5-sulfonate- d 6 as an internal standard. All NMR spectra were collected using a Varian Direct Drive 600-MHz NMR spectrometer equipped with a 5-mm triple resonance salt-tolerant cold probe. Chemical shifts were referenced to the 1H or 13C methyl signal in DSS-d6 at 0 ppm. The 1D 1 H NMR spectra of all samples were processed, assigned, and analyzed using Chenomx NMR Suite 8.3 with quantification based on spectral intensities relative to the internal standard as described previously [ 36 , 44 ]. Candidate metabolites present in each of the complex mixtures were determined by matching the chemical shift, J-coupling, and intensity information of experimental NMR signals against the NMR signals of standard metabolites in the Chenomx library. Compounds were assigned a rank and assign confidence to metabolites (RANCM) value according to the amount of spectral information used to identify the compound ( Additional File 7 ) [ 45 ]. For many metabolites, including aspartate, asparagine, sucrose, acetate, methanol, and glucose, we utilized 2D NMR to corroborate the 1D data, providing more confidence to an assignment. The two-dimensional 1H-1H total correlation spectroscopy (TOCSY) spectra were collected using the Varian TOCSY pulse sequence with a TOCSY mixing time of 80 ms (MLEV-17). Spectral widths were 12 ppm in both directions with 256 increments acquired in the indirect dimension and 64 transients per increment. The relaxation delay was 1.5 s during which presaturation of the water signal was applied and the acquisition time was 143 ms during which 2048 total points were acquired. The 2D 1H-13C heteronuclear single-quantum correlation spectroscopy (HSQC) spectra were acquired using the Varian gHSQCAD pulse sequence with a 1JCH of 146 Hz. Spectral widths were 12 ppm and 160 ppm for the direct and indirect dimensions, respectively, with 256 increments acquired in the indirect dimension and 128 transients per increment. The relaxation delay was 1.5 s during which presaturation of the water signal was applied. The acquisition time was 143 ms in which 13C composite pulse decoupling (wurst140) was applied and 2048 total points were acquired. NMR-identified metabolites discussed in the text were present in 30% of the samples. Results And Discussion The HUM-V genome database enabled metaproteomic characterization of river sediment microbiomes Here we created the Hyporheic Uncultured Microbial and Viral (HUM-V) genomic catalog from Columbia River HZ sediments. We leveraged this resource for metaproteomic peptide recruitment, enabling identification of the community members and their gene expression in these sediments. We reconstructed 655 bacterial and archaeal metagenome assembled genomes (MAGs); 102 were medium or high-quality genomes based on the Genome Consortium Standards [ 19 ] ( Additional File 2 ). These genomes were dereplicated into 55 genomic representatives to form the bacterial and archaeal portion of the HUM-V microbial genome database. These dereplicated HUM-V MAGs were distributed across 9 Bacterial and 2 Archaeal phyla. In terms of new genomic discoveries, 1 genome represented a new order within the Actinobacteriota, and 12 genomes represented 6 new genera from archaeal and bacterial phyla including members of the Thermoplasmatota, Acidobacteria, Actinobacteriota, CSP1-3, Proteobacteria, and Desulfobacterota (Fig. 1 b, Fig. 2 a). From the same metagenomic assemblies we reconstructed and reported viral metagenome assembled genomes (vMAGs), making this one of only a handful of genome-resolved studies that include viral genomes derived from rivers [ 46 – 48 ], and to our knowledge, the first study to complement these with bacterial and archaeal genomes. We reconstructed 2,482 vMAGs that dereplicated into 111 dereplicated viral populations > 10kb in size ( Additional File 6 ). Given their sparse sampling from river corridors, only 5 of the HUM-V viral genomes had taxonomic assignments using established viral taxonomies from standard reference databases. To better understand if the remaining 95% (n = 105) of viral genomes were completely novel or had been previously detected in similar ecosystems, we repeated the analyses, this time adding 1,861 viral genomes we reconstructed or pulled from public metagenomes from four freshwater sites in North and South America (Fig. 2 c, Additional File 6 ). Of the 105 remaining viral genomes in HUM-V, 15% (n = 17) clustered with these freshwater derived sequences, indicating a portion of this viral community is shared across diverse geographic and freshwater systems. Of the remaining viral genomes, 23% (n = 26) clustered only with genomes recovered in this data set, indicating multiple samplings of the same virus spatially at this site, while 57% (n = 63) of the viral genomes we sampled were singletons (i.e., only sampled from these sediments once). These results hint at the possible cosmopolitan and endemic viral lineages that warrant further exploration. HUM-V recruited viral and microbial peptides from our HZ sediment metaproteomic dataset (n = 33 lateral and depth resolved samples) (Fig. 2 b d, Additional File 5 ). Across all sediment samples, microbial genomes recruited 13,102 total peptides to ~ 1,300 proteins in HUM-V, with 68% of these proteins uniquely assigned to a single microbial genome. For viruses and microbes alike, the most abundant genomes were not necessarily the most actively expressing proteins. The most abundantly ranked microbial members included the Nitrospiraceae genus NS7, Binatia, and Nitrososphaeraceae genus TA-21 (Fig. 2 b), yet only the Nitrososphaeraceae had high proteomic recruitment (15%). Similarly, some low abundance members (e.g., members of the Actinobacteria) accounted for a majority of the uniquely assigned proteome relative abundance (49%) (Fig. 2 b). Like our microbial dataset, 66% of the viral genomes encoded genes that uniquely recruited peptides (Fig. 2 d). This exceeded prior viral metaproteome recruitment from other environmental systems (e.g., wastewater, saliva, rumen (0.4–15%, [ 49 – 51 ]), thus we infer a relatively large portion of the viral community was active at the time of sampling. While microbial and viral activity did not appear to be structured by transect, sediment depth, or geochemical conditions, these two assemblages were coordinated to one another ( Additional File 3: Figure S6 ). Explaining this lack of geochemical or spatial structuring, it is possible that the microbial heterogeneity in these samples occurred over a finer spatial resolution (pore or biofilm scale) than the bulk 10 cm depths sampled or that these HZ sediment microbiomes are metabolically robust to the small, but significant changes in chemistry measured across spatial gradients ( Additional File 3: Figure S6, Additional File 3: Figure S7, Additional File 7 ). Microbial cross feeding of organic carbon is likely sustained by aerobic respiration It is well recognized that microbial carbon oxidation in HZ sediments largely contributes to river respiration, yet the microbial food webs underpinning this process have yet to be documented. Consistent with resazurin (raz) data (see Additional File 3: supplementary methods) that indicated these sediments were oxygenated and supported aerobic microbial respiration ( Additional File 3: Figure S7 ) [ 52 ], all but one of the microbial genomes recovered from this site encoded aerobic respiration machinery, including a complete electron transport chain and a cytochrome oxidase ( Additional File 3: Figure S1 ). Proteomic evidence for aerobic respiration (cytochrome c oxidase aa3 ) was detected from nearly all samples, but only assigned to few members of the Nitrososphaeraceae. However, given limitations with detecting membrane cytochromes [ 53 ], we consider it likely this metabolism was more active than was captured in proteomic data, as we failed to find any evidence for other anaerobic metabolisms (e.g., methanogenesis). While the overall carbon content of these sediments was low (< 10 mg/g) ( Additional File 7 ), our FTICR-MS analysis indicated that plant litter could be an important substrate, as lignin-like compounds were the most abundant biochemical class detected ( Additional File 3: Figure S5 , Additional File 8 ). In support of this, from our metagenomes, 38% of the HUM-V genomes encoded genes for potentially degrading phenolic/aromatic monomers, while 10% could degrade the larger, more recalcitrant polymers ( Additional File 2 ). Gene expression of carbohydrate-active enzymes (CAZymes) also supported the degradation of plant polymers like starch and cellulose via extracellular glucoamylase (GH15) and endo-glucanase (GH5) from an actinobacterial genome (Microm_1) and the Nitrososphaeraceae (Nitroso_2), respectively (Fig. 3 ). In summary, many types of chemical and biological data reveal that heterotrophic, aerobic metabolism in these low carbon sediments is likely maintained by inputs from decomposition. Given the capacity for plant polymer decomposition (e.g., lignin, cellulose, and starch) across HUM-V genomes, we next tracked the microbial fate of the degradation products of these metabolisms, including sugar monomers, short chain fatty acids, and carbon dioxide (Fig. 3 , Additional File 2, Additional File 5 ). Metabolites detected by NMR included sugars (e.g., glucose, sucrose, and trehalose), which could be the result of depolymerization of plant derived polymers, and we confirmed the CAZYmes to use these substrates were also expressed in situ . Additionally, NMR also detected organic acids (acetate, butyrate, lactate, pyruvate, propionate) and alcohols (ethanol, methanol, isopropanol), with proteomics supporting the usage of acetate and methanol by Anaeromyxobacter MAG (Anaerom_1) and archaeal Woeseia (Woese_1), respectively. Here our metabolite and proteomic data demonstrated that plant biomass degradation supports sequential metabolic handoffs that lead to carbon dioxide production. Carbon dioxide production and consumption is widely encoded by HUM-V microorganisms In addition to carbon dioxide being generated from the heterotrophic metabolisms described above, our proteomics revealed that carbon dioxide could arise by the aerobic oxidation of carbon monoxide (CO). Genes for aerobic CO dehydrogenases (from Actinobacteria, Binatia, and CSP-1 genomes) were among the most expressed in these sediments. Analogous to findings from soil systems, it is possible that atmospheric carbon monoxide is a major energy source supporting persistent aerobic heterotrophic bacteria in deprived, or dynamic organic carbon environments [ 54 ]. Based on the genomic inventory of these HUM-V genomes, we posit that Binatia, CSP1-3, and Micromonosporaceae are capable of carboxydotrophy, while Actino_1 is a carboxydovore, using CO metabolism as supplemental energy or possible carbon source during starvation [ 54 ]. Since heterotrophic respiration and carbon monoxide oxidation would generate carbon dioxide in these sediments, we next tracked microorganisms that could use this carbon source autotrophically (Fig. 3 , Additional File 2, Additional File 5 ). The ability to fix carbon was prevalent, encoded by 75% of HUM-V microbial genomes. In fact, this metabolism was represented by multiple fixation pathways from 18 different lineages, demonstrating both functional and taxonomic redundancy. Specifically, this includes 4 different pathways (e.g., Calvin-Benson-Bassham cycle, reductive TCA cycle, 3-HydroxyPropionate /4-HydroxyButyrate cycle, 3-Hydroxypropionate bi-cycle) from members of nitrifying lineages (Thaumarchaeota and Nitrospirota) (discussed below), as well as from organisms with heterotrophic capabilities like Binatia, CSP1_3, Proteobacteria, Woeseiaceae, and Acidobacteria ( Additional File 3: Figure S1 , Additional File 2 ). Collectively our multi-omics data suggest that sediment microbial respiration is likely decoupled from river respiration, since some microbially produced carbon dioxide would be lost to supporting autotrophy. Our research further resolves the carbon economy in HZ sediments, implying that the net effect of carbon dioxide emissions from rivers could depend on the balance between carbon dioxide production from heterotrophy and carbon monoxide, as well as consumption by autotrophs. Microbial metaproteomics supports theoretical inferences derived from geochemistry The ratio of total element carbon (C) and total nitrogen (N) (e.g., C/N) is a geochemical indicator often used to assess the possible microbial metabolisms that can be supported in an ecosystem [ 55 , 56 ]. Here the C/N ratios of these sediments were relatively low to other sediments at 6.4 ± 1.1 across the samples ( Additional File 7 ). Biogeochemical theory posits that C/N values less than 15 would indicate rapid microbial mineralization of organic nitrogen to release inorganic nitrogen [ 57 ]. This theory also states that C/N ratios less than 10 may indicate ammonium is released to the surrounding environment, allowing sufficient concentrations to simultaneously support the assimilatory needs of heterotrophs and energy needs of nitrifiers, allowing for their co-occurrence [ 56 ]. Our multi-omics data offered a new opportunity to substantiate these geochemical inferences by profiling the possible substrates and microbial activity of nitrogen mineralizers and nitrifiers in river sediments. Given the prevalence of ammonium in all 33 sediment samples (0.28–11.22 µg gram − 1 ) ( Additional File 3: Figure S8, Additional File 7 ), we next examined our metaproteomic data for peptidases, genes that could contribute to the mineralization of organic nitrogen into amino acids and free ammonium. Hinting at the relevance of this metabolism, the gene expression of peptidases (n = 31) was 3 times more abundant and prevalent than glycoside hydrolase genes modulating organic C transformations ( Additional File 5 ). In support of active microbial N mineralization, hydrophobic, polar, and hydrophilic amino acids were prevalent (more so than sugars) in the H 1 -NMR characterized metabolites ( Additional File 3: Figure S8 ). We focused our analyses on the putative extracellular peptidases, as these were most likely to shape organic nitrogen pools in the sediment. We categorized expressed peptidase families as either amino acid releasing (end terminus cleaving, e.g., M28) or peptide releasing (endocleaving, e.g., S08A, M43B, M36, MO4) (Fig. 4 , Additional File 5 ). Linking these expressed peptidases to our genomes, members of the Actinobacteria, Thermoproteota, and Methylomirabilota, and Binatia are likely candidates for driving the mineralization of organic N. We then profiled amino acid transporters that were expressed, revealing uptake of branched chain amino acids, glutamate, osmoprotectants, spermidine/putrescine, and peptides (Fig. 4 ). This profiling indicated synergy and competitions for this organic N resource in these sediments. We propose that in HZ sediments extracellular peptidases are a shared public good whose cost of production is assumed by certain individuals with benefits to the entire community [ 58 ]. In some cases, taxa that mineralized organic N were consumers of the resulting products, as genomes in the Actinobacteria and Binatia expressed external peptidases genes and the genes for transporting the organic N products (Fig. 4 , linkages shown). In other instances, members of the Proteobacteria, Thermoplasmatota, and CSP1-3 could be functioning as cheater cells that expressed only genes for intracellular transport and benefitting from peptidases produced by others. Our findings reinforce that cooperative interactions based on cross-feeding and public goods are likely at the core of many processes relevant to organic carbon (Fig. 3 ) and nitrogen (Fig. 4 ) cycling in these sediments. Consistent with established conceptual geochemical theory, we showed the lower C:N ratios (< 10) of these sediments not only supported mineralization which could be a source of free ammonium in these sediments, but also nitrification. Supporting this, ammonium was detected in all sediments (average concentration 2.6 µg/gram of sediment) ( Additional File 3: Figure S8, Additional File 7 ). Proteomics confirmed ammonium (NH 4 + ) oxidation to nitrite was performed by Archaeal Nitrososphaeraceae (formerly Thaumarchaeota), with ammonia monooxygenase proteins being one of the most prevalent and highly expressed functional proteins (top 5%) across this dataset ( Additional File 5 ). The next step in nitrification, nitrite oxidation to nitrate was inferred from nitrite oxidoreductase peptides assigned to 5 genomes belonging to 2 new species (Nitro_40CM-3_1, Nitro_NS7_3, Nitro_NS7_4, Nitro_NS7_5, and Nitro_NS7_14) ( Additional File 3: Figure S9 , Additional File 5 , see sheet metabolism info). Both nitrifying lineages had the capacity for carbon dioxide fixation with the reductive tricarboxylic acid (TCA) cycle (e.g., ATP-citrate lyase) in Nitrospiraceae genomes, and 3-HydroxyPropionate/4-HydroxyButyrate (3HP/4HB) encoded by the Nitrososphaeraceae. We did not detect genomic evidence for comammox or anammox and thus aerobic, chemolithoautotrophic nitrification supported by a metabolic partnership between bacteria and archaea occurred in the presence of heterotrophs as predicted by C/N ratios. Similarly, others have reported the prominence of nitrifying lineages from the archaeal thaumarcheotal Thermoproteota and bacterial Nitrospirota both by 16S rRNA [ 59 ] and genome-resolved metagenomics [ 12 , 13 ] in HZ sediments. Here we nearly doubled the genomic sampling of these river nitrifiers, assigning unique gene expression patterns to 3 and 17 genomes from Nitrososphaeraceae and Nitrospiraceae respectively, including the first genomic sampling of new genera and species (Fig. 2 ). Our co-expression data indicate that metabolic handoffs between archaeal ammonia oxidizers and bacterial nitrite oxidizers may be an unaccounted-for biogenic source of nitrate in these sediments ( Additional File 5 ). This suggests the activity of nitrifiers could be an underappreciated modulator of nitrous oxide fluxes from oligotrophic HZ sediments, both through their indirect stimulation of denitrifiers and their own contributions to this greenhouse flux [ 60 ]. Taken together, the archaeal-bacterial nitrifying mutualism outlined here appears well adapted to the low nutrient conditions present in many HZ sediments, warranting future research on the variables that constrain nitrification rates (i.e., ammonium availability, dissolved oxygen, pH) and their role as driver of nitrogen fluxes from these systems [ 61 ]. Denitrification is encoded by novel and taxonomically diverse lineages in HZ sediments Beyond the possible biogenic sources of nitrate, we identified from nitrification, these HZ sediments receive significant allochthonous nitrate from groundwater. When river stage decreases, groundwater discharges through the HZ sediments, bringing nitrate concentrations to over 20 mg/L [ 2 , 62 ]. In support of an important influence of nitrate from either source, HUM-V genomes with the capacity for nitrate reduction spanned diverse taxonomies, with NarG or NapX encoded in 11 genomes from the Actinobacteriota, Binatia, Gammaproteobacteria, and Myxococcota ( Additional File 3: Figure S1 ). However, our proteomic evidence for nitrate reduction was detected in less than 10% of the 33 sediment samples, with unique peptides assigned to Binatia NarG from a single sample. Based on gene expression data, we inventoried other steps in the denitrification pathway. Nitrite was reduced via nitrifier and denitrifier reduction to nitric oxide from archaeal ammonia oxidizers of the Nitrososphaeraceae active in 79% of metaproteome samples, and from Gammaproteobacterial Burkholderia in a single sample, respectively. The role of nitrite reduction by Nitrososphaeraceae is still under investigation but could be used for detoxification [ 63 ]. Genes for converting nitric oxide to nitrous oxide were not detected in proteomics, but we did find evidence that the Desulfobacterota genome (Desulf_UBA2774_1) expressed the nos gene for reducing nitrous oxide to nitrogen gas. Phylogenetic analysis suggest this organism used a "Clade II” nos sequence type adapted for low atmospheric concentrations of nitrous oxide ( Additional File 3: Figure S2 ), and consistent with our genome metabolic summary did so without encoding other steps of the denitrification pathway [ 64 ]. Notably, the capacity for denitrification exists beyond those detected in proteomics, as Binatia encoded dissimilatory nitrite reduction to ammonium (DNRA) and the potential for nitrous oxide production via nor was encoded by two Gammaproteobacteria (Steroid-FEN-1191_1, Steroid_1) and a member of the Myxococcota (Anaerom_1). In summary, our data adds to the growing realization that complete denitrification by single microorganism is likely the exception rather than the rule in natural systems [ 65 ], including the HZ [ 66 ]. In support of this, none of the genomes reconstructed here encoded a complete denitrification pathway for reducing nitrate to nitrous oxide or dinitrogen gas ( Additional File 3: Figure S1 ). Similarly, our proteomics data hinted that separate microbial members likely catalyzed each step of the denitrification pathway ( Additional File 3: Figure S4 ). This suggests cross-organism inorganic nitrogen exchange would be necessary for nitrogen gas flux, such that physical processes (e.g., advection, diffusion) or the spatial colocalization of microorganisms, as well as organic carbon availability, may have disproportionate impacts on flux of nitrous oxide and dinitrogen from these sediments. HUM-V identifies new microbial and viral players in hyporheic zone carbon and nitrogen cycling The creation of a genome database expanded upon prior amplicon-based surveys, allowing us to assign new metabolic functions to microbes and even viruses in hyporheic sediments. While HUM-V contains genomes from phyla (CSP1-3, Eisenbacteria) and classes (Binatia, MOR-1 in Acidobacteria) composed entirely of uncultivated members (Fig. 1 d), here we focus our analysis on the Binatia, as we recovered 7 genomes (one which included a complete 16S rRNA gene), they recruited peptides, and they also played key roles in carbon and nitrogen cycling. Using the 16S rRNA gene (from Binatia_7), we inventoried the distribution of closely related species to our HUM-V genomes (> 97% similarity) in the Sequence Read Archive (SRA) samples, to uncover the ecological distribution of these organisms from soils, as well as a wide variety of terrestrial, terrestrial-aquatic, marine samples (Fig. 5 ), indicating the processes uncovered by proteomics here are likely applicable to a wide range of ecosystems. A recent comparative genomics analysis on Binatota MAGs provided a first assessment of their metabolic potential, indicating genes for methylotrophy, alkane degradation, and pigment production were distributed across the phylum [ 67 ]. These HUM-V genomes belong to a class and family denoted UBA9968. Contrary to their prior metabolic inventory, HUM-V UBA9968 MAGs do not encode the potential for methanol oxidation, and we identified a new role in the decomposition of aromatic compounds from plant biomass (phenylpropionic acid, phenylacetic acid, salicylic acid), and xenobiotics (phthalic acid) (Fig. 5 ). We provide the first proteomic evidence for any members of the Binatia, supporting their roles in aerobically oxidizing carbon monoxide, producing extracellular peptidases, and in denitrification. Together these findings illustrate the power of HUM-V paired proteomes to illuminate new roles for members of uncultivated, previously enigmatic lineages in HZ carbon and nitrogen cycling. The relatively high proteomic recruitment of viruses sampled in HUM-V (Fig. 2 d) suggested important viral contributions in these sediments. In silico analysis assigned a putative host to 29% of the 111 viral genomes linking 18 microbial genomes that belong to bacterial members in Acidobacteriota, Actinobacteriota, CSP1-3, Eisenbacteria Methylomirabilota, Myxococcota, Nitrospirota, and Proteobacteria ( Additional File 3: Figure S10 , Additional file 2 , Additional file 6 ). Analysis of the metaproteomes for these phage-impacted microorganisms revealed these hosts expressed genes for nitrification (Nitrospiraceae) as well as carbon monoxide oxidation and nitrogen mineralization (Actinobacteria) (Fig. 6 ). Additionally, HUM-V phage genomes encode auxiliary metabolic genes with the potential to enhance microbial metabolism of carbon (CAZymes), sulfur (sulfate adenyl transferase), and nitrogen (amidase to cleave ammonium) ( Additional File 3: Figure S11 , see Additional File 3 supplemental text). We also show viral abundances were better predictors of total carbon and nitrogen percentages relative to microbial genome abundances ( Additional File 3: Figure S12, Additional File S9 , see Additional File 3 supplemental text). Together, these HUM-V enabled results indicate viral infections may contribute to river sediment functioning and raise the question to whether enhanced viral interrogation might provide a means to improved ecosystem or biogeochemical models in these systems. Conclusions To our knowledge this study represents one of the first genome-resolved microbial and viral enabled proteomic studies in river sediments. Using genome-resolved proteomics with complementary metabolites (detected by NMR, FTICR-MS), and geochemistry we begin to illuminate the microbial contributions to processes well known to occur but previously poorly defined mechanistically in river sediments (e.g., nitrogen mineralization). We also show how multi-omic tools can uncover previously enigmatic processes which may directly impact river respiration (e.g., carbon monoxide oxidation). While river carbon and nitrogen budgets are often quantified by direct measurements of inputs and the concentration of inorganic and organic compounds exported from rivers, what is missing today is an appreciation for the microbially and virally mediated sources and sinks for key intermediates (e.g., carbon dioxide, ammonium, nitrate), the degree to which these compounds are recycled and exchanged, and the underlying microbial metabolic lifestyles that catalyze this interconnected carbon and nitrogen biogeochemistry. Here, we have created a conceptual framework that elaborates on these missing ideas. Empowered by our individual process-based metaproteomic analyses (Figs. 3 – 6 ), we created a conceptual model outlining the microbial conversions of carbon and nitrogen in these hyporheic sediments (Fig. 7 ). Heterotrophic oxidation of organic carbon derived from plant (and likely microbial) biomass supported by oxygen and nitrogen respiring populations produce carbon dioxide. In addition, metaproteomics divulged that aerobic carbon monoxide oxidation may also be a source of carbon dioxide. Like organic carbon, the organic nitrogen in microbial and plant biomass could be mineralized to release ammonium in these sediments. Together inorganic pools of nitrogen (ammonium) and carbon (carbon dioxide) sustain the coordinated activity of nitrifying populations. Together our findings put forth an integrated framework that advances microbial roles in hyporheic carbon and nitrogen transformations, yielding insights that could inform research strategies to reduce existing predictive uncertainties in river corridor models. Abbreviations HUM-V: Hyporheic Uncultured Microbial and Viral (HUM-V) database MAG: Metagenome assembled genome vMAG: Viral metagenome assembled genome HZ: Hyporheic zone FTICR-MS: Fourier-transform ion cyclotron resonance mass spectrometry NMR: Nuclear magnetic resonance DNA: Deoxyribonucleic acid rRNA: Ribosomal ribonucleic acid tRNA: Transfer ribonucleic acids JGI: Joint Genome Institute EMSL: Environmental Molecular Sciences Laboratory OSU: The Ohio State University NCBI: National Center for Biotechnology Information SRA: Sequence Read Archive Gbp: Giga base pair IDBA-UD: Iterative De Bruijn Graph De Novo Assembler – Uneven Depth metaSPAdes: metagenomic St. Petersburg genome Assembler GTDB: Genome Taxonomy Database ANI: Average Nucleotide Identity DRAM: Distilled and Refined Annotation of MAGs IMG/MER: Integrated Microbial Genomes / Microbiomes Expert Review PFAM database: Protein Family Database MUSCLE: Multiple Sequence Comparison by Log-Expectation MAFFT: Multiple Alignment using Fast Fourier Transform nar : respiratory nitrate reductase nxr : nitrite oxidoreductase nos : nitric oxide synthase DNRA: Dissimilatory Nitrate Reduction to Ammonium MIUViG: Minimum information about an Uncultivated Virus Genome ICTV: International Committee on Taxonomy of Viruses PHYRE2: Protein Homology / AnalogY Recognition Engine 2.0 sPLS: sparse Partial Least Squares Regression MS: Mass Spectrometry MS / MS: Tandem mass spectrometry LC-MS/MS: Liquid Chromatography Tandem Mass Spectrometry FDR: False Detection Rate PSM: Peptide Spectrum Matches NSAF: Normalized spectral abundance frequency MHz: Megahertz DSS-d6: 3-(Trimethylsilyl)-1-propanesulfonic acid-d6 sodium salt ppm: parts per million RANCM: Rank and AssigN Confidence to Metabolites TOCSY: Total Correlation Spectroscopy HSQC: Heteronuclear single-quantum correlation spectroscopy gHSQCAD: Gradient-enhanced Heteronuclear Single Quantum Coherence with Adiabatic Pulses CAZymes: Carbohydrate-Active Enzymes GH: Glycoside Hydrolase CO: Carbon Monoxide TCA: Tricarboxylic acid C: Carbon N: Nitrogen C:N: Carbon / Nitrogen Ratio NH 4 +: Ammonium 3HP/4HB: 3-HydroxyPropionate/4-HydroxyButyrate ATP: Adenosine triphosphate Declarations Ethics approval and consent to participate: Not applicable. Consent for publication: Not applicable. Availability of data and material: The datasets supporting the conclusions of this article are publicly available. Sequencing data are available in NCBI under bioproject PRJNA576070, with MAGs deposited under biosamples SAMN18867633-SAMN18867734 and 16S rRNA amplicon sequences under accession numbers SRX9312157-SRX9312180, and vMAGs have been temporarily deposited in Zenodo doi 10.5281/zenodo.5124937. Viral genomes fasta file is publicly available within the following repository link: https://doi.org/10.5281/zenodo.5124937. Metaproteomics data are deposited in the MassIVE database under accession MSV000087330. Metabolomics data are publicly available and deposited in Zenodo doi https://doi.org/10.5281/zenodo.5076253. Additional datasets supporting the conclusions of this article are included within the article (and its additional files). Competing interests: The authors declare they have no competing interests. Funding: This work was supported by the Subsurface Biogeochemical Research (SBR) program (DE-SC0018170); the National Sciences Foundation Division of Biological Infrastructure [#1759874]; and the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, Environmental System Science (ESS) program through subcontract from the River Corridor Scientific Focus Area project at Pacific Northwest National Laboratory. J.R.R. is funded by the National Science Foundation (NRT-DESE) [1450032], support for A Trans-Disciplinary Graduate Training Program in Biosensing and Computational Biology at Colorado State University. The NMR data, FTICR-MS data and MS-proteomics data in this work was collected using instrumentation in the Environmental Molecular Science Laboratory (grid.436923.9), a DOE Office of Science User Facility sponsored by the Office of Biological and Environmental Research and located at Pacific Northwest National Laboratory. Pacific Northwest National Lab is operated by Battelle for the DOE under Contract DE-AC05-76RL01830. Metagenomic sequencing for this research was performed by the Joint Genome Institute via a large-scale sequencing award (Award 1781) and at the Genomics Shared Resource Core at The Ohio State University Comprehensive Cancer Center supported by P30 CA016058. Authors’ contributions: Using the CRediT contributor roles,author contributions can be defined as follows:Conceptualization,JRR, MAB, JCS, and KCW; Data curation, JRR, MAB, and BBM; Formal analysis, JRR, MAB, BBM, GJS, LMS, RAD; Funding acquisition, JCS, KCW; Project administration, RAD, JCS, KCW; Investigation, JRR, MAB, BBM, GJS, SOP, CDN, DWH; Supervision, MSL, EBG, DWH, JCS, KCW; Writing- original draft, JRR, MAB, BBM, GJS, KCW; Writing- review and editing, JRR, MAB, BBM, GJS, JCS, KCW. Acknowledgements: The authors would like to thank Tyson Claffey and Richard Wolfe for Colorado State University server management; Sandy Shew for management of computing resources retained from The Ohio State University Unity cluster; Dr. Pearlly Yan at the Genomics Shared Resource Core at The Ohio State University Comprehensive Cancer Center for management of metagenomic sequencing; and Dr. J John for the continuous support. Author's information: J.R.R. and M.A.B. contributed equally to this work. References Boulton AJ, Findlay S, Marmonier P, Stanley EH, Valett HM. The functional significance of the hyporheic zone in streams and rivers. Annu Rev Ecol Syst. Annual Reviews 4139 El Camino Way, PO Box 10139, Palo Alto, CA 94303-0139, USA; 1998;29:59–81. Stegen JC, Johnson T, Fredrickson JK, Wilkins MJ, Konopka AE, Nelson WC, et al. Influences of organic carbon speciation on hyporheic corridor biogeochemistry and microbial ecology. Nat Commun. Nature Publishing Group; 2018;9:1–11. Newcomer ME, Hubbard SS, Fleckenstein JH, Maier U, Schmidt C, Thullner M, et al. Influence of hydrological perturbations and riverbed sediment characteristics on hyporheic zone respiration of CO2 and N2. J Geophys Res Biogeosciences. Wiley Online Library; 2018;123:902–22. Trimmer M, Grey J, Heppell CM, Hildrew AG, Lansdown K, Stahl H, et al. River bed carbon and nitrogen cycling: state of play and some new directions. Sci Total Environ. Elsevier; 2012;434:143–58. Villa JA, Smith GJ, Ju Y, Renteria L, Angle JC, Arntzen E, et al. Methane and nitrous oxide porewater concentrations and surface fluxes of a regulated river. Sci Total Environ. Elsevier; 2020;715:136920. Hu M, Chen D, Dahlgren RA. Modeling nitrous oxide emission from rivers: a global assessment. Glob Chang Biol. Wiley Online Library; 2016;22:3566–82. Beaulieu JJ, Tank JL, Hamilton SK, Wollheim WM, Hall RO, Mulholland PJ, et al. Nitrous oxide emission from denitrification in stream and river networks. Proc Natl Acad Sci. National Acad Sciences; 2011;108:214–9. Caruso A, Boano F, Ridolfi L, Chopp DL, Packman A. Biofilm-induced bioclogging produces sharp interfaces in hyporheic flow, redox conditions, and microbial community structure. Geophys Res Lett. Wiley Online Library; 2017;44:4917–25. Naegeli MW, Uehlinger U. Contribution of the hyporheic zone to ecosystem metabolism in a prealpine gravel-bed-river. J North Am Benthol Soc. North American Benthological Society; 1997;16:794–804. Battin TJ, Kaplan LA, Newbold JD, Hendricks SP. A mixing model analysis of stream solute dynamics and the contribution of a hyporheic zone to ecosystem function. Freshw Biol. Wiley Online Library; 2003;48:995–1014. Zhang M, Daraz U, Sun Q, Chen P, Wei X. Denitrifier abundance and community composition linked to denitrification potential in river sediments. Environ Sci Pollut Res. Springer; 2021;1–12. Liu S, Wang H, Chen L, Wang J, Zheng M, Liu S, et al. Comammox Nitrospira within the Yangtze River continuum: community, biogeography, and ecological drivers. ISME J. Nature Publishing Group; 2020;14:2488–504. Pinto OHB, Silva TF, Vizzotto CS, Santana RH, Lopes FAC, Silva BS, et al. Genome-resolved metagenomics analysis provides insights into the ecological role of Thaumarchaeota in the Amazon River and its plume. BMC Microbiol. BioMed Central; 2020;20:1–11. Graham EB, Crump AR, Kennedy DW, Arntzen E, Fansler S, Purvine SO, et al. Multi’omics comparison reveals metabolome biochemistry, not microbiome composition or gene expression, corresponds to elevated biogeochemical function in the hyporheic zone. Sci Total Environ. Elsevier; 2018;642:742–53. Peng Y, Leung HCM, Yiu SM, Chin FYL. IDBA-UD: a de novo assembler for single-cell and metagenomic sequencing data with highly uneven depth. Bioinformatics. Narnia; 2012;28:1420–8. Nurk S, Meleshko D, Korobeynikov A, Pevzner PA. metaSPAdes: a new versatile metagenomic assembler. Genome Res. Cold Spring Harbor Lab; 2017;27:824–34. Kang DD, Li F, Kirton E, Thomas A, Egan R, An H, et al. MetaBAT 2: An adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ. PeerJ Inc.; 2019;2019:e7359. Wu M, Scott AJ. Phylogenomic analysis of bacterial and archaeal sequences with AMPHORA2. Bioinformatics. Oxford University Press; 2012;28:1033–4. Bowers RM, Kyrpides NC, Stepanauskas R, Harmon-Smith M, Doud D, Reddy TBK, et al. Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea. Nat. Biotechnol. Nature Publishing Group; 2017. p. 725–31. Olm MR, Brown CT, Brooks B, Banfield JF. dRep: a tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication. ISME J. Nature Publishing Group; 2017;11:2864–8. Shaffer M, Borton MA, McGivern BB, Zayed AA, La Rosa SL, Solden LM, et al. DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Res. :gkaa621. Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics. 2018;36:1925–1927. Edgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. Oxford University Press; 2004;32:1792–7. Katoh K, Misawa K, Kuma K, Miyata T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. Oxford University Press; 2002;30:3059–66. Price MN, Dehal PS, Arkin AP. FastTree: computing large minimum evolution trees with profiles instead of a distance matrix. Mol Biol Evol. Oxford University Press; 2009;26:1641–50. Castelle CJ, Hug LA, Wrighton KC, Thomas BC, Williams KH, Wu D, et al. Extraordinary phylogenetic diversity and metabolic versatility in aquifer sediment. Nat Commun. Nature Publishing Group; 2013;4:2120. Yu NY, Wagner JR, Laird MR, Melli G, Rey S, Lo R, et al. PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. Bioinformatics. 2010;26:1608–15. Armenteros JJA, Tsirigos KD, Sønderby CK, Petersen TN, Winther O, Brunak S, et al. SignalP 5.0 improves signal peptide predictions using deep neural networks. Nat Biotechnol. Nature Publishing Group; 2019;37:420–3. Bendtsen JD, Kiemer L, Fausbøll A, Brunak S. Non-classical protein secretion in bacteria. BMC Microbiol. BioMed Central; 2005;5:1–13. Roux S, Enault F, Hurwitz BL, Sullivan MB. VirSorter: mining viral signal from microbial genomic data. PeerJ. PeerJ Inc.; 2015;3:e985. Roux S, Adriaenssens EM, Dutilh BE, Koonin E V., Kropinski AM, Krupovic M, et al. Minimum information about an uncultivated virus genome (MIUVIG). Nat Biotechnol. Nature Publishing Group; 2019;37:29–37. Merchant N, Lyons E, Goff S, Vaughn M, Ware D, Micklos D, et al. The iPlant collaborative: cyberinfrastructure for enabling data to discovery for the life sciences. PLoS Biol. Public Library of Science San Francisco, CA USA; 2016;14:e1002342. Jang H Bin, Bolduc B, Zablocki O, Kuhn JH, Roux S, Adriaenssens EM, et al. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat Biotechnol. Nature Publishing Group; 2019;37:632–9. Kelley LA, Mezulis S, Yates CM, Wass MN, Sternberg MJE. The Phyre2 web portal for protein modeling, prediction and analysis. Nat Protoc. Nature Publishing Group; 2015;10:845–58. Ahlgren NA, Ren J, Lu YY, Fuhrman JA, Sun F. Alignment-free oligonucleotide frequency dissimilarity measure improves prediction of hosts from metagenomically-derived viral sequences. Nucleic Acids Res. Oxford University Press; 2017;45:39–53. Borton MA, Hoyt DW, Roux S, Daly RA, Welch SA, Nicora CD, et al. Coupled laboratory and field investigations resolve microbial interactions that underpin persistence in hydraulically fractured shales. Proc Natl Acad Sci. National Academy of Sciences; 2018;115:E6585–E6594. Daly RA, Borton MA, Wilkins MJ, Hoyt DW, Kountz DJ, Wolfe RA, et al. Microbial metabolisms in a 2.5-km-deep ecosystem created by hydraulic fracturing in shales. Nat Microbiol. Nature Publishing Group; 2016;1:16146. Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. Nature Publishing Group; 2012;9:357. Lê Cao K-A, Rossouw D, Robert-Granié C, Besse P. A sparse PLS for variable selection when integrating omics data. Stat Appl Genet Mol Biol. De Gruyter; 2008;7. Nicora CD, Burnum-Johnson KE, Nakayasu ES, Casey CP, White III RA, Chowdhury TR, et al. The MPLEx protocol for multi-omic analyses of soil samples. J Vis Exp JoVE. MyJoVE Corporation; 2018; Elias JE, Gygi SP. Target-decoy search strategy for mass spectrometry-based proteomics. Proteome Bioinforma. Springer; 2010. p. 55–71. Kim S, Pevzner PA. MS-GF+ makes progress towards a universal database search tool for proteomics. Nat Commun. Nature Publishing Group; 2014;5:5277. McGivern BB, Tfaily MM, Borton MA, Kosina SM, Daly RA, Nicora CD, et al. Decrypting bacterial polyphenol metabolism in an anoxic wetland soil. Nat Commun. Nature Publishing Group; 2021;12:1–16. Weljie AM, Newton J, Mercier P, Carlson E, Slupsky CM. Targeted profiling: quantitative analysis of 1H NMR metabolomics data. Anal Chem. ACS Publications; 2006;78:4430–42. Joesten WC, Kennedy MA. RANCM: a new ranking scheme for assigning confidence levels to metabolite assignments in NMR-based metabolomics studies. Metabolomics. Springer; 2019;15:5. Moon K, Jeon JH, Kang I, Park KS, Lee K, Cha C-J, et al. Freshwater viral metagenome reveals novel and functional phage-borne antibiotic resistance genes. Microbiome. Springer; 2020;8:1–15. Wolf YI, Silas S, Wang Y, Wu S, Bocek M, Kazlauskas D, et al. Doubling of the known set of RNA viruses by metagenomic analysis of an aquatic virome. Nat Microbiol. Nature Publishing Group; 2020;5:1262–70. Peduzzi P. Virus ecology of fluvial systems: a blank spot on the map? Biol Rev. Wiley Online Library; 2016;91:937–49. Püttker S, Kohrs F, Benndorf D, Heyer R, Rapp E, Reichl U. Metaproteomics of activated sludge from a wastewater treatment plant--A pilot study. Proteomics. Wiley Online Library; 2015;15:3596–601. Rudney JD, Xie H, Rhodus NL, Ondrey FG, Griffin TJ. A metaproteomic analysis of the human salivary microbiota by three-dimensional peptide fractionation and tandem mass spectrometry. Mol Oral Microbiol. Wiley Online Library; 2010;25:38–49. Solden LM, Naas AE, Roux S, Daly RA, Collins WB, Nicora CD, et al. Interspecies cross-feeding orchestrates carbon degradation in the rumen ecosystem. Nat Microbiol. Nature Publishing Group; 2018;3:1274. Graham EB, Crump AR, Resch CT, Fansler S, Arntzen E, Kennedy DW, et al. Deterministic influences exceed dispersal effects on hydrologically-connected microbiomes. Environ Microbiol. Wiley Online Library; 2017;19:1552–67. Yang F, Bogdanov B, Strittmatter EF, Vilkov AN, Gritsenko M, Shi L, et al. Characterization of Purified c-Type Heme-Containing Peptides and Identification of c-Type Heme-Attachment Sites in Shewanella o neidenis Cytochromes Using Mass Spectrometry. J Proteome Res. ACS Publications; 2005;4:846–54. Cordero PRF, Bayly K, Leung PM, Huang C, Islam ZF, Schittenhelm RB, et al. Atmospheric carbon monoxide oxidation is a widespread mechanism supporting microbial survival. ISME J. Nature Publishing Group; 2019;13:2868–81. Xia X, Zhang S, Li S, Zhang L, Wang G, Zhang L, et al. The cycle of nitrogen in river systems: sources, transformation, and flux. Environ Sci Process \& Impacts. Royal Society of Chemistry; 2018;20:863–91. Strauss EA, Lamberti GA. Regulation of nitrification in aquatic sediments by organic carbon. Limnol Oceanogr. Wiley Online Library; 2000;45:1854–9. Brust GE. Management strategies for organic vegetable fertility. Saf Pract Org food. Elsevier; 2019. p. 193–212. Cavaliere M, Feng S, Soyer OS, Jiménez JI. Cooperation in microbial communities and their biotechnological applications. Environ Microbiol. Wiley Online Library; 2017;19:2949–63. Gibbons SM, Jones E, Bearquiver A, Blackwolf F, Roundstone W, Scott N, et al. Human and environmental impacts on river sediment microbial communities. PLoS One. Public Library of Science; 2014;9:e97435. Wrage N, Velthof GL, Van Beusichem ML, Oenema O. Role of nitrifier denitrification in the production of nitrous oxide. Soil Biol Biochem. Elsevier; 2001;33:1723–32. Strauss EA, Richardson WB, Bartsch LA, Cavanaugh JC, Bruesewitz DA, Imker H, et al. Nitrification in the Upper Mississippi River: patterns, controls, and contribution to the NO3- budget. J North Am Benthol Soc. 2004;23:1–14. Stegen JC, Fredrickson JK, Wilkins MJ, Konopka AE, Nelson WC, Arntzen E V, et al. Groundwater--surface water mixing shifts ecological assembly processes and stimulates organic carbon turnover. Nat Commun. Nature Publishing Group; 2016;7:1–12. Bartossek R, Nicol GW, Lanzen A, Klenk H-P, Schleper C. Homologues of nitrite reductases in ammonia-oxidizing archaea: diversity and genomic context. Environ Microbiol. Wiley Online Library; 2010;12:1075–88. Jones CM, Spor A, Brennan FP, Breuil M-C, Bru D, Lemanceau P, et al. Recently identified microbial guild mediates soil N 2 O sink capacity. Nat Clim Chang. Nature Publishing Group; 2014;4:801–5. Kuypers MMM, Marchant HK, Kartal B. The microbial nitrogen-cycling network. Nat Rev Microbiol. Nature Publishing Group; 2018;16:263–76. Nelson WC, Graham EB, Crump AR, Fansler SJ, Arntzen E V, Kennedy DW, et al. Distinct temporal diversity profiles for nitrogen cycling genes in a hyporheic microbiome. PLoS One. Public Library of Science San Francisco, CA USA; 2020;15:e0228165. Murphy CL, Sheremet A, Dunfield PF, Spear JR, Stepanauskas R, Woyke T, et al. Genomic Analysis of the Yet-Uncultured Binatota Reveals Broad Methylotrophic, Alkane-Degradation, and Pigment Production Capacities. MBio. Am Soc Microbiol; 2021;12. Supplementary Files AdditionalFile1.xlsx Additional File 1, xlsx: Title of data: Metagenome information. Description: Excel sheet containing metagenome read counts, accession numbers, biosamples and 16S rRNA data. AdditionalFile2.xlsx Additional File 2, xlsx: Title of data: Metagenome assembled genome information. Description: Excel sheet containing metagenome assembled genome completion information, accession numbers, Read mapping information, and metabolism summary from DRAM. AdditionalFile3.docx Additional File 3, .docx: Title: Supplementary methods, text, and figures. Description: Word document containing all supplementary methods and text. This file also contains all supplementary figures referenced within the main manuscript along with their respective legends. AdditionalFile4.zip Additional File 4, zip: Title of data: Phylogenetic trees for Nxr and Nos genes. Description: Compressed zip file containing fasta, phylogenetic tree, and excel files corresponding to Nxr and Nos genes present in HUM-V AdditionalFile5.xlsx Additional File 5, xlsx: Title: Metaproteomic data. Description: Excel sheet containing MAG Unique peptides mapped, MAG unique peptide NSAF relative abundance, MAG Per genome NSAF Relative abundance, and vMAG unique peptides mapped. AdditionalFile6.xlsx Additional File 6, xlsx: Title: Viral metagenome assembled genomes data. Description: Excel sheet containing MIUViG information, accession numbers, annotations, read mapping information, vContact2 results, downloaded MetaG information for vContact2 analyses, virhostmatcher2 results. AdditionalFile7.xlsx Additional File 7, xlsx: Title: Geochemistry and NMR. Description: Excel spreadsheet containing geochemistry data and NMR metabolite data. AdditionalFile8.xlsx Additional File 8, xlsx: Title: FTICR-MS Data. Description: Excel spreadsheet containing FTICR-MS data. AdditionalFile9.xlsx Additional File 9, xlsx: Title: sPLS data and correlations results. Description: Excel spreadsheet containing data corresponding to sparse partial least squares regressions (sPLS) and correlations between MAGs and vMAGs to geochemistry values from Additional File 7. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-746574","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research","associatedPublications":[],"authors":[{"id":43018798,"identity":"65a5901c-a26e-4aa8-9332-9ce1b631c461","order_by":0,"name":"Josué A. Rodríguez-Ramos","email":"","orcid":"https://orcid.org/0000-0002-2049-2765","institution":"Colorado State University Department of Soil and Crop Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Josué","middleName":"A.","lastName":"Rodríguez-Ramos","suffix":""},{"id":43018799,"identity":"ba419c31-410e-4cce-93a6-5485f0c8a32d","order_by":1,"name":"Mikayla A. Borton","email":"","orcid":"","institution":"Pacific Northwest National Laboratory Biological Sciences Division","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mikayla","middleName":"A.","lastName":"Borton","suffix":""},{"id":43018800,"identity":"ff1af1f9-b0e6-46ae-bb36-dfc20d6aebf4","order_by":2,"name":"Bridget B. McGivern","email":"","orcid":"","institution":"Colorado State University Department of Soil and Crop Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bridget","middleName":"B.","lastName":"McGivern","suffix":""},{"id":43018801,"identity":"b73ab978-8541-4ccf-8593-068c17575d47","order_by":3,"name":"Garrett J. Smith","email":"","orcid":"","institution":"Radboud University: Radboud Universiteit","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Garrett","middleName":"J.","lastName":"Smith","suffix":""},{"id":43018802,"identity":"3e7fe59c-8754-4cb5-85d3-158716010303","order_by":4,"name":"Lindsey M. Solden","email":"","orcid":"","institution":"Colorado State University Department of Soil and Crop Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lindsey","middleName":"M.","lastName":"Solden","suffix":""},{"id":43018803,"identity":"a057ecda-dc81-4fcc-a6e3-3b7d73c9f91a","order_by":5,"name":"Michael Shaffer","email":"","orcid":"","institution":"Colorado State University Department of Soil and Crop Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Michael","middleName":"","lastName":"Shaffer","suffix":""},{"id":43018804,"identity":"6cbf9545-3d7b-4062-9b84-0cdfd230a21d","order_by":6,"name":"Rebecca A. Daly","email":"","orcid":"","institution":"Colorado State University Department of Soil and Crop Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rebecca","middleName":"A.","lastName":"Daly","suffix":""},{"id":43018805,"identity":"18eaf710-75c6-44d6-8875-a6c0e36e7458","order_by":7,"name":"Samuel O. Purvine","email":"","orcid":"","institution":"Pacific Northwest National Laboratory","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Samuel","middleName":"O.","lastName":"Purvine","suffix":""},{"id":43018806,"identity":"04135233-7a46-4ead-b862-6e6407c9c6dd","order_by":8,"name":"Carrie D. Nicora","email":"","orcid":"","institution":"Pacific Northwest National Laboratory Biological Sciences Division","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Carrie","middleName":"D.","lastName":"Nicora","suffix":""},{"id":43018807,"identity":"4f3924da-0bbe-425f-a5df-46a43019fa72","order_by":9,"name":"Elizabeth K. Eder","email":"","orcid":"","institution":"Pacific Northwest National Laboratory","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Elizabeth","middleName":"K.","lastName":"Eder","suffix":""},{"id":43018808,"identity":"00b9eedb-dadc-4cd2-b45e-e9ce779e8f85","order_by":10,"name":"Mary Lipton","email":"","orcid":"","institution":"Pacific Northwest National Laboratory","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mary","middleName":"","lastName":"Lipton","suffix":""},{"id":43018809,"identity":"5f478420-39ec-4b0b-987e-b25f6a416cf9","order_by":11,"name":"David W. Hoyt","email":"","orcid":"","institution":"Pacific Northwest National Laboratory","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"David","middleName":"W.","lastName":"Hoyt","suffix":""},{"id":43018810,"identity":"72e83e14-6aed-41d3-bf92-64ca9b5a90a2","order_by":12,"name":"James C. Stegen","email":"","orcid":"","institution":"Pacific Northwest National Laboratory","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"James","middleName":"C.","lastName":"Stegen","suffix":""},{"id":43018811,"identity":"ce2665e2-5478-4c7e-b1c6-5eb60f71b926","order_by":13,"name":"Kelly C. Wrighton","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJ0lEQVRIie3QMUvEMBTA8VcL7RLt+o6D9iukFA4F8bO0CJ0EB0FuEEwpxOXQueBXcHARx5TA3VJ1Fbqc+yE9DsRB1MYqosRzFcx/aNP2/QgNgMn0B0OwM6EWDhB12wRwP79SPbHYB7EYQApgt0+im15OoCPyd9I7ypgkIHfXvLPJdH55E3i2LRfNwZMPbn6BGtIn5RvZ4HifZEVVh0XupCjGNAIy3tcRHxNFaupgFearvI6pJAMQDk0Y7gy0JLj7Qq4ViRrxTA9ZMNOSPlrvxBspIhShWHIaAxIt6Y0SVp7Sl3YXEhYF3+7+5eo4CjlJ99Y1BCeynM6GKQ1OKtrM+Vbgublshg9+u5Dnt7pTblsh6vwx/vba+WFcZT2qqyeWjJhMJtO/7hU0I2OE0mhxCwAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-0434-4217","institution":"Colorado State University Department of Soil and Crop Sciences","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Kelly","middleName":"C.","lastName":"Wrighton","suffix":""}],"badges":[],"createdAt":"2021-07-24 10:54:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-746574/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-746574/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":12048592,"identity":"58626e3d-ee42-432c-9f7d-8aae8f898122","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":448698,"visible":true,"origin":"","legend":"Overview of hyporheic zone sampling and the microbial genomes included in the HUM-V database a) Samples were collected from two transects, each with 3 sediment cores, with each core sectioned into 10-centimeter segments from 0-60 centimeters in depth and paired with metaproteomics and geochemistry. b) Schematic of the data types available for each of the depth samples within a core. Black-filled circles indicate depth samples for a particular data type, while the open circles denote missing analyses (due to limited sample availability). c) Summarized catalog of the total samples for each analysis. d) The total recovered genomes (# of MAGs), taxonomic string, inferred genome completion (Comp., %), and contamination (Cont., %) for the dereplicated microbial genomes retained in HUM-V.","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/22c44ad6e9f74b0765119a09.png"},{"id":12048593,"identity":"19b7d73f-216a-494a-9582-2f66f4aecb01","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":463439,"visible":true,"origin":"","legend":"HUM-V database coupled to metaproteomics reveals active members in the hyporheic zone microbiome. a) Stacked bar graph indicates the taxonomic novelty of the de-replicated MAGs colored by phylum and stacked according to the first empty position within the taxonomic string provided by the Genome Taxonomy Database GTDB-Tk. Each color represents a MAG phylum according to the MAG Phylum legend, a coloring maintained across this manuscript. b) Butterfly plot reports the summed genomic relative abundance across all samples (left side) and the normalized mean proteomic relative abundance (right side) for dereplicated MAGs (55 total, 49 shown), with bars colored by phylum. A butterfly plot with all MAGs is shown in Additional File 3: Figure S3. MAGs that contain a partial or complete 16S rRNA sequence are denoted with and asterisk (*). Non-unique peptide assignment is defined in the methods and is shown with grey bars. c) Similarity network of the few vMAGs from our study (black) that clustered to viruses belonging to the default RefSeq, ICTV and NCBI Taxonomy databases (gray), as well as clustering of our vMAGs to other freshwater, publicly available dataset we mined (Pink, Purple, Orange, and Turquoise). The remaining clusters of viruses that were novel (e.g., did not cluster with prior viral genomes) are shown, with the full network file including singletons shown in Additional File 6. d) Butterfly plot showing summed genomic relative abundance (left side) and total peptides recruited for each vMAG population (total 111, 58 shown). Bars are colored by clustering of vMAGs from this study with (i) viruses of known taxonomy in RefSeq, ICTV and NCBI Taxonomy (dark grey), (ii) novel genera, both only from this study and ubiquitous (black), and no clustering from any database (light grey, singletons).","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/0fba04c7b2a29a975bcb9846.png"},{"id":12048644,"identity":"2c635785-db43-4ca1-83a7-e04a7cc20d8c","added_by":"auto","created_at":"2021-08-02 21:33:05","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":208022,"visible":true,"origin":"","legend":"Metaproteomics and metabolomics reveal microbial metabolic handoffs that support carbon cycling in river sediments. Detected metabolites are given in boxes, with NMR-detected compounds listed in red, polymers from FTICR-MS in orange, undetected metabolites in black. These polymers were inferred from FTICR-MS assigned biochemical classes and the specificity of CAZymes detected in metaproteome, where starch and cellulose were within the “polysaccharide-like” class and glycoproteins were in the “amino sugar-like” class. MAG-resolved metaproteome information is indicated by solid arrows, with MAG shape colored by phylum. Red arrows indicate processes leading to CO2 production, while black arrows indicate other microbial carbon transforming genes expressed in the proteome. Shaded bold arrows indicate chemical connections, where (1) grey indicates a metabolite was detected along with putative downstream products (e.g., sucrose conversion to glucose) but metaproteomic lacked evidence for the transformation or (2) red indicates a metabolite not measured but metaproteomic evidence supported transformation (e.g., CO conversion to CO2).","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/c198250114502708835efcae.png"},{"id":12048594,"identity":"1baf3757-4a7e-4981-88f0-7bb7636f42ef","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":256567,"visible":true,"origin":"","legend":"Organic nitrogen mineralization and cellular transport are active microbial processes in river sediments. Bubble plots indicate the expressed genes that were uniquely assigned to specific genome including (a) extracellular peptidases and (b) cellular transporters for organic nitrogen. Unique peptides detected in at least 3 samples are reported as bubbles and colored by phylum. Table on the right shows putative amino acids cleaved or transported by respective peptidases or transporters, shades of color (green or grey) denote peptides that are cleaved into amino acids that could be transported, providing linkages between extracellular organic nitrogen transformation and transport of nitrogen into the cell. White boxes indicate an organic nitrogen transporter that recruited peptides but could not be linked to outputs of specific peptidases. ","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/168d5695ab6cbfa364e111bd.png"},{"id":12048645,"identity":"f7ae2c28-b275-49aa-8a54-ee0bb762a7da","added_by":"auto","created_at":"2021-08-02 21:33:05","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":206136,"visible":true,"origin":"","legend":"Uncultured Binatia are widely dispersed across ecosystems and express nitrogen and carbon genes in situ. (a) 16S rRNA comparisons between the HUM-V Binatia genomes and other microbiome datasets reveal that these uncultivated HUM-V species are prevalent in terrestrial and aquatic microbiomes. Graph shows the % of samples in the SRA database with reads assigned to Binatia. (b) The seven HUM-V genomes reconstructed here encode diverse pathways for transforming phenolic compounds. Pathways are shown with boxes corresponding to each enzyme in the pathway, colored by the number of Binatia MAGs that encode each step. Gene information shown here is reported in detail in Additional File 5. (c) Metabolic genome cartoon for the major functions encoded by the Binatia MAGs. Dotted arrows are functions encoded in the metagenome, with black arrows corresponding to metaproteome-detected enzymes. (BCAA, Branched chain amino acids). ","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/03b9b4697e2e4e375b5a606c.png"},{"id":12048601,"identity":"6d6da55d-bb93-460e-bd07-eb4bef62b7a4","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":108622,"visible":true,"origin":"","legend":"Evidence that viruses could impact microbial host metabolism and river geochemistry. a) Stacked bar chart of the total number of vMAGs (n=32) that have putative host linkages. Each bar represents a phylum and lines within bars indicate the linkages for specific genomes within each phylum. For example, there are three genomes within the Actinobacteriota phylum that collectively have 12 viral linkages and of the three genomes that have linkages, one host has 10 viruses linked, while the other two hosts have 1 virus linked. b) Genome cartoons of microbial metabolisms for two representative genomes that could be predated by vMAGs, with the genes shown in black text boxes denoting processes detected in proteomics. These two microorganisms were selected as examples because they were active members in shaping carbon and nitrogen metabolism in these river sediments but could be impacted by viral predation; other virus-host relationships are reported in Additional File 3: Figure S10. c) Heatmap reports correlations between a subset of vMAGs with rectangle colors denoting the putative phyla for the respective host. Correlations between these vMAGs and ecosystem geochemistry (NH4 µg/gram, %N, %C) are reported with significant correlation coefficients denoted by purple-green shading according to the legend. Red asterisks (*) indicate the vMAG relative abundance predicted a key environmental variable by sparse partial least squares (sPLS) regression. Note a subset of these predicted vMAGs are shown in 6b.","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/ed0eb5b23dd7414e0f59e30b.png"},{"id":12048603,"identity":"e962ee6f-d480-4326-9645-cf2203f62e3b","added_by":"auto","created_at":"2021-08-02 21:30:06","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":2094600,"visible":true,"origin":"","legend":"Conceptual model uncovering microbes and processes contributing to carbon and nitrogen cycling in river sediments. Integration of multi-omic data uncovered the microbial and viral effects on carbon and nitrogen cycling in river sediments. Black arrows signify microbial transformations uncovered in our metaproteomic data. Specific processes (e.g., mineralization, nitrification, CO-oxidation, denitrification, and aerobic respiration) are highlighted in beige boxes, with microorganisms inferred to carry out the specific process denoted by overlaid cell shapes colored by phylum. Prior to this research, little was known about the specific enzymes and organisms responsible for river organic nitrogen mineralization and CO-oxidation, thus this research adds new content to microbial roles in carbon and nitrogen transformations in these systems. Possible (biotic, atmospheric, and aquatic) carbon and nitrogen sources are shown by purple, green, and blue arrows respectively. Inorganic carbon and nitrogen sources are shown by black squares (aqueous) and black circles (gaseous) with white text and dashed arrows indicating possible gasses that could be released to the atmosphere. Processes that could be impacted by viruses are marked with grey viral symbols.","description":"","filename":"Figure7.png","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/cc5d9b0aba5085e425b21777.png"},{"id":14145482,"identity":"d4befd2e-d506-4b85-9e7b-f0a9c25033b4","added_by":"auto","created_at":"2021-09-30 11:57:37","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4560392,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/d9927b4d-e1ab-43f9-8fb7-954ee1cc1f0e.pdf"},{"id":12048598,"identity":"83171995-628f-4607-a1b5-18967f368fd5","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":17006,"visible":true,"origin":"","legend":"Additional File 1, xlsx: Title of data: Metagenome information. Description: Excel sheet containing metagenome read counts, accession numbers, biosamples and 16S rRNA data.","description":"","filename":"AdditionalFile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/7881824e3cb9461cc6c050fb.xlsx"},{"id":12048595,"identity":"9096e705-6a5f-4c8f-b227-203d3d937b38","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1422239,"visible":true,"origin":"","legend":"Additional File 2, xlsx: Title of data: Metagenome assembled genome information. Description: Excel sheet containing metagenome assembled genome completion information, accession numbers, Read mapping information, and metabolism summary from DRAM. ","description":"","filename":"AdditionalFile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/1f02fb96d1ff2641e4d4b1d1.xlsx"},{"id":12048643,"identity":"3fc1390a-cefc-4944-ad89-5667d96ce62a","added_by":"auto","created_at":"2021-08-02 21:33:05","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":2314642,"visible":true,"origin":"","legend":"Additional File 3, .docx: Title: Supplementary methods, text, and figures. Description: Word document containing all supplementary methods and text. This file also contains all supplementary figures referenced within the main manuscript along with their respective legends.","description":"","filename":"AdditionalFile3.docx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/47a161c84612e275673951d2.docx"},{"id":12048606,"identity":"0ee87fa8-c582-486b-a209-4f20fa4bfd29","added_by":"auto","created_at":"2021-08-02 21:30:06","extension":"zip","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":6963446,"visible":true,"origin":"","legend":"Additional File 4, zip: Title of data: Phylogenetic trees for Nxr and Nos genes. Description: Compressed zip file containing fasta, phylogenetic tree, and excel files corresponding to Nxr and Nos genes present in HUM-V","description":"","filename":"AdditionalFile4.zip","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/e73ac007f30dc0a1e2d9d792.zip"},{"id":12048597,"identity":"45fee1f7-7083-4bc6-b403-3f124a7c729e","added_by":"auto","created_at":"2021-08-02 21:30:05","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":909574,"visible":true,"origin":"","legend":"Additional File 5, xlsx: Title: Metaproteomic data. Description: Excel sheet containing MAG Unique peptides mapped, MAG unique peptide NSAF relative abundance, MAG Per genome NSAF Relative abundance, and vMAG unique peptides mapped.","description":"","filename":"AdditionalFile5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/da2818ff7b8698b6ab32f966.xlsx"},{"id":12048647,"identity":"d6b6cad5-3783-4c74-8ad2-848079ea598c","added_by":"auto","created_at":"2021-08-02 21:33:06","extension":"xlsx","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":894761,"visible":true,"origin":"","legend":"Additional File 6, xlsx: Title: Viral metagenome assembled genomes data. Description: Excel sheet containing MIUViG information, accession numbers, annotations, read mapping information, vContact2 results, downloaded MetaG information for vContact2 analyses, virhostmatcher2 results.","description":"","filename":"AdditionalFile6.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/a3d15b93e8fd0fd35f2216a9.xlsx"},{"id":12048893,"identity":"4d949b51-1801-4dc5-af39-644705f0b397","added_by":"auto","created_at":"2021-08-02 21:36:06","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":28771,"visible":true,"origin":"","legend":"Additional File 7, xlsx: Title: Geochemistry and NMR. Description: Excel spreadsheet containing geochemistry data and NMR metabolite data. ","description":"","filename":"AdditionalFile7.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/1054a197bdfb96597353c994.xlsx"},{"id":12048607,"identity":"4108c37a-5412-49c0-bad1-bd32b67758f1","added_by":"auto","created_at":"2021-08-02 21:30:06","extension":"xlsx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":4263317,"visible":true,"origin":"","legend":"Additional File 8, xlsx: Title: FTICR-MS Data. Description: Excel spreadsheet containing FTICR-MS data.","description":"","filename":"AdditionalFile8.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/1fe984ad74dd5f64af8b9f67.xlsx"},{"id":12048895,"identity":"36fd2794-7955-4b95-80f8-b6c8f26b2e00","added_by":"auto","created_at":"2021-08-02 21:36:07","extension":"xlsx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":468992,"visible":true,"origin":"","legend":"Additional File 9, xlsx: Title: sPLS data and correlations results. Description: Excel spreadsheet containing data corresponding to sparse partial least squares regressions (sPLS) and correlations between MAGs and vMAGs to geochemistry values from Additional File 7.","description":"","filename":"AdditionalFile9.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-746574/v1/af34fd7469997bf185de3a3d.xlsx"}],"financialInterests":"","formattedTitle":"\u003cp\u003eMicrobial Genome-Resolved Metaproteomic Analyses Frame Intertwined Carbon and Nitrogen Cycles in River Hyporheic Sediments\u003c/p\u003e","fulltext":[{"header":"Background","content":"\u003cp\u003eThe hyporheic zone (HZ) acts as a transitional space between river and groundwater compartments in the river corridor where the bidirectional supply of nutrients and organic carbon stimulate microbial activity [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Characterized as the permanently saturated interface between the river surface channel and underlying sediments, the HZ is considered a biogeochemical hotspot for microbial cycling of carbon and nitrogen [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. These zones have been reported to support microbial heterotrophic respiration, denitrification, and nitrification, as well as the consumption and production of greenhouse gases (such as nitrous oxide and carbon dioxide) [\u003cspan additionalcitationids=\"CR5 CR6\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. In addition to harboring diverse energy metabolisms, these sediments may act as an important sink of carbon and nitrogen in terms of microbial subsurface biomass [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Overall estimates of ecosystem respiration have revealed that the HZ accounts for 40 to 90% of total river respiration [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], highlighting that a substantial amount of respiration is associated with hyporheic microbial activities.\u003c/p\u003e \u003cp\u003eDespite the importance of microbial metabolism to river corridor biogeochemistry, the ability to partition metabolic handoffs between organisms, the linked use of carbon and nitrogen by individual organisms, and the mineralization of detritus have yet to be holistically interrogated. Metagenomic studies in river sediments have not fully inventoried carbon and nitrogen cycling metabolisms, instead focusing on specific aspects of the nitrogen cycle (e.g., genes in denitrification [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]). Moreover, most of these studies were not genome resolved, hindering the assignment of biogeochemical processes to specific microorganisms, and culture independent genomic reconstructions from river sediments are limited to a handful of studies [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], all of which focused exclusively on nitrification. Thus, little is known about uncultured microbial communities in river sediments, with the enzymes, interconnected chemical reactions, and microbial metabolic lifestyles mediating carbon and nitrogen transformations in river sediments not currently resolvable from existing HZ microbiome datasets.\u003c/p\u003e \u003cp\u003eHere we address this knowledge gap, creating a genome-resolved inventory of the microbial and viral members in HZ sediments collected from the Columbia River in Washington State, USA (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). This resource was used to recruit metaproteomic data, providing a first of its kind, comprehensive inventory of the active microbial organisms and their enzymatic machinery in river systems. We contextualized these biological findings, using chemical scaffolding provided from metabolomics and geochemistry. Reconstructing the expressed metabolic capabilities of numerous lineages enabled us to resolve microbial contributions to biogeochemistry in these sediments. Our proteome enabled road map outlined the metabolic circuitry coupling carbon and nitrogen biogeochemistry in these HZ sediments, providing a framework to develop hydrobiogeochemical models informed by biochemical mechanisms and ecological interactions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eExperimental Design\u003c/h2\u003e \u003cp\u003eTo investigate the microbial processes involved in biogeochemical cycling in HZ sediments, we leveraged previously collected geochemical, metagenomic, metaproteomic, and Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and combined it with new nuclear magnetic resonance (NMR) metabolomic characterization and additional metagenomic sequencing. Samples were collected from the hyporheic zone of the Columbia River (46\u0026deg;22\u0026rsquo;15.80\u0026Prime;N, 119\u0026deg;16\u0026rsquo;31.52\u0026Prime;W) in March 2015 as previously described [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Briefly, liquid nitrogen frozen sediment profiles (0-60cm) were collected along two transects separated by approximately 170 meters (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). At each transect, three sediment cores up to 60 cm in depth were collected at 5-meter intervals perpendicular to the river flow. All cores were collected during conditions in which the sediments were fully saturated. Each core was sectioned into 10 cm segments from 0-60-centimeter depths and stored at -80\u0026ordm;C. Analyses were carried out at 10-centimeter increments, with the exception of one core that was pooled from 0\u0026ndash;30 centimeters to have sufficient input masses. Collectively, this robust, paired multi-omic dataset was made up of previously reported metagenomes (n\u0026thinsp;=\u0026thinsp;33, 3-4Gbp), metaproteomes (n\u0026thinsp;=\u0026thinsp;33), FTICR-MS metabolomes (n\u0026thinsp;=\u0026thinsp;33), and geochemical characterizations (n\u0026thinsp;=\u0026thinsp;33), as well as new metagenomes (n\u0026thinsp;=\u0026thinsp;10, 10-25Gbp) and NMR metabolomes (n\u0026thinsp;=\u0026thinsp;17) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). Genome recovery from these samples was improved by increasing the metagenomic sequencing depth per sample (from an average of 3.8 to 25.3 Gbp for selected samples) and employing a hybrid of co-assembly and single assembly methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eDNA extraction and sequencing\u003c/h2\u003e \u003cp\u003eAs described previously [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], deoxyribonucleic acid (DNA) was extracted from the sediments using the MoBio PowerSoil kit (MoBio Laboratories, Inc., Carlsbad, CA) following manufacturer's instructions, with the addition of a 2-hour proteinase-K incubation at 55\u0026deg;C prior to bead-beating to facilitate cell lysis. Purified genomic DNA was sent to the Joint Genome Institute (JGI, n\u0026thinsp;=\u0026thinsp;33) under Joint Genome Institute / Environmental Molecular Sciences Laboratory (EMSL) proposal 1781 or to the Genomics Shared Resource facility at The Ohio State University (OSU, n\u0026thinsp;=\u0026thinsp;10), producing 43 metagenomes with average sequencing depth of 4 (JGI) and 25 Gbp (OSU) per sample, totaling 377Gbp. DNA submitted to JGI were prepared for sequencing using an Illumina Library creation kit (KAPA Biosystems), and then solid-phase reversible immobilization size selection. DNA submitted to OSU were prepared for sequencing with a Nextera XT library System followed by solid-phase reversible immobilization size selection. Libraries at both facilities were quantified to ensure input thresholds, and then sequenced using an Illumina HiSeq 2500 platform. Deeper sequencing was performed at OSU to enhance MAG recovery, resulting in an increase of 252 Gbp of additional sequencing for ten samples, increasing the sequencing depth per sample by at least 3-fold (15.37\u0026ndash;49.24 Gbp per sample) \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, \u003cb\u003eAdditional File 1)\u003c/b\u003e. \u003cb\u003eAdditional File 1\u003c/b\u003e details all sequencing information, including National Center for Biotechnology Information (NCBI) accession numbers.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eMetagenome assembly and binning\u003c/h2\u003e \u003cp\u003eRaw reads were trimmed for length and quality using Sickle v1.33 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/najoshi/sickle\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e and then subsequently assembled using Iterative De Bruijn Graph De Novo Assembler \u0026ndash; Uneven Depth (IDBA-UD) 1.1.0 [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] with an initial kmer of 40 or metagenomic St. Petersburg genome Assembler (metaSPAdes) 3.13.0 [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] with default parameters. To further increase genomic recovery, for the ten samples that had shallow and deep sequencing, metagenomic reads were coassembled using IDBA-UD 1.1.0 with an initial kmer of 40. All assemblies, including co-assemblies, were then individually binned using Metabat2 [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] with default parameters to obtain Metagenome Assembled Genomes (MAGs).\u003c/p\u003e \u003cp\u003eFor each bin, genome completion was estimated based on the presence of core gene sets (highly conserved genes that occur in single copy) for Bacteria (n\u0026thinsp;=\u0026thinsp;31 genes) and Archaea (n\u0026thinsp;=\u0026thinsp;104 genes) using Amphora2 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Bins were discarded for further analysis if completion was \u0026lt;\u0026thinsp;70% or contamination was \u0026gt;\u0026thinsp;10% to select for only medium to high quality bins [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. This resulted in 102 MAGs that were then dereplicated using dRep [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] with default parameters and resulted in a final set of 55 MAGs (\u0026gt;\u0026thinsp;99% ANI). To further assess bin quality, we used the Distilled and Refined Annotation of MAGs (DRAM) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] annotation pipeline to identify ribosomal ribonucleic acids (rRNAs) and transfer ribonucleic acids (tRNAs). The 102 MAGs detailed here are deposited on NCBI under the BioProject ID PRJNA576070, with genome quality information reported in \u003cb\u003eAdditional File 2.\u003c/b\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003ePhylogenetic and metabolic analysis of Metagenome Assembled Genomes (MAG)\u003c/h2\u003e \u003cp\u003eMedium and high-quality MAGs were taxonomically classified using the Genome Taxonomy Database (GTDB) Toolkit v1.3.0 on September 2020 [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Novel taxonomy was identified as the first taxonomic level with no designation using GTDB taxonomy. For example, MAGs whose GTDB taxonomy string was not designated after the family level (e.g., g__) were identified as novel genera. Of the MAGs that had multiple representatives sharing taxonomy strings up to the Family level (Binatia and CSP1-3), we used average nucleotide identity (ANI) to determine whether they belonged to the same genus. MAG scaffolds were annotated using the DRAM pipeline [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. The raw annotations for each genome are deposited in the Zenodo repository under doi \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5281/zenodo.5128772\u003c/span\u003e\u003c/span\u003e and can be accessed here: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.5128772\u003c/span\u003e\u003c/span\u003e. \u003cb\u003eAdditional File 2\u003c/b\u003e shows the metabolic summary of genomes (product DRAM output) and output is also displayed in \u003cb\u003eAdditional File 3: Figure S1\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eTarget metabolic marker genes of interest recovered in bins were used to query the Integrated Microbial Genomes / Microbiomes Expert Review (IMG/M ER) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://img.jgi.doe.gov/cgi-bin/mer/main.cgi\u003c/span\u003e\u003c/span\u003e) and NCBI (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov\u003c/span\u003e\u003c/span\u003e) databases using BLASTp, or retrieved from the Protein Family (PFAM) database by protein family. Returned amino acid sequences were compiled with other known genes not retrieved via sequence homology, and de-replicated to make a reference sequence database. Sequences from the metagenomes were then aligned to the reference sequences using Multiple Sequence Comparison by Log-Expectation (MUSCLE) version 3.8.31 [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] or Multiple Alignment using Fast Fourier Transform (MAFFT) version 7.427 [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Alignments were manually curated to remove end and other gap regions. These alignments were then used to construct phylogenetic trees using FastTree version 2.1.11 [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] with default settings.\u003c/p\u003e \u003cp\u003eAn additional phylogenetic analysis was performed on genes annotated as respiratory nitrate reductase (\u003cem\u003enar\u003c/em\u003e) and nitrite oxidoreductase (\u003cem\u003enxr\u003c/em\u003e) to resolve novel Binatia role in nitrogen cycling. Specifically, sequences from [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] were downloaded and combined with \u003cem\u003enar\u003c/em\u003e and \u003cem\u003enxr\u003c/em\u003e amino acid sequences from dereplicated bins, aligned using MUSCLE, version 3.8.31, and run through ProtPipeliner, a Python script developed in-house for generation of phylogenetic trees (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/TheWrightonLab\u003c/span\u003e\u003c/span\u003e). Phylogenetic trees are shown in \u003cb\u003eAdditional File 3: Figure S2\u003c/b\u003e and \u003cb\u003eAdditional File 4\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eFor polyphenol and organic polymer degradation, we used functional annotation in addition to predicted secretion to assess functional potential. To determine if the predicted genes encoded a secreted protein, we used pSortb [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] and SignalP [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] to predict location; if those methods did not detect a signal peptide, the amino acid sequence was queried to SecretomeP and a SecP score\u0026thinsp;\u0026gt;\u0026thinsp;0.5 [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] was used as a threshold to report non-canonical secretion signals. Metabolic information for each MAG discussed in this manuscript are available in \u003cb\u003eAdditional File 2\u003c/b\u003e and \u003cb\u003eAdditional File 5\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eViral Analyses\u003c/h2\u003e \u003cp\u003eMetagenomic assemblies (n\u0026thinsp;=\u0026thinsp;43) were screened for DNA viral sequences using VirSorter v1.0.3 with the ViromeDB database option [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], retaining viral contigs ranked 1, 2, 4 or 5 with greater than 10kb in genome length as stated by the Minimum Information about an Uncultivated Virus Genome (MIUViG) standards [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. To determine an approximate species level taxonomy for viral scaffolds, they were clustered into viral metagenome assembled genomes (vMAGs) at 95% ANI across 85% of the shortest contig using ClusterGenomes 5.1 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/simroux/ClusterGenome\u003c/span\u003e\u003c/span\u003e) [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. After clustering, vMAGs were manually confirmed to be viral by assessing the total of viral genes with regards to non-viral genes in the genome, where genomes containing more than 18% of non-viral genes were discarded (J flag, DRAM [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]).This resulted in 111 vMAGs that were deposited on NCBI under the BioProject ID PRJNA576070 \u003cb\u003eAdditional File 6.\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo determine taxonomic affiliation, vMAGs were clustered to viruses belonging to standard viral reference taxonomy databases NCBI Bacterial and Archaeal Viral RefSeq V85 with the International Committee on Taxonomy of Viruses (ICTV) and NCBI Taxonomy using the network-based protein classification software vContact2 v0.9.8 [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Default methods were used. To determine geographic distribution of viruses in freshwater ecosystems, we also included viruses mined from publicly available freshwater metagenomes in vContact2 analyses: 1) East River, CO (PRJNA579838) 2) A previous study from the Columbia River, WA (PRJNA375338) 3) Prairie Potholes, ND (PRJNA365086) and 4) the Amazon River (PRJNA237344).\u003c/p\u003e \u003cp\u003eViral contigs were annotated with DRAM-v [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], with annotations for each viral genome reported in \u003cb\u003eAdditional File 6\u003c/b\u003e. Genes that were identified by DRAM-v as being possible auxiliary metabolic genes (categories 1\u0026ndash;3) were subjected to protein modeling using Protein Homology / AnalogY Recognition Engine (PHYRE2) in order to improve the accuracy of annotation [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. To identify likely vMAG hosts, oligonucleotide frequencies between virus (n\u0026thinsp;=\u0026thinsp;111) and non-dereplicated hosts (n\u0026thinsp;=\u0026thinsp;102) were analyzed using VirHostMatcher using a threshold of d2* measurements of \u0026lt;\u0026thinsp;0.25 [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. Per genome results of VirHostMatcher predictions are reported in \u003cb\u003eAdditional File 6\u003c/b\u003e. The lowest d2* value for each viral contig\u0026thinsp;\u0026lt;\u0026thinsp;0.25 was used. Viruses and hosts could not be linked by matching clustered regularly interspaced short palindromic repeats (CRISPR) spacers in host genomes to vMAG genomes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eGenome relative abundance calculations\u003c/h2\u003e \u003cp\u003eRelative abundance for each MAG and vMAG was estimated using in-house scripts available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/TheWrightonLab\u003c/span\u003e\u003c/span\u003e [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. Briefly, all metagenomic reads were concatenated for each sample, rarified to 3Gbp, and multi-mapped to 55 unique MAGs via Bowtie2 [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. For MAGs, a minimum scaffold coverage of 75% and depth of 3x required for read recruitment at 7 mismatches. For vMAGs, reads were mapped using Bowtie2 [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] at a maximum mismatch of 15, a minimum contig coverage of 75% and a minimum depth coverage of 2x. Relative abundances for each MAG and vMAG were calculated as their coverage proportion from the sum of the whole coverage of all bins for each set of metagenomic reads. To identify correlations between MAGs and vMAGs to geochemistry and predictive capability, we also mapped the subset of deep sequencing (n\u0026thinsp;=\u0026thinsp;10) and rarefied to the lowest deep metagenome available (4.8Gbp). Correlations and sparse Partial Least Squares Regression (sPLS) predictions (PLS R package [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]) were done using mapping data pertaining to only the 10 deeply sequenced metagenomes. Genome relative abundances per sample for MAGs and vMAGs are reported in \u003cb\u003eAdditional File 2\u003c/b\u003e and \u003cb\u003eAdditional File 6\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eMetaproteome generation and peptide mapping\u003c/h2\u003e \u003cp\u003eSediment samples were prepared for metaproteome analysis as previously reported in Graham et al. 2018 [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] and the protocol outlined by Nicora et al [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. For protein identification, spectra were searched against two files that included (i) 55 dereplicated MAG and (ii) 111 clustered vMAGs amino acid sequences. Exact sequence duplicates were removed, and 16 commonly observed contaminants (e.g., tryptic fragments, human keratins, and serum albumin precursors) were included. The tandem mass spectrometry (MS/MS) spectra from all liquid chromatography tandem mass spectrometry\u003c/p\u003e \u003cp\u003e(LC-MS/MS) datasets were converted to ASCII text (.dta format) using MSConvert (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://proteowizard.sourceforge.net/tools/msconvert.html\u003c/span\u003e\u003c/span\u003e) which more precisely assigns the charge and parent mass values to an MS/MS spectrum. The data files were then interrogated via target-decoy approach [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e] using MSGF+ [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e] using a\u0026thinsp;\u0026plusmn;\u0026thinsp;20 ppm parent mass tolerance, partially tryptic digestion enzyme settings, and a variable posttranslational modification of oxidized Methionine. All MS/MS search results for each dataset were collated into tab separated ASCII text files listing the best scoring identification for each spectrum. Collated search results were further combined into a single result file. These results were imported into a Microsoft SQL Server database. Results were filtered to N1% false detection rate (FDR) using an MSGF\u0026thinsp;+\u0026thinsp;supplied Q-Value that assesses reversed sequence decoy identifications for a given MSGF score across each dataset. Using the protein references as a grouping term, unique peptides belonging to each protein were counted, as were all peptide spectrum matches (PSMs) belonging to all peptides for that protein (i.e., a protein level observation count value). PSM observation counts were reported for each sample that was analyzed. Crosstabulation tables were created to enumerate protein level PSM observations for each sample, allowing low-precision quantitative comparisons to be made.\u003c/p\u003e \u003cp\u003eMicrobial metaproteomes were converted to normalized spectral abundance frequency (NSAF) values and subsequently divided into unique, non-unique specialized, and non-unique categories, while viral metaproteomes were analyzed using peptide counts only from unique hits due to low recruitment [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Peptide recruitment for each MAG amino acid sequence per sample is reported in \u003cb\u003eAdditional File 5\u003c/b\u003e. Hits were divided into 3 categories: (1) uniques (peptide hits to a single protein), (2) non-unique specialized (peptide hits to multiple amino acid sequences that all had same annotation and MAG taxonomy), (3) non-unique (peptide hits to multiple amino acid sequences with different annotation or from MAGs with different taxonomy) [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. This designation was necessary as several hits could not be resolved to the MAG level due to functional conservation across closely related genomes in the HUM-V database. Data in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e showcases (1) and (2) categories, with the entire dataset shown in \u003cb\u003eAdditional File 3: Figure S3\u003c/b\u003e. Including the non-unique specialized hits assigned an additional 14% of the proteome (grey bar, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb) and confirmed we did not underrepresent the gene expression from genomically well-sampled strains (e.g., Nitrospiraceae). Metaproteome hits for MAGs were used for further metabolic analyses if they were detected in at least three samples. Annotations for the entire metaproteomic dataset are shown in \u003cb\u003eAdditional File 3: Figure S4\u003c/b\u003e.\u003c/p\u003e \u003ch2\u003eGeochemical measurements, FTICR-MS characterization of organic matter, and NMR detected metabolites.\u003c/h2\u003e \u003cp\u003eAs previously reported [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], total nitrogen, total carbon, and total sulfur were determined using Elementar vario EL cube (Elementar Co., Germany), with details in the Supplementary Information (\u003cb\u003eAdditional File 7\u003c/b\u003e). To characterize organic matter, we used FTICR-MS to analyze sediments as previously reported [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], with details in the Supplementary Information (\u003cb\u003eAdditional File 3: Figure S5, Additional File 8\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eTo identify the metabolites available to microorganisms in this river system, we performed \u003csup\u003e1\u003c/sup\u003eH Nuclear Magnetic Resonance (NMR) spectroscopy on sediment pore water. Sediment samples were mixed with 200, 300, or 600 \u0026micro;L of MilliQ water depending on the sediment mass (\u003cb\u003eAdditional File 7\u003c/b\u003e) and centrifuged to remove the sediment. Supernatant (180 \u0026micro;L) was then diluted by 10% (vol/vol) with 5 mM 2,2-dimethyl-2-silapentane-5-sulfonate-\u003cem\u003ed\u003c/em\u003e\u003csub\u003e6\u003c/sub\u003e as an internal standard. All NMR spectra were collected using a Varian Direct Drive 600-MHz NMR spectrometer equipped with a 5-mm triple resonance salt-tolerant cold probe. Chemical shifts were referenced to the 1H or 13C methyl signal in DSS-d6 at 0 ppm. The 1D \u003csup\u003e1\u003c/sup\u003eH NMR spectra of all samples were processed, assigned, and analyzed using Chenomx NMR Suite 8.3 with quantification based on spectral intensities relative to the internal standard as described previously [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. Candidate metabolites present in each of the complex mixtures were determined by matching the chemical shift, J-coupling, and intensity information of experimental NMR signals against the NMR signals of standard metabolites in the Chenomx library. Compounds were assigned a rank and assign confidence to metabolites (RANCM) value according to the amount of spectral information used to identify the compound (\u003cb\u003eAdditional File 7\u003c/b\u003e) [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFor many metabolites, including aspartate, asparagine, sucrose, acetate, methanol, and glucose, we utilized 2D NMR to corroborate the 1D data, providing more confidence to an assignment. The two-dimensional 1H-1H total correlation spectroscopy (TOCSY) spectra were collected using the Varian TOCSY pulse sequence with a TOCSY mixing time of 80 ms (MLEV-17). Spectral widths were 12 ppm in both directions with 256 increments acquired in the indirect dimension and 64 transients per increment. The relaxation delay was 1.5 s during which presaturation of the water signal was applied and the acquisition time was 143 ms during which 2048 total points were acquired. The 2D 1H-13C heteronuclear single-quantum correlation spectroscopy (HSQC) spectra were acquired using the Varian gHSQCAD pulse sequence with a 1JCH of 146 Hz. Spectral widths were 12 ppm and 160 ppm for the direct and indirect dimensions, respectively, with 256 increments acquired in the indirect dimension and 128 transients per increment. The relaxation delay was 1.5 s during which presaturation of the water signal was applied. The acquisition time was 143 ms in which 13C composite pulse decoupling (wurst140) was applied and 2048 total points were acquired. NMR-identified metabolites discussed in the text were present in 30% of the samples.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results And Discussion","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eThe HUM-V genome database enabled metaproteomic characterization of river sediment microbiomes\u003c/h2\u003e \u003cp\u003eHere we created the Hyporheic Uncultured Microbial and Viral (HUM-V) genomic catalog from Columbia River HZ sediments. We leveraged this resource for metaproteomic peptide recruitment, enabling identification of the community members and their gene expression in these sediments.\u003c/p\u003e \u003cp\u003eWe reconstructed 655 bacterial and archaeal metagenome assembled genomes (MAGs); 102 were medium or high-quality genomes based on the Genome Consortium Standards [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] (\u003cb\u003eAdditional File 2\u003c/b\u003e). These genomes were dereplicated into 55 genomic representatives to form the bacterial and archaeal portion of the HUM-V microbial genome database. These dereplicated HUM-V MAGs were distributed across 9 Bacterial and 2 Archaeal phyla. In terms of new genomic discoveries, 1 genome represented a new order within the Actinobacteriota, and 12 genomes represented 6 new genera from archaeal and bacterial phyla including members of the Thermoplasmatota, Acidobacteria, Actinobacteriota, CSP1-3, Proteobacteria, and Desulfobacterota (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003eFrom the same metagenomic assemblies we reconstructed and reported viral metagenome assembled genomes (vMAGs), making this one of only a handful of genome-resolved studies that include viral genomes derived from rivers [\u003cspan additionalcitationids=\"CR47\" citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e], and to our knowledge, the first study to complement these with bacterial and archaeal genomes. We reconstructed 2,482 vMAGs that dereplicated into 111 dereplicated viral populations\u0026thinsp;\u0026gt;\u0026thinsp;10kb in size (\u003cb\u003eAdditional File 6\u003c/b\u003e). Given their sparse sampling from river corridors, only 5 of the HUM-V viral genomes had taxonomic assignments using established viral taxonomies from standard reference databases. To better understand if the remaining 95% (n\u0026thinsp;=\u0026thinsp;105) of viral genomes were completely novel or had been previously detected in similar ecosystems, we repeated the analyses, this time adding 1,861 viral genomes we reconstructed or pulled from public metagenomes from four freshwater sites in North and South America (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec, \u003cb\u003eAdditional File 6\u003c/b\u003e). Of the 105 remaining viral genomes in HUM-V, 15% (n\u0026thinsp;=\u0026thinsp;17) clustered with these freshwater derived sequences, indicating a portion of this viral community is shared across diverse geographic and freshwater systems. Of the remaining viral genomes, 23% (n\u0026thinsp;=\u0026thinsp;26) clustered only with genomes recovered in this data set, indicating multiple samplings of the same virus spatially at this site, while 57% (n\u0026thinsp;=\u0026thinsp;63) of the viral genomes we sampled were singletons (i.e., only sampled from these sediments once). These results hint at the possible cosmopolitan and endemic viral lineages that warrant further exploration.\u003c/p\u003e \u003cp\u003eHUM-V recruited viral and microbial peptides from our HZ sediment metaproteomic dataset (n\u0026thinsp;=\u0026thinsp;33 lateral and depth resolved samples) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb\u003cb\u003ed, Additional File 5\u003c/b\u003e). Across all sediment samples, microbial genomes recruited 13,102 total peptides to ~\u0026thinsp;1,300 proteins in HUM-V, with 68% of these proteins uniquely assigned to a single microbial genome. For viruses and microbes alike, the most abundant genomes were not necessarily the most actively expressing proteins. The most abundantly ranked microbial members included the Nitrospiraceae genus NS7, Binatia, and Nitrososphaeraceae genus TA-21 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb), yet only the Nitrososphaeraceae had high proteomic recruitment (15%). Similarly, some low abundance members (e.g., members of the Actinobacteria) accounted for a majority of the uniquely assigned proteome relative abundance (49%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb). Like our microbial dataset, 66% of the viral genomes encoded genes that uniquely recruited peptides (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed). This exceeded prior viral metaproteome recruitment from other environmental systems (e.g., wastewater, saliva, rumen (0.4\u0026ndash;15%, [\u003cspan additionalcitationids=\"CR50\" citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]), thus we infer a relatively large portion of the viral community was active at the time of sampling. While microbial and viral activity did not appear to be structured by transect, sediment depth, or geochemical conditions, these two assemblages were coordinated to one another (\u003cb\u003eAdditional File 3: Figure S6\u003c/b\u003e). Explaining this lack of geochemical or spatial structuring, it is possible that the microbial heterogeneity in these samples occurred over a finer spatial resolution (pore or biofilm scale) than the bulk 10 cm depths sampled or that these HZ sediment microbiomes are metabolically robust to the small, but significant changes in chemistry measured across spatial gradients (\u003cb\u003eAdditional File 3: Figure S6, Additional File 3: Figure S7, Additional File 7\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eMicrobial cross feeding of organic carbon is likely sustained by aerobic respiration\u003c/h2\u003e \u003cp\u003eIt is well recognized that microbial carbon oxidation in HZ sediments largely contributes to river respiration, yet the microbial food webs underpinning this process have yet to be documented. Consistent with resazurin (raz) data (see Additional File 3: supplementary methods) that indicated these sediments were oxygenated and supported aerobic microbial respiration (\u003cb\u003eAdditional File 3: Figure S7\u003c/b\u003e) [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e], all but one of the microbial genomes recovered from this site encoded aerobic respiration machinery, including a complete electron transport chain and a cytochrome oxidase (\u003cb\u003eAdditional File 3: Figure S1\u003c/b\u003e). Proteomic evidence for aerobic respiration (cytochrome c oxidase \u003cem\u003eaa3\u003c/em\u003e) was detected from nearly all samples, but only assigned to few members of the Nitrososphaeraceae. However, given limitations with detecting membrane cytochromes [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e], we consider it likely this metabolism was more active than was captured in proteomic data, as we failed to find any evidence for other anaerobic metabolisms (e.g., methanogenesis).\u003c/p\u003e \u003cp\u003eWhile the overall carbon content of these sediments was low (\u0026lt;\u0026thinsp;10 mg/g) (\u003cb\u003eAdditional File 7\u003c/b\u003e), our FTICR-MS analysis indicated that plant litter could be an important substrate, as lignin-like compounds were the most abundant biochemical class detected (\u003cb\u003eAdditional File 3: Figure S5\u003c/b\u003e, \u003cb\u003eAdditional File 8\u003c/b\u003e). In support of this, from our metagenomes, 38% of the HUM-V genomes encoded genes for potentially degrading phenolic/aromatic monomers, while 10% could degrade the larger, more recalcitrant polymers (\u003cb\u003eAdditional File 2\u003c/b\u003e). Gene expression of carbohydrate-active enzymes (CAZymes) also supported the degradation of plant polymers like starch and cellulose via extracellular glucoamylase (GH15) and endo-glucanase (GH5) from an actinobacterial genome (Microm_1) and the Nitrososphaeraceae (Nitroso_2), respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). In summary, many types of chemical and biological data reveal that heterotrophic, aerobic metabolism in these low carbon sediments is likely maintained by inputs from decomposition.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGiven the capacity for plant polymer decomposition (e.g., lignin, cellulose, and starch) across HUM-V genomes, we next tracked the microbial fate of the degradation products of these metabolisms, including sugar monomers, short chain fatty acids, and carbon dioxide (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, \u003cb\u003eAdditional File 2, Additional File 5\u003c/b\u003e). Metabolites detected by NMR included sugars (e.g., glucose, sucrose, and trehalose), which could be the result of depolymerization of plant derived polymers, and we confirmed the CAZYmes to use these substrates were also expressed \u003cem\u003ein situ\u003c/em\u003e. Additionally, NMR also detected organic acids (acetate, butyrate, lactate, pyruvate, propionate) and alcohols (ethanol, methanol, isopropanol), with proteomics supporting the usage of acetate and methanol by Anaeromyxobacter MAG (Anaerom_1) and archaeal Woeseia (Woese_1), respectively. Here our metabolite and proteomic data demonstrated that plant biomass degradation supports sequential metabolic handoffs that lead to carbon dioxide production.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eCarbon dioxide production and consumption is widely encoded by HUM-V microorganisms\u003c/h2\u003e \u003cp\u003eIn addition to carbon dioxide being generated from the heterotrophic metabolisms described above, our proteomics revealed that carbon dioxide could arise by the aerobic oxidation of carbon monoxide (CO). Genes for aerobic CO dehydrogenases (from Actinobacteria, Binatia, and CSP-1 genomes) were among the most expressed in these sediments. Analogous to findings from soil systems, it is possible that atmospheric carbon monoxide is a major energy source supporting persistent aerobic heterotrophic bacteria in deprived, or dynamic organic carbon environments [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. Based on the genomic inventory of these HUM-V genomes, we posit that Binatia, CSP1-3, and Micromonosporaceae are capable of carboxydotrophy, while Actino_1 is a carboxydovore, using CO metabolism as supplemental energy or possible carbon source during starvation [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSince heterotrophic respiration and carbon monoxide oxidation would generate carbon dioxide in these sediments, we next tracked microorganisms that could use this carbon source autotrophically (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, \u003cb\u003eAdditional File 2, Additional File 5\u003c/b\u003e). The ability to fix carbon was prevalent, encoded by 75% of HUM-V microbial genomes. In fact, this metabolism was represented by multiple fixation pathways from 18 different lineages, demonstrating both functional and taxonomic redundancy. Specifically, this includes 4 different pathways (e.g., Calvin-Benson-Bassham cycle, reductive TCA cycle, 3-HydroxyPropionate /4-HydroxyButyrate cycle, 3-Hydroxypropionate bi-cycle) from members of nitrifying lineages (Thaumarchaeota and Nitrospirota) (discussed below), as well as from organisms with heterotrophic capabilities like Binatia, CSP1_3, Proteobacteria, Woeseiaceae, and Acidobacteria (\u003cb\u003eAdditional File 3: Figure S1\u003c/b\u003e, \u003cb\u003eAdditional File 2\u003c/b\u003e). Collectively our multi-omics data suggest that sediment microbial respiration is likely decoupled from river respiration, since some microbially produced carbon dioxide would be lost to supporting autotrophy. Our research further resolves the carbon economy in HZ sediments, implying that the net effect of carbon dioxide emissions from rivers could depend on the balance between carbon dioxide production from heterotrophy and carbon monoxide, as well as consumption by autotrophs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eMicrobial metaproteomics supports theoretical inferences derived from geochemistry\u003c/h2\u003e \u003cp\u003eThe ratio of total element carbon (C) and total nitrogen (N) (e.g., C/N) is a geochemical indicator often used to assess the possible microbial metabolisms that can be supported in an ecosystem [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e, \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. Here the C/N ratios of these sediments were relatively low to other sediments at 6.4\u0026thinsp;\u0026plusmn;\u0026thinsp;1.1 across the samples (\u003cb\u003eAdditional File 7\u003c/b\u003e). Biogeochemical theory posits that C/N values less than 15 would indicate rapid microbial mineralization of organic nitrogen to release inorganic nitrogen [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]. This theory also states that C/N ratios less than 10 may indicate ammonium is released to the surrounding environment, allowing sufficient concentrations to simultaneously support the assimilatory needs of heterotrophs and energy needs of nitrifiers, allowing for their co-occurrence [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. Our multi-omics data offered a new opportunity to substantiate these geochemical inferences by profiling the possible substrates and microbial activity of nitrogen mineralizers and nitrifiers in river sediments.\u003c/p\u003e \u003cp\u003eGiven the prevalence of ammonium in all 33 sediment samples (0.28\u0026ndash;11.22 \u0026micro;g gram\u003csup\u003e\u0026minus;\u0026thinsp;1\u003c/sup\u003e) (\u003cb\u003eAdditional File 3: Figure S8, Additional File 7\u003c/b\u003e), we next examined our metaproteomic data for peptidases, genes that could contribute to the mineralization of organic nitrogen into amino acids and free ammonium. Hinting at the relevance of this metabolism, the gene expression of peptidases (n\u0026thinsp;=\u0026thinsp;31) was 3 times more abundant and prevalent than glycoside hydrolase genes modulating organic C transformations (\u003cb\u003eAdditional File 5\u003c/b\u003e). In support of active microbial N mineralization, hydrophobic, polar, and hydrophilic amino acids were prevalent (more so than sugars) in the H\u003csup\u003e1\u003c/sup\u003e-NMR characterized metabolites (\u003cb\u003eAdditional File 3: Figure S8\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWe focused our analyses on the putative extracellular peptidases, as these were most likely to shape organic nitrogen pools in the sediment. We categorized expressed peptidase families as either amino acid releasing (end terminus cleaving, e.g., M28) or peptide releasing (endocleaving, e.g., S08A, M43B, M36, MO4) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, \u003cb\u003eAdditional File 5\u003c/b\u003e). Linking these expressed peptidases to our genomes, members of the Actinobacteria, Thermoproteota, and Methylomirabilota, and Binatia are likely candidates for driving the mineralization of organic N. We then profiled amino acid transporters that were expressed, revealing uptake of branched chain amino acids, glutamate, osmoprotectants, spermidine/putrescine, and peptides (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). This profiling indicated synergy and competitions for this organic N resource in these sediments.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe propose that in HZ sediments extracellular peptidases are a shared public good whose cost of production is assumed by certain individuals with benefits to the entire community [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. In some cases, taxa that mineralized organic N were consumers of the resulting products, as genomes in the Actinobacteria and Binatia expressed external peptidases genes and the genes for transporting the organic N products (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, linkages shown). In other instances, members of the Proteobacteria, Thermoplasmatota, and CSP1-3 could be functioning as cheater cells that expressed only genes for intracellular transport and benefitting from peptidases produced by others. Our findings reinforce that cooperative interactions based on cross-feeding and public goods are likely at the core of many processes relevant to organic carbon (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) and nitrogen (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) cycling in these sediments.\u003c/p\u003e \u003cp\u003eConsistent with established conceptual geochemical theory, we showed the lower C:N ratios (\u0026lt;\u0026thinsp;10) of these sediments not only supported mineralization which could be a source of free ammonium in these sediments, but also nitrification. Supporting this, ammonium was detected in all sediments (average concentration 2.6 \u0026micro;g/gram of sediment) (\u003cb\u003eAdditional File 3: Figure S8, Additional File 7\u003c/b\u003e). Proteomics confirmed ammonium (NH\u003csub\u003e4\u003c/sub\u003e\u003csup\u003e+\u003c/sup\u003e) oxidation to nitrite was performed by Archaeal Nitrososphaeraceae (formerly Thaumarchaeota), with ammonia monooxygenase proteins being one of the most prevalent and highly expressed functional proteins (top 5%) across this dataset (\u003cb\u003eAdditional File 5\u003c/b\u003e). The next step in nitrification, nitrite oxidation to nitrate was inferred from nitrite oxidoreductase peptides assigned to 5 genomes belonging to 2 new species (Nitro_40CM-3_1, Nitro_NS7_3, Nitro_NS7_4, Nitro_NS7_5, and Nitro_NS7_14) (\u003cb\u003eAdditional File 3: Figure S9\u003c/b\u003e, \u003cb\u003eAdditional File 5\u003c/b\u003e, see sheet metabolism info). Both nitrifying lineages had the capacity for carbon dioxide fixation with the reductive tricarboxylic acid (TCA) cycle (e.g., ATP-citrate lyase) in Nitrospiraceae genomes, and 3-HydroxyPropionate/4-HydroxyButyrate (3HP/4HB) encoded by the Nitrososphaeraceae. We did not detect genomic evidence for comammox or anammox and thus aerobic, chemolithoautotrophic nitrification supported by a metabolic partnership between bacteria and archaea occurred in the presence of heterotrophs as predicted by C/N ratios.\u003c/p\u003e \u003cp\u003eSimilarly, others have reported the prominence of nitrifying lineages from the archaeal thaumarcheotal Thermoproteota and bacterial Nitrospirota both by 16S rRNA [\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e] and genome-resolved metagenomics [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] in HZ sediments. Here we nearly doubled the genomic sampling of these river nitrifiers, assigning unique gene expression patterns to 3 and 17 genomes from Nitrososphaeraceae and Nitrospiraceae respectively, including the first genomic sampling of new genera and species (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Our co-expression data indicate that metabolic handoffs between archaeal ammonia oxidizers and bacterial nitrite oxidizers may be an unaccounted-for biogenic source of nitrate in these sediments (\u003cb\u003eAdditional File 5\u003c/b\u003e). This suggests the activity of nitrifiers could be an underappreciated modulator of nitrous oxide fluxes from oligotrophic HZ sediments, both through their indirect stimulation of denitrifiers and their own contributions to this greenhouse flux [\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e]. Taken together, the archaeal-bacterial nitrifying mutualism outlined here appears well adapted to the low nutrient conditions present in many HZ sediments, warranting future research on the variables that constrain nitrification rates (i.e., ammonium availability, dissolved oxygen, pH) and their role as driver of nitrogen fluxes from these systems [\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eDenitrification is encoded by novel and taxonomically diverse lineages in HZ sediments\u003c/h2\u003e \u003cp\u003eBeyond the possible biogenic sources of nitrate, we identified from nitrification, these HZ sediments receive significant allochthonous nitrate from groundwater. When river stage decreases, groundwater discharges through the HZ sediments, bringing nitrate concentrations to over 20 mg/L [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e]. In support of an important influence of nitrate from either source, HUM-V genomes with the capacity for nitrate reduction spanned diverse taxonomies, with NarG or NapX encoded in 11 genomes from the Actinobacteriota, Binatia, Gammaproteobacteria, and Myxococcota (\u003cb\u003eAdditional File 3: Figure S1\u003c/b\u003e). However, our proteomic evidence for nitrate reduction was detected in less than 10% of the 33 sediment samples, with unique peptides assigned to Binatia NarG from a single sample.\u003c/p\u003e \u003cp\u003eBased on gene expression data, we inventoried other steps in the denitrification pathway. Nitrite was reduced via nitrifier and denitrifier reduction to nitric oxide from archaeal ammonia oxidizers of the Nitrososphaeraceae active in 79% of metaproteome samples, and from Gammaproteobacterial Burkholderia in a single sample, respectively. The role of nitrite reduction by Nitrososphaeraceae is still under investigation but could be used for detoxification [\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e]. Genes for converting nitric oxide to nitrous oxide were not detected in proteomics, but we did find evidence that the Desulfobacterota genome (Desulf_UBA2774_1) expressed the \u003cem\u003enos\u003c/em\u003e gene for reducing nitrous oxide to nitrogen gas. Phylogenetic analysis suggest this organism used a \"Clade II\u0026rdquo; \u003cem\u003enos\u003c/em\u003e sequence type adapted for low atmospheric concentrations of nitrous oxide (\u003cb\u003eAdditional File 3: Figure S2\u003c/b\u003e), and consistent with our genome metabolic summary did so without encoding other steps of the denitrification pathway [\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e]. Notably, the capacity for denitrification exists beyond those detected in proteomics, as Binatia encoded dissimilatory nitrite reduction to ammonium (DNRA) and the potential for nitrous oxide production via \u003cem\u003enor\u003c/em\u003e was encoded by two Gammaproteobacteria (Steroid-FEN-1191_1, Steroid_1) and a member of the Myxococcota (Anaerom_1).\u003c/p\u003e \u003cp\u003eIn summary, our data adds to the growing realization that complete denitrification by single microorganism is likely the exception rather than the rule in natural systems [\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e], including the HZ [\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e]. In support of this, none of the genomes reconstructed here encoded a complete denitrification pathway for reducing nitrate to nitrous oxide or dinitrogen gas (\u003cb\u003eAdditional File 3: Figure S1\u003c/b\u003e). Similarly, our proteomics data hinted that separate microbial members likely catalyzed each step of the denitrification pathway (\u003cb\u003eAdditional File 3: Figure S4\u003c/b\u003e). This suggests cross-organism inorganic nitrogen exchange would be necessary for nitrogen gas flux, such that physical processes (e.g., advection, diffusion) or the spatial colocalization of microorganisms, as well as organic carbon availability, may have disproportionate impacts on flux of nitrous oxide and dinitrogen from these sediments.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eHUM-V identifies new microbial and viral players in hyporheic zone carbon and nitrogen cycling\u003c/h2\u003e \u003cp\u003eThe creation of a genome database expanded upon prior amplicon-based surveys, allowing us to assign new metabolic functions to microbes and even viruses in hyporheic sediments. While HUM-V contains genomes from phyla (CSP1-3, Eisenbacteria) and classes (Binatia, MOR-1 in Acidobacteria) composed entirely of uncultivated members (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ed), here we focus our analysis on the Binatia, as we recovered 7 genomes (one which included a complete 16S rRNA gene), they recruited peptides, and they also played key roles in carbon and nitrogen cycling. Using the 16S rRNA gene (from Binatia_7), we inventoried the distribution of closely related species to our HUM-V genomes (\u0026gt;\u0026thinsp;97% similarity) in the Sequence Read Archive (SRA) samples, to uncover the ecological distribution of these organisms from soils, as well as a wide variety of terrestrial, terrestrial-aquatic, marine samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e), indicating the processes uncovered by proteomics here are likely applicable to a wide range of ecosystems.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA recent comparative genomics analysis on Binatota MAGs provided a first assessment of their metabolic potential, indicating genes for methylotrophy, alkane degradation, and pigment production were distributed across the phylum [\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e]. These HUM-V genomes belong to a class and family denoted UBA9968. Contrary to their prior metabolic inventory, HUM-V UBA9968 MAGs do not encode the potential for methanol oxidation, and we identified a new role in the decomposition of aromatic compounds from plant biomass (phenylpropionic acid, phenylacetic acid, salicylic acid), and xenobiotics (phthalic acid) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). We provide the first proteomic evidence for any members of the Binatia, supporting their roles in aerobically oxidizing carbon monoxide, producing extracellular peptidases, and in denitrification. Together these findings illustrate the power of HUM-V paired proteomes to illuminate new roles for members of uncultivated, previously enigmatic lineages in HZ carbon and nitrogen cycling.\u003c/p\u003e \u003cp\u003eThe relatively high proteomic recruitment of viruses sampled in HUM-V (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed) suggested important viral contributions in these sediments. \u003cem\u003eIn silico\u003c/em\u003e analysis assigned a putative host to 29% of the 111 viral genomes linking 18 microbial genomes that belong to bacterial members in Acidobacteriota, Actinobacteriota, CSP1-3, Eisenbacteria Methylomirabilota, Myxococcota, Nitrospirota, and Proteobacteria (\u003cb\u003eAdditional File 3: Figure S10\u003c/b\u003e, \u003cb\u003eAdditional file 2\u003c/b\u003e, \u003cb\u003eAdditional file 6\u003c/b\u003e). Analysis of the metaproteomes for these phage-impacted microorganisms revealed these hosts expressed genes for nitrification (Nitrospiraceae) as well as carbon monoxide oxidation and nitrogen mineralization (Actinobacteria) (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Additionally, HUM-V phage genomes encode auxiliary metabolic genes with the potential to enhance microbial metabolism of carbon (CAZymes), sulfur (sulfate adenyl transferase), and nitrogen (amidase to cleave ammonium) (\u003cb\u003eAdditional File 3: Figure S11\u003c/b\u003e, see Additional File 3 supplemental text). We also show viral abundances were better predictors of total carbon and nitrogen percentages relative to microbial genome abundances (\u003cb\u003eAdditional File 3: Figure S12, Additional File S9\u003c/b\u003e, see Additional File 3 supplemental text). Together, these HUM-V enabled results indicate viral infections may contribute to river sediment functioning and raise the question to whether enhanced viral interrogation might provide a means to improved ecosystem or biogeochemical models in these systems.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusions","content":"\u003cp\u003eTo our knowledge this study represents one of the first genome-resolved microbial and viral enabled proteomic studies in river sediments. Using genome-resolved proteomics with complementary metabolites (detected by NMR, FTICR-MS), and geochemistry we begin to illuminate the microbial contributions to processes well known to occur but previously poorly defined mechanistically in river sediments (e.g., nitrogen mineralization). We also show how multi-omic tools can uncover previously enigmatic processes which may directly impact river respiration (e.g., carbon monoxide oxidation). While river carbon and nitrogen budgets are often quantified by direct measurements of inputs and the concentration of inorganic and organic compounds exported from rivers, what is missing today is an appreciation for the microbially and virally mediated sources and sinks for key intermediates (e.g., carbon dioxide, ammonium, nitrate), the degree to which these compounds are recycled and exchanged, and the underlying microbial metabolic lifestyles that catalyze this interconnected carbon and nitrogen biogeochemistry. Here, we have created a conceptual framework that elaborates on these missing ideas.\u003c/p\u003e \u003cp\u003eEmpowered by our individual process-based metaproteomic analyses (Figs.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e), we created a conceptual model outlining the microbial conversions of carbon and nitrogen in these hyporheic sediments (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e). Heterotrophic oxidation of organic carbon derived from plant (and likely microbial) biomass supported by oxygen and nitrogen respiring populations produce carbon dioxide. In addition, metaproteomics divulged that aerobic carbon monoxide oxidation may also be a source of carbon dioxide. Like organic carbon, the organic nitrogen in microbial and plant biomass could be mineralized to release ammonium in these sediments. Together inorganic pools of nitrogen (ammonium) and carbon (carbon dioxide) sustain the coordinated activity of nitrifying populations. Together our findings put forth an integrated framework that advances microbial roles in hyporheic carbon and nitrogen transformations, yielding insights that could inform research strategies to reduce existing predictive uncertainties in river corridor models.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eHUM-V:\u0026nbsp;Hyporheic Uncultured Microbial and Viral (HUM-V) database\u003c/p\u003e\n\u003cp\u003eMAG: Metagenome assembled genome\u003c/p\u003e\n\u003cp\u003evMAG: Viral metagenome assembled genome\u003c/p\u003e\n\u003cp\u003eHZ: Hyporheic zone\u003c/p\u003e\n\u003cp\u003eFTICR-MS: Fourier-transform ion cyclotron resonance mass spectrometry\u003c/p\u003e\n\u003cp\u003eNMR: Nuclear magnetic resonance\u003c/p\u003e\n\u003cp\u003eDNA: Deoxyribonucleic acid\u003c/p\u003e\n\u003cp\u003erRNA: Ribosomal ribonucleic acid\u003c/p\u003e\n\u003cp\u003etRNA: Transfer ribonucleic acids\u003c/p\u003e\n\u003cp\u003eJGI: Joint Genome Institute\u003c/p\u003e\n\u003cp\u003eEMSL: Environmental Molecular Sciences Laboratory\u003c/p\u003e\n\u003cp\u003eOSU: The Ohio State University\u003c/p\u003e\n\u003cp\u003eNCBI: National Center for Biotechnology Information\u003c/p\u003e\n\u003cp\u003eSRA: Sequence Read Archive\u003c/p\u003e\n\u003cp\u003eGbp: Giga base pair\u003c/p\u003e\n\u003cp\u003eIDBA-UD: Iterative De Bruijn Graph De Novo Assembler \u0026ndash; Uneven Depth\u003c/p\u003e\n\u003cp\u003emetaSPAdes: metagenomic St. Petersburg genome Assembler\u003c/p\u003e\n\u003cp\u003eGTDB: Genome Taxonomy Database\u003c/p\u003e\n\u003cp\u003eANI: Average Nucleotide Identity\u003c/p\u003e\n\u003cp\u003eDRAM: Distilled and Refined Annotation of MAGs\u003c/p\u003e\n\u003cp\u003eIMG/MER: Integrated Microbial Genomes / Microbiomes Expert Review\u003c/p\u003e\n\u003cp\u003ePFAM database: Protein Family Database\u003c/p\u003e\n\u003cp\u003eMUSCLE: Multiple Sequence Comparison by Log-Expectation\u003c/p\u003e\n\u003cp\u003eMAFFT: Multiple Alignment using Fast Fourier Transform\u003c/p\u003e\n\u003cp\u003e\u003cem\u003enar\u003c/em\u003e: respiratory nitrate reductase\u003c/p\u003e\n\u003cp\u003e\u003cem\u003enxr\u003c/em\u003e: nitrite oxidoreductase\u003c/p\u003e\n\u003cp\u003e\u003cem\u003enos\u003c/em\u003e: nitric oxide synthase\u003c/p\u003e\n\u003cp\u003eDNRA: Dissimilatory Nitrate Reduction to Ammonium\u003c/p\u003e\n\u003cp\u003eMIUViG: Minimum information about an Uncultivated Virus Genome\u003c/p\u003e\n\u003cp\u003eICTV: International Committee on Taxonomy of Viruses\u003c/p\u003e\n\u003cp\u003ePHYRE2: Protein Homology / AnalogY Recognition Engine 2.0\u003c/p\u003e\n\u003cp\u003esPLS: sparse Partial Least Squares Regression\u003c/p\u003e\n\u003cp\u003eMS: Mass Spectrometry\u003c/p\u003e\n\u003cp\u003eMS / MS: Tandem mass spectrometry\u003c/p\u003e\n\u003cp\u003eLC-MS/MS: Liquid Chromatography Tandem Mass Spectrometry\u003c/p\u003e\n\u003cp\u003eFDR: False Detection Rate\u003c/p\u003e\n\u003cp\u003ePSM: Peptide Spectrum Matches\u003c/p\u003e\n\u003cp\u003eNSAF: Normalized spectral abundance frequency\u003c/p\u003e\n\u003cp\u003eMHz: Megahertz\u003c/p\u003e\n\u003cp\u003eDSS-d6: 3-(Trimethylsilyl)-1-propanesulfonic acid-d6 sodium salt\u003c/p\u003e\n\u003cp\u003eppm: parts per million\u003c/p\u003e\n\u003cp\u003eRANCM: Rank and AssigN Confidence to Metabolites\u003c/p\u003e\n\u003cp\u003eTOCSY: Total Correlation Spectroscopy\u003c/p\u003e\n\u003cp\u003eHSQC: Heteronuclear single-quantum correlation spectroscopy\u003c/p\u003e\n\u003cp\u003egHSQCAD: Gradient-enhanced Heteronuclear Single Quantum Coherence with Adiabatic Pulses\u003c/p\u003e\n\u003cp\u003eCAZymes: Carbohydrate-Active Enzymes\u003c/p\u003e\n\u003cp\u003eGH: Glycoside Hydrolase\u003c/p\u003e\n\u003cp\u003eCO: Carbon Monoxide\u003c/p\u003e\n\u003cp\u003eTCA: Tricarboxylic acid\u003c/p\u003e\n\u003cp\u003eC: Carbon\u003c/p\u003e\n\u003cp\u003eN: Nitrogen\u003c/p\u003e\n\u003cp\u003eC:N: Carbon / Nitrogen Ratio\u003c/p\u003e\n\u003cp\u003eNH\u003csub\u003e4\u003c/sub\u003e\u003csup\u003e+:\u003c/sup\u003e Ammonium\u003c/p\u003e\n\u003cp\u003e3HP/4HB: 3-HydroxyPropionate/4-HydroxyButyrate\u003c/p\u003e\n\u003cp\u003eATP: Adenosine triphosphate\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cem\u003eEthics approval and consent to participate:\u0026nbsp;\u003c/em\u003eNot applicable.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eConsent for publication:\u0026nbsp;\u003c/em\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAvailability of data and material:\u0026nbsp;\u003c/em\u003eThe datasets supporting the conclusions of this article are publicly available. Sequencing data are available in NCBI under bioproject PRJNA576070, with MAGs deposited under biosamples SAMN18867633-SAMN18867734 and 16S rRNA amplicon sequences under accession numbers SRX9312157-SRX9312180, and vMAGs have been temporarily deposited in Zenodo doi 10.5281/zenodo.5124937. Viral genomes fasta file is publicly available within the following repository link: https://doi.org/10.5281/zenodo.5124937. Metaproteomics data are deposited in the MassIVE database under accession MSV000087330. Metabolomics data are publicly available and deposited in Zenodo doi https://doi.org/10.5281/zenodo.5076253. Additional datasets supporting the conclusions of this article are included within the article (and its additional files).\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eCompeting interests:\u0026nbsp;\u003c/em\u003eThe authors declare they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eFunding:\u003c/em\u003e This work was supported by the Subsurface Biogeochemical Research (SBR) program (DE-SC0018170); the National Sciences Foundation Division of Biological Infrastructure [#1759874]; and\u0026nbsp;the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, Environmental System Science (ESS) program through subcontract from the River Corridor Scientific Focus Area project at Pacific Northwest National Laboratory. J.R.R. is funded by the National Science Foundation (NRT-DESE) [1450032], support for A Trans-Disciplinary Graduate Training Program in Biosensing and Computational Biology at Colorado State University.\u0026nbsp;The NMR data, FTICR-MS data and MS-proteomics data in this work was collected using instrumentation in the Environmental Molecular Science Laboratory (grid.436923.9), a DOE Office of Science User Facility sponsored by the Office of Biological and Environmental Research and located at Pacific Northwest National Laboratory. Pacific Northwest National Lab is operated by Battelle for the DOE under Contract DE-AC05-76RL01830. Metagenomic sequencing for this research was performed by the Joint Genome Institute via a large-scale sequencing award (Award 1781) and at the Genomics Shared Resource Core at The Ohio State University Comprehensive Cancer Center supported by P30 CA016058.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAuthors\u0026rsquo; contributions:\u0026nbsp;\u003c/em\u003eUsing the CRediT contributor roles,author contributions can be defined as follows:Conceptualization,JRR, MAB, JCS, and KCW; \u0026nbsp;Data curation, JRR, MAB, and BBM; Formal analysis, JRR, MAB, BBM, GJS, LMS, RAD; Funding acquisition, JCS, KCW; Project administration, RAD, JCS, KCW; Investigation, JRR, MAB, BBM, GJS, SOP, CDN, DWH; Supervision, MSL, EBG, DWH, JCS, KCW; Writing- original draft, JRR, MAB, BBM, GJS, KCW; Writing- review and editing, \u0026nbsp; JRR, MAB, BBM, GJS, JCS, KCW.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAcknowledgements:\u0026nbsp;\u003c/em\u003eThe authors would like to thank Tyson Claffey and Richard Wolfe for Colorado State University server management; Sandy Shew for management of computing resources retained from The Ohio State University Unity cluster; Dr. Pearlly Yan at the Genomics Shared Resource Core at The Ohio State University Comprehensive Cancer Center for management of metagenomic sequencing; and Dr. J John for the continuous support.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAuthor\u0026apos;s information:\u0026nbsp;\u003c/em\u003eJ.R.R. and M.A.B. contributed equally to this work.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBoulton AJ, Findlay S, Marmonier P, Stanley EH, Valett HM. The functional significance of the hyporheic zone in streams and rivers. Annu Rev Ecol Syst. Annual Reviews 4139 El Camino Way, PO Box 10139, Palo Alto, CA 94303-0139, USA; 1998;29:59\u0026ndash;81. \u003c/li\u003e\n\u003cli\u003eStegen JC, Johnson T, Fredrickson JK, Wilkins MJ, Konopka AE, Nelson WC, et al. Influences of organic carbon speciation on hyporheic corridor biogeochemistry and microbial ecology. Nat Commun. Nature Publishing Group; 2018;9:1\u0026ndash;11. \u003c/li\u003e\n\u003cli\u003eNewcomer ME, Hubbard SS, Fleckenstein JH, Maier U, Schmidt C, Thullner M, et al. Influence of hydrological perturbations and riverbed sediment characteristics on hyporheic zone respiration of CO2 and N2. J Geophys Res Biogeosciences. Wiley Online Library; 2018;123:902\u0026ndash;22. \u003c/li\u003e\n\u003cli\u003eTrimmer M, Grey J, Heppell CM, Hildrew AG, Lansdown K, Stahl H, et al. River bed carbon and nitrogen cycling: state of play and some new directions. Sci Total Environ. Elsevier; 2012;434:143\u0026ndash;58. \u003c/li\u003e\n\u003cli\u003eVilla JA, Smith GJ, Ju Y, Renteria L, Angle JC, Arntzen E, et al. Methane and nitrous oxide porewater concentrations and surface fluxes of a regulated river. Sci Total Environ. Elsevier; 2020;715:136920. \u003c/li\u003e\n\u003cli\u003eHu M, Chen D, Dahlgren RA. Modeling nitrous oxide emission from rivers: a global assessment. Glob Chang Biol. Wiley Online Library; 2016;22:3566\u0026ndash;82. \u003c/li\u003e\n\u003cli\u003eBeaulieu JJ, Tank JL, Hamilton SK, Wollheim WM, Hall RO, Mulholland PJ, et al. Nitrous oxide emission from denitrification in stream and river networks. Proc Natl Acad Sci. National Acad Sciences; 2011;108:214\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eCaruso A, Boano F, Ridolfi L, Chopp DL, Packman A. Biofilm-induced bioclogging produces sharp interfaces in hyporheic flow, redox conditions, and microbial community structure. Geophys Res Lett. Wiley Online Library; 2017;44:4917\u0026ndash;25. \u003c/li\u003e\n\u003cli\u003eNaegeli MW, Uehlinger U. Contribution of the hyporheic zone to ecosystem metabolism in a prealpine gravel-bed-river. J North Am Benthol Soc. North American Benthological Society; 1997;16:794\u0026ndash;804. \u003c/li\u003e\n\u003cli\u003eBattin TJ, Kaplan LA, Newbold JD, Hendricks SP. A mixing model analysis of stream solute dynamics and the contribution of a hyporheic zone to ecosystem function. Freshw Biol. Wiley Online Library; 2003;48:995\u0026ndash;1014. \u003c/li\u003e\n\u003cli\u003eZhang M, Daraz U, Sun Q, Chen P, Wei X. Denitrifier abundance and community composition linked to denitrification potential in river sediments. Environ Sci Pollut Res. Springer; 2021;1\u0026ndash;12. \u003c/li\u003e\n\u003cli\u003eLiu S, Wang H, Chen L, Wang J, Zheng M, Liu S, et al. Comammox Nitrospira within the Yangtze River continuum: community, biogeography, and ecological drivers. ISME J. Nature Publishing Group; 2020;14:2488\u0026ndash;504. \u003c/li\u003e\n\u003cli\u003ePinto OHB, Silva TF, Vizzotto CS, Santana RH, Lopes FAC, Silva BS, et al. Genome-resolved metagenomics analysis provides insights into the ecological role of Thaumarchaeota in the Amazon River and its plume. BMC Microbiol. BioMed Central; 2020;20:1\u0026ndash;11. \u003c/li\u003e\n\u003cli\u003eGraham EB, Crump AR, Kennedy DW, Arntzen E, Fansler S, Purvine SO, et al. Multi\u0026rsquo;omics comparison reveals metabolome biochemistry, not microbiome composition or gene expression, corresponds to elevated biogeochemical function in the hyporheic zone. Sci Total Environ. Elsevier; 2018;642:742\u0026ndash;53. \u003c/li\u003e\n\u003cli\u003ePeng Y, Leung HCM, Yiu SM, Chin FYL. IDBA-UD: a de novo assembler for single-cell and metagenomic sequencing data with highly uneven depth. Bioinformatics. Narnia; 2012;28:1420\u0026ndash;8. \u003c/li\u003e\n\u003cli\u003eNurk S, Meleshko D, Korobeynikov A, Pevzner PA. metaSPAdes: a new versatile metagenomic assembler. Genome Res. Cold Spring Harbor Lab; 2017;27:824\u0026ndash;34. \u003c/li\u003e\n\u003cli\u003eKang DD, Li F, Kirton E, Thomas A, Egan R, An H, et al. MetaBAT 2: An adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ. PeerJ Inc.; 2019;2019:e7359. \u003c/li\u003e\n\u003cli\u003eWu M, Scott AJ. Phylogenomic analysis of bacterial and archaeal sequences with AMPHORA2. Bioinformatics. Oxford University Press; 2012;28:1033\u0026ndash;4. \u003c/li\u003e\n\u003cli\u003eBowers RM, Kyrpides NC, Stepanauskas R, Harmon-Smith M, Doud D, Reddy TBK, et al. Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea. Nat. Biotechnol. Nature Publishing Group; 2017. p. 725\u0026ndash;31. \u003c/li\u003e\n\u003cli\u003eOlm MR, Brown CT, Brooks B, Banfield JF. dRep: a tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through de-replication. ISME J. Nature Publishing Group; 2017;11:2864\u0026ndash;8. \u003c/li\u003e\n\u003cli\u003eShaffer M, Borton MA, McGivern BB, Zayed AA, La Rosa SL, Solden LM, et al. DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Res. :gkaa621. \u003c/li\u003e\n\u003cli\u003eChaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics. 2018;36:1925\u0026ndash;1927. \u003c/li\u003e\n\u003cli\u003eEdgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. Oxford University Press; 2004;32:1792\u0026ndash;7. \u003c/li\u003e\n\u003cli\u003eKatoh K, Misawa K, Kuma K, Miyata T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. Oxford University Press; 2002;30:3059\u0026ndash;66. \u003c/li\u003e\n\u003cli\u003ePrice MN, Dehal PS, Arkin AP. FastTree: computing large minimum evolution trees with profiles instead of a distance matrix. Mol Biol Evol. Oxford University Press; 2009;26:1641\u0026ndash;50. \u003c/li\u003e\n\u003cli\u003eCastelle CJ, Hug LA, Wrighton KC, Thomas BC, Williams KH, Wu D, et al. Extraordinary phylogenetic diversity and metabolic versatility in aquifer sediment. Nat Commun. Nature Publishing Group; 2013;4:2120. \u003c/li\u003e\n\u003cli\u003eYu NY, Wagner JR, Laird MR, Melli G, Rey S, Lo R, et al. PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. Bioinformatics. 2010;26:1608\u0026ndash;15. \u003c/li\u003e\n\u003cli\u003eArmenteros JJA, Tsirigos KD, S\u0026oslash;nderby CK, Petersen TN, Winther O, Brunak S, et al. SignalP 5.0 improves signal peptide predictions using deep neural networks. Nat Biotechnol. Nature Publishing Group; 2019;37:420\u0026ndash;3. \u003c/li\u003e\n\u003cli\u003eBendtsen JD, Kiemer L, Fausb\u0026oslash;ll A, Brunak S. Non-classical protein secretion in bacteria. BMC Microbiol. BioMed Central; 2005;5:1\u0026ndash;13. \u003c/li\u003e\n\u003cli\u003eRoux S, Enault F, Hurwitz BL, Sullivan MB. VirSorter: mining viral signal from microbial genomic data. PeerJ. PeerJ Inc.; 2015;3:e985. \u003c/li\u003e\n\u003cli\u003eRoux S, Adriaenssens EM, Dutilh BE, Koonin E V., Kropinski AM, Krupovic M, et al. Minimum information about an uncultivated virus genome (MIUVIG). Nat Biotechnol. Nature Publishing Group; 2019;37:29\u0026ndash;37. \u003c/li\u003e\n\u003cli\u003eMerchant N, Lyons E, Goff S, Vaughn M, Ware D, Micklos D, et al. The iPlant collaborative: cyberinfrastructure for enabling data to discovery for the life sciences. PLoS Biol. Public Library of Science San Francisco, CA USA; 2016;14:e1002342. \u003c/li\u003e\n\u003cli\u003eJang H Bin, Bolduc B, Zablocki O, Kuhn JH, Roux S, Adriaenssens EM, et al. Taxonomic assignment of uncultivated prokaryotic virus genomes is enabled by gene-sharing networks. Nat Biotechnol. Nature Publishing Group; 2019;37:632\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eKelley LA, Mezulis S, Yates CM, Wass MN, Sternberg MJE. The Phyre2 web portal for protein modeling, prediction and analysis. Nat Protoc. Nature Publishing Group; 2015;10:845\u0026ndash;58. \u003c/li\u003e\n\u003cli\u003eAhlgren NA, Ren J, Lu YY, Fuhrman JA, Sun F. Alignment-free oligonucleotide frequency dissimilarity measure improves prediction of hosts from metagenomically-derived viral sequences. Nucleic Acids Res. Oxford University Press; 2017;45:39\u0026ndash;53. \u003c/li\u003e\n\u003cli\u003eBorton MA, Hoyt DW, Roux S, Daly RA, Welch SA, Nicora CD, et al. Coupled laboratory and field investigations resolve microbial interactions that underpin persistence in hydraulically fractured shales. Proc Natl Acad Sci. National Academy of Sciences; 2018;115:E6585\u0026ndash;E6594. \u003c/li\u003e\n\u003cli\u003eDaly RA, Borton MA, Wilkins MJ, Hoyt DW, Kountz DJ, Wolfe RA, et al. Microbial metabolisms in a 2.5-km-deep ecosystem created by hydraulic fracturing in shales. Nat Microbiol. Nature Publishing Group; 2016;1:16146. \u003c/li\u003e\n\u003cli\u003eLangmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. Nature Publishing Group; 2012;9:357. \u003c/li\u003e\n\u003cli\u003eL\u0026ecirc; Cao K-A, Rossouw D, Robert-Grani\u0026eacute; C, Besse P. A sparse PLS for variable selection when integrating omics data. Stat Appl Genet Mol Biol. De Gruyter; 2008;7. \u003c/li\u003e\n\u003cli\u003eNicora CD, Burnum-Johnson KE, Nakayasu ES, Casey CP, White III RA, Chowdhury TR, et al. The MPLEx protocol for multi-omic analyses of soil samples. J Vis Exp JoVE. MyJoVE Corporation; 2018; \u003c/li\u003e\n\u003cli\u003eElias JE, Gygi SP. Target-decoy search strategy for mass spectrometry-based proteomics. Proteome Bioinforma. Springer; 2010. p. 55\u0026ndash;71. \u003c/li\u003e\n\u003cli\u003eKim S, Pevzner PA. MS-GF+ makes progress towards a universal database search tool for proteomics. Nat Commun. Nature Publishing Group; 2014;5:5277. \u003c/li\u003e\n\u003cli\u003eMcGivern BB, Tfaily MM, Borton MA, Kosina SM, Daly RA, Nicora CD, et al. Decrypting bacterial polyphenol metabolism in an anoxic wetland soil. Nat Commun. Nature Publishing Group; 2021;12:1\u0026ndash;16. \u003c/li\u003e\n\u003cli\u003eWeljie AM, Newton J, Mercier P, Carlson E, Slupsky CM. Targeted profiling: quantitative analysis of 1H NMR metabolomics data. Anal Chem. ACS Publications; 2006;78:4430\u0026ndash;42. \u003c/li\u003e\n\u003cli\u003eJoesten WC, Kennedy MA. RANCM: a new ranking scheme for assigning confidence levels to metabolite assignments in NMR-based metabolomics studies. Metabolomics. Springer; 2019;15:5. \u003c/li\u003e\n\u003cli\u003eMoon K, Jeon JH, Kang I, Park KS, Lee K, Cha C-J, et al. Freshwater viral metagenome reveals novel and functional phage-borne antibiotic resistance genes. Microbiome. Springer; 2020;8:1\u0026ndash;15. \u003c/li\u003e\n\u003cli\u003eWolf YI, Silas S, Wang Y, Wu S, Bocek M, Kazlauskas D, et al. Doubling of the known set of RNA viruses by metagenomic analysis of an aquatic virome. Nat Microbiol. Nature Publishing Group; 2020;5:1262\u0026ndash;70. \u003c/li\u003e\n\u003cli\u003ePeduzzi P. Virus ecology of fluvial systems: a blank spot on the map? Biol Rev. Wiley Online Library; 2016;91:937\u0026ndash;49. \u003c/li\u003e\n\u003cli\u003eP\u0026uuml;ttker S, Kohrs F, Benndorf D, Heyer R, Rapp E, Reichl U. Metaproteomics of activated sludge from a wastewater treatment plant--A pilot study. Proteomics. Wiley Online Library; 2015;15:3596\u0026ndash;601. \u003c/li\u003e\n\u003cli\u003eRudney JD, Xie H, Rhodus NL, Ondrey FG, Griffin TJ. A metaproteomic analysis of the human salivary microbiota by three-dimensional peptide fractionation and tandem mass spectrometry. Mol Oral Microbiol. Wiley Online Library; 2010;25:38\u0026ndash;49. \u003c/li\u003e\n\u003cli\u003eSolden LM, Naas AE, Roux S, Daly RA, Collins WB, Nicora CD, et al. Interspecies cross-feeding orchestrates carbon degradation in the rumen ecosystem. Nat Microbiol. Nature Publishing Group; 2018;3:1274. \u003c/li\u003e\n\u003cli\u003eGraham EB, Crump AR, Resch CT, Fansler S, Arntzen E, Kennedy DW, et al. Deterministic influences exceed dispersal effects on hydrologically-connected microbiomes. Environ Microbiol. Wiley Online Library; 2017;19:1552\u0026ndash;67. \u003c/li\u003e\n\u003cli\u003eYang F, Bogdanov B, Strittmatter EF, Vilkov AN, Gritsenko M, Shi L, et al. Characterization of Purified c-Type Heme-Containing Peptides and Identification of c-Type Heme-Attachment Sites in Shewanella o neidenis Cytochromes Using Mass Spectrometry. J Proteome Res. ACS Publications; 2005;4:846\u0026ndash;54. \u003c/li\u003e\n\u003cli\u003eCordero PRF, Bayly K, Leung PM, Huang C, Islam ZF, Schittenhelm RB, et al. Atmospheric carbon monoxide oxidation is a widespread mechanism supporting microbial survival. ISME J. Nature Publishing Group; 2019;13:2868\u0026ndash;81. \u003c/li\u003e\n\u003cli\u003eXia X, Zhang S, Li S, Zhang L, Wang G, Zhang L, et al. The cycle of nitrogen in river systems: sources, transformation, and flux. Environ Sci Process \\\u0026amp; Impacts. Royal Society of Chemistry; 2018;20:863\u0026ndash;91. \u003c/li\u003e\n\u003cli\u003eStrauss EA, Lamberti GA. Regulation of nitrification in aquatic sediments by organic carbon. Limnol Oceanogr. Wiley Online Library; 2000;45:1854\u0026ndash;9. \u003c/li\u003e\n\u003cli\u003eBrust GE. Management strategies for organic vegetable fertility. Saf Pract Org food. Elsevier; 2019. p. 193\u0026ndash;212. \u003c/li\u003e\n\u003cli\u003eCavaliere M, Feng S, Soyer OS, Jim\u0026eacute;nez JI. Cooperation in microbial communities and their biotechnological applications. Environ Microbiol. Wiley Online Library; 2017;19:2949\u0026ndash;63. \u003c/li\u003e\n\u003cli\u003eGibbons SM, Jones E, Bearquiver A, Blackwolf F, Roundstone W, Scott N, et al. Human and environmental impacts on river sediment microbial communities. PLoS One. Public Library of Science; 2014;9:e97435. \u003c/li\u003e\n\u003cli\u003eWrage N, Velthof GL, Van Beusichem ML, Oenema O. Role of nitrifier denitrification in the production of nitrous oxide. Soil Biol Biochem. Elsevier; 2001;33:1723\u0026ndash;32. \u003c/li\u003e\n\u003cli\u003eStrauss EA, Richardson WB, Bartsch LA, Cavanaugh JC, Bruesewitz DA, Imker H, et al. Nitrification in the Upper Mississippi River: patterns, controls, and contribution to the NO3- budget. J North Am Benthol Soc. 2004;23:1\u0026ndash;14. \u003c/li\u003e\n\u003cli\u003eStegen JC, Fredrickson JK, Wilkins MJ, Konopka AE, Nelson WC, Arntzen E V, et al. Groundwater--surface water mixing shifts ecological assembly processes and stimulates organic carbon turnover. Nat Commun. Nature Publishing Group; 2016;7:1\u0026ndash;12. \u003c/li\u003e\n\u003cli\u003eBartossek R, Nicol GW, Lanzen A, Klenk H-P, Schleper C. Homologues of nitrite reductases in ammonia-oxidizing archaea: diversity and genomic context. Environ Microbiol. Wiley Online Library; 2010;12:1075\u0026ndash;88. \u003c/li\u003e\n\u003cli\u003eJones CM, Spor A, Brennan FP, Breuil M-C, Bru D, Lemanceau P, et al. Recently identified microbial guild mediates soil N 2 O sink capacity. Nat Clim Chang. Nature Publishing Group; 2014;4:801\u0026ndash;5. \u003c/li\u003e\n\u003cli\u003eKuypers MMM, Marchant HK, Kartal B. The microbial nitrogen-cycling network. Nat Rev Microbiol. Nature Publishing Group; 2018;16:263\u0026ndash;76. \u003c/li\u003e\n\u003cli\u003eNelson WC, Graham EB, Crump AR, Fansler SJ, Arntzen E V, Kennedy DW, et al. Distinct temporal diversity profiles for nitrogen cycling genes in a hyporheic microbiome. PLoS One. Public Library of Science San Francisco, CA USA; 2020;15:e0228165. \u003c/li\u003e\n\u003cli\u003eMurphy CL, Sheremet A, Dunfield PF, Spear JR, Stepanauskas R, Woyke T, et al. Genomic Analysis of the Yet-Uncultured Binatota Reveals Broad Methylotrophic, Alkane-Degradation, and Pigment Production Capacities. MBio. Am Soc Microbiol; 2021;12. \u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"microbiome, viruses, metagenomics, river sediment, denitrification, nitrification, carboxydotrophs, Thaumarchaeota, peptidases, Binatia","lastPublishedDoi":"10.21203/rs.3.rs-746574/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-746574/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cem\u003eBackground:\u003c/em\u003e\u003c/p\u003e\u003cp\u003eRivers serve as a nexus for nutrient transfer between terrestrial and marine ecosystems and as such, have a significant impact on global carbon and nitrogen cycles. In river ecosystems, the sediments found within the hyporheic zone are microbial hotspots that can account for a significant portion of ecosystem respiration and have profound impacts on system biogeochemistry. Despite this, studies using genome-resolved analyses linking microbial and viral communities to nitrogen and carbon biogeochemistry are limited.\u003c/p\u003e\u003cp\u003e\u003cem\u003eResults:\u003c/em\u003e\u003c/p\u003e\u003cp\u003eHere, we characterized the microbial and viral communities of Columbia River hyporheic zone sediments to reveal the metabolisms that actively cycle carbon and nitrogen. Using genome-resolved metagenomics, we created the Hyporheic Uncultured Microbial and Viral (HUM-V) database, containing a dereplicated database of 55 microbial Metagenome-Assembled Genomes (MAGs), representing 12 distinct phyla. We also sampled 111 viral Metagenome Assembled Genomes (vMAGs) from 26 distinct and novel genera. The HUM-V recruited metaproteomes from these same samples, providing the first inventory of microbial gene expression in hyporheic zone sediments. Combining this data with metabolite data, we generated a conceptual model where heterotrophic and autotrophic metabolisms co-occur to drive an integrated carbon and nitrogen cycle, revealing microbial sources and sinks for carbon dioxide and ammonium in these sediments. We uncovered the metabolic handoffs underpinning these processes including mutualistic nitrification by Thermoproteota (formerly Thaumarchaeota) and Nitrospirota, as well as identified possible cooperative and cheating behavior impacting nitrogen mineralization. Finally, by linking vMAGs to microbial genome hosts, we reveal possible viral controls on microbial nitrification and organic carbon degradation.\u003c/p\u003e\u003cp\u003e\u003cem\u003eConclusions:\u003c/em\u003e\u003c/p\u003e\u003cp\u003eOur multi-omics analyses provide new mechanistic insight into coupled carbon-nitrogen cycling in the hyporheic zone. This is a key step in developing predictive hydrobiogeochemical models that account for microbial cross-feeding and viral influences over potential and expressed microbial metabolisms. Furthermore, the publicly available HUM-V genome resource can be queried and expanded by researchers working in other ecosystems to assess the transferability of our results to other parts of the globe.\u003c/p\u003e","manuscriptTitle":"Microbial Genome-Resolved Metaproteomic Analyses Frame Intertwined Carbon and Nitrogen Cycles in River Hyporheic Sediments","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-08-02 21:30:03","doi":"10.21203/rs.3.rs-746574/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"59f2c9e6-2082-40ee-bc80-79b21e7c5748","owner":[],"postedDate":"August 2nd, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":6153420,"name":"Applied \u0026 Industrial Microbiology"},{"id":6153421,"name":"General Microbiology"}],"tags":[],"updatedAt":"2021-09-30T11:57:27+00:00","versionOfRecord":[],"versionCreatedAt":"2021-08-02 21:30:03","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-746574","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-746574","identity":"rs-746574","version":["v1"]},"buildId":"rHA-KDH7Qsr4HCuvH75dn","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.