A systematically biosynthetic investigation of lactic acid bacteria reveals diverse antagonistic bacteriocins that potentially shape the human microbiome

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study systematically investigated lactic acid bacteria genomes and human metagenomes, revealing diverse bacteriocins, particularly class II bacteriocins in the vagina, that potentially contribute to microbiome homeostasis.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-17 · read from full text

This study systematically surveyed lactic acid bacteria using antiSMASH on 31,977 LAB genomes (SAGs and MAGs) to map biosynthetic gene clusters for secondary metabolites, then used analyses of 748 human-associated metagenomes (including metatranscriptomics) to assess their presence and activity in human niches. Across 130,051 BGCs grouped into 2,849 gene cluster families, most clusters were species- or strain-specific and many were uncharacterized, with machine-learning–predicted antagonistic bacteriocins proposed to shape microbial community structure. The authors report that class II bacteriocins are especially abundant and enriched in vaginal microbiomes, and that omics-guided and experimental validation support a role for these antagonistic peptides in regulating vaginal microbial communities and contributing to microbiome homeostasis. A key caveat is that conclusions about function are largely based on genomic/metagenomic prediction of bacteriocins and their antagonistic potential rather than complete chemical characterization of all identified BGCs. This paper is centrally about endometriosis and/or adenomyosis only tangentially; it does not specifically discuss endometriosis or adenomyosis, but it is included because it examines vaginal microbiome homeostasis and antimicrobial bacteriocins relevant to pelvic reproductive tract microbial environments.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background: Lactic acid bacteria (LAB) produce various bioactive secondary metabolites (SMs), which endow LAB with a protective role for the host. However, the biosynthetic potentials of LAB-derived SMs remain elusive, particularly in their diversity, abundance, and distribution in the human microbiome. Thus, it is still unknown to what extent LAB-derived SMs are involved in microbiome homeostasis. Results: : Here, we systematically investigate the biosynthetic potential of LAB from 31,977 LAB genomes, identifying 130,051 BGCs of 2,849 gene cluster families (GCFs). Most of these GCFs are species-specific or even strain-specific and uncharacterized yet. Analyzing 748 human-associated metagenomes, we gain an insight into the profile of LAB BGCs, which are highly diverse and niche-specific in the human microbiome. We discover that most LAB BGCs may encode bacteriocins with pervasive antagonistic activities predicted by machine learning models, potentially playing protective roles in the human microbiome. Class II bacteriocins, one of the most abundant and diverse LAB SMs, are particularly enriched and predominant in the vaginal microbiomes. Together with experimental validation, our metagenomic and metatranscriptomic analysis show that antagonistic class II bacteriocins potentially regulate microbial communities in the vagina, thereby contributing to microbiome homeostasis. Conclusions: : Our study systematically investigates LAB biosynthetic potential and their profile in the human microbiome, linking them to the antagonistic contributions to microbiome homeostasis via omics analysis. These discoveries of the diverse and prevalent antagonistic SMs are expected to stimulate the mechanism study of LAB’s protective roles for the microbiome and host, highlighting the potential of LAB and their bacteriocins as therapeutic alternatives.
Full text 198,018 characters · extracted from preprint-html · click to expand
A systematically biosynthetic investigation of lactic acid bacteria reveals diverse antagonistic bacteriocins that potentially shape the human microbiome | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A systematically biosynthetic investigation of lactic acid bacteria reveals diverse antagonistic bacteriocins that potentially shape the human microbiome Dengwei Zhang, Jian Zhang, Shanthini Kalimuthu, Jing Liu, Zhiman Song, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1868011/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 27 Apr, 2023 Read the published version in Microbiome → Version 1 posted 9 You are reading this latest preprint version Abstract Background: Lactic acid bacteria (LAB) produce various bioactive secondary metabolites (SMs), which endow LAB with a protective role for the host. However, the biosynthetic potentials of LAB-derived SMs remain elusive, particularly in their diversity, abundance, and distribution in the human microbiome. Thus, it is still unknown to what extent LAB-derived SMs are involved in microbiome homeostasis. Results: Here, we systematically investigate the biosynthetic potential of LAB from 31,977 LAB genomes, identifying 130,051 BGCs of 2,849 gene cluster families (GCFs). Most of these GCFs are species-specific or even strain-specific and uncharacterized yet. Analyzing 748 human-associated metagenomes, we gain an insight into the profile of LAB BGCs, which are highly diverse and niche-specific in the human microbiome. We discover that most LAB BGCs may encode bacteriocins with pervasive antagonistic activities predicted by machine learning models, potentially playing protective roles in the human microbiome. Class II bacteriocins, one of the most abundant and diverse LAB SMs, are particularly enriched and predominant in the vaginal microbiomes. Together with experimental validation, our metagenomic and metatranscriptomic analysis show that antagonistic class II bacteriocins potentially regulate microbial communities in the vagina, thereby contributing to microbiome homeostasis. Conclusions: Our study systematically investigates LAB biosynthetic potential and their profile in the human microbiome, linking them to the antagonistic contributions to microbiome homeostasis via omics analysis. These discoveries of the diverse and prevalent antagonistic SMs are expected to stimulate the mechanism study of LAB’s protective roles for the microbiome and host, highlighting the potential of LAB and their bacteriocins as therapeutic alternatives. Lactic acid bacteria Biosynthetic gene clusters Secondary metabolites Bacteriocins Human microbiome Vaginal microbiome Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Lactic acid bacteria (LAB) are Gram-positive, microaerophilic bacteria, which have drawn extensive attention due to their fundamental roles in different biological processes [1]. These bacteria feature lactic acid production in carbohydrate metabolism, which is important in food fermentation [2]. Due to general safety features, they have also been engineered to produce food ingredients and pharmaceutical agents and deliver therapeutic molecules [1–4]. More importantly, a growing body of evidence reveals the health-promoting effects of consuming certain LAB strains, rendering them promising candidates for probiotics [5]. Their probiotic actions may arise from multifaceted mechanisms, such as gut microflora regulation, bioactive metabolite production, and immune system modulation, which endow LAB with a protective role for the host [6]. Despite numerous studies focusing on characterizing LAB as probiotics for microflora regulation, how they impact microbiome homeostasis and host physiology is still not fully understood. Fundamentally, metabolic crosstalk within the kingdoms or with the host is the basis by which microbes, including LAB, engage with microbiome homeostasis. A myriad of metabolites produced by microbes, especially secondary metabolites (SMs, e.g. antibiotics and pigments), is among the mediums to achieve a high degree of crosstalk. Secondary metabolite-mediated interactions such as mutualism and antagonism are essential in maintaining microbiome homeostasis [7,8]. Many LAB members, including Lactobacillus, Streptococcus , and Lactococcus , produce bioactive SMs, ranging from bacteriocins nisin and lactocillin to tetramic acid reutericyclin [9–11]. Bacteriocins are ribosomally synthesized antimicrobial peptides with antibacterial potential and are generally divided into three classes (class I, post‐translationally modified peptides (e.g., nisin and lactocillin); class II, small unmodified peptides (e.g., amylovorin L and crispacin A); class III, large, heat-labile peptides) [12,13]. Moreover, bacteriocins, the most extensively studied SMs of LAB, have been recently revealed to play potential roles in shaping the microbiota by modulating microbial composition and inhibiting pathogens [14]. However, these studies of LAB SMs have largely focused on structures, biosynthesis, mechanisms of action, or therapeutic potential in preventing infection in a case-by-case manner [15,16]. The landscape of LAB SMs, particularly their diversity, prevalence, and potential roles in the human microbiome, remains elusive. Therefore, it is still unknown to what extent LAB-derived secondary metabolites are actively involved in microbiome homeostasis. With the development of bioinformatics techniques, the recent explosion of sequenced bacterial genomes and metagenomes provides fresh opportunities for large-scale biosynthetic analysis at both single species and community levels. Here we harnessed recent advances in biosynthetic and metagenomic analysis to investigate the untapped biosynthetic potential of LAB SMs systematically. Leveraging this comprehensive analysis of 31,977 LAB genomes and 748 human microbiome metagenomes, we gained previously undescribed insights into the biosynthetic capacity of LAB SMs and their diversity, abundance, and distribution in the human microbiome. We found that most BGCs may encode antagonistic SMs uncharacterized yet, exemplified by a new class II bacteriocin termed crispacin 407, potentially playing protective roles in the human microbiome. To our best knowledge, this is the first largest survey of LAB biosynthetic potential and their profile in the human microbiome, linking them to the antagonistic contributions to microbiome homeostasis via omics analysis. Our omics-guided findings of the diverse and prevalent antagonistic bacteriocins in the human microbiome, particularly the vaginal microbiome, provide insight into antagonistic interactions linked to microbiome homeostasis and highlight the probiotic potential of LAB with antagonistic SMs. Results Genomic analysis reveals the landscape of SM biosynthetic potential of LAB Given that environment or foods are possible LAB sources for the gut microbiome, to profile LAB SMs in the human microbiome, we first collected LAB genome data from different sources to comprehensively investigate the biosynthetic potential of LAB SMs. Publicly available bacterial single amplified genomes (SAGs) and metagenome-assembled genomes (MAGs) of LAB were gathered from three databases (RefSeq [17], PATRIC [18], and IMG/M [19]) and two previous studies [20,21], resulting in 40,879 SAGs and 4,575 MAGs in total (Supplementary Table 1). Genomes were then de-duplicated, and their taxonomical classifications were verified and unified using GTDB taxonomy. As a result, 31,977 LAB genomes (27,549 SAGs and 4,428 MAGs, Supplementary Table 2), spanning six families containing 56 genera, were retained for global biosynthetic analysis of LAB SMs (Supplementary Fig. 1a). Using a rule-based BGC detection tool, antiSMASH 6.0 [22], we identified 130,051 BGCs from 30,718 genomes (Fig. 1a, Supplementary Fig. 1b, Supplementary Table 3), including 1,333 nonribosomal peptide synthetase clusters (NRPS, 1.0%), 25,278 polyketides synthase clusters (PKS, 19.7%), 98,810 ribosomally encoded and post-translationally modified peptides (RiPPs, 76.0%), 1,629 terpene (1.3%) and 2,984 BGCs (2.3%) encoding other types of metabolites (Supplementary Fig. 2). The BGCs per genome ranged from 0 to 14, with an average of 4.07. Among the most abundant RiPPs, RiPP-like (formerly annotated as bacteriocin by antiSMASH) topped its list with 72,471 (55.7% of total BGCs). Using BiG-SLiCE [23] to extract bacteriocin biosynthesis-related domains from 72,471 RiPP-like BGCs (Supplementary Fig. 3 and Supplementary Table 4), we identified 60,497 class II bacteriocins (RiPP-like BGCs that contain class II bacteriocins-related domains identified by BiG-SLiCE) (46.5%), which are the most abundant LAB SMs. Also worth noting was that 99.8% (25,224/25,278) of PKS were type III PKS (T3PKS) (Fig. 1a). To gain insight into the phylogenetic distribution of BGC in LAB genera, we examined 26,983 BGC-containing SAGs, excluding MAGs due to their incompleteness. The biosynthetic capacity varied considerably at the families, genera, or species level (Fig. 1b, Supplementary Fig. 4). The RiPPs and T3PKS BGCs dominated all LAB genera except for Tetragenococcus , while terpene BGCs and NRPS BGCs were sporadically distributed in those genera (Fig. 1, Supplementary Figs. 5, 6). Among 55 genera, we found a median of ≥1 RiPPs per genome in 28 genera and one T3PKS in 33 genera. While a high proportion of RiPPs has been reported in Firmicutes [24], T3PKS dominating in certain LAB genera might encode specialized metabolites that have basic biological functions. Of note, despite being small in genome size, Streptococcaceae generally harbored more abundant BGCs than other families (Supplementary Fig. 4a), with a median of five BGC per genome, exemplified by genera Lactococcus and Streptococcus . Contrastingly, 33 genera only harbored a median of <=1 BGC per genome, indicating the limited biosynthetic capacity of the LAB majority (Fig. 1b). By comparing 53 LAB genera and 3,805 non-LAB genera (164,417 genomes, Supplementary Table 5), we found comparatively limited biosynthetic capacity in LAB and a significantly strong correlation (Spearman rho = 0.712, P < 0.001) between bacterial biosynthetic potential and their genome size (Supplementary Fig. 7). The small genomes and reduced biosynthetic capacities in LAB might correlate with their adaptation to nutritionally-rich niches. Most LAB BGCs are species-specific and uncharacterized yet Although BGCs highly vary in gene content, grouping them into families (GCFs) or clans (GCCs) based on architectural relationships of biosynthetic elements is an effective way to uncover the similarity of their encoding products in terms of the chemical features and biological functions [25]. To gain insight into the novelty and diversity of 130,051 LAB BGCs, we extracted BGC features (biosynthetic domains) using BiG-SLiCE [23] and grouped them based on an all-to-all cosine distance among BGCs [26]. The 129,878 BGCs with features were classified into 2,849 GCFs and 112 GCCs, with a distance threshold of 0.2 and 0.8, respectively (Fig. 2, Supplementary Fig. 8). We further compared the 112 GCCs to the reference known BGCs described in the 'Minimum Information about a Biosynthetic Gene' (MIBiG) repository [27]. Notably, only three clans, linear azol(in)e-containing peptides (LAP, GCC_29), RiPP-like (GCC_84), and class II lanthipeptides (GCC_110), were closely similar to known BGCs (average cosine distances < 0.2), leaving the vast majority unknown. This highlights the huge knowledge gap of LAB SMs and demonstrates the potential for discovering novel chemistry from the LAB. Of note, the majority of NRPS (73.2%), terpene (99.3%) and T3PKS (99.3%) were clustered into one respective clan. In contrast, RiPP BGCs, contributing 83 GCCs with 1,818 GCFs (RiPP proportion > 80% in GCCs/GCFs), were highly diverse due to the diversity of their post-translational modification (PTM) enzyme genes and adjacent genes (Fig. 2a). Among them, 621 GCFs of 23 RiPP-like GCCs encoded class II bacteriocins, representing one of the most diverse LAB SMs. The other 60 RiPP GCCs mainly encoded class I bacteriocins, including lanthipeptide, LAP, lassopeptide, and rSAM-modified RiPPs. In prokaryotic genome evolution, the conserved genes cross genera are more likely to contribute to essential ecological processes, whereas species- or even strain-specific genes often arise from natural selection, thus enhancing niche adaptation or host fitness [28]. In this context, we next examined the distribution and diversity of genus- or species-specific BGCs. While GCC clustering shows the distribution and novelty of LAB SMs, a fine resolution of GCF clustering can offer an insight into the diversity of BGCs that are predicted to encode similar natural products [29]. We found that the majority of GCFs were genus-specific (92.6%, 2,637/2,849) and species-specific (75.8%, 2,159/2,849). Remarkably, 1,165 GCFs (40.9%) contained only one BGC harbored by a specific strain (Fig. 2b). In contrast, only 7% (212) were cross-genus GCFs, including 142 RiPPs (present in 2-20 genera), 17 NRPS (in 2-3 genera), 21 T3PKS (in 2-9 genera), and 9 terpenes (in 2-7 genera) (Supplementary Figs. 9, 10). Among these 142 cross-genus RiPP GCFs, 62 are class II bacteriocins. Owing to this high GCF diversity between genera, we did not observe a phylogenetic relationship in GCF presence/absence (Supplementary Fig. 8). These taxa-specific BGCs usually encode specialized SMs and provide a competitive edge to the producer for niche adaption. Considering the wide presence of LAB in different niches [30], a high proportion of species- and strain-specific BGCs might result from niche selection. LAB BGCs are diverse and niche-specific in the human microbiome Numerous studies have revealed the variable prevalence of LAB species in the human microbiome [21], raising the question of to what extent their SMs vary in different body sites for niche adaption. Thus, we next explored the profile of LAB BGCs in the healthy human microbiome by re-visiting 748 metagenomes of six body sites from the Human Microbiome Project (HMP) [31]. These sites included aerobic (anterior nares, representing skin microbiome), microaerobic (supragingival plaque, buccal mucosa, tongue dorsum, and posterior fornix, representing oral and vaginal microbiome), and anaerobic (stool, representing gut microbiome) environments (Supplementary Table 6). In line with a previous larger-scale study [21], LAB exhibited variable abundance and prevalence in different body sites (Fig. 3a). Of note, the genus Lactobacillus dominated in the vagina with a median abundance of 99.0% [interquartile range (IQR), 91.0%-99.8%], whereas Streptococcus was moderately abundant but highly prevalent in six body sites. To profile the diversity, abundance, and distribution of LAB BGCs in the human microbiome, we de-duplicated 130,051 BGCs to 24,222 representative BGCs and mapped metagenomic reads to 24,222 nonredundant BGCs. The number of LAB BGCs detected in six body sites varied considerably, with the highest in the oral cavity and the lowest in the skin (Supplementary Fig. 11a), probably due to the variable abundance of LAB. From 748 metagenomes, we detected 5,687 BGCs of 610 GCFs, including 71 T3PKS and 312 RiPPs with 92 class II bacteriocins (Supplementary Fig. 11b). The GCF accumulation curve indicated that more GCFs would be detected in those body sites as more samples were included, revealing the huge diversity of LAB SMs in the human microbiome (Fig. 3b). The three oral sites were the richest in GCFs (averaging 29, 38, and 54). Compared to the oral cavity, the vagina harbors a lower diversity of GCFs (averaging 12) but a significantly higher abundance of LAB BGCs (Figs. 3c, d). Particularly, the vaginal microbiome harbored a high abundance of class II bacteriocins, lassopeptide, lanthipeptide, and LAP. Of note, influenced by sequencing depth, the diversity and abundance of LAB BGCs in the human microbiome may be underestimated. Of 610 detected GCFs, ~52% were niche-specific in one of six sites, which accounted for 18% - 38% of GCFs in a particular site (Fig. 3e). We also observed that those niche-specific GCFs were generally species-specific (Chi-squared test, P < 0.001), but not genus-specific ( P = 0.026) nor strain-specific ( P = 0.013) (Fig. 3f). This result indicated that niche-specific GCFs were derived from different species residing in distinct niches, which may provide a competitive advantage to the niche adaptation of their hosts. Our genomic and metagenomic analysis of biosynthetic potential revealed that the LAB SMs are diverse and variably prevalent in the human microbiome. Machine learning models reveal that most BGCs may encode antagonistic SMs Given the abundance and prevalence of LAB BGCs in the human microbiome, we next want to study the potential bioactivities of BGC-encoding SMs. The bioactivity of SMs encoded by BGCs was recently predicted using machine learning strategies based on chemical fingerprints of predicted compound structure, protein family (PFAM) domains, and other genetic features [32–34]. Here, we adapted four common machine learning classifiers (logistic regression, elastic net regression, random forest, and support vector machines) to predict the bioactivities of LAB-derived SMs. For the training data (950 known BGCs, Supplementary Table 7, Supplementary Fig. 12), ten-fold cross-validation revealed that the random forest classifier outperformed others with an average area under the receiver operating characteristic curve (AUROC) being 0.76, 0.80, and 0.82, for antibacterial, antifungal, and antitumor or cytotoxic, respectively (Fig. 4a, Supplementary Fig. 13). The performance of the random forest classifier was comparable with previously reported methods [32,33]. Using the random forest classifier, 129,878 LAB BGCs with features were predicted to encode different bioactive SMs, comprising antibacterial (n=123,587, 95.2%), cytotoxic (n=2,268, 1.8%), antibacterial-antifungal (n=79, 0.1 %), antibacterial-cytotoxic (n=2,548, 2.0%), unknown (1,395, 1.1%). Most BGCs, regardless of BGC classes, were predicted to be antibacterial (Fig. 4b). Of note, 97.8% of RiPPs (96,466/98,637) were predicted to exhibit antibacterial activity, implying that bacteriocins were plenteous in LAB. With predicted antibacterial activity, most RiPPs with known post-modifications were classified as class I bacteriocins, while most RiPPs-like (83.4%) were class II bacteriocins. Almost 100% of class II bacteriocins (60,494/60,497) were captured as antibacterial SMs, contributing to 48.1% of LAB-derived antibacterials. Those antibacterial BGCs dominated almost all LAB genera (Supplementary Fig. 14), possibly conferring a competitive edge in the microbial community. Compared to BGCs (n=1,121,156) identified from non-LAB genomes, LAB-derived BGCs potentially encoded a significantly higher proportion of antibacterial SMs, indicating a higher antagonistic potential of LAB SMs (Wilcoxon rank-sum test, P < 0.001) (Fig. 4c, Supplementary Fig. 15a, b). Low percentage of LAB-derived BGCs encoding putative cytotoxic or antifungal SMs were found in specific species (Supplementary Fig. 15c, d). For example, a certain family of LAP (GCF_199) possibly conferring cytotoxic activity were distributed in ten Streptococcus species, especially Streptococcus pyogenes , in which common pathogenicity feature endowed by conserved LAP had been reported [35]. In the six body sites, almost all BGCs potentially encoded antimicrobials (Fig. 4d), possibly mediating bacterial antagonism for maintaining microbiome homeostasis. We speculated that LAB harboring antimicrobial SMs, especially class II bacteriocin, could protect the host from pathogen invasion and maintain the microbiome homeostasis via antagonistic interaction [7]. Underexplored class II bacteriocins are widely distributed in the human microbiome The findings of the antagonistic potential of class II bacteriocins and their variable prevalence and predominance in the human microbiome raise the question of what extent the class II bacteriocins may link to microbiome homeostasis. We next attempted to group class II bacteriocins into subfamily with similar biological functions based on precursor sequence space and investigate their profiling in the human microbiome in detail . To fully reveal the chemical diversity of class II bacteriocin, we first adapted two approaches for the identification of precursor peptides, with hmmsearch [36] to search Pfam domains of precursors of class II bacteriocin (Supplementary Table 4) and with BAGEL4 [37] which is a tool specifically designed for bacteriocin mining. We combined two approaches to identify 187,649 precursors from class II bacteriocin BGCs (Fig. 5a, Supplementary Table 8). We then grouped 187,649 putative precursors into 2,005 clusters with a threshold of 50% sequence identity (Supplementary Fig. 16). The sequence lengths of those representative precursors were approximately normally distributed, with a center of ~55 amino acids (Fig. 5b). The accumulation curve showed that the precursor diversity increased with the number of genomes included, indicating more class II bacteriocins will be disclosed with more genomes sequenced (Fig. 5c). Moreover, we found that only 188 clusters were similar to 333 known class II bacteriocins (identity>90%, coverage>95%), leaving the vast majority (1,817/2,005) underexplored (Supplementary Table 9). Of note, while the rule-based method hmmsearch and BAGEL4 enable a high likelihood of positive detection, at the same time, they probably underestimate the real biosynthetic potentials of bacteriocins. In line with the taxa-specificity of the GCFs, most class II bacteriocin precursors family were genus-specific (n=1,862, 92.9%) and species-specific (n=1,327, 66.2%), with 33.7% of precursor family being even strain-specific (n=675) (Fig. 5d). We next examined their profile in the human microbiome. Of 644 clusters detected in six body sites, about 31.8% of clusters were niche-specific in one of six sites (Supplementary Fig. 17a). Moreover, the profiles of class II bacteriocins in different body sites were distinct as revealed by t-SNE plot, further supporting the niche-specificity of class II bacteriocin in the human microbiome (Fig. 5e). Their profiles in vagina and skin showed great individual variations, whereas the class II bacteriocins in other body sites were relatively conserved with being clustered together. Additionally, those class II bacteriocins were sporadically present in the skin and gut, whereas some class II bacteriocins were particularly enriched in the oral cavity and vagina with a high prevalence and abundance (Fig. 5f, Supplementary Fig. 17b). Probably due to the individual variations in the vagina, a member of subfamilies of class II bacteriocins exhibited a relatively smaller prevalence in the vagina than in the oral. Both the GCFs profile (Fig. 3d) and precursors profile (Figs. 5e, f) in the human microbiome suggested that class II bacteriocins are particularly enriched in the vaginal microbiome. Considering the vagina has simple communities with the lowest alpha diversity than other body sites [38], we reasoned that those enriched and predominant class II bacteriocins might play prominent roles in regulating microbial community in the vagina. Multi-omics analysis revealed class II bacteriocins potentially contributing to vaginal microbiome homeostasis To examine the class II bacteriocins that may account for the homeostasis of the vaginal microbiome, we first constructed the association network between class II bacteriocins and bacterial species at the metagenomic level. We found 23 precursor clusters correlated negatively with various species, indicating their antagonistic potential in regulating the vaginal microbiome (Fig. 6a). In particular, 21 clusters were negatively correlated with Lactobacillus iners , which is more conducive to the occurrence of abnormal vaginal microflora and thus a potential new therapeutic target for bacterial vaginosis treatment. Additionally, 21 of 23 clusters were also found to be inversely correlated to the Shannon index (Spearman rho < -0.4, adjusted P < 0.05) (Fig. 6b). Lower bacteria diversity in the microbiome with these detected class II bacteriocins suggested their regulative role in shaping the microbiome. To confirm whether these antagonistic bacteriocins are biologically functional, we next inspected their expression profile in the 180 metatranscriptomic datasets (Supplementary Table 6) and found that most of them were actively transcribed in the vaginal microbiome of healthy individuals (Fig. 6c). Three of the 21 clusters grouped with known class II bacteriocins, including Amylovorin L (cluster_342 and cluster_346, a two-component class IIb bacteriocin) from Lactobacillus amylovorus DCE 471 [39] and gassericin T (cluster_94) from Lactobacillus gasseri SBT2055 [40]. The findings of known class II bacteriocins with protective roles by omics-based associated analysis further validated the effectiveness of our approach in discovering the regulatory bacteriocins in the microbiome. The other 18 clusters of class II bacteriocins were also prevalent and actively transcribed in the vagina microbiome but uncharacterized yet. We next sought to validate the antagonistic potential of those uncharacterized bacteriocins experimentally. For proof of principle, we selected two precursor clusters (cluster_467 and cluster_468) with high abundance and a short peptide length that make their chemical synthesis practical. Those two precursors were located on BGCs (e.g, bgc120802) identified from 16 genomes of L. crispatus (Supplementary Fig. 18). The bgc120802 harbors specific class II bacteriocins-related genes, including a two-component regulator system (histidine kinase and response regulator), ABC transporter, and immunity protein. The precursor sequences from 16 BGCs were identical and featured a canonical double-glycine leader (Fig. 6d, Supplementary Fig. 18). We thus synthesized the core peptides (Supplementary Fig. 19), namely crispacin 467 (27 amino acids) and crispacin 468 (30 amino acids), respectively, and validated their antagonistic activity toward bacteria and fungi. The antimicrobial assay showed that the crispacin 467 exhibited a narrow-spectrum antibacterial activity against phylogenetical-closely related strain L. delbrueckii subsp. bulgaricus with a minimum inhibitory concentration of 12.5 μg/mL, while inhibitory effects of crispacin 468 were not observed, nor a synergy of them (Fig. 6e). The crispacin 468 might exhibit antimicrobial activity against other species beyond the tested strains. Taken together, we believe that the bacteriocin producers arm the vagina microbiome with diverse antagonistic bacteriocins, potentially preventing pathogen invasion and stabilizing the microbial community. Though how LAB employs SMs to shape their microbiome communities is not yet fully understood, our omics-guided discovery of new bacteriocins from LAB provides an alternative way to the discovery of new antibacterial therapeutics for microbiome dysbiosis. Discussion Despite increasing evidence revealing the health-promoting effects of LAB in the human microbiome, how they interplay with other microbes and influence microbiome homeostasis in the human is still understudied. Previous biosynthetic analysis of LAB in a limited dataset focuses on particular metabolites, proposing their protective roles to host [41,42]. However, the landscape of LAB SMs, particularly their profiles and potential roles in the human microbiome, remains elusive. In this study, we conducted a comprehensive omic analysis for LAB BGCs, significantly enhancing our understanding of the diversity and distribution of LAB BGCs. We found 129,878 BGCs of 2,849 GCFs from 31,977 LAB genomes, most of which were species-specific and encoded diverse uncharacterized SMs. We further investigated human metagenomes of six body sites to disclose the BGCs profile of LAB in the human microbiome, revealing that the diverse LAB SMs are diverse and variably prevalent in the human microbiome. Of note, BGCs of class II bacteriocins were particularly enriched and predominant in the vaginal microbiome. The niche specificity of GCFs in the human microbiome together with their specific- or even strain-specificity, suggested that the LAB SMs may provide a competitive advantage to the niche adaptation of their producing hosts. To profile BGCs in the human microbiome, we grouped them into families and clans based on architectural relationships, a well-accepted approach to study the BGC similarity and prioritize novel BGCs for natural product discovery [25,43]. Although informative, the GCF grouping will be affected by the imperfect BGC boundary prediction of antiSMASH [22]. Most LAB SMs were predicted to be antibacterial using machine learning models, indicating their potential regulating roles in the human microbiome and putative protective roles for the host. However, due to limited training data, our machine learning model can only predict limited bioactivities (antibacterial, antifungal, and antitumor), underestimating another biological potential of LAB SMs. Although such evidence does not exclude the possibility of other biological functions, we believe that antagonistic LAB SMs, particularly bacteriocins, potentially provide competitive edges to their producers and regulate the microbiome community. Applying metagenomics and metatranscriptomics analysis, we underscored 21 class II bacteriocins actively expressed in the vaginal microbiome and negatively correlated with individual bacteria species. Together with their negative association with the α-diversity of the vaginal microbiome, we can envision that these bacteriocins play prominent roles in regulating homeostasis. For proof of principle, we identified a novel class II bacteriocin produced by L. crispatus , namely crispacin 467, which exhibited characteristic narrow-spectrum antibacterial activity against phylogenetical-closely related strains. Our results suggested that crispacin 467 may arm L. crispatus with a protective role in vaginal health [44]. Although previous analyses have disclosed the bacteriocin biosynthetic genes and antibacterial activity in L. crispatus [45,46], little is known regarding their antagonistic bacteriocin except for crispacin A [47], not to mention their potential roles in the microbiome. A myriad of class II bacteriocins are harbored in Lactobacillus and the entire LAB of the human microbiome, providing enormous potential for new antimicrobial discovery. While the precursors of 21 bacteriocins were prevalent and transcribed in the vaginal microbiome, whether they are produced in situ in the vagina still needs to be examined by metabolomics. Meanwhile, how LAB employ SMs to shape their microbiome communities needs to be further explored in the future using in vivo mouse models or in vitro polymicrobial models. LAB, especially genus Lactobacillus , dominate the existing probiotics that confer a health benefit on the host when administered adequately [48]. The beneficial effects of probiotics result from diverse mechanisms, among which is bacteriocin production. Antagonistic bacteriocins that assist the producer colonization and provide protective roles for the host are important for probiotics to offer beneficial effects. Our findings reinforced the understanding of pervasive bacteriocins of LAB, which are also widespread in the human microbiome, particularly in the vagina. Those LAB and their bacteriocins present in healthy individuals are promising priorities for microbiome-based therapeutics [49]. For example, with the potential to modulate the vaginal microbiota, probiotics containing Lactobacillus spp. have been applied to treat bacterial vaginosis [44]. Moreover, the diverse antagonistic bacteriocins could be borrowed and arm genetically engineered beneficial LAB probiotics with a therapeutic property [50], preventing the host from the pathogen invasion directly or indirectly (via microbiota- and/or immune modulation). Continued investigation of the biosynthetic capacity and ecological roles will help to facilitate the translation of LAB and their antagonistic SMs into clinical application. In summary, our study provides a global insight into the biosynthetic potentials of LAB SMs and a starting point for the omics-guided discovery of antagonistic SMs that potentially regulate microbiome homeostasis. Class II bacteriocin predominant in vaginal microbiome but negatively associated with its bacterial diversity is experimentally validated to play antagonistic roles in microbial communities. To the best of our knowledge, our study is the first to systematically unveil LAB SM biosynthetic potentials and their profile in the healthy human microbiome. However, the analysis presented here cannot be considered exhaustive. The machine learning strategies employed to predict the bioactivity of SMs remain refined by knowledge accumulation of LAB SMs and their biosynthesis and bioactivity. Additionally, how LAB employ SMs to shape their microbiome communities in the human niche remains to be studied. Nevertheless, our systematic investigation of the biosynthetic potential of LAB provides a good starting point for the omics-guided discovery of SMs with therapeutic potential from the human microbiome. In addition to enhancing our understanding of the profile of LAB SMs and their potential regulating roles in the human microbiome, the discovery of antagonistic bacteriocins opens up exciting opportunities for future research on various probiotic applications of LABs. Methods Data acquisition As defined early, lactic acid bacteria include 14 genera, comprising Lactobacillus , Lactococcus , Leuconostoc , Pediococcus , Streptococcus , Aerococcus , Alloiococcus , Carnobacterium , Dolosigranulum , Enterococcus , Oenococcus , Tetragenococcus , Vagococcus , and Weissella [51]. Particularly, genus Lactobacillus has been reclassified recently [52], extending to 25 genera consisting of Lactobacillus , Paralactobacillus, Amylolactobacillus, Acetilactobacillus , Agrilactobacillus , Apilactobacillus , Bombilactobacillus , Companilactobacillus , Dellaglioa , Fructilactobacillus , Furfurilactobacillus , Holzapfelia , Lacticaseibacillus , Lactiplantibacillus , Lapidilactobacillus , Latilactobacillus , Lentilactobacillus , Levilactobacillus , Ligilactobacillus , Limosilactobacillus , Liquorilactobacillus , Loigolactobacilus, Paucilactobacillus , Schleiferilactobacillus , and Secundilactobacillus . Filtered with taxonomy, genomes from these 38 genera were then retrieved from NCBI reference sequences (RefSeq) database [17] (as of Aug. 2021, including SAGs only), PATRIC database (including SAGs) [18], IMG/M database (including SAGs and MAGs) [19]. Besides them, genomes from two previous studies focusing on the human gut microbiome (including SAGs and MAGs) [20] and food-originated LAB (including MAGs) [21] were also included. To avoid the reference genome redundancy, genomes from RefSeq were compared to themselves and those from other sources using Mash v2.3 [53]. Genomes with a Mash distance of 0 were considered identical. Only the one with a minimal number of contigs was retained. As potential misclassification might be present, GTDB-Tk v1.7.0 [54] was further used to confirm and unify taxonomic annotation against GTDB-Tk reference data version r202 [55]. There is a slight difference between NCBI taxonomy and GTDB taxonomy [56]. Under GTDB taxonomy, five genera are sub-divided: Carnobacterium , Enterococcus , Lactococcus , Vagococcus , and Weissella . Finally, a total of 56 genera belonging to six families (Lactobacillaceae, Aerococcaceae, Streptococcaceae, Vagococcaceae, Enterococcaceae, and Carnobacteriaceae) were considered as members of LAB in this study. Biosynthetic gene cluster analysis Biosynthetic gene clusters for each genome were annotated by antiSMASH 6.0 [22] with default parameters. In total, 31,977 LAB genomes and 164,417 non-LAB genomes (the intersection between RefSeq and GTDB repository version r202) [57] were included for BGC annotation. This resulted in 130,051 BGCs from 30,718 LAB genomes and 1,122,204 BGCs from 155,540 non-LAB genomes. No BGCs were annotated in 1,259 LAB genomes and 8,877 non-LAB genomes. Clustering BGCs into families and clans BiG-SLiCE [23], a tool to cluster sizable BGCs, contains two BGC features (biosynthetic-Pfam and sub-Pfam domains). Those BGC features are sufficient to distinguish distinct BGC classes. As a previous study described [26], all features of LAB BGCs and 1910 experimentally validated BGCs from the MIBiG 2.0 repository were extracted by BiG-SLiCE v1.1.0, and subsequently used to compute all-to-all cosine distances between BGCs using Python suite SciPy version 1.6.2 [58]. The cosine distances were next subject to hierarchical clustering with average linkage, grouping BGCs into families (GCFs, distances < 0.2) and clans (GCCs, distances < 0.8) by Python 3.8 with Scikit-learn version 0.24.2 [59]. Metagenomics and metatranscriptomics analysis The raw metagenomic sequencing reads of 748 HMP samples [31] and the raw metatranscriptomic data of 180 vaginal samples [60] were acquired from NCBI SRA (Sequence Read Archive) [61] under project accession number PRJNA48479 and PRJNA797778, respectively. Fastp 0.21.1 [62] with default parameters was adopted for detecting and removing low-quality sequencing reads. High-quality metagenomic sequencing reads were subjected to kneaddata (https://github.com/biobakery/kneaddata) for discarding reads belonging to the human host, through searching against the human reference genome (GRCh38.p13) from GENCODE [63]; high-quality metatranscriptomic reads were also subjected to SortMeRNA v4.3.4 [64] for removing reads derived from ribosomal RNAs. Following that, MetaPhlAn v3.0.13 [65] was used for taxonomic profiling. Prior to assessing the abundance of BGCs in metagenomics and metatranscriptomic data, we used a modified script from BiG-MAP [66] to de-duplicate 130,051 BGCs. To reduce the computational load, we de-duplicated them within each GCF, at a 0.8 nucleotide identity threshold, leading to 24,222 non-redundant BGCs, the nucleotide sequences of which were used to generate the reference database. Next, the non-host metagenomic and metatranscriptomic reads were mapped to this BGC reference using Bowtie 2 v2.3.5.1 [67], with a parameter of “-k 1”. We then utilized featureCounts v2.0.3 [68] (with parameters of “-T 30 -f -p -B -C -t CDS -g ID -M -O --fracOverlap 0.2”) to assign sequencing reads to the BGC genes. When calculating the abundance of a BGC, we only considered the core and additional biosynthetic genes, excluding the other genes such as transporters, regulators, transposases, and so forth. For each BGC, a corresponding GTF (General Transfer Format) annotation file was generated by antiSMASH. We retrieved the biosynthetic-related genes (tag “biosynthetic” for the core biosynthetic genes and “biosynthetic-additional” for the additional biosynthetic genes) according to the “gene_kind” tag in GTF files. A BGC was considered present in a metagenomic sample when fulfilling the following criteria: (1) the percentage of biosynthetic-related genes detected is over 50% of total biosynthetic-related genes in a BGC; (2) at least one core biosynthetic gene was found in a BGC. The abundance of a BGC was computed via the equation (1): N i represents the number of reads mapped on a biosynthetic-related gene; L i represents the gene length; k represents the number of biosynthetic-related genes in a BGC; N represents the total number of high-quality non-host reads in a metagenome/metatranscriptome sample. Prediction of secondary metabolite activity To predict the activity of BGC-encoding compounds, we used mlr v2.19.0 [69] to perform machine learning. The training dataset comprising 950 MIBiG BGCs with known activities (antibacterial, antifungal, antitumor or cytotoxic, or other activities) was gathered by Walker et al. [33]. BGC features of those known BGCs were extracted by BiG-SLiCE [23]. Prior to training models, we removed BGC features present in < 10 BGCs. Rather than multiclass classification, binary classification was adopted for each activity class since a molecule might have multiple functions. Four two-class classifiers (namely logistic regression, elastic net regression, random forest, and support vector machines) were adopted for binary classification of the activities of BGC products. In order to obtain the honest performance of four classifiers, we measured their accuracies using 10-fold cross-validation. Moreover, 3-fold cross-validation was adopted in each test to tune the hyperparameters, generating 30 instances for each classifier. The average AUROC was used to evaluate the performance of four classifiers. The function generateThreshVsPerfData was used to generate data on threshold vs. performances, which was further adapted for plotting the ROC curve. Using the random forest model, 129,878 LAB-derived BGCs and 1,121,156 non-LAB-derived BGCs containing BGC features were subject to activity prediction. Chord diagram showing the association between BGC classes and predicted activities of their products was plotted using R package circlize v0.4.13 [70]. Sankey diagram showing the association between species and BGC classes was done by package networkD3 v0.4 [71]. Precursor of class II bacteriocins In order to pinpoint the precursors of class II bacteriocins, we first used Prodigal-short [72] to identify all small ORFs. We then used hmmsearch [36] to search class II bacteriocins-related domains (provided in Supplementary Table 4) against ORFs of all RiPP-like BGCs. The hits with a threshold of E-value < 0.01 were considered as the precursors of class II bacteriocins. Meanwhile, BAGEL4 [37] was also adopted for searching class II bacteriocins from RiPP-like BGCs. They detected 128,599 and 90,101 putative precursors, respectively, with 30,764 in common. We discarded 287 sequences that were larger than 150 AAs, retaining 187,649 sequences for further analysis. Those sequences were then grouped into clusters using Cd-hit [73], with the parameters of “-n 2 -p 1 -c 0.5 -d 200 -M 50000 -l 5 -s 0.95 –aL 0.95 –g 1”. The sequences with an identity of > 50% will be grouped into one cluster, as proteins with > 50% identity generally share a common function [74]. To collect the known class II bacteriocins, we queried NCBI PubMed with the keyword “class II bacteriocin”. Meanwhile, we also included the sequences gathered by Yi et al. [15] as well as the sequence deposited in the BAGEL4 database [37]. In total, 333 sequences of class II bacteriocins were obtained (Supplementary Table 9). As the curated 333 sequences might be the mature peptides, a local sequence aligner, DIAMOND v2.0.15 [75], was utilized to compare 333 known class II bacteriocins to 187,649 precursor sequences with the parameter of “--id 90 --query-cover 95 --masking 0”. The known class II bacteriocins showed an alignment of identity > 90% and coverage > 95% with 1,775 precursor sequences belonging to 188 clusters that were thus regarded as homologous. For 21 selected precursor clusters, we identified the global identity relative to the known class II bacteriocins using the Needleman-Wunsch algorithm in the function “needleall” of EMBOSS software package [76]. The alignment of precursors was done by MAFFT v7.490 [77] with the parameter of “--maxiterate 1000 --localpair”, and then was visualized using Jalview software [78]. To conveniently inspect the gene organizations of BGCs harboring precursors of cluster_467 and cluster_468, we adopted BiG-SCAPE [25] for exploring their architectures. Precursor abundance in metagenome/metatranscriptome samples was computed via the equation (2): Here, N i represents the number of reads mapped on a precursor gene; L i represents the gene length; N represents the total number of high-quality non-host reads in a metagenome/metatranscriptome sample. The abundance of a precursor cluster is the sum of the abundance of precursors in this cluster. Phylogenetic tree construction GTDB repository version r202 [57] contains 822 representative genomes of 56 LAB genera. The genome with the largest N50 length in each genus was selected as a proxy for its corresponding genus. Consistent with the previous approach to constructing bacterial reference trees [57], the multiple sequence alignment of the concatenation of 120 phylogenetically informative marker genes of 56 representative genomes was used to infer the phylogenetic tree. IQ-TREE version 2.1.4-beta [79] was adapted for constructing maximum likelihood (ML) phylogenetic trees, with 1,000 ultrafast bootstrap replicates. In-built ModelFinder [80] identified the best-fit model as LG+F+R8. Inferred phylogeny was visualized using iTOL [81]. Peptide synthesis The two deducted core peptides of cluster_467 and cluster_468 were chemically synthesized by Sangon Biotech (Shanghai, China). Their molecular weights were confirmed by mass spectrometry, and their required purity was ≥ 90%, determined by high-performance liquid chromatography. The synthesized peptide powder was stored at −80 °C and dissolved in sterilized double-distilled water to 2 mg/mL upon use. Bacterial and fungal strains A total of 14 bacterial strains and two fungal strains were used in this study. Their growth conditions are as follows: two bacterial strains ( Escherichia coli DH5α, Staphylococcus aureus B04) were incubated in Luria–Bertani (LB) culture medium at 37 ℃ under 180 rpm rotation; two strains ( Chromobacterium violaceum , Bacillus subtilis 168) were incubated in LB medium at 30 ℃ with shaking at 180 rpm; Lactococcus lactis subsp. cremoris MG1363 was grown statically at 30 ℃ in M17 medium; eight bacterial strains (including Streptococcus mitis , Enterococcus faecium , Enterococcus faecalis , Lactobacillus delbrueckii subsp. Bulgaricus , Lactobacillus crispatus ATCC 33820, Lactobacillus acidophillus , Lactobacillus casei , Lactobacillus fermentum ) were incubated statically in MRS medium at 37 ℃; Gardnerella vaginalis was maintained in Colombia blood agar at 37 ℃ in anaerobic conditions, and the suspension culture was grown by taking a loop full of colonies from the agar plate and incubating in Brain heart infusion broth (BHI), at 37 ℃ in anaerobic conditions; two fungal strains ( Candida albicans SC5314, Candida albicans ATCC 10231) were grown in RPMI media at 37 ℃ with shaking at 150 rpm. Determination of minimum inhibitory concentrations The minimum inhibitory concentrations (MICs) of the two peptides [individually and in combination (1:1 ration)] against bacterial and fungal strains were performed by broth microdilution. Tested bacterial strains were inoculated overnight in the corresponding culture medium (LB, M17, BHI, or MRS) and at respective growth conditions. The optical density at 600 nm (OD 600 ) of bacterial cultures was determined to estimate the bacterial concentration. The bacteria cultures were diluted to ~5 × 10 5 CFU/mL using the respective broth. 100 μL aliquots of bacterial suspensions were transferred into 96-well plates containing two-fold serial dilutions of peptides (ranging from 200 μg/ml to 0.19 μg/ml). After incubating for 24 h, bacterial growth was assessed by determining OD 600 . Besides, MIC against the fungal strains was determined according to the CLSI M27-A3 guidelines [82]. Briefly, C. albicans strains were cultured overnight in RPMI medium and grown fungal suspensions were centrifuged at 5,000 rpm for 10 min and the pellet was resuspended and washed twice with 1× PBS to remove the dead cells. The fungal inoculum was standardized to 1 × 10 6 CFU/mL using a spectrophotometer and added to the well plate containing varying concentrations of the peptide (200 μg/mL to 0.19 μg/mL). The media without the peptide served as a control. The plates were then incubated at 37°C for 24 h with shaking at 80 rpm, and the absorbance was measured at 520 nm using SpectraMax 340 tunable microplate reader (Molecular Devices, San Jose, CA, USA). The MIC value was determined as the lowest concentration of the peptides where no bacterial or fungal growth was detected. All assays were conducted in triplicate on three independent occasions. Statistical analysis and visualization The accumulations of GCFs detected in metagenomes as well as clusters of class II bacteriocin precursors were computed with function specaccum in R package vegan v2.5-7 [83]. R package UpSetR v1.4.0 [84] was adopted for visualizing the intersection of GCFs or precursor clusters detected in different body sites. Chi-squared test and wilcoxon rank-sum test (two sided) were done by function chisq.test and wilcox.test in R, respectively. The alpha diversity (Shannon index) of the vaginal microbiome was calculated with R package vegan v2.5-7 [83]. To visualize the distribution of class II bacteriocins detected in six body sites, a dimensionality reduction was performed using t-distributed stochastic neighbor embedding (t-SNE), which was done by R package Rtsne v0.15 [85]. Spearman's correlations between precursor clusters vs. bacterial species and between clusters vs. Shannon index were computed with the function corr.test in R package psych v2.1.9 [86], and P values were adjusted with the “BH” method [87]. The heat maps in this study were plotted using package pheatmap v1.0.12 [88]. Cytoscape 3.9.0 [89] was used to visualize the network of similarity of class II bacteriocins and the network of species-precursor correlation. Without a specific statement, other figures were generated using ggplot2 v3.3.5 [90]. All statistical analyses were finished in R v4.1.2. Declarations Ethics approval and Consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials The bacterial genomes are publicly available in NCBI Assembly RefSeq database (https://www.ncbi.nlm.nih.gov/assembly), PATRIC database (https://docs.patricbrc.org/user_guides/ftp.html), and IMG/M database (https://img.jgi.doe.gov/cgi-bin/m/main.cgi). The genomes from human gut are available in the European Nucleotide Archive under study accession ERP116715, and genomes from food metagenomes are available at http://www.tfm.unina.it/DATA001-2020-Pasolli. Genomes can be obtained through the accession numbers provided in Supplementary Tables 1 and 5. The raw data for HMP metagenomes and vaginal metatranscriptomes are deposited in NCBI-SRA under the BioProjects PRJNA48479 (https://www.ncbi.nlm.nih.gov/bioproject/48479) and PRJNA797778 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA797778), respectively. The samples used in this study are provided in Supplementary Table 6. Competing interests The authors declare no competing interests. Funding This work is partially funded by a Shenzhen Basic Research General Programme (JCYJ20210324122211031) and two Hong Kong Research Grants Council General research grants (HKU27107320 and HKU17115322). Authors' contributions YL and DZ conceived of the study, participated in its design and coordination, and drafted the manuscript. DZ and JZ gathered publicly available data used in this study. DZ and SK performed MIC determination. JL, ZS, BH, and PC performed bacterial culture and metabolic analysis. DZ, JZ, ZZ, and YL performed data analysis and interpretation. CF and PN provided advice. YL was involved in the overall supervision of the project. All authors read, revised, and approved the final manuscript. Acknowledgments The authors would like to thank Dr. Mingqiang Qiao and Wanjin Qiao at Nankai University for providing L. lactis strain. References Carr FJ, Chill D, Maida N. The lactic acid bacteria: A literature survey. Critical Reviews in Microbiology. 2002;28:281–370. Leroy F, de Vuyst L. Lactic acid bacteria as functional starter cultures for the food fermentation industry. Trends in Food Science and Technology. Elsevier; 2004;15:67–78. Teusink B, Smid EJ. Modelling strategies for the industrial exploitation of lactic acid bacteria [Internet]. Nature Reviews Microbiology. Nature Publishing Group; 2006 [cited 2022 Jul 5]. p. 46–56. Available from: https://www.nature.com/articles/nrmicro1319 Wells JM, Mercenier A. Mucosal delivery of therapeutic and prophylactic molecules using lactic acid bacteria. Nature Reviews Microbiology [Internet]. Nature Publishing Group; 2008 [cited 2022 Jul 5];6:349–62. Available from: https://www.nature.com/articles/nrmicro1840 Saez-Lara MJ, Gomez-Llorente C, Plaza-Diaz J, Gil A. The Role of Probiotic Lactic Acid Bacteria and Bifidobacteria in the Prevention and Treatment of Inflammatory Bowel Disease and Other Related Diseases: A Systematic Review of Randomized Human Clinical Trials. BioMed Research International [Internet]. Hindawi Publishing Corporation; 2015 [cited 2022 Jul 5];2015:1–15. Available from: http://www.hindawi.com/journals/bmri/2015/505878/ Ren C, Faas MM, de Vos P. Disease managing capacities and mechanisms of host effects of lactic acid bacteria. Critical Reviews in Food Science and Nutrition [Internet]. Taylor & Francis; 2021 [cited 2021 Nov 12];61:1365–93. Available from: https://www.tandfonline.com/doi/abs/10.1080/10408398.2020.1758625 García-Bayona L, Comstock LE. Bacterial antagonism in host-associated microbial communities. Science (1979). 2018;361. Braga RM, Dourado MN, Araújo WL. Microbial interactions: ecology in a molecular perspective. Brazilian Journal of Microbiology [Internet]. Sociedade Brasileira de Microbiologia; 2016;47:86–98. Available from: http://dx.doi.org/10.1016/j.bjm.2016.10.005 Delves-Broughton J, Blackburn P, Evans RJ, Hugenholtz J. Applications of the bacteriocin, nisin. Antonie van Leeuwenhoek 1996 69:2 [Internet]. Springer; 1996 [cited 2022 May 11];69:193–202. Available from: https://link.springer.com/article/10.1007/BF00399424 Donia MS, Cimermancic P, Schulze CJ, Wieland Brown LC, Martin J, Mitreva M, et al. A systematic analysis of biosynthetic gene clusters in the human microbiome reveals a common family of antibiotics. Cell [Internet]. Elsevier Inc.; 2014;158:1402–14. Available from: http://dx.doi.org/10.1016/j.cell.2014.08.032 Höltzel A, Gänzle MG, Nicholson GJ, Hammes WP, Jung G. The First Low Molecular Weight Antibiotic from Lactic Acid Bacteria: Reutericyclin, a New Tetramic Acid. Angew Chem Int Ed Engl [Internet]. Angew Chem Int Ed Engl; 2000 [cited 2022 May 11];39:2766–8. Available from: https://pubmed.ncbi.nlm.nih.gov/10934421/ Acedo JZ, Chiorean S, Vederas JC, van Belkum MJ. The expanding structural variety among bacteriocins from Gram-positive bacteria. FEMS Microbiology Reviews [Internet]. Oxford Academic; 2018 [cited 2021 Dec 29];42:805–28. Available from: https://academic.oup.com/femsre/article/42/6/805/5063573 Cotter PD, Ross RP, Hill C. Bacteriocins — a viable alternative to antibiotics? Nature Reviews Microbiology [Internet]. Nature Publishing Group; 2013;11:95–105. Available from: http://www.nature.com/articles/nrmicro2937 Heilbronner S, Krismer B, Brötz-Oesterhelt H, Peschel A. The microbiome-shaping roles of bacteriocins [Internet]. Nature Reviews Microbiology. Nature Publishing Group; 2021 [cited 2021 Dec 6]. p. 726–39. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/s41579-021-00569-w Yi Y, Li P, Zhao F, Zhang T, Shan Y, Wang X, et al. Current status and potentiality of class II bacteriocins from lactic acid bacteria: structure, mode of action and applications in the food industry. Trends in Food Science and Technology [Internet]. Elsevier Ltd; 2022;120:387–401. Available from: https://doi.org/10.1016/j.tifs.2022.01.018 «d Ennahar S, Sashihara T, Sonomoto K, Ishizaki A. Class IIa bacteriocins: biosynthesis, structure and activity. FEMS Microbiology Reviews [Internet]. Oxford Academic; 2000 [cited 2021 Dec 24];24:85–106. Available from: https://academic.oup.com/femsre/article/24/1/85/526454 O’Leary NA, Wright MW, Brister JR, Ciufo S, Haddad D, McVeigh R, et al. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research [Internet]. Oxford Academic; 2016 [cited 2022 Feb 19];44:D733–45. Available from: https://academic.oup.com/nar/article/44/D1/D733/2502674 Wattam AR, Abraham D, Dalay O, Disz TL, Driscoll T, Gabbard JL, et al. PATRIC, the bacterial bioinformatics database and analysis resource. Nucleic Acids Res [Internet]. Nucleic Acids Res; 2014 [cited 2022 Feb 21];42. Available from: https://pubmed.ncbi.nlm.nih.gov/24225323/ Chen IMA, Chu K, Palaniappan K, Ratner A, Huang J, Huntemann M, et al. The IMG/M data management and analysis system v.6.0: new tools and advanced capabilities. Nucleic Acids Research [Internet]. Oxford Academic; 2021 [cited 2022 Feb 19];49:D751–63. Available from: https://academic.oup.com/nar/article/49/D1/D751/5943189 Almeida A, Nayfach S, Boland M, Strozzi F, Beracochea M, Shi ZJ, et al. A unified catalog of 204,938 reference genomes from the human gut microbiome. Nature Biotechnology [Internet]. Nature Publishing Group; 2021 [cited 2021 Oct 12];39:105–14. Available from: https://www.nature.com/articles/s41587-020-0603-3 Pasolli E, de Filippis F, Mauriello IE, Cumbo F, Walsh AM, Leech J, et al. Large-scale genome-wide analysis links lactic acid bacteria from food with the gut microbiome. Nature Communications [Internet]. Nature Publishing Group; 2020 [cited 2021 Dec 12];11:1–12. Available from: https://www.nature.com/articles/s41467-020-16438-8 Blin K, Shaw S, Kloosterman AM, Charlop-Powers Z, van Wezel GP, Medema MH, et al. antiSMASH 6.0: improving cluster detection and comparison capabilities. Nucleic Acids Research [Internet]. Oxford Academic; 2021 [cited 2021 Jul 23];49:W29–35. Available from: https://academic.oup.com/nar/article/49/W1/W29/6274535 Kautsar SA, van der Hooft JJJ, de Ridder D, Medema MH. BiG-SLiCE: A highly scalable tool maps the diversity of 1.2 million biosynthetic gene clusters. Gigascience. Oxford University Press; 2021;10:1–17. Nayfach S, Roux S, Seshadri R, Udwary D, Varghese N, Schulz F, et al. A genomic catalog of Earth’s microbiomes. Nature Biotechnology. Nature Research; 2021;39:499–509. Navarro-Muñoz JC, Selem-Mojica N, Mullowney MW, Kautsar SA, Tryon JH, Parkinson EI, et al. A computational framework to explore large-scale biosynthetic diversity. Nature Chemical Biology [Internet]. Nature Publishing Group; 2020 [cited 2021 Sep 18];16:60–8. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/s41589-019-0400-9 Paoli L, Ruscheweyh H-J, Forneris CC, Hubrich F, Kautsar S, Bhushan A, et al. Biosynthetic potential of the global ocean microbiome. Nature [Internet]. Nature Publishing Group; 2022 [cited 2022 Jun 23];1–8. Available from: https://www.nature.com/articles/s41586-022-04862-3 Kautsar SA, Blin K, Shaw S, Navarro-Muñoz JC, Terlouw BR, van der Hooft JJJ, et al. MIBiG 2.0: A repository for biosynthetic gene clusters of known function. Nucleic Acids Research. 2020; Brockhurst MA, Harrison E, Hall JPJ, Richards T, McNally A, MacLean C. The Ecology and Evolution of Pangenomes. Current Biology. Cell Press; 2019;29:R1094–103. Gavriilidou A, Kautsar SA, Zaburannyi N, Krug D, Müller R, Medema MH, et al. Compendium of specialized metabolite biosynthetic diversity encoded in bacterial genomes. Nature Microbiology [Internet]. Nature Publishing Group; 2022 [cited 2022 May 3];7:726–35. Available from: https://www.nature.com/articles/s41564-022-01110-2 George F, Daniel C, Thomas M, Singer E, Guilbaud A, Tessier FJ, et al. Occurrence and dynamism of lactic acid bacteria in distinct ecological niches: A multifaceted functional health perspective. Frontiers in Microbiology. Frontiers Media S.A.; 2018. p. 2899. Methé BA, Nelson KE, Pop M, Creasy HH, Giglio MG, Huttenhower C, et al. A framework for human microbiome research. Nature 2012 486:7402 [Internet]. Nature Publishing Group; 2012 [cited 2022 Mar 16];486:215–21. Available from: https://www.nature.com/articles/nature11209 Hannigan GD, Prihoda D, Palicka A, Soukup J, Klempir O, Rampula L, et al. A deep learning genome-mining strategy for biosynthetic gene cluster prediction. Nucleic Acids Research [Internet]. Oxford Academic; 2019 [cited 2019 Oct 8];47:e110–e110. Available from: https://academic.oup.com/nar/article/47/18/e110/5545735 Walker AS, Clardy J. A Machine Learning Bioinformatics Method to Predict Biological Activity from Biosynthetic Gene Clusters. Journal of Chemical Information and Modeling [Internet]. American Chemical Society; 2021 [cited 2019 Oct 1];61:2560–71. Available from: https://pubs.acs.org/doi/full/10.1021/acs.jcim.0c01304 Skinnider MA, Johnston CW, Gunabalasingam M, Merwin NJ, Kieliszek AM, MacLellan RJ, et al. Comprehensive prediction of secondary metabolite structure and biological activity from microbial genome sequences. Nature Communications [Internet]. Nature Publishing Group; 2020 [cited 2019 Sep 30];11:1–9. Available from: https://www.nature.com/articles/s41467-020-19986-1 Lee SW, Mitchell DA, Markley AL, Hensler ME, Gonzalez D, Wohlrab A, et al. Discovery of a widely distributed toxin biosynthetic gene cluster. Proc Natl Acad Sci U S A [Internet]. 2008 [cited 2022 Mar 10];105:5879–84. Available from: www.pnas.org/cgi/content/full/ Eddy SR. Accelerated Profile HMM Searches. PLOS Computational Biology [Internet]. Public Library of Science; 2011 [cited 2022 May 23];7:e1002195. Available from: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1002195 van Heel AJ, de Jong A, Song C, Viel JH, Kok J, Kuipers OP. BAGEL4: A user-friendly web server to thoroughly mine RiPPs and bacteriocins. Nucleic Acids Research [Internet]. Oxford Academic; 2018 [cited 2022 Jan 18];46:W278–81. Available from: https://academic.oup.com/nar/article/46/W1/W278/5000017 Structure, function and diversity of the healthy human microbiome. Nature [Internet]. 2012;486:207–14. Available from: http://www.nature.com/articles/nature11234 Foulquié Moreno MR, Baert B, Denayer S, Cornelis P, de Vuyst L. Characterization of the amylovorin locus of Lactobacillus amylovorus DCE 471, producer of a bacteriocin active against Pseudomonas aeruginosa, in combination with colistin and pyocins. FEMS Microbiology Letters [Internet]. Oxford Academic; 2008 [cited 2022 Jun 13];286:199–206. Available from: https://academic.oup.com/femsle/article/286/2/199/591282 Kawai Y, Saitoh B, Takahashi O, Kitazawa H, Saito T, Nakajima H, et al. Primary Amino Acid and DNA Sequences of Gassericin T, a Lactacin F-Family Bacteriocin Produced by Lactobacillus gasseri SBT2055. Bioscience, Biotechnology, and Biochemistry [Internet]. Oxford Academic; 2000 [cited 2022 Jun 13];64:2201–8. Available from: https://academic.oup.com/bbb/article/64/10/2201/5945713 Alvarez-Sieiro P, Montalbán-López M, Mu D, Kuipers OP. Bacteriocins of lactic acid bacteria: extending the family. Applied Microbiology and Biotechnology. 2016. p. 2939–51. Mukesh Kumar M, Dhanasekaran D. Biosynthetic Gene Cluster Analysis in Lactobacillus Species Using antiSMASH. Advances in Probiotics [Internet]. Elsevier; 2021 [cited 2022 Apr 28]. p. 113–20. Available from: https://linkinghub.elsevier.com/retrieve/pii/B9780128229095000071 Doroghazi JR, Albright JC, Goering AW, Ju K-S, Haines RR, Tchalukov KA, et al. A roadmap for natural product discovery based on large-scale genomics and metabolomics. Nature Chemical Biology [Internet]. Nature Publishing Group; 2014 [cited 2021 Nov 15];10:963–8. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/nchembio.1659 France M, Alizadeh M, Brown S, Ma B, Ravel J. Towards a deeper understanding of the vaginal microbiota. Nature Microbiology. Springer US; 2022. p. 367–78. Fontana F, Alessandri G, Lugli GA, Mancabelli L, Longhi G, Anzalone R, et al. Probiogenomics Analysis of 97 Lactobacillus crispatus Strains as a Tool for the Identification of Promising Next-Generation Probiotics. Microorganisms [Internet]. 2020;9:73. Available from: https://www.mdpi.com/2076-2607/9/1/73 Argentini C, Fontana F, Alessandri G, Lugli GA, Mancabelli L, Ossiprandi MC, et al. Evaluation of Modulatory Activities of Lactobacillus crispatus Strains in the Context of the Vaginal Microbiota. Microbiology Spectrum. American Society for Microbiology; 2022; Tahara T, Kanatani K. Isolation and partial characterization of crispacin A, a cell-associated bacteriocin produced by Lactobacillus crispatus JCM 2009. FEMS Microbiology Letters [Internet]. Oxford Academic; 2006 [cited 2022 May 27];147:287–90. Available from: https://academic.oup.com/femsle/article/147/2/287/532692 Suez J, Zmora N, Segal E, Elinav E. The pros, cons, and many unknowns of probiotics [Internet]. Nature Medicine. 2019 [cited 2021 Dec 11]. p. 716–29. Available from: https://doi.org/10.1038/s41591-019-0439-x Sorbara MT, Pamer EG. Microbiome-based therapeutics. Nature Reviews Microbiology [Internet]. Nature Publishing Group; 2022 [cited 2022 Jan 17];20:365–80. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/s41579-021-00667-9 Plavec TV, Berlec A. Engineering of lactic acid bacteria for delivery of therapeutic proteins and peptides. Applied Microbiology and Biotechnology [Internet]. Springer Verlag; 2019;103:2053–66. Available from: http://link.springer.com/10.1007/s00253-019-09628-y Mokoena MP. Lactic Acid Bacteria and Their Bacteriocins: Classification, Biosynthesis and Applications against Uropathogens: A Mini-Review. Molecules [Internet]. MDPI AG; 2017;22:1255. Available from: http://www.mdpi.com/1420-3049/22/8/1255 Zheng J, Wittouck S, Salvetti E, Franz CMAP, Harris HMB, Mattarelli P, et al. A taxonomic note on the genus Lactobacillus: Description of 23 novel genera, emended description of the genus Lactobacillus Beijerinck 1901, and union of Lactobacillaceae and Leuconostocaceae. International Journal of Systematic and Evolutionary Microbiology [Internet]. Microbiology Society; 2020;70:2782–858. Available from: https://www.microbiologyresearch.org/content/journal/ijsem/10.1099/ijsem.0.004107 Ondov BD, Treangen TJ, Melsted P, Mallonee AB, Bergman NH, Koren S, et al. Mash: Fast genome and metagenome distance estimation using MinHash. Genome Biology [Internet]. BioMed Central Ltd.; 2016 [cited 2022 Feb 19];17:1–14. Available from: https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0997-x Chaumeil PA, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics [Internet]. Oxford Academic; 2020 [cited 2022 Feb 19];36:1925–7. Available from: https://academic.oup.com/bioinformatics/article/36/6/1925/5626182 Parks DH, Chuvochina M, Rinke C, Mussig AJ, Chaumeil P-A, Hugenholtz P. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Research [Internet]. Oxford Academic; 2022 [cited 2022 Feb 25];50:D785–94. Available from: https://academic.oup.com/nar/article/50/D1/D785/6370255 Parks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil PA, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nature Biotechnology 2018 36:10 [Internet]. Nature Publishing Group; 2018 [cited 2022 Feb 25];36:996–1004. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/nbt.4229 Parks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil PA, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nature Biotechnology 2018 36:10 [Internet]. Nature Publishing Group; 2018 [cited 2022 Mar 1];36:996–1004. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/nbt.4229 Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nature Methods 2020 17:3 [Internet]. Nature Publishing Group; 2020 [cited 2022 Mar 19];17:261–72. Available from: https://www.nature.com/articles/s41592-019-0686-2 Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research [Internet]. 2011 [cited 2022 Mar 19];12:2825–30. Available from: http://scikit-learn.sourceforge.net. France MT, Fu L, Rutt L, Yang H, Humphrys MS, Narina S, et al. Insight into the ecology of vaginal bacteria through integrative analyses of metagenomic and metatranscriptomic data. Genome Biology 2022 23:1 [Internet]. BioMed Central; 2022 [cited 2022 Mar 5];23:1–26. Available from: https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02635-9 Leinonen R, Sugawara H, Shumway M. The Sequence Read Archive. Nucleic Acids Research [Internet]. Oxford University Press; 2011 [cited 2022 Mar 19];39:D19. Available from: /pmc/articles/PMC3013647/ Chen S, Zhou Y, Chen Y, Gu J. Fastp: An ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018. p. i884–90. Frankish A, Diekhans M, Ferreira AM, Johnson R, Jungreis I, Loveland J, et al. GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Research [Internet]. Oxford Academic; 2019 [cited 2022 Mar 19];47:D766–73. Available from: https://academic.oup.com/nar/article/47/D1/D766/5144133 Kopylova E, Noé L, Touzet H. SortMeRNA: fast and accurate filtering of ribosomal RNAs in metatranscriptomic data. Bioinformatics [Internet]. Oxford Academic; 2012 [cited 2022 Jun 18];28:3211–7. Available from: https://academic.oup.com/bioinformatics/article/28/24/3211/246053 Beghini F, McIver LJ, Blanco-Míguez A, Dubois L, Asnicar F, Maharjan S, et al. Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with biobakery 3. Elife. eLife Sciences Publications Ltd; 2021;10. Pascal Andreu V, Augustijn HE, van den Berg K, van der Hooft JJJ, Fischbach MA, Medema MH. BiG-MAP: an Automated Pipeline To Profile Metabolic Gene Cluster Abundance and Expression in Microbiomes. mSystems [Internet]. American Society for Microbiology; 2021 [cited 2022 Mar 19];6. Available from: https://journals.asm.org/doi/abs/10.1128/mSystems.00937-21 Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nature Methods. 2012;9:357–9. Liao Y, Smyth GK, Shi W. FeatureCounts: An efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014; Bischl B, Lang M, Kotthoff L, Schiffner J, Richter J, Studerus E, et al. mlr: Machine Learning in R. Journal of Machine Learning Research [Internet]. 2016 [cited 2022 Jun 18];17:1–5. Available from: https://github.com/mlr-org/mlr Gu Z, Gu L, Eils R, Schlesner M, Brors B. circlize implements and enhances circular visualization in R. Bioinformatics [Internet]. Oxford Academic; 2014 [cited 2022 Jun 20];30:2811–2. Available from: https://academic.oup.com/bioinformatics/article/30/19/2811/2422259 Allaire JJ, Ellis P, Gandrud C, Kuo K, Lewis BW, Owen J, et al. Package “networkD3.” 2017 [cited 2022 Jun 20]; Available from: https://github.com/christophergandrud/networkD3/issues Santos-Aberturas J, Chandra G, Frattaruolo L, Lacret R, Pham TH, Vior NM, et al. Uncovering the unexplored diversity of thioamidated ribosomal peptides in Actinobacteria using the RiPPER genome mining tool. Nucleic Acids Research [Internet]. Oxford Academic; 2019 [cited 2022 Jun 18];47:4624–37. Available from: https://academic.oup.com/nar/article/47/9/4624/5420534 Li W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics [Internet]. Oxford Academic; 2006 [cited 2022 Jun 18];22:1658–9. Available from: https://academic.oup.com/bioinformatics/article/22/13/1658/194225 Sangar V, Blankenberg DJ, Altman N, Lesk AM. Quantitative sequence-function relationships in proteins based on gene ontology. BMC Bioinformatics [Internet]. BioMed Central; 2007 [cited 2022 May 23];8:1–15. Available from: https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-8-294 Buchfink B, Reuter K, Drost HG. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nature Methods 2021 18:4 [Internet]. Nature Publishing Group; 2021 [cited 2022 Jun 20];18:366–8. Available from: https://www.nature.com/articles/s41592-021-01101-x Rice P, Longden L, Bleasby A. EMBOSS: The European Molecular Biology Open Software Suite. Trends in Genetics [Internet]. Elsevier; 2000 [cited 2022 Jun 20];16:276–7. Available from: http://www.cell.com/article/S0168952500020242/fulltext Katoh K, Standley DM. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. Molecular Biology and Evolution [Internet]. Oxford Academic; 2013 [cited 2022 Jun 20];30:772–80. Available from: https://academic.oup.com/mbe/article/30/4/772/1073398 Waterhouse AM, Procter JB, Martin DMA, Clamp M, Barton GJ. Jalview Version 2—a multiple sequence alignment editor and analysis workbench. Bioinformatics [Internet]. Oxford Academic; 2009 [cited 2022 Jun 20];25:1189–91. Available from: https://academic.oup.com/bioinformatics/article/25/9/1189/203460 Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Molecular Biology and Evolution [Internet]. Oxford Academic; 2020 [cited 2022 Feb 28];37:1530–4. Available from: https://academic.oup.com/mbe/article/37/5/1530/5721363 Kalyaanamoorthy S, Minh BQ, Wong TKF, von Haeseler A, Jermiin LS. ModelFinder: Fast model selection for accurate phylogenetic estimates. Nature Methods [Internet]. Nat Methods; 2017 [cited 2021 Jul 25];14:587–9. Available from: https://pubmed.ncbi.nlm.nih.gov/28481363/ Letunic I, Bork P. Interactive Tree of Life (iTOL) v4: Recent updates and new developments. Nucleic Acids Research [Internet]. Nucleic Acids Res; 2019 [cited 2021 Jul 25];47. Available from: https://pubmed.ncbi.nlm.nih.gov/30931475/ Clinical Laboratory Standards Institute. Reference Method for Broth Dilution Antifungal susceptibility testing of yeast. Clinical Laboratory Standards Institute [Internet]. 2008 [cited 2022 Jun 20];22. Available from: www.clsi.org. Oksanen J. Vegan: ecological diversity [Internet]. R Package Version 2.4-4. 2017. p. 11. Available from: https://cran.r-project.org/package=vegan Conway JR, Lex A, Gehlenborg N. UpSetR: an R package for the visualization of intersecting sets and their properties. Bioinformatics [Internet]. Oxford Academic; 2017 [cited 2022 Jun 20];33:2938–40. Available from: https://academic.oup.com/bioinformatics/article/33/18/2938/3884387 Package “Rtsne” Title T-Distributed Stochastic Neighbor Embedding using a Barnes-Hut Implementation. 2022 [cited 2022 Jun 20]; Available from: https://github.com/jkrijthe/Rtsne Revelle W. Package “psych” - Procedures for Psychological, Psychometric and Personality Research. R Package [Internet]. 2015 [cited 2022 Jun 20];1–358. Available from: https://www.scholars.northwestern.edu/en/publications/psych-procedures-for-personality-and-psychological-research Benjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological). 1995;57:289–300. Kolde R, others. Pheatmap: pretty heatmaps. R package version. 2012;1:726. Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, et al. Cytoscape: A software Environment for integrated models of biomolecular interaction networks. Genome Research. 2003; Villanueva RAM, Chen ZJ. ggplot2: elegant graphics for data analysis. Taylor \& Francis; 2019. Supplementary Tables Supplementary Tables S1-S9 are not available with this version Additional Declarations No competing interests reported. Supplementary Files SupplementaryFigures.docx Cite Share Download PDF Status: Published Journal Publication published 27 Apr, 2023 Read the published version in Microbiome → Version 1 posted Editorial decision: Major revision 22 Feb, 2023 Reviews received at journal 27 Jan, 2023 Reviews received at journal 19 Dec, 2022 Reviewers agreed at journal 12 Dec, 2022 Reviewers agreed at journal 12 Dec, 2022 Reviewers invited by journal 07 Dec, 2022 Editor assigned by journal 21 Jul, 2022 Submission checks completed at journal 19 Jul, 2022 First submitted to journal 17 Jul, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1868011","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":122438673,"identity":"4a24c5ee-3845-4132-9b07-c0c6c17b0d42","order_by":0,"name":"Dengwei Zhang","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Dengwei","middleName":"","lastName":"Zhang","suffix":""},{"id":122438676,"identity":"02e946c1-00c6-48cf-8c5d-1519d1f7bd76","order_by":1,"name":"Jian Zhang","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jian","middleName":"","lastName":"Zhang","suffix":""},{"id":122438677,"identity":"fb86538c-1c40-4b14-a48f-19571efa579b","order_by":2,"name":"Shanthini Kalimuthu","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shanthini","middleName":"","lastName":"Kalimuthu","suffix":""},{"id":122438678,"identity":"ccd56902-30f2-4441-b330-dbb8d9cf18f6","order_by":3,"name":"Jing Liu","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jing","middleName":"","lastName":"Liu","suffix":""},{"id":122438680,"identity":"c7886f85-11cb-4527-abc7-65705b74a2ed","order_by":4,"name":"Zhiman Song","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zhiman","middleName":"","lastName":"Song","suffix":""},{"id":122438682,"identity":"9fe1ea98-e02b-4d86-84c7-adba9542302e","order_by":5,"name":"Bei-bei He","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bei-bei","middleName":"","lastName":"He","suffix":""},{"id":122438683,"identity":"70460543-a5aa-4677-9666-82fb68381f7e","order_by":6,"name":"Peiyan Cai","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Peiyan","middleName":"","lastName":"Cai","suffix":""},{"id":122438685,"identity":"bf22b0d4-fd0c-4b4e-978b-f7fd8390f2d4","order_by":7,"name":"Zheng Zhong","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zheng","middleName":"","lastName":"Zhong","suffix":""},{"id":122438686,"identity":"d77934f5-0c1b-4601-826a-a8a6f1a834a7","order_by":8,"name":"Chenchen Feng","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Chenchen","middleName":"","lastName":"Feng","suffix":""},{"id":122438687,"identity":"cb3ffc3e-bb46-42ba-9a47-02b303a7910f","order_by":9,"name":"Prasanna Neelakantan","email":"","orcid":"","institution":"University of Hong Kong","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Prasanna","middleName":"","lastName":"Neelakantan","suffix":""},{"id":122438688,"identity":"d2bd8a49-491e-47cf-862e-84070d836506","order_by":10,"name":"Yong-Xin Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAnUlEQVRIiWNgGAWjYHACxgMJFQcMGBh4mInXcyDhDMlaGNtI0WJwI/fAgYfz7hgbHOA9bECcljPnEg4kbntmZnCALzmBOC3HewyAWg7bGBzgMT5AnJbDPEAtc0jSAral4bAZSAtxDpME+SXh2GFjycM8xsR5n+9G7sGHP2oOG/Yd7zGWIEqLwgEeKIvoiJRv4CGsaBSMglEwCkY4AADNmTemk2xEegAAAABJRU5ErkJggg==","orcid":"","institution":"University of Hong Kong","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Yong-Xin","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2022-07-18 02:59:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1868011/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1868011/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s40168-023-01540-y","type":"published","date":"2023-04-27T20:36:26+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":24368016,"identity":"d68a3ad1-0bfd-46c0-b48d-d7c0a8dbddd5","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":279495,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of secondary metabolite biosynthetic capacity in LAB.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, Overall BGCs identified from 31,977 LAB genomes. The numbers outside the brackets indicate BGC count, and numbers represent the corresponding percentage and the count of genera in which BGCs are present. \u003cstrong\u003eb\u003c/strong\u003e, Layers are as follows: ①, the maximum likelihood phylogenetic tree based on 120 concatenate marker genes of 56 representative LAB genomes. LAB in this study covers 6 families and 56 genera under GTDB taxonomy; ②, the count of genomes included in this study, with being transformed with log2; ③, log2-transformed BGC count; ④, average BGC abundance in LAB genera; ⑤, average RiPPs abundance in LAB genera. In RiPPs term, subtypes not shown here or a combination of \u0026gt;1 subtypes are clustered into “Others”. rSAM-Modified RiPPs consist of RaS-RiPP, ranthipeptide and sactipeptide. Figures (a) and (b) share a common legend.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/589b643e31e95ff1c1085473.png"},{"id":24368015,"identity":"0dbca615-e7db-4234-ba0b-29b43643e8ff","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":297380,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLAB BGCs are diverse and taxa-specific.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, A total of 129,878 BGCs were grouped into 2,849 GCFs and 112 GCCs. The innermost dendrogram is the hierarchical clustering of 112 GCCs, on the basis of their average cosine distance to MIBiG BGCs. The next outer layer is the proportion of different genera. Genera with genome count \u0026lt; 200 are grouped into “Others”. The following layer is the proportion of BGC classes, which is proportionate to point size. The two outer layers refer to log2-transformed BGC count and average distance to MiBiG BGCs. The triangles denote clans dominated by particular class II bacteriocins-related domains (proportion \u0026gt; 80% in one clan). The predominant bacteriocins-related domains are shown in Supplementary Fig. 10. \u003cstrong\u003eb\u003c/strong\u003e, The bar plot shows the number of GCFs present in different genera (left), species (medial), and genomes (right).\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/78b6dc9be68cc43766c7ae9a.png"},{"id":24368020,"identity":"79a82c40-9cac-4198-9d12-58d940f34889","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":163258,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLAB BGCs in the human microbiome are variable and niche-specific.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, The prevalence and abundance of LAB genera in the six body sites microbial communities. \u003cstrong\u003eb\u003c/strong\u003e, GCF accumulation curve, reflecting the number of GCFs detected will increase as more samples are included. The numbers of metagenomes of six body sites are as follows: anterior nares, 66; stool, 100; posterior fornix, 169; supragingival plaque, 281; buccal mucosa, 66; tongue dorsum, 66. \u003cstrong\u003ec\u003c/strong\u003e, Number of GCFs detected in six body sites. Data are mean ± standard deviation, with only the upper error bar being shown. \u003cstrong\u003ed\u003c/strong\u003e, Boxplot showing the abundance of different BGCs detected in six body sites. \u003cstrong\u003ee\u003c/strong\u003e, The pie chart shows the proportion of 610 GCFs detected in how many sites. Corresponding percentages are shown in the brackets. The bar plot on the left refers to the number of GCFs in each site. The bar plot on the top depicts the number of GCFs of each intersection. Connecting lines are drawn if an intersection is present in more than one site. \u003cstrong\u003ef\u003c/strong\u003e, Numbers of genus/species/genome-specific GCFs and cross-genus/species/genome GCFs that were detected in one site (niche-specific) or more than one sites (cross-niche). The Chi-squared test gave \u003cem\u003eP\u003c/em\u003e values. *, 0.01 \u0026lt; \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05; ***, \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/65b3995dea05952be3f160c6.png"},{"id":24368017,"identity":"526fbfdb-f7eb-4705-98e9-acc40f1248e7","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":226658,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePutative compound activity of LAB BGCs.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, Performance of four machine learning classifiers [logistic regression, elastic net regression, support vector machines (SVM), and random forest] in determining compound activities using 10-fold cross-validation. The receiver operating characteristic (ROC) curves were based on aggregated performances of 10-fold cross-validation. Average AUROC was shown. \u003cstrong\u003eb\u003c/strong\u003e, Chord diagram showing the predicted activity of 129,878 BGCs. The scale was the proportion of each BGC class or predicted activity. The number shown in brackets refers to the BGC count and percentage relative to overall BGCs. Antibacterial-antifungal and antibacterial-cytotoxic represent BGCs encoding bifunctional SMs. \u003cstrong\u003ec\u003c/strong\u003e, Proportion of antibacterial SMs encoded by BGCs from LAB and non-LAB species. The proportion of antibacterial activity was calculated from randomly selected 10,000 BGCs of LAB or non-LAB, with be resampled 1,000 times. Data are mean ± standard deviation. \u003cem\u003eP\u003c/em\u003e value was given by Wilcoxon rank-sum test (two sided), with “***” denoting \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001. \u003cstrong\u003ed\u003c/strong\u003e, The proportion of putative activities of BGCs detected in six body sites. AN, anterior nares; St, stool; PF, posterior fornix; SP, supragingival plaque; BM, buccal mucosa; TD, tongue dorsum.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/14356acf9c34515baf9ea781.png"},{"id":24368019,"identity":"0489234f-f242-4651-b51f-a5ef44ff97cf","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":173380,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eClass II bacteriocins are structurally diverse and variably prevalent in human microbiome. a\u003c/strong\u003e, The number of putative precursors of class II bacteriocins detected by hmmsearch and BAGEL4. They identified 128,369 and 90,026 putative precursors, respectively, with 30,746 sequences in common. \u003cstrong\u003eb\u003c/strong\u003e, Distribution of the length of 2,005 representative precursors, which Cd-hit designated. \u003cstrong\u003ec\u003c/strong\u003e, Rarefaction curve of clusters of class II bacteriocin precursors. The Red line shows that 1,775 sequences belonging to 188 clusters were highly similar to known class II bacteriocins (identity\u0026gt;90%, coverage\u0026gt;95%). \u003cstrong\u003ed\u003c/strong\u003e, Number of clusters in different genera (left), species (middle), and genomes (right). \u003cstrong\u003ee\u003c/strong\u003e, t-SNE plot reveals the distinct profile of class II bacteriocins in different body sites. Each dot represents one metagenome sample. \u003cstrong\u003ef\u003c/strong\u003e, The prevalence and average abundance of 644 precursor clusters detected in six body sites. Each dot denotes one precursor cluster. The abundance in the individual is shown in Supplementary Fig. 17. The red numbers are the number of precursor clusters with a log2-transformed average abundance \u0026gt; -5 (shown in red dash line) and the number of precursor clusters detected in each body site.\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/458f67472899a6f4428362a5.png"},{"id":24368018,"identity":"3de48439-31a9-4816-bedf-eb441f857943","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":169663,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAntagonistic class II bacteriocins potentially play a regulating role in the vagina microbiome.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e, Correlation network between precursor clusters of class II bacteriocins and bacterial species in the vaginal microbiome. 23 clusters are negatively correlated with species in the community level, with spearman’s rho \u0026lt; -0.3 and adjusted\u003cem\u003e P\u003c/em\u003e \u0026lt; 0.05 shown in the network. The number in the node denotes the precursor cluster number. Av, \u003cem\u003eAtopobium vaginae\u003c/em\u003e; Dm, \u003cem\u003eDialister micraerophilus\u003c/em\u003e; Lc, \u003cem\u003eLactobacillus crispatus\u003c/em\u003e; Li, \u003cem\u003eLactobacillus iners\u003c/em\u003e; Lp, \u003cem\u003eLactobacillus paragasseri\u003c/em\u003e; Vb, \u003cem\u003eVeillonellaceae bacterium\u003c/em\u003e DNF00626. \u003cstrong\u003eb\u003c/strong\u003e, Spearman correlation between precursor clusters and alpha diversity (Shannon index). The dashed line denotes the correlation coefficient cutoff \u0026lt; -0.4 and adjust \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05. Points refer to 240 precursor clusters detected in the vaginal metagenomes, 21 of which were significantly associated with the alpha diversity of the vaginal microbiome. \u003cstrong\u003ec\u003c/strong\u003e, Global sequence identity to known class II bacteriocins (upper), sequence length (middle), abundance and prevalence of 21 clusters (bottom) in the vaginal metagenome (MG, n=169) and metatranscriptome (MT, n=180). \u003cstrong\u003ed\u003c/strong\u003e, Gene organization of bgc120802 and precursor sequences of cluster_467 and cluster_468. Putative double-glycines leader peptides are in grey. \u003cstrong\u003ee\u003c/strong\u003e, The minimum inhibitory concentration of chemically synthesized bacteriocins. NA: not available, no inhibitory effect detected with 200 μg/ml.\u003c/p\u003e\u003cp\u003e\u0026nbsp;\u003c/p\u003e","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/afbd224d3f3fcc57e044002f.png"},{"id":44727787,"identity":"0af71503-981d-4f75-87a3-8c8b2a28e68d","added_by":"auto","created_at":"2023-10-16 20:55:32","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1894945,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/d8984427-840d-4244-9e28-ebfe6b1ea8bf.pdf"},{"id":24368021,"identity":"ff81796d-51fe-426f-aaab-b22963a6d0bf","added_by":"auto","created_at":"2022-07-26 19:16:16","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":11643850,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-1868011/v1/c6fd629d1e45c59433ee8650.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"A systematically biosynthetic investigation of lactic acid bacteria reveals diverse antagonistic bacteriocins that potentially shape the human microbiome","fulltext":[{"header":"Introduction","content":"\u003cp\u003eLactic acid bacteria (LAB) are Gram-positive, microaerophilic bacteria, which have drawn extensive attention due to their fundamental roles in different biological processes\u0026nbsp;[1]. These bacteria feature lactic acid production in carbohydrate metabolism, which is important in food fermentation\u0026nbsp;[2]. Due to general safety features, they have also been engineered to produce food ingredients and pharmaceutical agents and deliver therapeutic molecules\u0026nbsp;[1\u0026ndash;4]. More importantly, a growing body of evidence reveals the health-promoting effects of consuming certain LAB strains, rendering them promising candidates for probiotics\u0026nbsp;[5]. Their probiotic actions may arise from multifaceted mechanisms, such as\u0026nbsp;gut microflora regulation, bioactive metabolite production, and immune system modulation, which endow LAB with a protective role for the host\u0026nbsp;[6]. Despite numerous studies focusing on characterizing LAB as probiotics for microflora regulation, how they impact microbiome homeostasis and host physiology is still not fully understood.\u003c/p\u003e\n\u003cp\u003eFundamentally, metabolic crosstalk within the kingdoms or with the host is the basis by which microbes, including LAB, engage with microbiome homeostasis. A myriad of metabolites produced by microbes, especially secondary metabolites (SMs, e.g. antibiotics and pigments), is among the mediums to achieve a high degree of crosstalk. Secondary metabolite-mediated interactions such as mutualism and antagonism are essential in maintaining microbiome homeostasis\u0026nbsp;[7,8]. Many LAB members, including \u003cem\u003eLactobacillus,\u003c/em\u003e \u003cem\u003eStreptococcus\u003c/em\u003e, and \u003cem\u003eLactococcus\u003c/em\u003e, produce bioactive SMs, ranging from bacteriocins nisin and\u0026nbsp;lactocillin\u0026nbsp;to tetramic acid reutericyclin\u0026nbsp;[9\u0026ndash;11]. Bacteriocins are ribosomally synthesized antimicrobial peptides with antibacterial potential and are generally divided into three classes (class I, post‐translationally modified peptides (e.g., nisin and lactocillin); class II, small unmodified peptides (e.g.,\u0026nbsp;amylovorin\u0026nbsp;L and crispacin A); class III, large, heat-labile peptides)\u0026nbsp;[12,13]. Moreover, bacteriocins, the most extensively studied SMs of LAB, have been recently revealed to play potential roles in shaping the microbiota by modulating microbial composition and inhibiting pathogens\u0026nbsp;[14].\u0026nbsp;However, these studies of LAB SMs have largely focused on structures, biosynthesis, mechanisms of action, or therapeutic potential in preventing infection in a case-by-case manner\u0026nbsp;[15,16].\u0026nbsp;The landscape of LAB SMs, particularly their diversity, prevalence, and potential roles in the human microbiome, remains elusive. Therefore, it is still unknown to what extent LAB-derived secondary metabolites are actively involved in microbiome homeostasis.\u003c/p\u003e\n\u003cp\u003eWith the development of bioinformatics techniques, the recent explosion of sequenced bacterial genomes and metagenomes provides fresh opportunities for large-scale biosynthetic analysis at both single species and community levels. Here we harnessed recent advances in biosynthetic and metagenomic analysis to investigate the untapped biosynthetic potential of LAB SMs systematically. Leveraging this comprehensive analysis of 31,977 LAB genomes and 748 human microbiome metagenomes, we gained previously undescribed insights into the biosynthetic capacity of LAB SMs and their diversity, abundance, and distribution in the human microbiome. We found that most BGCs may encode antagonistic SMs uncharacterized yet, exemplified by a new class II bacteriocin termed\u0026nbsp;crispacin 407, potentially playing protective roles in the human microbiome. To our best knowledge, this is the first largest survey of LAB biosynthetic potential and their profile in the human microbiome, linking them to the antagonistic contributions to microbiome homeostasis via omics analysis. Our omics-guided findings of the diverse and prevalent antagonistic bacteriocins in the human microbiome, particularly the vaginal microbiome, provide insight into antagonistic interactions linked to microbiome homeostasis and highlight the probiotic potential of LAB with antagonistic SMs.\u0026nbsp;\u003c/p\u003e\n"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eGenomic analysis reveals the landscape of SM biosynthetic potential of LAB\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGiven that environment or foods are possible LAB sources for the gut microbiome, to profile LAB SMs in the human microbiome, we first collected LAB genome data from different sources to comprehensively investigate the biosynthetic potential of LAB SMs.\u0026nbsp;Publicly available bacterial single amplified genomes (SAGs) and metagenome-assembled genomes (MAGs) of LAB were gathered from three databases (RefSeq\u0026nbsp;[17], PATRIC\u0026nbsp;[18], and IMG/M\u0026nbsp;[19]) and two previous studies\u0026nbsp;[20,21], resulting in 40,879 SAGs and 4,575 MAGs in total (Supplementary Table 1). Genomes were then de-duplicated, and their taxonomical classifications were verified and unified using GTDB taxonomy. As a result, 31,977 LAB genomes (27,549 SAGs and 4,428 MAGs, Supplementary Table 2), spanning six families containing 56 genera, were retained for global biosynthetic analysis of LAB SMs (Supplementary Fig. 1a). Using a rule-based BGC detection tool, antiSMASH 6.0\u0026nbsp;[22], we identified 130,051 BGCs from 30,718 genomes (Fig. 1a, Supplementary Fig. 1b, Supplementary Table 3), including 1,333 nonribosomal peptide synthetase clusters (NRPS, 1.0%), 25,278 polyketides synthase clusters (PKS, 19.7%), 98,810 ribosomally encoded and post-translationally modified peptides (RiPPs, 76.0%), 1,629 terpene (1.3%) and 2,984 BGCs (2.3%) encoding other types of metabolites (Supplementary Fig. 2). The BGCs per genome ranged from 0 to 14, with an average of 4.07. Among the most abundant RiPPs, RiPP-like (formerly annotated as bacteriocin by antiSMASH) topped its list with 72,471 (55.7% of total BGCs). Using BiG-SLiCE\u0026nbsp;[23]\u0026nbsp;to extract bacteriocin biosynthesis-related domains from 72,471 RiPP-like BGCs (Supplementary Fig. 3 and Supplementary Table 4), we identified 60,497 class II bacteriocins (RiPP-like BGCs that contain class II bacteriocins-related domains identified by BiG-SLiCE) (46.5%), which are the most abundant LAB SMs.\u0026nbsp;Also worth noting was that 99.8% (25,224/25,278) of PKS were type III PKS (T3PKS) (Fig. 1a).\u003c/p\u003e\n\u003cp\u003eTo gain insight into the phylogenetic distribution of BGC in LAB genera, we examined 26,983 BGC-containing SAGs, excluding MAGs due to their incompleteness. The biosynthetic capacity varied considerably at the families, genera, or species level (Fig. 1b, Supplementary Fig. 4). The RiPPs and T3PKS BGCs dominated all LAB genera except for \u003cem\u003eTetragenococcus\u003c/em\u003e, while terpene BGCs and NRPS BGCs were sporadically distributed in those genera (Fig. 1, Supplementary Figs. 5, 6). Among 55 genera, we found a median of \u0026ge;1 RiPPs per genome in 28 genera and one T3PKS in 33 genera. While a high proportion of RiPPs has been reported in Firmicutes [24], T3PKS dominating in certain LAB genera might encode specialized metabolites that have basic biological functions. Of note, despite being small in genome size, Streptococcaceae generally harbored more abundant BGCs than other families (Supplementary Fig. 4a), with a median of five BGC per genome, exemplified by genera \u003cem\u003eLactococcus\u003c/em\u003e and \u003cem\u003eStreptococcus\u003c/em\u003e. Contrastingly, 33 genera only harbored a median of \u0026lt;=1 BGC per genome, indicating the limited biosynthetic capacity of the LAB majority (Fig. 1b). By comparing 53 LAB genera and 3,805 non-LAB genera (164,417 genomes, Supplementary Table 5), we found comparatively limited biosynthetic capacity in LAB and a significantly strong correlation (Spearman \u003cem\u003erho\u003c/em\u003e = 0.712, \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001) between bacterial biosynthetic potential and their genome size (Supplementary Fig. 7). The small genomes and reduced biosynthetic capacities in LAB might correlate with their adaptation to nutritionally-rich niches.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMost LAB BGCs are species-specific and uncharacterized yet\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAlthough BGCs highly vary in gene content, grouping them into families (GCFs) or clans (GCCs) based on architectural relationships of biosynthetic elements is an effective way to uncover the similarity of their encoding products in terms of the chemical features and biological functions\u0026nbsp;[25]. To gain insight into the novelty and diversity of 130,051 LAB BGCs, we extracted BGC features (biosynthetic domains) using BiG-SLiCE\u0026nbsp;[23]\u003csup\u003e\u0026nbsp;\u003c/sup\u003eand grouped them based on an all-to-all cosine distance among BGCs [26]. The 129,878 BGCs with features were classified into 2,849 GCFs and 112 GCCs, with a distance threshold of 0.2 and 0.8, respectively (Fig. 2, Supplementary Fig. 8). We further compared the 112 GCCs to the reference known BGCs described in the \u0026apos;Minimum Information about a Biosynthetic Gene\u0026apos; (MIBiG) repository [27]. Notably, only three clans, linear azol(in)e-containing peptides (LAP, GCC_29), RiPP-like (GCC_84), and class II lanthipeptides (GCC_110), were closely similar to known BGCs (average cosine distances \u0026lt; 0.2), leaving the vast majority unknown. This highlights the huge knowledge gap of LAB SMs and demonstrates the potential for discovering novel chemistry from the LAB. Of note, the majority of NRPS (73.2%), terpene (99.3%) and T3PKS (99.3%) were clustered into one respective clan. In contrast, RiPP BGCs, contributing 83 GCCs with 1,818 GCFs (RiPP proportion \u0026gt; 80% in GCCs/GCFs), were highly diverse due to the diversity of their post-translational modification (PTM) enzyme genes and adjacent genes (Fig. 2a). Among them, 621 GCFs of 23 RiPP-like GCCs encoded class II bacteriocins, representing one of the most diverse LAB SMs. The other 60 RiPP GCCs mainly encoded class I bacteriocins, including lanthipeptide, LAP, lassopeptide, and rSAM-modified RiPPs.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;In prokaryotic genome evolution, the conserved genes cross genera are more likely to contribute to essential ecological processes, whereas species- or even strain-specific genes often arise from natural selection, thus enhancing niche adaptation or host fitness [28]. In this context, we next examined the distribution and diversity of genus- or species-specific BGCs. While GCC clustering shows the distribution and novelty of LAB SMs, a fine resolution of GCF clustering can offer an insight into the diversity of BGCs that are predicted to encode similar natural products [29].\u003cem\u003e\u0026nbsp;\u003c/em\u003eWe found that the majority of GCFs were genus-specific (92.6%, 2,637/2,849) and species-specific (75.8%, 2,159/2,849). Remarkably, 1,165 GCFs (40.9%) contained only one BGC harbored by a specific strain (Fig. 2b). In contrast, only 7% (212) were cross-genus GCFs, including 142 RiPPs (present in 2-20 genera), 17 NRPS (in 2-3 genera), 21 T3PKS (in 2-9 genera), and 9 terpenes (in 2-7 genera) (Supplementary Figs. 9, 10). Among these 142 cross-genus RiPP GCFs, 62 are class II bacteriocins. Owing to this high GCF diversity between genera, we did not observe a phylogenetic relationship in GCF presence/absence (Supplementary Fig. 8). These taxa-specific BGCs usually encode specialized SMs and provide a competitive edge to the producer for niche adaption. Considering the wide presence of LAB in different niches [30], a high proportion of species- and strain-specific BGCs might result from niche selection.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLAB BGCs are diverse and niche-specific in the human microbiome\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNumerous studies have revealed the variable prevalence of LAB species in the human microbiome [21], raising the question of to what extent their SMs vary in different body sites for niche adaption. Thus, we next explored the profile of LAB BGCs in the healthy human microbiome by re-visiting 748 metagenomes of six body sites from the Human Microbiome Project (HMP) [31]. These sites included aerobic (anterior nares, representing skin microbiome), microaerobic (supragingival plaque, buccal mucosa, tongue dorsum, and posterior fornix, representing oral and vaginal microbiome), and anaerobic (stool, representing gut microbiome) environments (Supplementary Table 6). In line with a previous larger-scale study [21], LAB exhibited variable abundance and prevalence in different body sites (Fig. 3a). Of note, the genus \u003cem\u003eLactobacillus\u0026nbsp;\u003c/em\u003edominated in the vagina with a median abundance of 99.0% [interquartile range (IQR), 91.0%-99.8%], whereas \u003cem\u003eStreptococcus\u003c/em\u003e was moderately abundant but highly prevalent in six body sites.\u003c/p\u003e\n\u003cp\u003eTo profile the diversity, abundance, and distribution of LAB BGCs in the human microbiome, we de-duplicated 130,051 BGCs to 24,222 representative BGCs and mapped metagenomic reads to 24,222 nonredundant BGCs. The number of LAB BGCs detected in six body sites varied considerably, with the highest in the oral cavity and the lowest in the skin (Supplementary Fig. 11a), probably due to the variable abundance of LAB. From 748 metagenomes, we detected 5,687 BGCs of 610 GCFs, including 71 T3PKS and 312 RiPPs with 92 class II bacteriocins (Supplementary Fig. 11b). The GCF accumulation curve indicated that more GCFs would be detected in those body sites as more samples were included, revealing the huge diversity of LAB SMs in the human microbiome (Fig. 3b). The three oral sites were the richest in GCFs (averaging 29, 38, and 54). Compared to the oral cavity, the vagina harbors a lower diversity of GCFs (averaging 12) but a significantly higher abundance of LAB BGCs (Figs. 3c, d). Particularly, the vaginal microbiome harbored a high abundance of class II bacteriocins, lassopeptide, lanthipeptide, and LAP. Of note, influenced by sequencing depth, the diversity and abundance of LAB BGCs in the human microbiome may be underestimated. Of 610 detected GCFs, ~52% were niche-specific in one of six sites, which accounted for 18% - 38% of GCFs in a particular site (Fig. 3e). We also observed that those niche-specific GCFs were generally species-specific (Chi-squared test, \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001), but not genus-specific (\u003cem\u003eP\u003c/em\u003e = 0.026) nor strain-specific (\u003cem\u003eP\u003c/em\u003e = 0.013) (Fig. 3f). This result indicated that niche-specific GCFs were derived from different species residing in distinct niches, which may provide a competitive advantage to the niche adaptation of their hosts. Our genomic and metagenomic analysis of biosynthetic potential revealed that the LAB SMs are diverse and variably prevalent in the human microbiome.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMachine learning models reveal that most BGCs may encode antagonistic SMs\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGiven the abundance and prevalence of LAB BGCs in the human microbiome, we next want to study the potential bioactivities of BGC-encoding SMs. The bioactivity of SMs encoded by BGCs was recently predicted using machine learning strategies based on chemical fingerprints of predicted compound structure, protein family (PFAM) domains, and other genetic features\u0026nbsp;[32\u0026ndash;34]. Here, we adapted four common machine learning classifiers (logistic regression, elastic net regression, random forest, and support vector machines) to predict the bioactivities of LAB-derived SMs. For the training data (950 known BGCs, Supplementary Table 7, Supplementary Fig. 12), ten-fold cross-validation revealed that the random forest classifier outperformed others with an average area under the receiver operating characteristic curve (AUROC) being 0.76, 0.80, and 0.82, for antibacterial, antifungal, and antitumor or cytotoxic, respectively (Fig. 4a, Supplementary Fig. 13). The performance of the random forest classifier was comparable with previously reported methods\u0026nbsp;[32,33].\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; Using the random forest classifier, 129,878 LAB BGCs with features were predicted to encode different bioactive SMs, comprising antibacterial (n=123,587, 95.2%), cytotoxic (n=2,268, 1.8%), antibacterial-antifungal (n=79, 0.1 %), antibacterial-cytotoxic (n=2,548, 2.0%), unknown (1,395, 1.1%). Most BGCs, regardless of BGC classes, were predicted to be antibacterial (Fig. 4b). Of note, 97.8% of RiPPs (96,466/98,637) were predicted to exhibit antibacterial activity, implying that bacteriocins were plenteous in LAB. With predicted antibacterial activity, most RiPPs with known post-modifications were classified as class I bacteriocins, while most RiPPs-like (83.4%) were class II bacteriocins. Almost 100% of class II bacteriocins (60,494/60,497) were captured as antibacterial SMs, contributing to 48.1% of LAB-derived antibacterials. Those antibacterial BGCs dominated almost all LAB genera (Supplementary Fig. 14), possibly conferring a competitive edge in the microbial community. Compared to BGCs (n=1,121,156) identified from non-LAB genomes, LAB-derived BGCs potentially encoded a significantly higher proportion of antibacterial SMs, indicating a higher antagonistic potential of LAB SMs (Wilcoxon rank-sum test, \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.001) (Fig. 4c, Supplementary Fig. 15a, b). Low percentage of LAB-derived BGCs encoding putative cytotoxic or antifungal SMs were found in specific species (Supplementary Fig. 15c, d). For example, a certain family of LAP (GCF_199) possibly conferring cytotoxic activity were distributed in ten \u003cem\u003eStreptococcus\u003c/em\u003e species, especially \u003cem\u003eStreptococcus pyogenes\u003c/em\u003e, in which common pathogenicity feature endowed by conserved\u003cem\u003e\u0026nbsp;\u003c/em\u003eLAP had been reported [35]. In the six body sites, almost all BGCs potentially encoded antimicrobials (Fig. 4d), possibly mediating bacterial antagonism for maintaining microbiome homeostasis. We speculated that LAB harboring antimicrobial SMs, especially class II bacteriocin, could protect the host from pathogen invasion and maintain the microbiome homeostasis via antagonistic interaction [7].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eUnderexplored class II bacteriocins are widely distributed in the human microbiome\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe findings of the antagonistic potential of class II bacteriocins and their variable prevalence and predominance in the human microbiome raise the question of what extent the class II bacteriocins may link to microbiome homeostasis. We next attempted to group class II bacteriocins into subfamily with similar biological functions based on precursor sequence space and investigate their profiling in the human microbiome in detail\u003cem\u003e.\u003c/em\u003e To fully reveal the chemical diversity of class II bacteriocin, we first adapted two approaches for the identification of precursor peptides, with hmmsearch [36] to search Pfam domains of precursors of class II bacteriocin (Supplementary Table 4) and with BAGEL4 [37] which is a tool specifically designed for bacteriocin mining. We combined two approaches to identify 187,649 precursors from class II bacteriocin BGCs (Fig. 5a, Supplementary Table 8). We then grouped 187,649 putative precursors into 2,005 clusters with a threshold of 50% sequence identity (Supplementary Fig. 16). The sequence lengths of those representative precursors were approximately normally distributed, with a center of ~55 amino acids (Fig. 5b). The accumulation curve showed that the precursor diversity increased with the number of genomes included, indicating more class II bacteriocins will be disclosed with more genomes sequenced (Fig. 5c). Moreover, we found that only 188 clusters were similar to 333 known class II bacteriocins (identity\u0026gt;90%, coverage\u0026gt;95%), leaving the vast majority (1,817/2,005) underexplored (Supplementary Table 9). Of note, while the rule-based method hmmsearch and BAGEL4 enable a high likelihood of positive detection, at the same time, they probably underestimate the real biosynthetic potentials of bacteriocins.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;In line with the taxa-specificity of the GCFs, most class II bacteriocin precursors family were genus-specific (n=1,862, 92.9%) and species-specific (n=1,327, 66.2%), with 33.7% of precursor family being even strain-specific (n=675) (Fig. 5d). We next examined their profile in the human microbiome. Of 644 clusters detected in six body sites, about 31.8% of clusters were niche-specific in one of six sites (Supplementary Fig. 17a). Moreover, the profiles of class II bacteriocins in different body sites were distinct as revealed by t-SNE plot, further supporting the niche-specificity of class II bacteriocin in the human microbiome (Fig. 5e). Their profiles in vagina and skin showed great individual variations, whereas the class II bacteriocins in other body sites were relatively conserved with being clustered together. Additionally, those class II bacteriocins were sporadically present in the skin and gut, whereas some class II bacteriocins were particularly enriched in the oral cavity and vagina with a high prevalence and abundance (Fig. 5f, Supplementary Fig. 17b). Probably due to the individual variations in the vagina, a member of subfamilies of class II bacteriocins exhibited a relatively smaller prevalence in the vagina than in the oral. Both the GCFs profile (Fig. 3d) and precursors profile (Figs. 5e, f) in the human microbiome suggested that class II bacteriocins are particularly enriched in the vaginal microbiome. Considering the vagina has simple communities with the lowest alpha diversity than other body sites [38], we reasoned that those enriched and predominant class II bacteriocins might play prominent roles in regulating microbial community in the vagina.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMulti-omics analysis revealed class II bacteriocins potentially contributing to vaginal microbiome homeostasis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo examine the class II bacteriocins that may account for the homeostasis of the vaginal microbiome, we first constructed the association network between class II bacteriocins and bacterial species at the metagenomic level. We found 23 precursor clusters correlated negatively with various species, indicating their antagonistic potential in regulating the vaginal microbiome (Fig. 6a). In particular, 21 clusters were negatively correlated with \u003cem\u003eLactobacillus iners\u003c/em\u003e, which is more conducive to the occurrence of abnormal vaginal microflora and thus a potential new therapeutic target for bacterial vaginosis treatment. Additionally, 21 of 23 clusters were also found to be inversely correlated to the Shannon index (Spearman \u003cem\u003erho\u003c/em\u003e \u0026lt; -0.4, adjusted \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05) (Fig. 6b). Lower bacteria diversity in the microbiome with these detected class II bacteriocins suggested their regulative role in shaping the microbiome. To confirm whether these antagonistic bacteriocins are biologically functional, we next inspected their expression profile in the 180 metatranscriptomic datasets (Supplementary Table 6) and found that most of them were actively transcribed in the vaginal microbiome of healthy individuals (Fig. 6c). Three of the 21 clusters grouped with known class II bacteriocins, including Amylovorin L (cluster_342 and cluster_346, a two-component class IIb bacteriocin) from \u003cem\u003eLactobacillus amylovorus\u003c/em\u003e DCE 471 [39] and gassericin T (cluster_94) from \u003cem\u003eLactobacillus gasseri\u0026nbsp;\u003c/em\u003eSBT2055\u0026nbsp;[40]. The findings of known class II bacteriocins with protective roles by omics-based associated analysis further validated the effectiveness of our approach in discovering the regulatory bacteriocins in the microbiome. The other 18 clusters of class II bacteriocins were also prevalent and actively transcribed in the vagina microbiome but uncharacterized yet.\u003c/p\u003e\n\u003cp\u003eWe next sought to validate the antagonistic potential of those uncharacterized bacteriocins experimentally. For proof of principle, we selected two precursor clusters (cluster_467 and cluster_468) with high abundance and a short peptide length that make their chemical synthesis practical. Those two precursors were located on BGCs (e.g, bgc120802) identified from 16 genomes of \u003cem\u003eL. crispatus\u0026nbsp;\u003c/em\u003e(Supplementary Fig. 18). The bgc120802 harbors specific class II bacteriocins-related genes, including a two-component regulator system (histidine kinase and response regulator), ABC transporter, and immunity protein. The precursor sequences from 16 BGCs were identical and featured a canonical double-glycine leader (Fig. 6d, Supplementary Fig. 18). We thus synthesized the core peptides (Supplementary Fig. 19), namely crispacin 467 (27 amino acids) and crispacin 468 (30 amino acids), respectively, and validated their antagonistic activity toward bacteria and fungi. The antimicrobial assay showed that the crispacin 467 exhibited a narrow-spectrum antibacterial activity against phylogenetical-closely related strain \u003cem\u003eL. delbrueckii\u003c/em\u003e subsp. \u003cem\u003ebulgaricus\u003c/em\u003e with\u003cem\u003e\u0026nbsp;\u003c/em\u003ea\u003cem\u003e\u0026nbsp;\u003c/em\u003eminimum inhibitory concentration of 12.5 \u0026mu;g/mL, while inhibitory effects of crispacin 468 were not observed, nor a synergy of them (Fig. 6e). The crispacin 468 might exhibit antimicrobial activity against other species beyond the tested strains. Taken together, we believe that the bacteriocin producers arm the vagina microbiome with diverse antagonistic bacteriocins, potentially preventing pathogen invasion and stabilizing the microbial community. Though how LAB employs SMs to shape their microbiome communities is not yet fully understood, our omics-guided discovery of new bacteriocins from LAB provides an alternative way to the discovery of new antibacterial therapeutics for microbiome dysbiosis.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eDespite increasing evidence revealing the health-promoting effects of LAB in the human microbiome, how they interplay with other microbes and influence microbiome homeostasis in the human is still understudied.\u0026nbsp;Previous biosynthetic analysis of LAB in a limited dataset focuses on particular metabolites, proposing their protective roles to host\u0026nbsp;[41,42]. However, the landscape of LAB SMs, particularly their profiles and potential roles in the human microbiome, remains elusive.\u0026nbsp;In this study, we conducted a comprehensive omic analysis for LAB BGCs, significantly enhancing our understanding of the diversity and distribution of LAB BGCs. We found\u0026nbsp;129,878 BGCs of\u0026nbsp;2,849 GCFs\u0026nbsp;from 31,977 LAB genomes, most of which were species-specific and encoded diverse uncharacterized SMs.\u0026nbsp;We further investigated human metagenomes of six body sites to disclose the BGCs profile of LAB in the human microbiome, revealing that the diverse LAB SMs are diverse and variably prevalent in the human microbiome. Of note, BGCs of class II bacteriocins were particularly enriched and predominant in the vaginal microbiome.\u0026nbsp;The niche specificity of GCFs in the human microbiome together with their specific- or even strain-specificity, suggested that the LAB SMs may\u0026nbsp;provide a competitive advantage to the niche adaptation of their producing hosts.\u0026nbsp;To profile BGCs in the human microbiome, we grouped them into families and clans based on architectural relationships, a well-accepted approach to study the BGC similarity and prioritize novel BGCs for natural product discovery\u0026nbsp;[25,43]. Although informative, the GCF grouping will be affected by the imperfect BGC boundary prediction of antiSMASH\u0026nbsp;[22].\u0026nbsp;Most LAB SMs were predicted to be antibacterial using machine learning models, indicating their potential regulating roles in the human microbiome and putative protective roles for the host. However,\u0026nbsp;due to limited training data, our machine learning model can only predict limited bioactivities (antibacterial, antifungal, and antitumor), underestimating another biological potential of LAB SMs. Although such evidence does not exclude the possibility of other biological functions, we believe that antagonistic LAB SMs, particularly bacteriocins, potentially provide competitive edges to their producers and regulate the microbiome community.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; Applying metagenomics and metatranscriptomics analysis, we underscored 21 class II bacteriocins actively expressed in the vaginal microbiome and negatively correlated with individual bacteria species. Together with their negative association with the \u0026alpha;-diversity of the vaginal microbiome, we can envision that these bacteriocins play prominent roles in regulating homeostasis. For proof of principle, we identified a novel class II bacteriocin produced by \u003cem\u003eL. crispatus\u003c/em\u003e, namely\u003cem\u003e\u0026nbsp;\u003c/em\u003ecrispacin 467, which exhibited characteristic narrow-spectrum antibacterial activity against phylogenetical-closely related strains. Our results suggested that crispacin 467 may arm \u003cem\u003eL. crispatus\u003c/em\u003e with a protective role in vaginal health [44]. Although previous analyses have disclosed the bacteriocin biosynthetic genes and antibacterial activity in \u003cem\u003eL. crispatus\u0026nbsp;\u003c/em\u003e[45,46], little is known regarding their antagonistic bacteriocin except for crispacin A [47], not to mention their potential roles in the microbiome. A myriad of class II bacteriocins are harbored in \u003cem\u003eLactobacillus\u003c/em\u003e and the entire LAB of the human microbiome, providing enormous potential for new antimicrobial discovery. While the precursors of 21 bacteriocins were prevalent and transcribed in the vaginal microbiome, whether they are produced in situ in the vagina still needs to be examined by metabolomics. Meanwhile, how LAB employ SMs to shape their microbiome communities needs to be further explored in the future using in vivo mouse models or in vitro polymicrobial models.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLAB, especially genus \u003cem\u003eLactobacillus\u003c/em\u003e, dominate the existing probiotics that confer a health benefit on the host when administered adequately\u0026nbsp;[48]. The beneficial effects of probiotics result from diverse mechanisms, among which is bacteriocin production. Antagonistic bacteriocins that assist the producer colonization and provide protective roles for the host are important for probiotics to offer beneficial effects. Our findings reinforced the understanding of pervasive bacteriocins of LAB, which are also widespread in the human microbiome, particularly in the vagina. Those LAB and their bacteriocins present in healthy individuals are promising priorities for microbiome-based therapeutics\u0026nbsp;[49]. For example, with the potential to modulate the vaginal microbiota, probiotics containing\u003cem\u003e\u0026nbsp;Lactobacillus\u003c/em\u003e spp. have been applied to treat bacterial vaginosis [44]. Moreover, the diverse antagonistic bacteriocins could be borrowed and arm genetically engineered beneficial LAB probiotics with a therapeutic property [50], preventing the host from the pathogen invasion directly or indirectly (via microbiota- and/or immune modulation). Continued investigation of the biosynthetic capacity and ecological roles will help to facilitate the translation of LAB and their antagonistic SMs into clinical application.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIn summary, our study provides a global insight into the biosynthetic potentials of LAB SMs and a starting point for the omics-guided discovery of antagonistic SMs that potentially regulate microbiome homeostasis. Class II bacteriocin predominant in vaginal microbiome but negatively associated with its bacterial diversity is experimentally validated to play antagonistic roles in microbial communities. To the best of our knowledge, our study is the first to systematically unveil LAB SM biosynthetic potentials and their profile in the healthy human microbiome. However, the analysis presented here cannot be considered exhaustive. The machine learning strategies employed to predict the bioactivity of SMs remain refined by knowledge accumulation of LAB SMs and their biosynthesis and bioactivity. Additionally, how LAB employ SMs to shape their microbiome communities in the human niche remains to be studied. Nevertheless, our systematic investigation of the biosynthetic potential of LAB provides a good starting point for the omics-guided discovery of SMs with therapeutic potential from the human microbiome. In addition to enhancing our understanding of the profile of LAB SMs and their potential regulating roles in the human microbiome, the discovery of antagonistic bacteriocins opens up exciting opportunities for future research on various probiotic applications of LABs.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eData acquisition\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAs defined early, lactic acid bacteria include 14 genera, comprising \u003cem\u003eLactobacillus\u003c/em\u003e, \u003cem\u003eLactococcus\u003c/em\u003e, \u003cem\u003eLeuconostoc\u003c/em\u003e, \u003cem\u003ePediococcus\u003c/em\u003e, \u003cem\u003eStreptococcus\u003c/em\u003e, \u003cem\u003eAerococcus\u003c/em\u003e, \u003cem\u003eAlloiococcus\u003c/em\u003e, \u003cem\u003eCarnobacterium\u003c/em\u003e, \u003cem\u003eDolosigranulum\u003c/em\u003e, \u003cem\u003eEnterococcus\u003c/em\u003e, \u003cem\u003eOenococcus\u003c/em\u003e, \u003cem\u003eTetragenococcus\u003c/em\u003e, \u003cem\u003eVagococcus\u003c/em\u003e, and \u003cem\u003eWeissella\u0026nbsp;\u003c/em\u003e[51]. Particularly, genus \u003cem\u003eLactobacillus\u0026nbsp;\u003c/em\u003ehas been reclassified recently\u0026nbsp;[52], extending to 25 genera consisting of \u003cem\u003eLactobacillus\u003c/em\u003e, \u003cem\u003eParalactobacillus,\u003c/em\u003e \u003cem\u003eAmylolactobacillus,\u003c/em\u003e \u003cem\u003eAcetilactobacillus\u003c/em\u003e, \u003cem\u003eAgrilactobacillus\u003c/em\u003e, \u003cem\u003eApilactobacillus\u003c/em\u003e, \u003cem\u003eBombilactobacillus\u003c/em\u003e, \u003cem\u003eCompanilactobacillus\u003c/em\u003e, \u003cem\u003eDellaglioa\u003c/em\u003e, \u003cem\u003eFructilactobacillus\u003c/em\u003e, \u003cem\u003eFurfurilactobacillus\u003c/em\u003e, \u003cem\u003eHolzapfelia\u003c/em\u003e, \u003cem\u003eLacticaseibacillus\u003c/em\u003e, \u003cem\u003eLactiplantibacillus\u003c/em\u003e, \u003cem\u003eLapidilactobacillus\u003c/em\u003e, \u003cem\u003eLatilactobacillus\u003c/em\u003e, \u003cem\u003eLentilactobacillus\u003c/em\u003e, \u003cem\u003eLevilactobacillus\u003c/em\u003e, \u003cem\u003eLigilactobacillus\u003c/em\u003e, \u003cem\u003eLimosilactobacillus\u003c/em\u003e, \u003cem\u003eLiquorilactobacillus\u003c/em\u003e, \u003cem\u003eLoigolactobacilus,\u003c/em\u003e \u003cem\u003ePaucilactobacillus\u003c/em\u003e, \u003cem\u003eSchleiferilactobacillus\u003c/em\u003e, and \u003cem\u003eSecundilactobacillus\u003c/em\u003e. Filtered with taxonomy, genomes from these 38 genera were then retrieved from NCBI reference sequences (RefSeq) database\u0026nbsp;[17]\u0026nbsp;(as of Aug. 2021, including SAGs only), PATRIC database (including SAGs)\u0026nbsp;[18], IMG/M database (including SAGs and MAGs)\u0026nbsp;[19]. Besides them, genomes from two previous studies focusing on the human gut microbiome (including SAGs and MAGs)\u0026nbsp;[20]\u0026nbsp;and food-originated LAB (including MAGs)\u0026nbsp;[21]\u0026nbsp;were also included. To avoid the reference genome redundancy, genomes from RefSeq were compared to themselves and those from other sources using Mash v2.3\u0026nbsp;[53]. Genomes with a Mash distance of 0 were considered identical. Only the one with a minimal number of contigs was retained. As potential misclassification might be present, GTDB-Tk v1.7.0\u0026nbsp;[54]\u0026nbsp;was further used to confirm and unify taxonomic annotation against GTDB-Tk reference data version r202\u0026nbsp;[55]. There is a slight difference between NCBI taxonomy and GTDB taxonomy\u0026nbsp;[56]. Under GTDB taxonomy, five genera are sub-divided: \u003cem\u003eCarnobacterium\u003c/em\u003e, \u003cem\u003eEnterococcus\u003c/em\u003e, \u003cem\u003eLactococcus\u003c/em\u003e, \u003cem\u003eVagococcus\u003c/em\u003e, and \u003cem\u003eWeissella\u003c/em\u003e. Finally, a total of 56 genera belonging to six families (Lactobacillaceae, Aerococcaceae, Streptococcaceae, Vagococcaceae, Enterococcaceae, and Carnobacteriaceae) were considered as members of LAB in this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBiosynthetic gene cluster analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBiosynthetic gene clusters for each genome were annotated by antiSMASH 6.0\u0026nbsp;[22]\u0026nbsp;with default parameters. In total, 31,977 LAB genomes and 164,417 non-LAB genomes (the intersection between RefSeq and GTDB repository version r202)\u0026nbsp;[57] were included for BGC annotation. This resulted in 130,051 BGCs from 30,718 LAB genomes and 1,122,204 BGCs from 155,540 non-LAB genomes. No BGCs were annotated in 1,259 LAB genomes and 8,877 non-LAB genomes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClustering BGCs into families and clans\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBiG-SLiCE\u0026nbsp;[23], a tool to cluster sizable BGCs, contains two BGC features (biosynthetic-Pfam and sub-Pfam domains). Those BGC features are sufficient to distinguish distinct BGC classes. As a previous study described\u0026nbsp;[26], all features of LAB BGCs and 1910 experimentally validated BGCs from the MIBiG 2.0 repository were extracted by BiG-SLiCE\u0026nbsp;v1.1.0, and subsequently used to compute all-to-all\u0026nbsp;cosine distances between BGCs using Python suite SciPy version 1.6.2\u0026nbsp;[58]. The cosine distances were next subject to hierarchical clustering with average linkage, grouping BGCs into families (GCFs, distances \u0026lt; 0.2) and clans (GCCs, distances \u0026lt; 0.8) by Python 3.8 with Scikit-learn version 0.24.2\u0026nbsp;[59].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMetagenomics and metatranscriptomics analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe raw metagenomic sequencing reads of 748 HMP samples [31] and the raw metatranscriptomic data of 180 vaginal samples [60] were acquired from NCBI SRA (Sequence Read Archive) [61] under project accession number PRJNA48479 and PRJNA797778, respectively. Fastp 0.21.1 [62] with default parameters was adopted for detecting and removing low-quality sequencing reads. High-quality metagenomic sequencing reads were subjected to kneaddata (https://github.com/biobakery/kneaddata) for discarding reads belonging to the human host, through searching against the human reference genome (GRCh38.p13) from GENCODE [63]; high-quality metatranscriptomic reads were also subjected to SortMeRNA v4.3.4 [64] for removing reads derived from ribosomal RNAs. Following that, MetaPhlAn v3.0.13 [65] was used for taxonomic profiling. Prior to assessing the abundance of BGCs in metagenomics and metatranscriptomic data, we used a modified script from BiG-MAP [66] to de-duplicate 130,051 BGCs. To reduce the computational load, we de-duplicated them within each GCF, at a 0.8 nucleotide identity threshold, leading to 24,222 non-redundant BGCs, the nucleotide sequences of which were used to generate the reference database. Next, the non-host metagenomic and metatranscriptomic reads were mapped to this BGC reference using Bowtie 2 v2.3.5.1 [67], with a parameter of \u0026ldquo;-k 1\u0026rdquo;. We then utilized featureCounts v2.0.3 [68] (with parameters of \u0026ldquo;-T 30 -f -p -B -C -t CDS -g ID -M -O --fracOverlap 0.2\u0026rdquo;) to assign sequencing reads to the BGC genes. When calculating the abundance of a BGC, we only considered the core and additional biosynthetic genes, excluding the other genes such as transporters, regulators, transposases, and so forth. For each BGC, a corresponding GTF (General Transfer Format) annotation file was generated by antiSMASH. We retrieved the biosynthetic-related genes (tag \u0026ldquo;biosynthetic\u0026rdquo; for the core biosynthetic genes and \u0026ldquo;biosynthetic-additional\u0026rdquo; for the additional biosynthetic genes) according to the \u0026ldquo;gene_kind\u0026rdquo; tag in GTF files. A BGC was considered present in a metagenomic sample when fulfilling the following criteria: (1) the percentage of biosynthetic-related genes detected is over 50% of total biosynthetic-related genes in a BGC; (2) at least one core biosynthetic gene was found in a BGC. The abundance of a BGC was computed via the equation (1):\u003c/p\u003e\n\u003cp\u003e\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAUsAAAA3CAYAAACFIGiQAAAPrklEQVR4nO2d3YsT1//H3/vjd1vb2XhVqBeZXFQsbCmzu9CawhbqpFakUiXZWopQcTsjFCzFpel65WNCsVCoyRYWpLQmitJeNGGzQrzIVHRNlwy17IU7g8ji1Yzb1T/gfC/kTGeSSTZPmsT9vGAgmYdzTjLnfOZ8ns4MMcYYCIIgiIb8X68bQBAEMQiQsCQIgmgCEpYEQRBNQMKSIAiiCUhYEgRBNMH/97oBBNFrNE0DAGzfvh2BQAC2bWN5eRkAsHPnzl42jegjaGZJbHoOHTqEcDiMubk5AMDq6ir27t2LQ4cO9bhlRD8xRHGWxKBhmiYePnyIV199FcFgsOPykskk1tfXcfnyZaysrAAAYrEYPvvsM+zevbvj8okXA5pZEgPH4uIi9u7di8XFxa6UVywW8dVXXwEAstksAGBhYcEjKE3TRDKZ7Ep9xGBCwpIYOGKxGERRxNjYGEzThKZpME2zrbJs24Zt2wgEAjhy5AguXrwIXdcxNjZWc+7LL7/cadOJAYaEJTFw2LaNtbU1BINBXL9+HeFwGN99911bZd2+fRu7du0CAHz++ecoFAq4cuUK3nvvPeccTdNw9epV7NixoyvtJwYTEpbEwLG8vIxQKOSozKVSCRcuXGi5HNu28cMPPzjfA4EAFEXBmTNn8Pbbb3vOPXfuHFZXVztrODHQkLAk+oJkMomhoaGGG+fmzZtYWVmBqqp47bXX2g7v4cLvr7/+cvZ9/fXXkGXZU+bOnTsxPDzsq5oTmwfyhhN9gWmakCQJY2NjmJ+f9xxTVRWpVAq8q0YiEXz55Zf4448/AKCtWWU7bXv06NEzrYfob2hmSfQFwWAQZ8+eRaFQQD6f9xw7efKk53uhUMDrr7+OPXv2IJvNIp/P11zTTRYXF2lWSZCwJPqHqakpyLKMTz/9FLZtO/sDgYAzqzRNE7IsIxgMYvfu3di1axf+/PPPZxoP+eDBA6ysrCAWi7XtdScGH1LDib5C13W8+eabiEajjgOn19i2jbm5Oezfv78rQfDEYELCkug7kskkpqenUSqVKDeb6BtIWBJ9h23bGB8fRygUqnH2cCKRCPbt24epqam263F72FuBhszmhFYdIvqWmZmZusfqCdFWIKFHtAI5eIi+4+jRozhy5EhdFbw67rIZNE1DLBaj/G6ibUhYEn1FNpuFYRg4fvx43XNyuRxkWXa+RyIR3yB2Lhg1TcPq6ipEUaxbpq7r0DQNuq5378cQLxSkhhN9g2mamJmZwcLCQs2xUCjkLJ929+5dT+72Rio5n6E+ePCg7jmzs7NIpVJQFOWZB7kTgwnNLIm+IRqN4tSpUzXhOfl8HoZhON+LxSLeeOONrtb9ySefAKgNgCcIDglLoi9IJpMol8uYnJysUac//PBDz7mFQgEAnKydjdRwzv3797G+vu5b/82bNyHLMgKBQM0xTdP6Juazn8jn84jFYohEIpsiWJ9Ch4iBY3R0FGtra7h27RpGRkY2PF/TNITDYc++6hjO0dFRHD582DcUSdM0PH78mFZNd6GqKh49eoQzZ85smkB9EpbEpse2bWzduhWVSgUjIyOwbRvnz5/H6dOnkc1m8eDBA8recZFMJvHTTz85NuTNwqZXw92qRK+wbRvZbBaRSKSn7disnD9/HoIg4MmTJ9A0DQcPHnRWRV9fX8f09HSPW9g5tm0jmUxu2L/y+TxCoRCGhoYQCoV8Q63OnTuHU6dOIZ/PI5lMNr2IiW3bmJ2dbVpl56/ycK8T0FNYC8iyzAD4boIgsGg0ykqlUt3rDcNg8XiciaLoXCdJEkun08wwDJZIJHyvK5VKLBqNMkEQnLoURXHKaxfLslipVGIAmCzLbZfTKZZlMcMwet6OzYosyzVbpVJhjD3ts6Io9riFnZFOp52x06h/5XI5BoBlMhnnOgCecVmpVBgAJooii8fjLB6PM0EQNhyHlUqFRaNR32OZTIbJsuwrOwzDYLIsM8uymvmpz5SWhCVjzBnUbjlrWRbLZDLODcnlcjXX8eOyLHuOcyFZfVM4iqI4xwzDcPbncjkmSRJrUd77IklSXwgpEpb9RyaTqTvI28GyLI8wrlenoihdqS8ej7N0Os0Y27ifi6JY81slSWKCIDjCik8u3GORC9V6WJbFRFH0FXi5XM4Z//UmWqVSiUmSVP9HPifakjT1hJT7qeOG/8GNOl0mk6l5OnFByZ901ViWxSRJ8ty4duCziV5DwrL/UBSlrsbTDrzPCoLg228zmUzDPt8Jjfo5H7tcsHK4IOMTHD6W3cKez0jroShKw5knL7ORVirLclfvQzu0ZbP0C68A4Hgm3TFxAJyX1f/44491y4zFYp7vmqYhlUpBkqSaY+52fP/993j48GHdcm3bhqqqGB4eduww9cJAdF137DXucAh3aIqmaQC8r0Hg+zRNg6qqiEQiTr28zurMENu28e233zrt8vuNuq4jFos59YyOjnrK4XbOZDIJXdeddvL6G5WlqmrNOclk0mlPdV2bFdM0USwW6/bBVgkEApifn4coiohGo557kM1mMTk5iUwm07X6muXWrVsAUPNSNh7PevfuXQBPA/xFUcTs7CyAp/34559/9mRUuTFNE6lUCh988EFH7du3bx/OnTvXURmd0lUHDxcugiA4+3hAcTQarStkOadPn3Y+X7p0CQBw4MCBhtfs3Lmz4TJeJ06cwMLCAsrlMizLgiRJmJycrBEUKysruHLlCi5evIhEIoFCoYBoNArgaYZIOp32nH/8+PGafcDT903bto0TJ05gz549yOVyMAwDZ8+e9Zx39OhRT7uGh4dryvr4448BAJZlwTAMrK2t4fDhwwCedtItW7agUChgaWkJs7OzmJmZcdp+9epVpxxd1zExMYGJiQkwxpDL5ZBKpXDixAnnHFVVcf/+fdy7d8952E1MTNQ1rteLbaze+INkUEkkEti3b1/DB32rcIEJwHmwaZqGyclJJBKJ5y4oAdSNP92yZUvNvmvXrsE0TQwNDWHr1q0AgF9//dX3+rm5OQDoeKm9HTt2YG1trbfxru1MR7mjx02lUnHUc7cK0cge6S7LvZVKJWd/o6l5M0SjUY/6z1UGd7l+6glvN/8tfqqC3z5ZlmvsK9Xlc1XLrYZZllWjhguC4FGL4vF4zf8OoEbFqf6/o9FojQ2MO+QYe3rvqk0n/H+qVsuI7mFZFhMEwVHLu2WnrEcjNbye3ZD38XZVYFmWmSAIDc9pRg3nvpJOHLqd0lFuuN/KL/WejPWeXPPz887q2IIg4N69exvOQFuBP4lM08Tc3BxSqVRT1+3fvx/T09MN84nr4dd+nnUCAL///jtEUfTE7fldw1+QpWkaLl26VLftPMzFTbFYdBajuHz5MhKJhG/ZvG2GYfjez+elire7tmQvYF0KTQ4EArhw4QImJychCMILmZNeKBTqquitwMeK+02cz5uO1HD21EEExhgqlQoURcH09DRGR0cd9Y0PZL/FETjc1jk2NuYIjVdeeQUA8M8//3TSRMc2KEkStm3bhl9++aWp655lAPK///6LUCi04XmmaSISieDYsWN49913awReN5Fl2XM/+VZvAHdbDferu1+3bqHrOlRVhSRJAJ6aQvoVvwfyZqNrNsuRkRFcuHABiqKgXC7j/PnzAID3338fAFAul1uyX3300UcAgN9++62jdh08eBALCwu4d+8epqamfG0wflQL+26zuLi4Yf2SJCEYDOLOnTsd27GWlpZq9rntP4VCwdc+WS/geH5+vinBQq+F8Me2bRw+fBiiKGJ+fh43btxAKpXqmcDkjpybN2969vPv1Y6fzUhbwrJRRL3buQM8naFxR8ixY8eajsaPxWKQJAmFQqGhUTebzTYUwoVCAQcOHGhZtb9+/TqA/4Q95/Hjx76fWyEYDGJtba1hu5eXl7G2tuashtMJsizj8uXLnvps28bff//tHAfgcfgA9QUl0Rm2bTuZNPPz8wgEAhgZGUEmk0EqlXI8zc+T8fFxCIKAYrHo2b+0tARRFNt+6Mmy7DFBtQuXGz1NOW3VyOkXlM73cyOxXwwZPyaKIstkMk6AKg9oh0+MIY9JA8AURfHEdlUqFaYoyobxaNyAzrNkeOxmOp12nBfcCO2OJRMEwVM2/92SJLFSqcQSiYTjcInH46xSqTjBt+4gXsb+i0vl+wzDYIIgMEEQHKM2d6jw8nncGzf6u51euVyO5XI55xz3/8YdRW4nEzeg80yLRCLBJEny/J/RaNS5P4lEgimKQjGfzwDep3mfrOZZxVnyvunuh9XwMcrHBW+LX5JJs/AyG2Xg8HoaOZF4H34W8afN0rV0Rz7QeBqiHzw9kQsPvsmyzOLxuO91lmWxdDrtuaaZ1EoOzxzinmUuqNwpVFzw8gwkLrCq4ZkKXJCWSiUmiiJLp9OOkHJv/AZX7+N18v9TFEWWy+UcjygXYlwYi6LISqWSI1Dj8Xjdsqv3cXK5nDNYqgUl/5956hoX0v2QYvaiYVnWhv9tOp3umlDgqYTV/UKWZd86EomE0wd4v+wEPsmoN1b92uYHT9nsZZ+kVYeIvkLXdTx58gQvvfRSU8uvdRvTNJ0kB7fqyU0YvWrXIKOqKgRB8MRRt0osFsNbb73V8HUjz5pNv+oQ0V9cuXIF4XC4LTtXI3t4s7byxcVFhMNhhMNhj8320KFDCIfDTqYL0TwnT57EwsJC2wsE67q+4XuZngs9m9MSRB3gUttKpRIrlUpN5f9zW2s1reZ38+Btd1n9spjDoMLNTq2q0c2YLZ4XJCyJvoLbuPhnbj9rNpOoWsi146jiTki4HBPpdLrnCzkMOtz/0KzgMwzD4wzuNSQsib4ik8k4MzjuOW4l5ZJfwyMl6nmd68HXTuWpiNwJ0qxDkXhxIQcP0Vfwd7tMTExA13WcPHmy5RhZ27YxPj4OALh9+3ZL1yeTSWzbtg2xWAyqquLOnTu4c+cOhoeHu56KSwwW5OAh+oqFhQUYhoFvvvkGjx49aks4Xb9+HYIgQBAEJ7mgWYrFIrZv3w4A2LNnD8rlMvL5PERRJEG5yaGZJdE3mKYJURRRqVRw69YtfPHFF7AsqyUhlc1moaoqyuUyXn75ZYyPj+PUqVNNp4sODw97FhkJhUIQBAG7du3qKPSFGHxoZkn0DYuLixAEASMjI06a6dzcHFRVbSrsRNd1zMzM4MaNGwgGgwgEArh27RpUVW1qXQK++LG7riNHjqBcLuOdd95p/4cRLwQkLIm+YX19HYqiAPhvTYGlpSVMTU01lRO8vLxc8y7xkZER3LhxA6urqxteXywWEQqFPIuc7N+/H7IsOzZQYvNCajhBEEQT0MySIAiiCUhYEgRBNAEJS4IgiCb4H0NLEon1HE+tAAAAAElFTkSuQmCC\"\u003e\u003c/p\u003e\n\u003cp\u003eN\u003csub\u003ei\u003c/sub\u003e represents the number of reads mapped on a biosynthetic-related gene; L\u003csub\u003ei\u003c/sub\u003e represents the gene length; k represents the number of biosynthetic-related genes in a BGC; N represents the total number of high-quality non-host reads in a metagenome/metatranscriptome sample.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePrediction of secondary metabolite activity\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo predict the activity of BGC-encoding compounds, we used mlr v2.19.0 [69] to perform machine learning. The training dataset comprising 950 MIBiG BGCs with known activities (antibacterial, antifungal, antitumor or cytotoxic, or other activities) was gathered by Walker \u003cem\u003eet al.\u0026nbsp;\u003c/em\u003e[33]. BGC features of those known BGCs were extracted by BiG-SLiCE [23]. Prior to training models, we removed BGC features present in \u0026lt; 10 BGCs. Rather than multiclass classification, binary classification was adopted for each activity class since a molecule might have multiple functions. Four two-class classifiers (namely logistic regression, elastic net regression, random forest, and support vector machines) were adopted for binary classification of the activities of BGC products. In order to obtain the honest performance of four classifiers, we measured their accuracies using 10-fold cross-validation. Moreover, 3-fold cross-validation was adopted in each test to tune the hyperparameters, generating 30 instances for each classifier. The average AUROC was used to evaluate the performance of four classifiers. The function \u003cem\u003egenerateThreshVsPerfData\u003c/em\u003e was used to generate data on threshold vs. performances, which was further adapted for plotting the ROC curve. Using the random forest model, 129,878 LAB-derived BGCs and 1,121,156 non-LAB-derived BGCs containing BGC features were subject to activity prediction. Chord diagram showing the association between BGC classes and predicted activities of their products was plotted using R package \u003cem\u003ecirclize\u0026nbsp;\u003c/em\u003ev0.4.13 [70]. Sankey diagram showing the association between species and BGC classes was done by package \u003cem\u003enetworkD3\u003c/em\u003e v0.4 [71].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePrecursor of class II bacteriocins\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn order to pinpoint the precursors of class II bacteriocins, we first used Prodigal-short\u0026nbsp;[72]\u0026nbsp;to identify all small ORFs. We then used hmmsearch\u0026nbsp;[36]\u0026nbsp;to search class II bacteriocins-related domains (provided in Supplementary Table 4) against ORFs of all RiPP-like BGCs. The hits with a threshold of E-value \u0026lt; 0.01 were considered as the precursors of class II bacteriocins. Meanwhile, BAGEL4\u0026nbsp;[37]\u0026nbsp;was also adopted for searching class II bacteriocins from RiPP-like BGCs. They detected 128,599 and 90,101 putative precursors, respectively, with 30,764 in common. We discarded 287 sequences that were larger than 150 AAs, retaining 187,649 sequences for further analysis. Those sequences were then grouped into clusters using Cd-hit\u0026nbsp;[73], with the parameters of \u0026ldquo;-n 2 -p 1 -c 0.5 -d 200 -M 50000 -l 5 -s 0.95 \u0026ndash;aL 0.95 \u0026ndash;g 1\u0026rdquo;. The sequences with an identity of \u0026gt; 50% will be grouped into one cluster, as proteins with \u0026gt; 50% identity generally share a common function\u0026nbsp;[74]. To collect the known class II bacteriocins, we queried NCBI PubMed with the keyword \u0026ldquo;class II bacteriocin\u0026rdquo;. Meanwhile, we also included the sequences gathered by Yi \u003cem\u003eet al.\u0026nbsp;\u003c/em\u003e[15]\u0026nbsp;as well as the sequence deposited in the BAGEL4 database\u0026nbsp;[37]. In total, 333 sequences of class II bacteriocins were obtained (Supplementary Table 9). As the curated 333 sequences might be the mature peptides, a local sequence aligner, DIAMOND v2.0.15\u0026nbsp;[75], was utilized to compare 333 known class II bacteriocins to 187,649 precursor sequences with the parameter of \u0026ldquo;--id 90 --query-cover 95 --masking 0\u0026rdquo;. The known class II bacteriocins showed an alignment of identity \u0026gt; 90% and coverage \u0026gt; 95% with 1,775 precursor sequences belonging to 188 clusters that were thus regarded as homologous. For 21 selected precursor clusters, we identified the global identity relative to the known class II bacteriocins using the Needleman-Wunsch algorithm in the function \u0026ldquo;needleall\u0026rdquo; of EMBOSS software package\u0026nbsp;[76]. The alignment of precursors was done by MAFFT v7.490\u0026nbsp;[77]\u0026nbsp;with the parameter of \u0026ldquo;--maxiterate 1000 --localpair\u0026rdquo;, and then was visualized using Jalview software\u0026nbsp;[78]. To conveniently inspect the gene organizations of BGCs harboring precursors of cluster_467 and cluster_468, we adopted BiG-SCAPE\u0026nbsp;[25]\u0026nbsp;for exploring their architectures.\u003c/p\u003e\n\u003cp\u003ePrecursor abundance in metagenome/metatranscriptome samples was computed via the equation (2):\u003c/p\u003e\n\u003cp\u003e\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAbgAAAAtCAYAAADcFn73AAAO80lEQVR4nO2dz2sbRxvHv3p5r05Yqaf+MEXrHkwKLnTllMYKNGCv2pTQgIPUNgdDio10KLQHh8iGXJwQiTiH4NoS2BBKg9YE4x4q4R9gQ6WaYIsiQYIPtfZQTE5ay67/gHkPfmc7K61t2ZYjRX4+sJCsZ3ee+bH77DzznZGDMcZAEARBEE3Gf+ptAEEQBEGcBuTgCIIgiKbkv/U2gCCajXw+j93dXbS0tKCjowMAYBgG1tfXAQBvv/023G53PU0kiDMBjeAIosbE43F4vV589NFHMAwDALC5uYlr167B6/Vid3e3zhYSxNnAQSITgqg9Ho8HhUIB4+PjCAQCAIBoNIqdnR3cu3evztYRxNmARnAEUWMymQw8Hg8CgQBGR0fN83/++Sc+//zzOlpGEGcLGsERRI2JRqM4f/48Lly4AK/Xi0KhALfbDafTia2tLTMdn5fr6uqqo7UE0bzQCI4gaszS0hK6u7vR1dUFWZYxNTWFfD4PWZYt6TY3NzE2NlYnK5ufVCqFQCAAn88HXdfrbQ5RB0hFSRA1ZnV11VRJ+v1+TE9Po7W1FT09PWYaXdexu7uL+/fv18vMpiYUCmFrawv3798nxeoZhkZwBFFDMpmMxZHdunULhUIBk5OTuHTpknl+dXUVXq8Xi4uL9TCzqYlGo1hYWICmaeTczjjk4OqEGD6pF4ZhQNM0+Hy+utrRLBiGYYYc+fIAt9sNVVWRzWZx8eJFM20gEICiKPjkk0/qYmujYRgGotHoof0wGo3C6XTC4XDA4/EglUpVpHnw4AFGRkaQSqUQjUZt09SLeDyOfD5fdfqhoSEKr54AcnB1wDAMnDt3DtPT0/U2BZ2dnZifn6+3GU3B5uYmtre3sb29jc3NTfP8999/j2AwCJfLZZ4zDAPZbNZcCH6Wicfj+OCDD3D79u0D0w0NDeHBgwdYXl5GsViEx+PB1atXkclkzDT5fB6lUgnDw8P4448/sLOzg5s3b2JoaOi0i3EghmHA5/Oht7fXbHNd1xEIBOBwOOBwOODz+SxlAYAff/wR4XD4SE6REGBE3VAUhamqWm8zGICGsOMskUwmG7LOi8UiU1WV5XK5fdMkEgkWDAZrkl84HGaxWIwxdvDzUCgUGAAzLbe1vO+m02kGgBUKBfNcLBZj9X7VqarK0um0+f9ischkWWbBYJBFIhEWDAYZAAagou55WrFMRHXQCK6OiF/0xNnixYsXZliu0UJQhmHgs88+s7VL0zR8/fXXuHz5ck3yunfvHgYGBgAc/DzwuUoxpOtyuaCqKubn582QMEfcLea9996ria3HRdM0GIZhWQ6yuLiImZkZjI+PY3BwEOPj44jFYgD2RrQiLpcL/f39CIVCr9XuZuBQB5fP5zE0NASn0wnDMMwhdVtbmyW2zedyotEoNE2D0+k0d3AA9uac2trazGs1TbPkYxgGQqGQGV/3+XzmsJwP4R0Oh5ne5/NVnDvIhnw+b8nf4/FU5M/LaRff13XdUg8+nw9Op7MipLBfeezKLNYxt02UNItl5PlEo9GKc5lMBqFQCD6fz8yX51ke2igvp9hGoj1i6MTj8VjuI9ZzPp837eT5i6RSKXg8HjgcDjidTkSjUdt64nkFAoGKezQjqqrC5XLhww8/bCghhMvlwtzcHGRZht/vt7QFd26JRMK235wmy8vLAFAR0v34448BwNznky/N4E7CMAz8/PPPUFX1NVprZXh4GDdu3LCcCwQCFWXhjl5cK8nhjpxClUfksCFeOp1mqqoyACwcDrNkMskSiQSTJMkMBRSLRZZMJhkA5vf7WSwWY+Fw2AwdJBIJpiiKmZYPx/mQvVgsMkVRWDAYZMVikRUKBSZJEpNl2bTD7/dXhBnEc4fZIEkSSyaTZpkkSbLcS8xftDGRSDDGGMvlcmZ+kUiEJZNJpiiKec9ygsGgGVYoFovmtcVi0UyjqiqTZZmFw2GWTqdZJBJhAJiiKGYaHl4Rwxvl59LpNJNl2SxDMpm01EV5ndm1hRjmkWWZ+f1+sy34ve3qORgMWmwXQ0i8n3A7w+GwpU55u/NreLscFLrjffGwQ6wv4ujwtlEUhRWLRTP0F4lETi1PVVX3bXve7uXwfie2dy6Xs/QT3pfrAX9Wqu2P4vNh97dahYbPClUFpnknEjsJ7/Bihe/XAJIkWeLHPHbOX76xWMzizBjbexGLTojbYGeXpUD72MCdMUfsRIlEoqJ8PO5tZ0M1sXC/329xLnYd3e6B5nlw+3g9i9fZnVNV1eIY7e7Py2nXFmI6SZIsjoo7JhH+wVN+TnwBSpJkqedcLmdxgrFYrKKteF4HzQERr4discgkSWKKojBJkk795VorB9dIcPuq6c/pdLriGRYRPzSJ6jjSHJwYI+/q6oIkSRVx+vfff9/y/0wmg1KpBFmWzTDUW2+9BQDIZrMAgNnZWbS1tVmu0zTNdqheDeU2AHsLbnt6eswwoRhiefLkCWRZtpTP5XLB7/ejVCpVhCGrCSlpmgZN08zQ5s2bN6uyvbe3FwDw999/V5VexG4OQ1RI/vrrr5Bl2WK/3TVbW1sYGBgwQ5/7LUY+f/58xbmlpSUA/7b7u+++a/6to6MDjDEzFDM7O4uJiQlLCJrnxUNOp42Yd7MctcLlcmF8fNx8TsfHx2t277MCfx6qUcuOjY1hcnJy37+3tbWZbUFUx4lEJp2dnVWnZXujRcuxsbFxkuyPhKZpGBkZwfDwcMX8IWAf97Z7gVcLn+tSFAWtra345ZdfqrruNOdjtre3Kz4k7NB1HT6fDz/88AMuX76MSCRyajZFIhHbvrHfHI84L3nQsd/caDl2eb/pR63I5/MIhUJQFAUAGlrk0NLSUm8TToSmabhz5w4tG6kxJ3JwRxED2C22FM+trq5W/F3X9apfVNUQCASwsbGB/v5+XL161ZJ/qVTatzzHeXi+/fZbLCws4K+//sLAwADOnTtX1XXchpM414Owq+fy/BVFgdvtxtra2onFBC9fvqw4J4ptnj17ZmvDfu0+NzdX1UueNjA+GYZh4LvvvoMsy5ibm8Py8jImJibq5uS4mKS8XxxlhNSopFIptLe3v9FlaFSO7eB0XUc2m0VfX9+B6Xgo8+7duxYHous6Xrx4AQC4cuUKSqVShcLu4cOHFS8q8R47OztV2yvee3BwEKqq4vHjxwCA69evAwCmpqYs1+zs7ECW5WN1vPn5edy4cePISwG4HLq7u9ty/p9//rH991Fwu922IVeR9fV1lEolfPPNN8fKg9Pe3g5JknDnzh1Lm+XzebPdrl+/jmw2W6EuffToEdrb20+UP3F8uEoY2PugcLlc6OjoQCKRwMTERIWM/XXAf2ZoZWXFcn51dRXBYPC121MtV65cAVDpmDmpVArvvPNOVe8YwzAqNuwmDqGaiTo+USqqDLkaj8MFBLIsVyiWuOpPkiQWDodZOBy2pOOT2TyPSCTCVFW1FYJw1V44HDaViYlEghUKhQNtEAUPXBnIBRFcMSYq/riiT1RJ8vxEAcZ+8Ml5rkTkasVYLGZer6qqrbpTLDdf4KooiqlW5EKMcDjMcrmcRRAjlltRFIt4hqtTxXJy8Qu/P69DLigQVbRcncnTiIIALlYR+wTvN7yuuapVbHdZls178XY/TaUecTDl6sly+HO4n9LvJPnyvrCf4pE/L7zvBoPBCgFbo8EFYXZq61gsxvx+P4tEIpbD7/fbimbE55KojiM7OPxfesudHWP/NqJ4lMOVkvxlVq4qEqW9sizbPkDcwciyzHK5nMURHmYDd2CioxURJfPiC59TLk8/7CXMJfJckcidi/iCz+Vy5kNql6dYd9xuXlZZllksFjMdi3jY1QW/b3k9J5NJUyHH24Q7UFmWWTqdNp0gX85QTX6cSCRils9Orl0oFMx2lSSJnFud4c/BQbL6WCxWMweXSCRsl36Uf+CKtolp3gS1rfgxzeHPmN1hp5TkH5WNqhZtVKr6wdNoNIrbt2/XdAKbIBoJXdfx6tUrAHvh1UbdZUa0Uwzf8xBYS0sLzeU0GJqmYXR0FGtra8e+Rzwex+zsLObm5mpoWfNDW3URBPbme71eL65du2bZKLmWHCTKqlawxX9mx+v1WkRSfX198Hq9eP78+YntJGpLIBCALMvH/lUDwzAwOTlJyzSOQVUOjiuVGm3PPIKoFfzlMT4+fmojoKmpKVsVYigUqhA47UcgEICqqlBVFb/99pt5/smTJ1AUxVxjSDQWP/30Ex4/fnysrbYePXqEycnJhtrS7U3hUAfncDjMxcKyLFcoHQmiGeAhvqOs7Twqg4OD0HXd4uRCoRB0Xcfg4GDV99nY2EBfXx8mJibMkd/Lly8r9jskGge+x+fz58+PNFCIx+O4desWhZ2PyaEOjpWtLzrKg0gQbworKysVu7yI1GrX/6dPn2JtbQ2hUAiapmFtbQ1Pnz6t+vpMJgNFUdDd3Q1JksxlJcvLy/j0009PbB9xugwMDBxpJHbU9IQVmoMjCOyF4Xt6eg5MU4vF9/xLfmFhAcPDw+Y6s2pZWVnBV199BZfLhUAggNHRUQDAwsICrR0kiDLIwREE9hbmi79xpuu6eUSjUbS2ttZMWbm4uAhJkiwjsGpZWloyHdmXX36JbDaLVCpVsZcqQRDk4AjCdv7t4cOHphz/2bNn+P3332uSl6ZpCIVCmJ6extzcHIaHh/f9nUA7VldXzfmYL774ArIs4+7du4eOPgniLEIOjjjTGIaBsbExSJKEV69eIZPJIB6PY2JiAl1dXXC73fB4PDX5Bet8Po/h4WEsLy/D7XbD5XJhZmYGoVCoqj1Xo9EonE6nZS6wv78f2WwWly5dOrF9BNFs/LfeBhBEPdnc3MT29jY6OzsxMjJing+Hw+a/FxYWaiK/X19fx8zMjEUR19HRgeXl5ap+HmhpaQltbW1YXV01hQe9vb1YWlrCxYsXT2wfQTQbVe1kQhBnFV3XIcsy7eJDEG8gFKIkiAPg83ChUOjYO1EQBFEfaARHEIcQj8dx4cIF+o05gnjDIAdHEARBNCUUoiQIgiCaEnJwBEEQRFNCDo4gCIJoSsjBEQRBEE0JOTiCIAiiKSEHRxAEQTQl5OAIgiCIpuR/6vOA3z0ekzYAAAAASUVORK5CYII=\"\u003e\u003c/p\u003e\n\u003cp\u003eHere, N\u003csub\u003ei\u003c/sub\u003e represents the number of reads mapped on a precursor gene; L\u003csub\u003ei\u003c/sub\u003e represents the gene length; N represents the total number of high-quality non-host reads in a metagenome/metatranscriptome sample. The abundance of a precursor cluster is the sum of the abundance of precursors in this cluster.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePhylogenetic tree construction\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGTDB repository version r202\u0026nbsp;[57]\u0026nbsp;contains 822 representative genomes of 56 LAB genera. The genome with the largest N50 length in each genus was selected as a proxy for its corresponding genus. Consistent with the previous approach to constructing bacterial reference trees\u0026nbsp;[57], the multiple sequence alignment of the concatenation of 120 phylogenetically informative marker genes of 56 representative genomes was used to infer the phylogenetic tree. IQ-TREE version 2.1.4-beta\u0026nbsp;[79]\u0026nbsp;was adapted for constructing\u0026nbsp;maximum likelihood (ML) phylogenetic trees, with 1,000 ultrafast bootstrap replicates. In-built\u0026nbsp;ModelFinder\u0026nbsp;[80]\u0026nbsp;identified the best-fit model as LG+F+R8. Inferred phylogeny was visualized using iTOL\u0026nbsp;[81].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePeptide synthesis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe two deducted core peptides of cluster_467 and cluster_468 were chemically synthesized by Sangon Biotech (Shanghai, China). Their molecular weights were confirmed by mass spectrometry, and their required purity was \u0026ge; 90%, determined by high-performance liquid chromatography. The synthesized peptide powder was stored at \u0026minus;80\u0026nbsp;\u0026deg;C\u0026nbsp;and\u0026nbsp;dissolved in sterilized\u0026nbsp;double-distilled water to 2 mg/mL upon use.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBacterial and fungal strains\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of 14 bacterial strains and two fungal strains were used in this study. Their growth conditions are as follows: two bacterial strains (\u003cem\u003eEscherichia coli\u003c/em\u003e DH5\u0026alpha;, \u003cem\u003eStaphylococcus aureus\u003c/em\u003e B04) were incubated in Luria\u0026ndash;Bertani (LB) culture medium at 37 ℃ under 180 rpm rotation; two strains (\u003cem\u003eChromobacterium violaceum\u003c/em\u003e, \u003cem\u003eBacillus subtilis\u003c/em\u003e 168) were incubated in LB medium at 30 ℃ with shaking at 180 rpm; \u003cem\u003eLactococcus lactis\u003c/em\u003e subsp. \u003cem\u003ecremoris\u003c/em\u003e MG1363 was grown statically at 30 ℃ in M17 medium; eight bacterial strains (including \u003cem\u003eStreptococcus mitis\u003c/em\u003e, \u003cem\u003eEnterococcus faecium\u003c/em\u003e, \u003cem\u003eEnterococcus faecalis\u003c/em\u003e, \u003cem\u003eLactobacillus delbrueckii\u003c/em\u003e subsp. \u003cem\u003eBulgaricus\u003c/em\u003e, \u003cem\u003eLactobacillus crispatus\u003c/em\u003e ATCC 33820, \u003cem\u003eLactobacillus acidophillus\u003c/em\u003e, \u003cem\u003eLactobacillus casei\u003c/em\u003e, \u003cem\u003eLactobacillus fermentum\u003c/em\u003e) were incubated statically in MRS medium at 37 ℃; \u003cem\u003eGardnerella vaginalis\u003c/em\u003e was maintained in Colombia blood agar at 37 ℃ in anaerobic conditions, and the suspension culture was grown by taking a loop full of colonies from the agar plate and incubating in Brain heart infusion broth (BHI), at 37 ℃ in anaerobic conditions; two fungal strains (\u003cem\u003eCandida albicans\u003c/em\u003e SC5314, \u003cem\u003eCandida albicans\u003c/em\u003e ATCC 10231) were grown in RPMI media at 37 ℃ with shaking at 150 rpm.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDetermination of minimum inhibitory concentrations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe minimum inhibitory concentrations (MICs) of the two peptides [individually and in combination (1:1 ration)] against bacterial and fungal strains were performed by broth microdilution. Tested bacterial strains were inoculated overnight in the corresponding culture medium (LB, M17, BHI, or MRS) and at respective growth conditions. The optical density at 600 nm (OD\u003csub\u003e600\u003c/sub\u003e) of bacterial cultures was determined to estimate the bacterial concentration. The bacteria cultures were diluted to ~5\u0026nbsp;\u0026times; 10\u003csup\u003e5\u003c/sup\u003e CFU/mL using the respective broth. 100 \u0026mu;L aliquots of bacterial suspensions were transferred into 96-well plates containing two-fold serial dilutions of peptides (ranging from 200 \u0026mu;g/ml to 0.19 \u0026mu;g/ml). After incubating for 24 h, bacterial growth was assessed by determining OD\u003csub\u003e600\u003c/sub\u003e. Besides, MIC against the fungal strains was determined according to the CLSI M27-A3 guidelines [82]. Briefly, \u003cem\u003eC. albicans\u003c/em\u003e strains were cultured overnight in RPMI medium and grown fungal suspensions were centrifuged at 5,000 rpm for 10 min and the pellet was resuspended and washed twice with 1\u0026times; PBS to remove the dead cells. The fungal inoculum was standardized to 1 \u0026times; 10\u003csup\u003e6\u003c/sup\u003e CFU/mL using a spectrophotometer and added to the well plate containing varying concentrations of the peptide (200 \u0026mu;g/mL to 0.19 \u0026mu;g/mL). The media without the peptide served as a control. The plates were then incubated at 37\u0026deg;C for 24 h with shaking at 80 rpm, and the absorbance was measured at 520 nm using SpectraMax 340 tunable microplate reader (Molecular Devices, San Jose, CA, USA). The MIC value was determined as the lowest concentration of the peptides where no bacterial or fungal growth was detected. All assays were conducted in triplicate on three independent occasions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical analysis and visualization\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe accumulations of GCFs detected in metagenomes as well as clusters of class II bacteriocin precursors were computed with function \u003cem\u003especaccum\u003c/em\u003e in R package \u003cem\u003evegan\u003c/em\u003e v2.5-7\u0026nbsp;[83]. R package \u003cem\u003eUpSetR\u003c/em\u003e v1.4.0\u0026nbsp;[84]\u0026nbsp;was adopted for visualizing the intersection of GCFs or precursor clusters detected in different body sites. Chi-squared test and wilcoxon rank-sum test (two sided) were done by function \u003cem\u003echisq.test\u003c/em\u003e and \u003cem\u003ewilcox.test\u003c/em\u003e in R, respectively. The alpha diversity (Shannon index) of the vaginal microbiome was\u0026nbsp;calculated with R package \u003cem\u003evegan\u003c/em\u003e v2.5-7\u0026nbsp;[83]. To visualize the distribution of class II bacteriocins detected in six body sites, a dimensionality reduction was performed using t-distributed stochastic neighbor embedding (t-SNE), which was done by R package \u003cem\u003eRtsne\u003c/em\u003e v0.15\u0026nbsp;[85]. Spearman\u0026apos;s correlations between precursor clusters vs. bacterial species and between clusters vs. Shannon index were computed with the function \u003cem\u003ecorr.test\u003c/em\u003e in R package \u003cem\u003epsych\u003c/em\u003e v2.1.9\u0026nbsp;[86], and\u0026nbsp;\u003cem\u003eP\u0026nbsp;\u003c/em\u003evalues were adjusted with the \u0026ldquo;BH\u0026rdquo; method\u0026nbsp;[87]. The heat maps in this study were plotted using package \u003cem\u003epheatmap\u003c/em\u003e v1.0.12 [88]. Cytoscape 3.9.0 [89] was used to visualize the network of similarity of class II bacteriocins and the network of species-precursor correlation. Without a specific statement, other figures were generated using ggplot2 v3.3.5 [90]. All statistical analyses were finished in R v4.1.2.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and Consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe bacterial genomes are publicly available in NCBI Assembly RefSeq database (https://www.ncbi.nlm.nih.gov/assembly), PATRIC database (https://docs.patricbrc.org/user_guides/ftp.html), and IMG/M database (https://img.jgi.doe.gov/cgi-bin/m/main.cgi). The genomes from human gut are available in the European Nucleotide Archive under study accession ERP116715, and genomes from food metagenomes are available at http://www.tfm.unina.it/DATA001-2020-Pasolli. Genomes can be obtained through the accession numbers provided in\u003c/p\u003e\n\u003cp\u003eSupplementary Tables 1 and 5. The raw data for HMP metagenomes and vaginal metatranscriptomes are deposited in NCBI-SRA under the BioProjects PRJNA48479 (https://www.ncbi.nlm.nih.gov/bioproject/48479) and PRJNA797778 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA797778), respectively. The samples used in this study are provided in Supplementary Table 6.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work is partially funded by a Shenzhen Basic Research General Programme (JCYJ20210324122211031) and two Hong Kong Research Grants Council General research grants (HKU27107320 and HKU17115322).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eYL and DZ conceived of the study, participated in its design and coordination, and drafted the manuscript. DZ and JZ gathered publicly available data used in this study. DZ and SK performed MIC determination. JL, ZS, BH, and PC performed bacterial culture and metabolic analysis. DZ, JZ, ZZ, and YL performed data analysis and interpretation. CF and PN provided advice. YL was involved in the overall supervision of the project. All authors read, revised, and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to thank Dr. Mingqiang Qiao and Wanjin Qiao at Nankai University for providing \u003cem\u003eL. lactis\u003c/em\u003e strain.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eCarr FJ, Chill D, Maida N. The lactic acid bacteria: A literature survey. Critical Reviews in Microbiology. 2002;28:281\u0026ndash;370.\u003c/li\u003e\n \u003cli\u003eLeroy F, de Vuyst L. Lactic acid bacteria as functional starter cultures for the food fermentation industry. Trends in Food Science and Technology. Elsevier; 2004;15:67\u0026ndash;78.\u003c/li\u003e\n \u003cli\u003eTeusink B, Smid EJ. Modelling strategies for the industrial exploitation of lactic acid bacteria [Internet]. Nature Reviews Microbiology. Nature Publishing Group; 2006 [cited 2022 Jul 5]. p. 46\u0026ndash;56. Available from: https://www.nature.com/articles/nrmicro1319\u003c/li\u003e\n \u003cli\u003eWells JM, Mercenier A. Mucosal delivery of therapeutic and prophylactic molecules using lactic acid bacteria. Nature Reviews Microbiology [Internet]. Nature Publishing Group; 2008 [cited 2022 Jul 5];6:349\u0026ndash;62. Available from: https://www.nature.com/articles/nrmicro1840\u003c/li\u003e\n \u003cli\u003eSaez-Lara MJ, Gomez-Llorente C, Plaza-Diaz J, Gil A. The Role of Probiotic Lactic Acid Bacteria and Bifidobacteria in the Prevention and Treatment of Inflammatory Bowel Disease and Other Related Diseases: A Systematic Review of Randomized Human Clinical Trials. BioMed Research International [Internet]. Hindawi Publishing Corporation; 2015 [cited 2022 Jul 5];2015:1\u0026ndash;15. Available from: http://www.hindawi.com/journals/bmri/2015/505878/\u003c/li\u003e\n \u003cli\u003eRen C, Faas MM, de Vos P. Disease managing capacities and mechanisms of host effects of lactic acid bacteria. Critical Reviews in Food Science and Nutrition [Internet]. Taylor \u0026amp; Francis; 2021 [cited 2021 Nov 12];61:1365\u0026ndash;93. Available from: https://www.tandfonline.com/doi/abs/10.1080/10408398.2020.1758625\u003c/li\u003e\n \u003cli\u003eGarc\u0026iacute;a-Bayona L, Comstock LE. Bacterial antagonism in host-associated microbial communities. Science (1979). 2018;361.\u003c/li\u003e\n \u003cli\u003eBraga RM, Dourado MN, Ara\u0026uacute;jo WL. Microbial interactions: ecology in a molecular perspective. Brazilian Journal of Microbiology [Internet]. Sociedade Brasileira de Microbiologia; 2016;47:86\u0026ndash;98. Available from: http://dx.doi.org/10.1016/j.bjm.2016.10.005\u003c/li\u003e\n \u003cli\u003eDelves-Broughton J, Blackburn P, Evans RJ, Hugenholtz J. Applications of the bacteriocin, nisin. Antonie van Leeuwenhoek 1996 69:2 [Internet]. Springer; 1996 [cited 2022 May 11];69:193\u0026ndash;202. Available from: https://link.springer.com/article/10.1007/BF00399424\u003c/li\u003e\n \u003cli\u003eDonia MS, Cimermancic P, Schulze CJ, Wieland Brown LC, Martin J, Mitreva M, et al. A systematic analysis of biosynthetic gene clusters in the human microbiome reveals a common family of antibiotics. Cell [Internet]. Elsevier Inc.; 2014;158:1402\u0026ndash;14. Available from: http://dx.doi.org/10.1016/j.cell.2014.08.032\u003c/li\u003e\n \u003cli\u003eH\u0026ouml;ltzel A, G\u0026auml;nzle MG, Nicholson GJ, Hammes WP, Jung G. The First Low Molecular Weight Antibiotic from Lactic Acid Bacteria: Reutericyclin, a New Tetramic Acid. Angew Chem Int Ed Engl [Internet]. Angew Chem Int Ed Engl; 2000 [cited 2022 May 11];39:2766\u0026ndash;8. Available from: https://pubmed.ncbi.nlm.nih.gov/10934421/\u003c/li\u003e\n \u003cli\u003eAcedo JZ, Chiorean S, Vederas JC, van Belkum MJ. The expanding structural variety among bacteriocins from Gram-positive bacteria. FEMS Microbiology Reviews [Internet]. Oxford Academic; 2018 [cited 2021 Dec 29];42:805\u0026ndash;28. Available from: https://academic.oup.com/femsre/article/42/6/805/5063573\u003c/li\u003e\n \u003cli\u003eCotter PD, Ross RP, Hill C. Bacteriocins \u0026mdash; a viable alternative to antibiotics? Nature Reviews Microbiology [Internet]. Nature Publishing Group; 2013;11:95\u0026ndash;105. Available from: http://www.nature.com/articles/nrmicro2937\u003c/li\u003e\n \u003cli\u003eHeilbronner S, Krismer B, Br\u0026ouml;tz-Oesterhelt H, Peschel A. The microbiome-shaping roles of bacteriocins [Internet]. Nature Reviews Microbiology. Nature Publishing Group; 2021 [cited 2021 Dec 6]. p. 726\u0026ndash;39. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/s41579-021-00569-w\u003c/li\u003e\n \u003cli\u003eYi Y, Li P, Zhao F, Zhang T, Shan Y, Wang X, et al. Current status and potentiality of class II bacteriocins from lactic acid bacteria: structure, mode of action and applications in the food industry. Trends in Food Science and Technology [Internet]. Elsevier Ltd; 2022;120:387\u0026ndash;401. Available from: https://doi.org/10.1016/j.tifs.2022.01.018\u003c/li\u003e\n \u003cli\u003e\u0026laquo;d Ennahar S, Sashihara T, Sonomoto K, Ishizaki A. Class IIa bacteriocins: biosynthesis, structure and activity. FEMS Microbiology Reviews [Internet]. Oxford Academic; 2000 [cited 2021 Dec 24];24:85\u0026ndash;106. Available from: https://academic.oup.com/femsre/article/24/1/85/526454\u003c/li\u003e\n \u003cli\u003eO\u0026rsquo;Leary NA, Wright MW, Brister JR, Ciufo S, Haddad D, McVeigh R, et al. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research [Internet]. Oxford Academic; 2016 [cited 2022 Feb 19];44:D733\u0026ndash;45. Available from: https://academic.oup.com/nar/article/44/D1/D733/2502674\u003c/li\u003e\n \u003cli\u003eWattam AR, Abraham D, Dalay O, Disz TL, Driscoll T, Gabbard JL, et al. PATRIC, the bacterial bioinformatics database and analysis resource. Nucleic Acids Res [Internet]. Nucleic Acids Res; 2014 [cited 2022 Feb 21];42. Available from: https://pubmed.ncbi.nlm.nih.gov/24225323/\u003c/li\u003e\n \u003cli\u003eChen IMA, Chu K, Palaniappan K, Ratner A, Huang J, Huntemann M, et al. The IMG/M data management and analysis system v.6.0: new tools and advanced capabilities. Nucleic Acids Research [Internet]. Oxford Academic; 2021 [cited 2022 Feb 19];49:D751\u0026ndash;63. Available from: https://academic.oup.com/nar/article/49/D1/D751/5943189\u003c/li\u003e\n \u003cli\u003eAlmeida A, Nayfach S, Boland M, Strozzi F, Beracochea M, Shi ZJ, et al. A unified catalog of 204,938 reference genomes from the human gut microbiome. Nature Biotechnology [Internet]. Nature Publishing Group; 2021 [cited 2021 Oct 12];39:105\u0026ndash;14. Available from: https://www.nature.com/articles/s41587-020-0603-3\u003c/li\u003e\n \u003cli\u003ePasolli E, de Filippis F, Mauriello IE, Cumbo F, Walsh AM, Leech J, et al. Large-scale genome-wide analysis links lactic acid bacteria from food with the gut microbiome. Nature Communications [Internet]. Nature Publishing Group; 2020 [cited 2021 Dec 12];11:1\u0026ndash;12. Available from: https://www.nature.com/articles/s41467-020-16438-8\u003c/li\u003e\n \u003cli\u003eBlin K, Shaw S, Kloosterman AM, Charlop-Powers Z, van Wezel GP, Medema MH, et al. antiSMASH 6.0: improving cluster detection and comparison capabilities. Nucleic Acids Research [Internet]. Oxford Academic; 2021 [cited 2021 Jul 23];49:W29\u0026ndash;35. Available from: https://academic.oup.com/nar/article/49/W1/W29/6274535\u003c/li\u003e\n \u003cli\u003eKautsar SA, van der Hooft JJJ, de Ridder D, Medema MH. BiG-SLiCE: A highly scalable tool maps the diversity of 1.2 million biosynthetic gene clusters. Gigascience. Oxford University Press; 2021;10:1\u0026ndash;17.\u003c/li\u003e\n \u003cli\u003eNayfach S, Roux S, Seshadri R, Udwary D, Varghese N, Schulz F, et al. A genomic catalog of Earth\u0026rsquo;s microbiomes. Nature Biotechnology. Nature Research; 2021;39:499\u0026ndash;509.\u003c/li\u003e\n \u003cli\u003eNavarro-Mu\u0026ntilde;oz JC, Selem-Mojica N, Mullowney MW, Kautsar SA, Tryon JH, Parkinson EI, et al. A computational framework to explore large-scale biosynthetic diversity. Nature Chemical Biology [Internet]. Nature Publishing Group; 2020 [cited 2021 Sep 18];16:60\u0026ndash;8. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/s41589-019-0400-9\u003c/li\u003e\n \u003cli\u003ePaoli L, Ruscheweyh H-J, Forneris CC, Hubrich F, Kautsar S, Bhushan A, et al. Biosynthetic potential of the global ocean microbiome. Nature [Internet]. Nature Publishing Group; 2022 [cited 2022 Jun 23];1\u0026ndash;8. Available from: https://www.nature.com/articles/s41586-022-04862-3\u003c/li\u003e\n \u003cli\u003eKautsar SA, Blin K, Shaw S, Navarro-Mu\u0026ntilde;oz JC, Terlouw BR, van der Hooft JJJ, et al. MIBiG 2.0: A repository for biosynthetic gene clusters of known function. Nucleic Acids Research. 2020;\u003c/li\u003e\n \u003cli\u003eBrockhurst MA, Harrison E, Hall JPJ, Richards T, McNally A, MacLean C. The Ecology and Evolution of Pangenomes. Current Biology. Cell Press; 2019;29:R1094\u0026ndash;103.\u003c/li\u003e\n \u003cli\u003eGavriilidou A, Kautsar SA, Zaburannyi N, Krug D, M\u0026uuml;ller R, Medema MH, et al. Compendium of specialized metabolite biosynthetic diversity encoded in bacterial genomes. Nature Microbiology [Internet]. Nature Publishing Group; 2022 [cited 2022 May 3];7:726\u0026ndash;35. Available from: https://www.nature.com/articles/s41564-022-01110-2\u003c/li\u003e\n \u003cli\u003eGeorge F, Daniel C, Thomas M, Singer E, Guilbaud A, Tessier FJ, et al. Occurrence and dynamism of lactic acid bacteria in distinct ecological niches: A multifaceted functional health perspective. Frontiers in Microbiology. Frontiers Media S.A.; 2018. p. 2899.\u003c/li\u003e\n \u003cli\u003eMeth\u0026eacute; BA, Nelson KE, Pop M, Creasy HH, Giglio MG, Huttenhower C, et al. A framework for human microbiome research. Nature 2012 486:7402 [Internet]. Nature Publishing Group; 2012 [cited 2022 Mar 16];486:215\u0026ndash;21. Available from: https://www.nature.com/articles/nature11209\u003c/li\u003e\n \u003cli\u003eHannigan GD, Prihoda D, Palicka A, Soukup J, Klempir O, Rampula L, et al. A deep learning genome-mining strategy for biosynthetic gene cluster prediction. Nucleic Acids Research [Internet]. Oxford Academic; 2019 [cited 2019 Oct 8];47:e110\u0026ndash;e110. Available from: https://academic.oup.com/nar/article/47/18/e110/5545735\u003c/li\u003e\n \u003cli\u003eWalker AS, Clardy J. A Machine Learning Bioinformatics Method to Predict Biological Activity from Biosynthetic Gene Clusters. Journal of Chemical Information and Modeling [Internet]. American Chemical Society; 2021 [cited 2019 Oct 1];61:2560\u0026ndash;71. Available from: https://pubs.acs.org/doi/full/10.1021/acs.jcim.0c01304\u003c/li\u003e\n \u003cli\u003eSkinnider MA, Johnston CW, Gunabalasingam M, Merwin NJ, Kieliszek AM, MacLellan RJ, et al. Comprehensive prediction of secondary metabolite structure and biological activity from microbial genome sequences. Nature Communications [Internet]. Nature Publishing Group; 2020 [cited 2019 Sep 30];11:1\u0026ndash;9. Available from: https://www.nature.com/articles/s41467-020-19986-1\u003c/li\u003e\n \u003cli\u003eLee SW, Mitchell DA, Markley AL, Hensler ME, Gonzalez D, Wohlrab A, et al. Discovery of a widely distributed toxin biosynthetic gene cluster. Proc Natl Acad Sci U S A [Internet]. 2008 [cited 2022 Mar 10];105:5879\u0026ndash;84. Available from: www.pnas.org/cgi/content/full/\u003c/li\u003e\n \u003cli\u003eEddy SR. Accelerated Profile HMM Searches. PLOS Computational Biology [Internet]. Public Library of Science; 2011 [cited 2022 May 23];7:e1002195. Available from: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1002195\u003c/li\u003e\n \u003cli\u003evan Heel AJ, de Jong A, Song C, Viel JH, Kok J, Kuipers OP. BAGEL4: A user-friendly web server to thoroughly mine RiPPs and bacteriocins. Nucleic Acids Research [Internet]. Oxford Academic; 2018 [cited 2022 Jan 18];46:W278\u0026ndash;81. Available from: https://academic.oup.com/nar/article/46/W1/W278/5000017\u003c/li\u003e\n \u003cli\u003eStructure, function and diversity of the healthy human microbiome. Nature [Internet]. 2012;486:207\u0026ndash;14. Available from: http://www.nature.com/articles/nature11234\u003c/li\u003e\n \u003cli\u003eFoulqui\u0026eacute; Moreno MR, Baert B, Denayer S, Cornelis P, de Vuyst L. Characterization of the amylovorin locus of Lactobacillus amylovorus DCE 471, producer of a bacteriocin active against Pseudomonas aeruginosa, in combination with colistin and pyocins. FEMS Microbiology Letters [Internet]. Oxford Academic; 2008 [cited 2022 Jun 13];286:199\u0026ndash;206. Available from: https://academic.oup.com/femsle/article/286/2/199/591282\u003c/li\u003e\n \u003cli\u003eKawai Y, Saitoh B, Takahashi O, Kitazawa H, Saito T, Nakajima H, et al. Primary Amino Acid and DNA Sequences of Gassericin T, a Lactacin F-Family Bacteriocin Produced by Lactobacillus gasseri SBT2055. Bioscience, Biotechnology, and Biochemistry [Internet]. Oxford Academic; 2000 [cited 2022 Jun 13];64:2201\u0026ndash;8. Available from: https://academic.oup.com/bbb/article/64/10/2201/5945713\u003c/li\u003e\n \u003cli\u003eAlvarez-Sieiro P, Montalb\u0026aacute;n-L\u0026oacute;pez M, Mu D, Kuipers OP. Bacteriocins of lactic acid bacteria: extending the family. Applied Microbiology and Biotechnology. 2016. p. 2939\u0026ndash;51.\u003c/li\u003e\n \u003cli\u003eMukesh Kumar M, Dhanasekaran D. Biosynthetic Gene Cluster Analysis in Lactobacillus Species Using antiSMASH. Advances in Probiotics [Internet]. Elsevier; 2021 [cited 2022 Apr 28]. p. 113\u0026ndash;20. Available from: https://linkinghub.elsevier.com/retrieve/pii/B9780128229095000071\u003c/li\u003e\n \u003cli\u003eDoroghazi JR, Albright JC, Goering AW, Ju K-S, Haines RR, Tchalukov KA, et al. A roadmap for natural product discovery based on large-scale genomics and metabolomics. Nature Chemical Biology [Internet]. Nature Publishing Group; 2014 [cited 2021 Nov 15];10:963\u0026ndash;8. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/nchembio.1659\u003c/li\u003e\n \u003cli\u003eFrance M, Alizadeh M, Brown S, Ma B, Ravel J. Towards a deeper understanding of the vaginal microbiota. Nature Microbiology. Springer US; 2022. p. 367\u0026ndash;78.\u003c/li\u003e\n \u003cli\u003eFontana F, Alessandri G, Lugli GA, Mancabelli L, Longhi G, Anzalone R, et al. Probiogenomics Analysis of 97 Lactobacillus crispatus Strains as a Tool for the Identification of Promising Next-Generation Probiotics. Microorganisms [Internet]. 2020;9:73. Available from: https://www.mdpi.com/2076-2607/9/1/73\u003c/li\u003e\n \u003cli\u003eArgentini C, Fontana F, Alessandri G, Lugli GA, Mancabelli L, Ossiprandi MC, et al. Evaluation of Modulatory Activities of Lactobacillus crispatus Strains in the Context of the Vaginal Microbiota. Microbiology Spectrum. American Society for Microbiology; 2022;\u003c/li\u003e\n \u003cli\u003eTahara T, Kanatani K. Isolation and partial characterization of crispacin A, a cell-associated bacteriocin produced by Lactobacillus crispatus JCM 2009. FEMS Microbiology Letters [Internet]. Oxford Academic; 2006 [cited 2022 May 27];147:287\u0026ndash;90. Available from: https://academic.oup.com/femsle/article/147/2/287/532692\u003c/li\u003e\n \u003cli\u003eSuez J, Zmora N, Segal E, Elinav E. The pros, cons, and many unknowns of probiotics [Internet]. Nature Medicine. 2019 [cited 2021 Dec 11]. p. 716\u0026ndash;29. Available from: https://doi.org/10.1038/s41591-019-0439-x\u003c/li\u003e\n \u003cli\u003eSorbara MT, Pamer EG. Microbiome-based therapeutics. Nature Reviews Microbiology [Internet]. Nature Publishing Group; 2022 [cited 2022 Jan 17];20:365\u0026ndash;80. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/s41579-021-00667-9\u003c/li\u003e\n \u003cli\u003ePlavec TV, Berlec A. Engineering of lactic acid bacteria for delivery of therapeutic proteins and peptides. Applied Microbiology and Biotechnology [Internet]. Springer Verlag; 2019;103:2053\u0026ndash;66. Available from: http://link.springer.com/10.1007/s00253-019-09628-y\u003c/li\u003e\n \u003cli\u003eMokoena MP. Lactic Acid Bacteria and Their Bacteriocins: Classification, Biosynthesis and Applications against Uropathogens: A Mini-Review. Molecules [Internet]. MDPI AG; 2017;22:1255. Available from: http://www.mdpi.com/1420-3049/22/8/1255\u003c/li\u003e\n \u003cli\u003eZheng J, Wittouck S, Salvetti E, Franz CMAP, Harris HMB, Mattarelli P, et al. A taxonomic note on the genus Lactobacillus: Description of 23 novel genera, emended description of the genus Lactobacillus Beijerinck 1901, and union of Lactobacillaceae and Leuconostocaceae. International Journal of Systematic and Evolutionary Microbiology [Internet]. Microbiology Society; 2020;70:2782\u0026ndash;858. Available from: https://www.microbiologyresearch.org/content/journal/ijsem/10.1099/ijsem.0.004107\u003c/li\u003e\n \u003cli\u003eOndov BD, Treangen TJ, Melsted P, Mallonee AB, Bergman NH, Koren S, et al. Mash: Fast genome and metagenome distance estimation using MinHash. Genome Biology [Internet]. BioMed Central Ltd.; 2016 [cited 2022 Feb 19];17:1\u0026ndash;14. Available from: https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0997-x\u003c/li\u003e\n \u003cli\u003eChaumeil PA, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics [Internet]. Oxford Academic; 2020 [cited 2022 Feb 19];36:1925\u0026ndash;7. Available from: https://academic.oup.com/bioinformatics/article/36/6/1925/5626182\u003c/li\u003e\n \u003cli\u003eParks DH, Chuvochina M, Rinke C, Mussig AJ, Chaumeil P-A, Hugenholtz P. GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Research [Internet]. Oxford Academic; 2022 [cited 2022 Feb 25];50:D785\u0026ndash;94. Available from: https://academic.oup.com/nar/article/50/D1/D785/6370255\u003c/li\u003e\n \u003cli\u003eParks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil PA, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nature Biotechnology 2018 36:10 [Internet]. Nature Publishing Group; 2018 [cited 2022 Feb 25];36:996\u0026ndash;1004. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/nbt.4229\u003c/li\u003e\n \u003cli\u003eParks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil PA, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nature Biotechnology 2018 36:10 [Internet]. Nature Publishing Group; 2018 [cited 2022 Mar 1];36:996\u0026ndash;1004. Available from: https://www-nature-com.eproxy.lib.hku.hk/articles/nbt.4229\u003c/li\u003e\n \u003cli\u003eVirtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nature Methods 2020 17:3 [Internet]. Nature Publishing Group; 2020 [cited 2022 Mar 19];17:261\u0026ndash;72. Available from: https://www.nature.com/articles/s41592-019-0686-2\u003c/li\u003e\n \u003cli\u003ePedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research [Internet]. 2011 [cited 2022 Mar 19];12:2825\u0026ndash;30. Available from: http://scikit-learn.sourceforge.net.\u003c/li\u003e\n \u003cli\u003eFrance MT, Fu L, Rutt L, Yang H, Humphrys MS, Narina S, et al. Insight into the ecology of vaginal bacteria through integrative analyses of metagenomic and metatranscriptomic data. Genome Biology 2022 23:1 [Internet]. BioMed Central; 2022 [cited 2022 Mar 5];23:1\u0026ndash;26. Available from: https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02635-9\u003c/li\u003e\n \u003cli\u003eLeinonen R, Sugawara H, Shumway M. The Sequence Read Archive. Nucleic Acids Research [Internet]. Oxford University Press; 2011 [cited 2022 Mar 19];39:D19. Available from: /pmc/articles/PMC3013647/\u003c/li\u003e\n \u003cli\u003eChen S, Zhou Y, Chen Y, Gu J. Fastp: An ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018. p. i884\u0026ndash;90.\u003c/li\u003e\n \u003cli\u003eFrankish A, Diekhans M, Ferreira AM, Johnson R, Jungreis I, Loveland J, et al. GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Research [Internet]. Oxford Academic; 2019 [cited 2022 Mar 19];47:D766\u0026ndash;73. Available from: https://academic.oup.com/nar/article/47/D1/D766/5144133\u003c/li\u003e\n \u003cli\u003eKopylova E, No\u0026eacute; L, Touzet H. SortMeRNA: fast and accurate filtering of ribosomal RNAs in metatranscriptomic data. Bioinformatics [Internet]. Oxford Academic; 2012 [cited 2022 Jun 18];28:3211\u0026ndash;7. Available from: https://academic.oup.com/bioinformatics/article/28/24/3211/246053\u003c/li\u003e\n \u003cli\u003eBeghini F, McIver LJ, Blanco-M\u0026iacute;guez A, Dubois L, Asnicar F, Maharjan S, et al. Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with biobakery 3. Elife. eLife Sciences Publications Ltd; 2021;10.\u003c/li\u003e\n \u003cli\u003ePascal Andreu V, Augustijn HE, van den Berg K, van der Hooft JJJ, Fischbach MA, Medema MH. BiG-MAP: an Automated Pipeline To Profile Metabolic Gene Cluster Abundance and Expression in Microbiomes. mSystems [Internet]. American Society for Microbiology; 2021 [cited 2022 Mar 19];6. Available from: https://journals.asm.org/doi/abs/10.1128/mSystems.00937-21\u003c/li\u003e\n \u003cli\u003eLangmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nature Methods. 2012;9:357\u0026ndash;9.\u003c/li\u003e\n \u003cli\u003eLiao Y, Smyth GK, Shi W. FeatureCounts: An efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014;\u003c/li\u003e\n \u003cli\u003eBischl B, Lang M, Kotthoff L, Schiffner J, Richter J, Studerus E, et al. mlr: Machine Learning in R. Journal of Machine Learning Research [Internet]. 2016 [cited 2022 Jun 18];17:1\u0026ndash;5. Available from: https://github.com/mlr-org/mlr\u003c/li\u003e\n \u003cli\u003eGu Z, Gu L, Eils R, Schlesner M, Brors B. circlize implements and enhances circular visualization in R. Bioinformatics [Internet]. Oxford Academic; 2014 [cited 2022 Jun 20];30:2811\u0026ndash;2. Available from: https://academic.oup.com/bioinformatics/article/30/19/2811/2422259\u003c/li\u003e\n \u003cli\u003eAllaire JJ, Ellis P, Gandrud C, Kuo K, Lewis BW, Owen J, et al. Package \u0026ldquo;networkD3.\u0026rdquo; 2017 [cited 2022 Jun 20]; Available from: https://github.com/christophergandrud/networkD3/issues\u003c/li\u003e\n \u003cli\u003eSantos-Aberturas J, Chandra G, Frattaruolo L, Lacret R, Pham TH, Vior NM, et al. Uncovering the unexplored diversity of thioamidated ribosomal peptides in Actinobacteria using the RiPPER genome mining tool. Nucleic Acids Research [Internet]. Oxford Academic; 2019 [cited 2022 Jun 18];47:4624\u0026ndash;37. Available from: https://academic.oup.com/nar/article/47/9/4624/5420534\u003c/li\u003e\n \u003cli\u003eLi W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics [Internet]. Oxford Academic; 2006 [cited 2022 Jun 18];22:1658\u0026ndash;9. Available from: https://academic.oup.com/bioinformatics/article/22/13/1658/194225\u003c/li\u003e\n \u003cli\u003eSangar V, Blankenberg DJ, Altman N, Lesk AM. Quantitative sequence-function relationships in proteins based on gene ontology. BMC Bioinformatics [Internet]. BioMed Central; 2007 [cited 2022 May 23];8:1\u0026ndash;15. Available from: https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-8-294\u003c/li\u003e\n \u003cli\u003eBuchfink B, Reuter K, Drost HG. Sensitive protein alignments at tree-of-life scale using DIAMOND. Nature Methods 2021 18:4 [Internet]. Nature Publishing Group; 2021 [cited 2022 Jun 20];18:366\u0026ndash;8. Available from: https://www.nature.com/articles/s41592-021-01101-x\u003c/li\u003e\n \u003cli\u003eRice P, Longden L, Bleasby A. EMBOSS: The European Molecular Biology Open Software Suite. Trends in Genetics [Internet]. Elsevier; 2000 [cited 2022 Jun 20];16:276\u0026ndash;7. Available from: http://www.cell.com/article/S0168952500020242/fulltext\u003c/li\u003e\n \u003cli\u003eKatoh K, Standley DM. MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. Molecular Biology and Evolution [Internet]. Oxford Academic; 2013 [cited 2022 Jun 20];30:772\u0026ndash;80. Available from: https://academic.oup.com/mbe/article/30/4/772/1073398\u003c/li\u003e\n \u003cli\u003eWaterhouse AM, Procter JB, Martin DMA, Clamp M, Barton GJ. Jalview Version 2\u0026mdash;a multiple sequence alignment editor and analysis workbench. Bioinformatics [Internet]. Oxford Academic; 2009 [cited 2022 Jun 20];25:1189\u0026ndash;91. Available from: https://academic.oup.com/bioinformatics/article/25/9/1189/203460\u003c/li\u003e\n \u003cli\u003eMinh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Molecular Biology and Evolution [Internet]. Oxford Academic; 2020 [cited 2022 Feb 28];37:1530\u0026ndash;4. Available from: https://academic.oup.com/mbe/article/37/5/1530/5721363\u003c/li\u003e\n \u003cli\u003eKalyaanamoorthy S, Minh BQ, Wong TKF, von Haeseler A, Jermiin LS. ModelFinder: Fast model selection for accurate phylogenetic estimates. Nature Methods [Internet]. Nat Methods; 2017 [cited 2021 Jul 25];14:587\u0026ndash;9. Available from: https://pubmed.ncbi.nlm.nih.gov/28481363/\u003c/li\u003e\n \u003cli\u003eLetunic I, Bork P. Interactive Tree of Life (iTOL) v4: Recent updates and new developments. Nucleic Acids Research [Internet]. Nucleic Acids Res; 2019 [cited 2021 Jul 25];47. Available from: https://pubmed.ncbi.nlm.nih.gov/30931475/\u003c/li\u003e\n \u003cli\u003eClinical Laboratory Standards Institute. Reference Method for Broth Dilution Antifungal susceptibility testing of yeast. Clinical Laboratory Standards Institute [Internet]. 2008 [cited 2022 Jun 20];22. Available from: www.clsi.org.\u003c/li\u003e\n \u003cli\u003eOksanen J. Vegan: ecological diversity [Internet]. R Package Version 2.4-4. 2017. p. 11. Available from: https://cran.r-project.org/package=vegan\u003c/li\u003e\n \u003cli\u003eConway JR, Lex A, Gehlenborg N. UpSetR: an R package for the visualization of intersecting sets and their properties. Bioinformatics [Internet]. Oxford Academic; 2017 [cited 2022 Jun 20];33:2938\u0026ndash;40. Available from: https://academic.oup.com/bioinformatics/article/33/18/2938/3884387\u003c/li\u003e\n \u003cli\u003ePackage \u0026ldquo;Rtsne\u0026rdquo; Title T-Distributed Stochastic Neighbor Embedding using a Barnes-Hut Implementation. 2022 [cited 2022 Jun 20]; Available from: https://github.com/jkrijthe/Rtsne\u003c/li\u003e\n \u003cli\u003eRevelle W. Package \u0026ldquo;psych\u0026rdquo; - Procedures for Psychological, Psychometric and Personality Research. R Package [Internet]. 2015 [cited 2022 Jun 20];1\u0026ndash;358. Available from: https://www.scholars.northwestern.edu/en/publications/psych-procedures-for-personality-and-psychological-research\u003c/li\u003e\n \u003cli\u003eBenjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological). 1995;57:289\u0026ndash;300.\u003c/li\u003e\n \u003cli\u003eKolde R, others. Pheatmap: pretty heatmaps. R package version. 2012;1:726.\u003c/li\u003e\n \u003cli\u003eShannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, et al. Cytoscape: A software Environment for integrated models of biomolecular interaction networks. Genome Research. 2003;\u003c/li\u003e\n \u003cli\u003eVillanueva RAM, Chen ZJ. ggplot2: elegant graphics for data analysis. Taylor \\\u0026amp; Francis; 2019.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Supplementary Tables","content":"\u003cp\u003eSupplementary Tables S1-S9 are not available with this version\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"microbiome","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mbio","sideBox":"Learn more about [Microbiome](http://microbiomejournal.biomedcentral.com/)","snPcode":"40168","submissionUrl":"https://submission.nature.com/new-submission/40168/3","title":"Microbiome","twitterHandle":"@MicrobiomeJ","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Lactic acid bacteria, Biosynthetic gene clusters, Secondary metabolites, Bacteriocins, Human microbiome, Vaginal microbiome","lastPublishedDoi":"10.21203/rs.3.rs-1868011/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1868011/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e Lactic acid bacteria (LAB) produce various bioactive secondary metabolites (SMs), which endow LAB with a protective role for the host. However, the biosynthetic potentials of LAB-derived SMs remain elusive, particularly in their diversity, abundance, and distribution in the human microbiome. Thus, it is still unknown to what extent LAB-derived SMs are involved in microbiome homeostasis.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eHere, we systematically investigate the biosynthetic potential of LAB from 31,977 LAB genomes, identifying 130,051 BGCs of 2,849 gene cluster families (GCFs). Most of these GCFs are species-specific or even strain-specific and uncharacterized yet. Analyzing 748 human-associated metagenomes, we gain an insight into the profile of LAB BGCs, which are highly diverse and niche-specific in the human microbiome. We discover that most LAB BGCs may encode bacteriocins with pervasive antagonistic activities predicted by machine learning models, potentially playing protective roles in the human microbiome. Class II bacteriocins, one of the most abundant and diverse LAB SMs, are particularly enriched and predominant in the vaginal microbiomes. Together with experimental validation, our metagenomic and metatranscriptomic analysis show that antagonistic class II bacteriocins potentially regulate microbial communities in the vagina, thereby contributing to microbiome homeostasis. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e Our study systematically investigates LAB biosynthetic potential and their profile in the human microbiome, linking them to the antagonistic contributions to microbiome homeostasis via omics analysis. These discoveries of the diverse and prevalent antagonistic SMs are expected to stimulate the mechanism study of LAB’s protective roles for the microbiome and host, highlighting the potential of LAB and their bacteriocins as therapeutic alternatives. \u003c/p\u003e","manuscriptTitle":"A systematically biosynthetic investigation of lactic acid bacteria reveals diverse antagonistic bacteriocins that potentially shape the human microbiome","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-07-26 19:16:14","doi":"10.21203/rs.3.rs-1868011/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-02-22T21:48:29+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-01-27T18:05:34+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-12-19T13:51:34+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"22b9a8c5-5495-4403-83d5-596db443425d","date":"2022-12-12T10:02:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"4c37b638-edae-40b4-9c1b-da3d67d3fe6f","date":"2022-12-12T07:45:12+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-12-07T14:20:45+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-07-21T04:30:32+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-07-20T00:58:54+00:00","index":"","fulltext":""},{"type":"submitted","content":"Microbiome","date":"2022-07-18T02:51:03+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"microbiome","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mbio","sideBox":"Learn more about [Microbiome](http://microbiomejournal.biomedcentral.com/)","snPcode":"40168","submissionUrl":"https://submission.nature.com/new-submission/40168/3","title":"Microbiome","twitterHandle":"@MicrobiomeJ","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1591d87c-8674-43d3-b1ac-9697690786e4","owner":[],"postedDate":"July 26th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T20:46:08+00:00","versionOfRecord":{"articleIdentity":"rs-1868011","link":"https://doi.org/10.1186/s40168-023-01540-y","journal":{"identity":"microbiome","isVorOnly":false,"title":"Microbiome"},"publishedOn":"2023-04-27 20:36:26","publishedOnDateReadable":"April 27th, 2023"},"versionCreatedAt":"2022-07-26 19:16:14","video":{"identity":"fcadc6b06af8cf2aaab84dae92f4ff35"},"vorDoi":"10.1186/s40168-023-01540-y","vorDoiUrl":"https://doi.org/10.1186/s40168-023-01540-y","workflowStages":[]},"version":"v1","identity":"rs-1868011","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1868011","identity":"rs-1868011","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0