Deciphering the microbial landscape of lower respiratory tract in health and respiratory diseases | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Deciphering the microbial landscape of lower respiratory tract in health and respiratory diseases cheng cheng, Yangqian Li, Suyan Wang, Haoyu Wang, Dan Liu, QingLan Wang, and 13 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7937892/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The microbiota plays an important role in maintaining lung homeostasis and the development of respiratory diseases. However, there have not been detailed characteristics of lung microbiome in health and respiratory diseases. Bronchoalveolar lavage fluid of 43 healthy individuals, 23 COVID patients, and 28 non-small cell lung cancerpatients was conducted to characterize microbial diversity and function by metagenomic sequencing. We identified 196 species in health lung after removing low abnudance species. The most abundant species were Staphylococcus aureus, Salmonella enterica, and Klebsiella pneumoniae. Keystone species were identified through co-network analysis, such as Granulicatella adiacens and Mogibacterium diversum. To further explore the microbial function in the healthy lung, we obtained 56 non-redundant metagenome-assembled genomes (MAGs) and subsequently classified them into nine phyla. Metabolism and genetic information processing exhibit high abundance, including bacterial secretion systems, protein export, and Staphylococcus aureus infection, indicating a close correlation between microbiota and the host. Furthermore, the microbiome pairwise distances of non-small cell lung cancer (NSCLC) participants were more similar to those of the healthy group than those of individuals with COVID-19. Rothia dentocariosa, Corynebacterium kefirresidentii, Sphingobium yanoikuyae were enriched in the healthy group, while Schaalia odontolytica, Actinomyces graevenitzii, and Rothia mucilaginosa was enriched in the NSCLC group. Overall, this study provides a comprehensive overview of the diversity and gene function of the lung microbiome, which will facilitate future studies of microbiota associated with respiratory diseases in humans. metagenome lower respiratory tract pulmonary microbiome COVID-19 lung cancer Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction The lung microbiota plays an important role in the human lung respiratory system[ 1 ]. Numerous studies have indicated that the diversity and composition of the respiratory microbiome in healthy individuals exhibit individual variability but are generally stable [ 2 – 6 ]. Using high-throughput sequencing, scientists can more accurately describe the composition of the lung microbiota in healthy individuals, including the abundance and functions of microorganisms [ 7 ]. Furthermore, the development of metagenomic binning technology has enabled the acquisition of nearly complete metagenome-assembled genomes (MAGs) on a large scale [ 8 ]. The microbiota composition in healthy human lungs has been partially characterized primarily through 16SrRNA amplicon sequencing [ 9 , 10 ]. However, there have been limited assemblies of MAGs specifically targeting the healthy lung microbiome. Due to clinical sample acquisition difficulty in healthy individuals and limited sequencing depth, the availability of sequenced genomes and functional information for most healthy individual lung microbes remains limited. In our previous studies, we showed pulmonary microbiome dysbiosis in lung cancer patients based on non-assembled metagenomic data[ 11 , 12 ]. A large-scale and high-depth metagenomic sequencing scan is needed to perform to characterize microbiota compositions and functions in healthy human lungs. Here we address this problem by constructing co-network of bacterium based on their relative abundance, reconstructing draft prokaryotic genomes from 43 healthy individuals. Furthermore, we revealed the microbiota alterations in lung diseases focused on COVID-19 and non-small cell lung cancer (NSCLC) using microbial profiling of this healthy cohort. Our study will provide more comprehensive baseline data for research on human pulmonary microbiota, and better understand its mechanisms of action in both health and disease states. Methods Sample collection and processing All details about the experimental samples are listed in Table S1 , including 43 healthy individuals, 23 COVID-19 patients, and 28 NSCLC patients. Healthy individuals were recruited from Suining Central Hospital, Sichuan Province, in 2021. They underwent CT imaging, pulmonary function tests, electrocardiograms, and routine blood tests, all showing normal liver and kidney function. COVID-19 patients were diagnosed with SARS-CoV-2 infection through qPCR or antigen testing. NSCLC patients were histologically and endoscopically confirmed by examination from three pathologists. The samples of those patients were obtained from the West China Hospital of Sichuan University. All patients were followed up to confirm the final diagnosis. In this study, samples were collected via bronchoscopy for distal alveolar lavage. We used only bronchoscope working channel washes, which were done twice with a minimum volume of 15 mL of 0.9% saline solution. Immediately dispense the collected BALF into 1.8 mL sterile freezing tubes and store at -80°C until processing. The collection process was carried out following aseptic procedures to prevent contamination from environmental, human commensal, and miscellaneous bacteria. DNA extraction and quality control DNA extraction from BALF was performed using the DNA Microbiome kit (51704, Qiagen, USA) according to the manufacturer’s instructions. The host cells are differentially lysed, followed by enzymatic digestion, to efficiently remove host DNA. Subsequently, complete cell lysis is achieved through physical and chemical methods, facilitating the isolation of bacterial DNA. Library construction and sequencing Using an NGS Nextera™ DNA Flex Library Prep kit (Illumina), the DNA libraries were conducted. By performing an enzymatic reaction called segmentation, DNA is fragmented, followed by the addition of adapter sequences and PCR amplification. An Agilent 4200 Bioanalyzer (Agilent Technologies, Santa Clara, USA) was used to assess the quality of the DNA libraries. After a quality check, Illumina PE150 sequencing was conducted by pooling multiple libraries based on the required effective concentration and target data amount. Sequencing was performed on the Illumina NovaSeq 6000. Data quality control For the quality control of raw data obtained from the Illumina sequencing platform, mocat2 (v2.1.3) was used to filter out low-quality reads, and reads aligned to the GRCh38 Homo sapiens reference genome were removed [ 34 ]. The remaining high-quality reads were further analyzed. Taxonomic classification Taxonomic classification was conducted with Kraken2 using the standard Kraken2 database ( https://benlangmead.github.io/aws-indexes/k2 ) containing archaea, bacteria, viruses, plasmids, human1, and UniVec_Core. Bracken [ 35 ] was then used to calculate the abundance of the identified species from the Kraken2 analysis. Diversity analysis Bacterial diversity was assessed by calculating the Shannon index using the vegan package (v2.5.7) in R. The Wilcoxon test was conducted to determine the statistical significance of differences among groups of healthy individuals. Differences associated with a P-value of 0.8 and statistically significant (P-value < 0.01) were visualized using the R package igraph (v2.0.2) with a layout based on the Fruchterman-Reingold algorithm. We also calculated degree and page rank values to facilitate the identification of keystone species. Network module detection was performed using the greedy algorithm. Microbiome-functional pathway analysis To identify microbial pathways, the human removed reads were analysed with HUMAnN (v3.7) [ 36 ] by using the databases of uniref 90[ 37 ] and pathway MetaCyc[ 38 ]. The “WGCNA” package[ 39 ] was used to conduct WGCNA (weighted gene co-expression network analysis). For correlation analysis, spearman correlations were calculated using function corr.test in R. Metagenomic de novo assembly and binning The Megahit module from metaWRAP-Binning was utilized to assemble clean data from the aforementioned processed Illumina sequencing data [ 40 ]. Metagenomic binning was subsequently conducted using MaxBin2, metaBAT2, and CONCOCT [ 40 ]. DAS Tool (v1.1.5) was used to aggregate multiple binning predictions into a new and enhanced bin set [ 41 ]. CheckM2 was used to evaluate the quality of the assembled bins, which were screened based on the criteria of completeness (> 50%) and contamination (< 10%)[ 42 ]. Then dRep (v3.0.0) was then used to remove duplicate bins with the parameter "-sa 0.95". After analysis, 56 MAGs were retained. Taxonomic assignment and phylogenetic analysis MAGs were taxonomically classified using the Genome Taxonomy Database Toolkit (GTDB-Tk) (v2) with "classify_wf" function and default parameters[ 43 ]. The Average Nucleotide Identity (ANI) values between the five new MAGs and genomes from two reference cohorts were calculated using fastANI (v1.33). All phylogenetic trees of the 56 MAGs were constructed using iqtree2 (v2.2.3) ( https://github.com/iqtree/iqtree2 ). Finally, iTOL ( https://itol.embl.de ) was used to generate a phylogenetic tree of the 56 MAGs. Function annotation of MAGs High-quality MAGs were utilized for gene prediction by Prodigal (v2.6.3) software ( https://github.com/hyattpd/Prodigal ) [ 44 ]. All complete genes were clustered at 90% protein identification using CD-HIT (v4.8.1) [ 45 ]. The KEGG (Kyoto Encyclopedia of Genes and Genomes) annotation results were extracted using KofamKOALA (v1.3.0) software [ 46 ]. Additionally, key virulence factors were identified using the Virulence Factor Database (VFDB) [ 47 ] via BLAST (v2.10.1+). Calculation of the abundances of MAGs and genes The quant_bins module from metaWRAP-Binning was used to quantify the abundance of each MAG in each sample. BWA MEM was used to align clean reads from each sample to gene catalogs after removing human contamination [ 48 ]. The abundances were normalized to fragments per kilobase of gene sequence per million reads mapped (FPKM). The abundances of microbial taxa, KEGG pathways, and Virulence Factors (VFs) were calculated by combining the abundances of all the members within each category. Microbiome analysis of three groups Principal Coordinates Analysis (PCoA) based on the Bray–Curtis distance matrix was utilized to visualize the compositional profiles of 212 species from samples belonging to three groups. The differences in microbiome composition between different phenotypes were calculated using permutational multivariate analysis of variance with distance matrices in the adonis2 function of the vegan package, with 999 permutations. Associations of specific microbial species with phenotypes were calculated using multivariate analysis by linear models (MaAsLin2) ( http://huttenhower.sph.harvard.edu/galaxy ) with healthy and lung-diseased samples as references. The p-value of associations was computed by MaAsLin2, and the false discovery rate (FDR) was calculated using the Benjamini–Hochberg correction. FDR < 0.01 was considered significant. Furthermore, the fold change was calculated using R (v4.3.2) with the Wilcoxon test. Results Distribution of species composition in the healthy human population We conducted meta-genomics profiling of bronchoalveolar lavage fluid (BALF) from 43 healthy individuals, including 11 males with smoking history (Fig. 1 a-b; Table S2 ). We obtained an average of 4,683,836,943 bases in each sample after quality control and removal of human-derived reads (Table S2 ). Using Kraken2 profiling, we identified 196 species with a relative abundance higher than 0.01 and present in more than 5% of individuals, accounting for 0.991 of the relative abundance per sample from healthy individuals (Table S3 -4). Those 196 species were assigned to eight phyla, with the dominant phyla being Firmicutes , Proteobacteria , and Actinobacteria (Fig. 1 c). In Ibironke et al.’s study, the three phyla mentioned above were identified as dominant phyla in the human respiratory microbiome of five individuals based on full-length 16SrRNA sequencing[ 13 ]. To identify microbial species may play a critical role in the maintaining of lung homeostasis, we examined our cohort for microbial taxa with high abundance and core microbes present in over 95% of individuals. We detected 15 high-abundance species as dominant species (Fig. 1 d-e). Meanwhile, we observed 32 core species (> 5% of individuals) out of 196 species. 12 out of the 32 core microbes are also dominant species, indicating their critical roles in the lung ecosystem. We further investigated the microbial component in our cohort through principal coordinate analysis (PCoA) and founded that there was no significantly separation based on gender, age, smoking history, or family cancer history (Figure S1 a). Furthermore, the alpha diversity of individuals grouped by the aforementioned clinical features was insignificant, indicating that differences in microbial communities between clinical groups were minor (Figure S1 b). To identify keystone species and explore relationships among species. Spearman’s correlation coefficients were calculated, and the ecological relationships across 155 different microbes (a relative abundance > 0.001 and present in more than 60% of individuals) were visualized using co-occurrence networks. There were 155 nodes and 1084 edges (Fig. 1 f). The top five species with the highest degree and page rank values were Granulicatella adiacens , Mogibacterium diversum , Streptococcus australis , Mogibacterium pumilum , and Streptococcus sp.116-D4 (Fig. 1 f). Two dominant species, Staphylococcus aureus and Salmonella enterica , have relatively low degree and page rank values (Figure S1 d). Additionally, 12 communities were detected (Figure S1 c). Firmicutes , Proteobacteria , and Actinobacteria comprised the three largest communities. Module one was dominated by Proteobacteria , module two by Bacteroidetes and Firmicutes , and module three by Firmicutes and Actinobacteria . The phylum of the species was significantly associated with the corresponding module (P value < 0.001, Fisher test), suggesting that species within the same phylum may perform similar roles in the healthy lung micro-ecology. The results indicate that the three major phyla ( Firmicutes , Proteobacteria , and Actinobacteria ), based on abundance and co-network analysis played a crucial role in lung homeostasis. Metabolic profiling of microbe from the healthy human lung To illustrate the metabolic composition of healthy individuals, we identified 374 pathways from the MetaCyc database. The top 15 relative abundances of metabolic terms are shown in Fig. 2 a-b. These dominant metabolic terms include aerobic respiration I, aminoimidazole ribonucleotide biosynthesis, adenosine nucleotides de novo biosynthesis, guanosine ribonucleotides biosynthesis, and others. We conducted weighted gene co-expression network analysis (WGCNA) to explore the relationship between the 374 pathways and basic clinical information. Two modules were obtained (Figure S2 a). However, we did not find modules significantly related to the basic clinical information (Figure S2 b). Ethanolamine Utilization and Pyruvate Fermentation to Isobutanol was enriched in the female group compared to the male group (Fig. 2 c). UMP biosynthesis II was positively correlated with age (Fig. 2 d). Our results suggest that the overall metabolic composition of microbes in BALF is not related to basic clinical information and indicates a stable state. Reconstruction of microbial genomes from the healthy human lung The workflow of binning is shown in Fig. 3 a. Firstly, assembly was performed based on single-sample and co-sample strategies after quality control on raw sequencing data and removing human contamination. The density of N50, the number of contigs, the length of the largest contigs, and the total length were shown in Fig. 3 b and Figure S3 a. The length of the largest contig and the total length of contigs from co-samples were larger than those from single samples. Microbial genomes representing individual bacterial species were then constructed from the assembled metagenomic sequencing data obtained from the 43 samples described above. We used the metaWRAP-Binning module to generate 844 bins from single-sample assemblies and 212 bins from mixed-sample assemblies. After dereplication, aggregation, and scoring filtering, we observed 121 bins. Furthermore, we obtained 84 metagenome-assembled genomes (MAGs) after quality assessment, meeting the criteria of > 50% completeness and 0.95). A final set of 56 non-redundant MAGs were obtained. Particularly, 14 MAGs were from co-assembly (Figure S3 b), indicating that co-assembly expands the number of detected MAGs. Among these 56 MAGs, 27 MAGs met the medium quality criteria (> 50% completeness and 90% completeness and < 5% contamination). Furthermore, various indexes of high-quality MAGs were significantly higher than the medium-quality ones, including completeness, contamination, contig number, and the length of the largest contig length (Figure S3 c). All MAGs exhibited a comparatively high prevalence in the 43 metagenomes (Figure S3 d), indicating that these MAGs were likely strains of core species. The 56 MAGs were subsequently classified into taxa using the Genome Taxonomy Database Toolkit (GTDB-Tk) (Table S5). Further analysis showed that 48 out of the 56 MAGs were identified at the species level (Fig. 3 c). As shown in Fig. 3 d, those MAGs covered nine bacterial phyla. Most MAGs belonged to Actinobacteriota (18 MAGs), followed by Firmicutes_A (12 MAGs), Firmicutes (eight MAGs), Proteobacteria (three MAGs), Fusobacteriota (two MAGs), Bacteroidota (two MAGs), Spirochaetota (one MAG), Patescibacteria (one MAG), and Campylobacterota (one MAG). Among those phyla, Actinobacteriota, Firmicutes, Proteobacteria, Fusobacteriota, Bacteroidota , and Spirochaetota were also detected in Kraken2, which belonged to the 196 species set. Moreover, we successfully identified Corynebacterium argentoratense and Haemophilus_A parahaemolyticus , which had low mean relative abundance (< 1%). New species identified in healthy individual lung Meanwhile, we uncovered some fascinating results regarding the identification of new bacterial species. The other eight MAGs not identified at the species level were defined as potential novel species and assigned to three phyla: Patescibacteria (three MAG), Firmicutes_A (three MAG), and Firmicutes (two MAG) (Fig. 3 d). The phylogenetic tree of the 56 MAGs was constructed based on 120 conserved proteins. The taxonomy of these MAGs at the phylum level was consistent with the phylogenetic tree (Fig. 3 e, Table S4 ). Particularly, three out of four MAGs in Proteobacteria were potential new species. Five out of the eight potential MAGs were highly qualified and had an ANI value less than 0.95. The five MAGs were further compared with bacterial genomes recently reported from two cohorts [ 14 , 15 ]. Among these bacterial species, ANI with these five MAGs was less than 0.95, indicating that they represent unknown species identified for the first time in this study. Functional characterizations of high-quality MAGs We analyzed the gene functions of 29 high-quality MAGs to gain a better understanding of the functions of the lung microbiota. At KEGG level 1, there were six identified pathways, including metabolism (109 pathways in level 2), human diseases (32 pathways in level 2), organismal systems (27 pathways in level 2), environmental information processing (13 pathways in level 2), cellular processes (12 pathways in level 2), and genetic information processing (12 pathways in level 2). In total, we identified 205 pathways at level 2. Among these, 118 pathways (57.56%) were consistently present in all samples, indicating that these pathways represent core functions of the lung microbiota. The most abundant pathways were depicted in Fig. 3 f. The largest number of pathways in level 2 were related to metabolism and genetic information processing, indicating their critical role in the lung micro-ecosystem. Particularly, the bacterial secretion system, protein export, and staphylococcus aureus infection were also detected, indicating the cross-talk between bacteria and the host. The observation of staphylococcus aureus infection is supported by the high relative abundance of staphylococcus aureus detected using Kraken2. The prevalence of viral factors from 29 highly qualified MAGs was assessed based on the Virulence Factor Database (VFDB). A total of 575 virulence factors were identified. The top 20 high-abundance virulence factors are shown in Fig. 3 g, including ABC transport-associated virulence genes. Taxonomic and functional profiles indicate minimal alterations in COVID-19 and NSCLC To explore the differences in microbial profiling between healthy individuals and lung-diseased patients, we selected COVID-19 and NSCLC as representatives of infectious disease and cancer, respectively. We identified 212 species from 23 COVID-19 patients, 28 NSCLC patients, and 43 individuals in the healthy group after quality control and removal of human reads. Overall, the lung microbiome compositions, based on the beta-diversity metrics of the three groups, were minimally different (Fig. 4 a). Moreover, COVID-19 samples were more scattered, while normal samples cluster closer together. The microbiome pairwise distances of NSCLC participants were more similar to those of the healthy group than those of individuals with COVID-19 (Fig. 4 b). Meanwhile, the top 15 species with high relative abundance in the healthy group changed more significantly in the COVID-19 group compared to NSCLC (Fig. 4 c). To further explore the microbiota changes between healthy and lung-diseased individuals, MaAsLin2 analysis was conducted based on species detected via Kraken2. We identified 47 high and 14 lower species in the comparison between the healthy group and the lung-diseased group, 23 high and 118 lower species in the COVID-19 versus healthy group comparison, and 44 high and 33 lower species in the NSCLC versus healthy group comparison (Table S6). Among the top 15 high-abundance bacteria in the healthy group, Escherichia coli , Tropheryma whipplei , and Aeromonas hydrophila were significantly more abundant in the healthy group compared to the lung disease group. Seven out of 32 representative species in the healthy group showed enrichment compared to the lung disease group, emphasizing their potential role in maintaining lung homeostasis. The top 50 significantly different microbial species were displayed in Fig. 4 d. Among those species, 20 were significantly enriched in the healthy group, while 30 were in the lung disease groups. Among those 50 species, Parvimonas micra also belonged to the 15 high-abundance species in the healthy group. Next, we calculated differences in the relative abundance of bacteria among the three groups with a threshold of absolute value of fold change > 1 and P-value < 0.01 using the Wilcoxon test (Table S7 and S8; Fig. 4 e-f). Rothia dentocariosa , Corynebacterium kefirresidentii, Sphingobium yanoikuyae were enriched in the healthy group, while Schaalia odontolytica was enriched in the NSCLC group. According to a case report, Rothia dentocariosa was cultured from BALF and caused opportunistic pulmonary infection [ 16 ]. Corynebacterium kefirresidentii was identified as a dominant member of the human skin microbiome through 16SrRNA sequencing [ 17 ]. Sphingobium yanoikuyae can degrade carcinogenic products and may reduce in gastric cancer [ 18 ]. Schaalia odontolytica was linked to resistance to neoadjuvant chemoradiotherapy in rectal cancer [ 19 ]. The relative abundance of MAGs among the three groups was also assessed (Table S9 and S10; Fig. 4 f). Actinomyces graevenitzii and Rothia mucilaginosa were significantly enriched in NSCLC. Actinomyces graevenitzii was isolated and cultured from the BALF of a patient with pulmonary actinomycosis [ 2 ]. Discussion This study provides a comprehensive dataset of microbiota, an exhaustive catalog of MAGs, identifies potential new species in healthy lungs, and explores microbial alterations in lung diseases. Microbial profiling of 196 non-low-abundance species from 43 healthy individuals was conducted. The 15 dominant species consist of communal pathogens, opportunistic pathogens, and non-pathogenic commensals. For instance, Staphylococcus aureus is detected in 30% of the population, mostly colonized in the nasal, throat, skin, and gastrointestinal tract [ 20 ], and has the potential to cause pneumonia. Klebsiella pneumoniae has emerged as a major clinical health threat due to the rise of multidrug-resistant strains [ 21 ]. Interestingly, Rothia mucilaginosa has an anti-inflammatory effect induced by pathogens or lipopolysaccharides in chronic lung disease [ 22 ]. Proteins produced by Streptococcus mitis and Streptococcus oralis from viral pneumonia patients significantly increased virus replication in lung epithelial cells [ 23 ]. And Streptococcus mitis was linked to preserved lung function and favorable survival [ 24 ]. Therefore, our results indicated that the dominant and keystone bacterial species consisted of both pathogens and non-pathogens, suggesting a complex microbial community in the lower respiratory tract of healthy lungs. It has been reported that pathogens, including Streptococcus pneumoniae and Haemophilus influenzae were found to be colonized in the airways of 20–50% of healthy individuals [ 25 ]. With an increasing number of bioinformatic methods being utilized to distinguish pathogens from background microorganisms, our findings may assist in identifying pathogens in the clinical diagnosis of pulmonary infections[ 26 ]. We also identified a significant amount of oral-associated microbiota in the lower respiratory tract (LTR) of healthy adult participant. In keystone species obtained through co-occurrence network analysis, Granulicatella adiacens belongs to microbiota in the oral cavity, urogenital tract, and intestinal tract [ 27 ] and can cause periodontitis [ 28 ]. Mogibacterium diversum was detected in human saliva samples [ 29 ]. Salmonella enterica has been isolated from stool samples of individuals with diarrhea in Grace[ 30 ]. Meanwhile, the dominant species Streptococcus mitis, Streptococcus oralis were also oral involved. It have been reported that oral bacteria are likely the source of lung microbiota in healthy individuals [ 31 ]. It was concluded that an increased presence of oral microbes entering the lungs was associated with reduced lung function and elevated levels of pro-inflammatory cytokines [ 32 ]. Our study provided exhaustive genomics datasets for the lung microbiota. The co-assembly strategy increased the recovery ratio of MAGs, for instance, four of 29 high-quality MAGs were obtained using the co-assembly method. Furthermore, one out of five new species originated from that method, indicating that co-assembly enhanced the potential for reconstructing novel genomes. It was reported that co-assembly is an effective approach for reconstructing low-abundance MAGs [ 33 ]. In this study, MAG6, MAG8, MAG9, and MAG12 were obtained through the co-assembly method, and their abundance was lower at the individual level compared to MAGs from the single-sample assembly. Moreover, in 29 high-quality MAGs, 15 were detected by Kraken2. Six were identified at the genus level, three were not identified at the genus level, and five new species were discovered. This indicates that assembly expanded the scope of microbiota identification. Meanwhile, Tropheryma whipplei (MAG1), Rothia mucilaginosa (MAG34), and Parvimonas micra (MAG56), which belonged to 29 MAGs, were detected among the top 15 high-abundance species according to Kraken2, highlighting their significant role in lung microecology. We observed that the microbiota of the diseased lung was altered compared to a healthy state, and the abundance of bacteria varied between the two different lung diseases. Interestingly, COVID-19 patients have a more distinct profile compared to NSCLS. For instance, individuals with COVID-19 had smaller Bray-Curtis distance values and experienced more microbial disturbance. Our results suggest that the microbial communities in infection-induced lung diseases may be more pronounced than in lung cancer, which requires further data for confirmation. There are several limitations in our study. First, the sample size was not large enough to elucidate the lung micro-ecology. Second, each sample needs a blank control experiment to eliminate environmental contamination. Third, some crucial species need to be validated through experiments. Overall, we present microbial abundance and MAGs, which will significantly enhance the capacity to conduct taxonomic grouping and metagenomic analyses for future studies on the pulmonary microbiota, as well as the associations between health and diseases. We also lay the groundwork for the future development of more personalized and precise management and therapy strategies for lung health. Declarations Funding This work was supported by National Natural Science Foundation of China (Nos. 92159302, 32370628, 32170592); The Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences(2022-I2M-CoV19-006); State Key Laboratory Special Fund (No. 2060204); Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences (No. 2023-12M-2-001); State Key Laboratory of Respiratory Health and Multimorbidity, State Key Laboratory Special Fund 2060204; the Science and Technology Project of Sichuan (Nos. 2022ZDZX0018, 2023NSFSC004, 2024NSFSC0402); Key R&D Support Plan of Chengdu Science and Technology Bureau (No. 2023-YF09-00039-SN) and 1.3.5 project for disciplines of excellence, West China Hospital, Sichuan University (No. ZYGD22009). Author statement WML and ZFW conceived the study design. XW, YRY, YQZ and RJX collected clinical information about samples. YQL LLX and NNC CPL, THT and JZ performed genome sequencing. CC carried out bioinformatic analysis. SYW, HYW, DL, QLW, YC and XH interpreted the data. CC, YQL, HYW and SYW wrote the article. WML, HYW and DL provided clinical insights. All authors discussed the results and reviewed the article. Acknowledgments We are grateful to all participants who volunteered for this study. Declaration of Interests The authors declare no competing interests. Data sharing statement The raw metagenomic sequencing data generated in this study was available in the National Genomics Data Center (NGDC) Genome Sequence Archive (GSA) database (https://bigd.big.ac.cn/gsa-human/) under accession number HRA008205. Ethics approval and consent to participate This study was conducted in accordance with the Declaration of Helsinki and approved by the ethics committees of all participating centers, reference number 2023.1175. Informed written consent was obtained from all patients at the participating institutions to permit the archiving of their biospecimens and their use in future studies. Basic and clinical information about the patients was collected during their hospital visits. Consent for publication Not applicable. References Natalini, J.G., S. Singh, and L.N. Segal, The dynamic lung microbiome in health and disease. Nat Rev Microbiol, 2023. 21 (4): p. 222-235. Yuan, Y., et al., Pulmonary Actinomyces graevenitzii Infection: Case Report and Review of the Literature. Front Med (Lausanne), 2022. 9 : p. 916817. Charlson, E.S., et al., Topographical continuity of bacterial populations in the healthy human respiratory tract. Am J Respir Crit Care Med, 2011. 184 (8): p. 957-63. Stearns, J.C., et al., Culture and molecular-based profiles show shifts in bacterial communities of the upper respiratory tract that occur with age. Isme j, 2015. 9 (5): p. 1268. Ahmed, B. and M.J. Cox, Comparison of the upper and lower airway microbiota in children with chronic lung diseases. 2018. 13 (8): p. e0201156. Pattaroni, C., et al., Early life inter-kingdom interactions shape the immunological environment of the airways. Microbiome, 2022. 10 (1): p. 34. !!! INVALID CITATION !!! Nayfach, S., et al., New insights from uncultivated genomes of the global human gut microbiome. Nature, 2019. 568 (7753): p. 505-510. Man, W.H., et al., Bacterial and viral respiratory tract microbiota and host characteristics in children with lower respiratory tract infections: a matched case-control study. Lancet Respir Med, 2019. 7 (5): p. 417-426. Han, W., N. Wang, and M. Han, Identification of microbial markers associated with lung cancer based on multi-cohort 16 s rRNA analyses: A systematic review and meta-analysis. 2023. 12 (18): p. 19301-19319. Jin, J., et al., Diminishing microbiome richness and distinction in the lower respiratory tract of lung cancer patients: A multiple comparative study design with independent validation. Lung Cancer, 2019. 136 : p. 129-135. Li, Y., et al., Dysbiosis of lower respiratory tract microbiome are associated with proinflammatory states in non-small cell lung cancer patients. Thorac Cancer, 2024. 15 (2): p. 111-121. Ibironke, O., et al., Species-level evaluation of the human respiratory microbiome. Gigascience, 2020. 9 (4). Zou, Y., et al., 1,520 reference genomes from cultivated human gut bacteria enable functional microbiome analyses. Nat Biotechnol, 2019. 37 (2): p. 179-185. Almeida, A., et al., A unified catalog of 204,938 reference genomes from the human gut microbiome. Nat Biotechnol, 2021. 39 (1): p. 105-114. Schiff, M.J. and M.H. Kaplan, Rothia dentocariosa pneumonia in an immunocompromised patient. Lung, 1987. 165 (5): p. 279-82. Swaney, M.H., et al., Sweat and Sebum Preferences of the Human Skin Microbiota. 2023. 11 (1): p. e0418022. Hu, Y.L., et al., The Gastric Microbiome Is Perturbed in Advanced Gastric Adenocarcinoma Identified Through Shotgun Metagenomics. Front Cell Infect Microbiol, 2018. 8 : p. 433. Huang, X., et al., Metagenomic Analysis of Intratumoral Microbiome Linking to Response to Neoadjuvant Chemoradiotherapy in Rectal Cancer. Int J Radiat Oncol Biol Phys, 2023. 117 (5): p. 1255-1269. Deinhardt-Emmer, S., et al., Virulence patterns of Staphylococcus aureus strains from nasopharyngeal colonization. J Hosp Infect, 2018. 100 (3): p. 309-315. Hu Y, Feng Y, Zong Z. Global distribution of Klebsiella pneumoniae producing extended-spectrum β -lactamases in neonates . Precis Clin Med. 2024. 7 (4):pbae031. Rigauts, C., J. Aizawa, and S.L. Taylor, R othia mucilaginosa is an anti-inflammatory bacterium in the respiratory tract of patients with chronic lung disease. 2022. 59 (5). Nishioka, K., et al., Proteins produced by Streptococcus species in the lower respiratory tract can modify antiviral responses against influenza virus in respiratory epithelial cells. Microbes Infect, 2021. 23 (1): p. 104764. O'Dwyer, D.N., et al., Commensal Oral Microbiota, Disease Severity and Mortality in Fibrotic Lung Disease. 2023. Crawford, E.D., et al., Proc Natl Acad Sci U S A. Langelier, C., et al., Metagenomic Sequencing Detects Respiratory Pathogens in Hematopoietic Cellular Transplant Patients. Am J Respir Crit Care Med, 2018. 197 (4): p. 524-528. Ruoff, K.L., Nutritionally variant streptococci. Clin Microbiol Rev, 1991. 4 (2): p. 184-90. Belstrøm, D., et al., Differences in bacterial saliva profile between periodontitis patients and a control cohort. J Clin Periodontol, 2014. 41 (2): p. 104-12. Baker, J.L., Using Nanopore Sequencing to Obtain Complete Bacterial Genomes from Saliva Samples. mSystems, 2022. 7 (5): p. e0049122. Maraki, S., et al., A 5-year study of the bacterial pathogens associated with acute diarrhoea on the island of Crete, Greece, and their resistance to antibiotics. Eur J Epidemiol, 2003. 18 (1): p. 85-90. Bassis, C.M., et al., Analysis of the upper respiratory tract microbiotas as the source of the lung and gastric microbiotas in healthy individuals. mBio, 2015. 6 (2): p. e00037. Zhang, J., et al., Differential Oral Microbial Input Determines Two Microbiota Pneumo-Types Associated with Health Status. 2022. 9 (32): p. e2203115. Stewart, R.D., et al., Assembly of 913 microbial genomes from metagenomic sequencing of the cow rumen. Nat Commun, 2018. 9 (1): p. 870. Kultima, J.R., et al., MOCAT2: a metagenomic assembly, annotation and profiling framework. Bioinformatics, 2016. 32 (16): p. 2520-3. Lu, J., et al., Bracken: Estimating species abundance in metagenomics data. PeerJ Computer Science, 2017. 3 : p. e104. Beghini, F., et al., Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3. 2021. 10 . Suzek, B.E., et al., UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics, 2007. 23 (10): p. 1282-8. Caspi, R., et al., The MetaCyc database of metabolic pathways and enzymes - a 2019 update. Nucleic Acids Res, 2020. 48 (D1): p. D445-d453. Langfelder, P. and S. Horvath, WGCNA: an R package for weighted correlation network analysis. BMC Bioinformatics, 2008. 9 : p. 559. Uritskiy, G.V., J. DiRuggiero, and J. Taylor, MetaWRAP-a flexible pipeline for genome-resolved metagenomic data analysis. Microbiome, 2018. 6 (1): p. 158. Sieber, C.M.K., et al., Recovery of genomes from metagenomes via a dereplication, aggregation and scoring strategy. Nat Microbiol, 2018. 3 (7): p. 836-843. Chklovski, A., et al., CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning. Nat Methods, 2023. 20 (8): p. 1203-1212. Chaumeil, P.A. and A.J. Mussig, GTDB-Tk v2: memory friendly classification with the genome taxonomy database. 2022. 38 (23): p. 5315-5316. Hyatt, D., et al., Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics, 2010. 11 : p. 119. Li, W. and A. Godzik, Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics, 2006. 22 (13): p. 1658-9. Aramaki, T., et al., KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold. Bioinformatics, 2020. 36 (7): p. 2251-2252. Liu, B., et al., VFDB 2022: a general classification scheme for bacterial virulence factors. Nucleic Acids Res, 2022. 50 (D1): p. D912-d917. Li, H. and R. Durbin, Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics, 2009. 25 (14): p. 1754-60. Additional Declarations No competing interests reported. Supplementary Files SupplementTable.xlsx FigureS1.pdf Figure S1. Overview of the micro-biomecomposition of healthy individuals, related to Figure 1. A. PCoA visualizing the beta-diversity of the 43 BALF samples from healthy volunteers grouped by gender, history of smoking, history of family-member cancer and ages. B. Alpha-diversity of individuals grouped by gender, age, smoking history, and family cancer history. P-value was calculated using a two-sided Wilcoxon test. C,D. Co-occurrence network of 155 species based on Spearman correlation coefficients among those species. Each node represents one species. Ecological relationships are represented by edges connecting two nodes. Grey lines indicate a positive correlation. In (C), Each node's size is proportional to the rank value; The background color indicates that 12 communications have been detected. In (D), the node color represent phyla; the node size is proportional to the relative abundance of the corresponding species. Abbreviations: PCoA, principal coordinate analysis. FigureS2.pdf Figure S2. WGCNA of metabolic composition in 43 healthy individuals, related to Figure 2. A. The dendrogram and modules of metabolic pathway detected by the WGCNA. B. Pearson correlation analysis of modules and basic clinical information including age, gender, smoking history and family history of cancer. Abbreviations: WGCNA, weighted gene co-expression network analysis. FigureS3.pdf Figure S3. Taxonomic annotation and phylogenetic tree of 56 MAGs, related to Figure 3. A. Distribution of contig information for single-sample assembly per sample, including the number of contigs, length of the largest contig, and total length of all contigs. The dashed lines represent values of co-assembly. B. The fraction of MAGs from the single-sample assembly and co-assembly. C. Quality metrics across high-quality (n = 29) and medium-quality (n = 27) MAGs. D. Heatmap showing the abundance of 29 high-quality MAGs. MAGs in the purple frame were from co-assembly. MAGs in red indicate novel species. MAGs in green indicate species ranked 18th based on relative abundance using Kranken2. Abbreviations: MAGs, metagenome-assembled genome. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7937892","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":544742871,"identity":"021cfbee-7fac-4c46-9df4-13966b9fbc1c","order_by":0,"name":"cheng cheng","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"cheng","middleName":"","lastName":"cheng","suffix":""},{"id":544742872,"identity":"ce2b0e4a-c42b-4c26-8911-85e07a66e2c6","order_by":1,"name":"Yangqian Li","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Yangqian","middleName":"","lastName":"Li","suffix":""},{"id":544742873,"identity":"5202bbde-2e9d-4e52-8dce-a88295c314a5","order_by":2,"name":"Suyan Wang","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Suyan","middleName":"","lastName":"Wang","suffix":""},{"id":544742874,"identity":"962f0816-d62b-42dc-be91-e7c49a637bf6","order_by":3,"name":"Haoyu Wang","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Haoyu","middleName":"","lastName":"Wang","suffix":""},{"id":544742875,"identity":"57fc8171-ccb1-40a6-9879-8f940309ab38","order_by":4,"name":"Dan Liu","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Dan","middleName":"","lastName":"Liu","suffix":""},{"id":544742876,"identity":"947cd34d-0590-4c94-96b0-75c44d0e6d47","order_by":5,"name":"QingLan Wang","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"QingLan","middleName":"","lastName":"Wang","suffix":""},{"id":544742877,"identity":"6b2f24e8-ed7b-4d03-b3d5-a324cbe87ccc","order_by":6,"name":"You Che","email":"","orcid":"","institution":"NIH","correspondingAuthor":false,"prefix":"","firstName":"You","middleName":"","lastName":"Che","suffix":""},{"id":544742878,"identity":"e10ef798-afa7-4998-8917-6866e04aa7ab","order_by":7,"name":"Linlin Xue","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Linlin","middleName":"","lastName":"Xue","suffix":""},{"id":544742879,"identity":"ef2e8b2d-aec5-4565-b7eb-9a17de7bb51c","order_by":8,"name":"Ningning Chao","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Ningning","middleName":"","lastName":"Chao","suffix":""},{"id":544742880,"identity":"d266256d-e7c0-4caf-86bc-cfe51d01f54d","order_by":9,"name":"Xuan He","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Xuan","middleName":"","lastName":"He","suffix":""},{"id":544742881,"identity":"4000a036-f718-4ed7-8ece-8bbc0276905c","order_by":10,"name":"Chengping Li","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Chengping","middleName":"","lastName":"Li","suffix":""},{"id":544742882,"identity":"6d49f4d2-befd-4e5d-81b9-f28437d5e1b8","order_by":11,"name":"Huohua Tian","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Huohua","middleName":"","lastName":"Tian","suffix":""},{"id":544742883,"identity":"58cb0840-275c-47ff-8092-f02faba7a302","order_by":12,"name":"Jing Zhou","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Jing","middleName":"","lastName":"Zhou","suffix":""},{"id":544742887,"identity":"5429c1a4-186d-4399-8615-7034f86bc09e","order_by":13,"name":"Xin Wang","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Xin","middleName":"","lastName":"Wang","suffix":""},{"id":544742888,"identity":"312e433a-8cfe-4be7-8b5e-346d1738090b","order_by":14,"name":"Yurui Yang","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Yurui","middleName":"","lastName":"Yang","suffix":""},{"id":544742889,"identity":"809c658c-6247-42a3-a001-09955b281515","order_by":15,"name":"Yuqi Zhu","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Yuqi","middleName":"","lastName":"Zhu","suffix":""},{"id":544742890,"identity":"f96dbfdc-f3ea-42a7-ad9f-0bd44d4bd6d2","order_by":16,"name":"Renjie Xu","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Renjie","middleName":"","lastName":"Xu","suffix":""},{"id":544742891,"identity":"a951a1ef-5d98-4180-bc40-9d625962c9d2","order_by":17,"name":"Zhoufeng Wang","email":"","orcid":"","institution":"Sichuan University","correspondingAuthor":false,"prefix":"","firstName":"Zhoufeng","middleName":"","lastName":"Wang","suffix":""},{"id":544742892,"identity":"eba90f71-fb7b-42db-835e-d9bb4954e3b2","order_by":18,"name":"Weimin Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAzklEQVRIiWNgGAWjYFACxgaGBAMJO34Ij5lYLRU2yZINxGsBgTNpjBsOEKvF4Hhz64aHbYeZjc8ffibBUGGd2MB+9gB+LWcOtt1IbDvMZ3YjzUyC4Ux6YgNPXgJeLWZA9SAtzGY3eNgkGNsOJzZI8Bjg13L/IVgL4+b+M0At/4jRcoOx7UYCyPsMOUAtDURosT8DdBgokCVupBlbJBxLN27jycGvRbL9+LObP0BR2X/44Y0PNday/exn8GtBBQlAzEaC+lEwCkbBKBgFOAAArYpJbrA6pq4AAAAASUVORK5CYII=","orcid":"","institution":"Sichuan University","correspondingAuthor":true,"prefix":"","firstName":"Weimin","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2025-10-24 07:38:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7937892/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7937892/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":96377064,"identity":"1c7a4d73-af59-4121-8d28-33380d653efa","added_by":"auto","created_at":"2025-11-20 11:32:08","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":132735,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/630ca157d1e9543e8bc4f151.docx"},{"id":96377067,"identity":"ee0937ab-b637-450a-9927-4dc0fcb2b3ec","added_by":"auto","created_at":"2025-11-20 11:32:08","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5552916,"visible":true,"origin":"","legend":"","description":"","filename":"Figure4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/6e20eaa51fb33657e782cdcd.pdf"},{"id":96453069,"identity":"ab061952-27ba-4195-b7c9-dcd39c820364","added_by":"auto","created_at":"2025-11-21 09:57:56","extension":"json","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":17741,"visible":true,"origin":"","legend":"","description":"","filename":"db1733a256df4cdc9c5a417e72788b5f.json","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/85b2ae888b895a4234b5819f.json"},{"id":96377068,"identity":"d67b086e-4216-4f04-a28f-dcb10d3a729f","added_by":"auto","created_at":"2025-11-20 11:32:08","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1340841,"visible":true,"origin":"","legend":"","description":"","filename":"FigureS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/9abb558b68bc5c3c017f9986.pdf"},{"id":96453116,"identity":"57795461-0c87-43cb-8867-c9dcf1d3ba71","added_by":"auto","created_at":"2025-11-21 09:58:12","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":514057,"visible":true,"origin":"","legend":"","description":"","filename":"FigureS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/87f417affb77a0bab1c9d336.pdf"},{"id":96377072,"identity":"4dcfef32-2ccc-4b06-bf46-e9ee413376f0","added_by":"auto","created_at":"2025-11-20 11:32:09","extension":"pdf","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1527611,"visible":true,"origin":"","legend":"","description":"","filename":"FigureS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/e3cc9f52bae5409f03566e65.pdf"},{"id":96452902,"identity":"1ee6dae9-1839-4e73-8852-91fa6d88c3ef","added_by":"auto","created_at":"2025-11-21 09:53:11","extension":"xlsx","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":163240,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementTable.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/59f0bfb7f02ea3538dd6df0a.xlsx"},{"id":96377074,"identity":"c8558c19-589a-4ea9-9784-65fcba260f81","added_by":"auto","created_at":"2025-11-20 11:32:09","extension":"xml","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110901,"visible":true,"origin":"","legend":"","description":"","filename":"db1733a256df4cdc9c5a417e72788b5f1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/33cf4e6a74f9fec44a35724a.xml"},{"id":96453078,"identity":"a05aefdd-eac6-412a-9b78-a90a41bf5f28","added_by":"auto","created_at":"2025-11-21 09:57:58","extension":"pdf","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3755546,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/893ff32a6604e34ae9709b22.pdf"},{"id":96452888,"identity":"42a38f2f-5d64-433a-8a4c-1cbfafadb827","added_by":"auto","created_at":"2025-11-21 09:52:36","extension":"pdf","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1057809,"visible":true,"origin":"","legend":"","description":"","filename":"Figure2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/30dc02a57c5919bbce95173c.pdf"},{"id":96377078,"identity":"750c098a-52a5-4bb7-b1fb-88b9e9b7ba4a","added_by":"auto","created_at":"2025-11-20 11:32:09","extension":"pdf","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":809879,"visible":true,"origin":"","legend":"","description":"","filename":"Figure3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/044f3e1346d397a52304d8d8.pdf"},{"id":96377076,"identity":"2816ad8f-75b1-4581-9515-57be25fa15e7","added_by":"auto","created_at":"2025-11-20 11:32:09","extension":"pdf","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5552916,"visible":true,"origin":"","legend":"","description":"","filename":"Figure4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/326b0d9e67f0eaa9d693248c.pdf"},{"id":96377073,"identity":"1e3ee363-06c6-4510-b30f-55d240823d0b","added_by":"auto","created_at":"2025-11-20 11:32:09","extension":"xml","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":107649,"visible":true,"origin":"","legend":"","description":"","filename":"db1733a256df4cdc9c5a417e72788b5f1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/19796e390c8ded46aeaee8a3.xml"},{"id":96377070,"identity":"06cdee96-ab3a-44a9-8514-227227114c1d","added_by":"auto","created_at":"2025-11-20 11:32:09","extension":"html","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":125375,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/e5d3f5fef3199dba7bb1c63f.html"},{"id":96453749,"identity":"35ffd107-855c-45a8-908f-2869bf0eed3e","added_by":"auto","created_at":"2025-11-21 10:01:29","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":110537,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of the microbiome composition of healthy individuals.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. A schematic representation of the workflow of metagenomic sequencing on healthy individuals.\u003c/p\u003e\n\u003cp\u003eB. The clinical information of 43 BALF samples from healthy volunteers including gender, age, history of smoking, and history of family-member cancer.\u003c/p\u003e\n\u003cp\u003eC. (1) Average relative abundances of bacterial phyla after filtering out rare species (those with a relative abundance lower than 0.01 or present in less than 5% of individuals). (2) Number of species present in each phylum.\u003c/p\u003e\n\u003cp\u003eD. Top 15 species-level abundances of all samples in the cohort are sorted by the mean relative abundance of bacteria at the species level in each person. The scale of 0–1 corresponds to 0–100% abundance. The color of the species name indicates the phyla, as shown in Figure 1c legend.\u003c/p\u003e\n\u003cp\u003eE. Top 15 species-level compositions of all samples in the cohort, sorted by the abundance of \u003cem\u003eStaphylococcus aureus\u003c/em\u003e. Each vertical line indicates one sample.\u003c/p\u003e\n\u003cp\u003eF. Co-occurrence network of 155 species based on Spearman correlation coefficients among those species. Each node represents one species. Ecological relationships are represented by edges connecting two nodes. Grey lines indicate a positive correlation. The node color represent phyla and the node size is proportional to the rank value..\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/01272483e0521087ab176b85.png"},{"id":96377056,"identity":"9c2a3e38-d920-4734-b863-9c8a4a08bdd5","added_by":"auto","created_at":"2025-11-20 11:32:08","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":41922,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of the metabolism composition of healthy individuals.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. Top 15 metabolic abundances of all samples in the cohort are sorted by the mean \u0026nbsp;\u0026nbsp;relative abundance of metabolism composition in each person. The scale of 0–1 corresponds to 0–100% abundance.\u003c/p\u003e\n\u003cp\u003eB. Top 15 metabolic abundance compositions of all samples in the cohort. Each vertical line indicates one sample.\u003c/p\u003e\n\u003cp\u003eC. Box plots illustrating variations in relative abundance for metabolic compositions between different groups.\u003c/p\u003e\n\u003cp\u003eD. UMP biosynthesis II was significantly correlated with Age. The shadings indicate the 95% confidence intervals.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/2cd388163fd8cd730641dfe1.png"},{"id":96452885,"identity":"46028e12-4186-4bb6-8a89-f394e829ff5b","added_by":"auto","created_at":"2025-11-21 09:52:27","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":78653,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTaxonomic annotation and phylogenetic tree of 56 metagenome-assembled genomes (MAGs).\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. A schematic diagram of the pipeline for constructing MAG from contigs based on the metagenomes of 43 healthy individuals.\u003c/p\u003e\n\u003cp\u003eB. Statistical distribution of log-scaled N50 for single-sample assembly per sample, with dashed lines indicating the log-scaled N50 value of co-assembly.\u003c/p\u003e\n\u003cp\u003eC. Taxonomic classification of 56 MAGs at various levels.\u003c/p\u003e\n\u003cp\u003eD. The number of MAGs at each phylum level.\u003c/p\u003e\n\u003cp\u003eE. Phylogenetic distribution of 56 MAGs. The inner-colored strips represent phyla. The outer rings indicate whether the quality of a MAG is high or medium. The star labeled MAGs represent new species.\u003c/p\u003e\n\u003cp\u003eF. Heatmap of the distribution of pathways. The vertical axis represents six different kinds of pathways in level 1 and 12 different kinds of pathways in level 2. Annotation results were obtained using KEGG.\u003c/p\u003e\n\u003cp\u003eG. Heatmap illustrating the distribution of virulence factors.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/b4cbcef38d24207a8d3704a5.png"},{"id":96377060,"identity":"02bb45ce-cd8f-4e9f-b4f9-46183e378a23","added_by":"auto","created_at":"2025-11-20 11:32:08","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":151064,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTaxonomic and functional profiles indicate changes in COVID-19 and NSCLC compared to healthy individuals.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. PCoA visualizing the beta-diversity of the three cohorts. Colors indicate different groups.\u003c/p\u003e\n\u003cp\u003eB. Bray-Curtis dissimilarity comparisons of three groups. The center line represents the median, the box limits indicate the upper and lower quartiles, and the whiskers show 1.5 times the interquartile ranges.\u003c/p\u003e\n\u003cp\u003eC. Histogram showing the relative abundance of the top 15 species with high relative abundance from the healthy group across three groups.\u003c/p\u003e\n\u003cp\u003eD. Principal microbial species associated with health status, COVID-19, and NSCLC. The top 50 microbial species were clustered based on taxonomy. Associations were color-coded based on the direction of effect (red for positive, blue for negative) and coefficient. Significant associations at FDR \u0026lt; 0.05 were indicated with a plus sign (for positive correlations) or a minus sign (for negative correlations).\u003c/p\u003e\n\u003cp\u003eE. Dot plots showing relative abundance of species significantly (|FoldChange| \u0026gt; 2 and Wilcox test P value \u0026lt; 0.01) enriched in COVID-19 and NSCLC.\u003c/p\u003e\n\u003cp\u003eF. Box plots illustrating variations in relative abundance for species and MAGs among three groups.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAbbreviations: \u003c/strong\u003ePCoA, Principal coordinate analysis; NSCLC, \u003ca href=\"https://pubmed.ncbi.nlm.nih.gov/31095926/\"\u003enon-small cell lung cancer\u003c/a\u003e.\u003c/p\u003e\n\u003cp\u003e1. Kultima JR, Coelho LP, Forslund K, Huerta-Cepas J, Li SS, Driessen M, Voigt AY, Zeller G, Sunagawa S, Bork P. MOCAT2: a metagenomic assembly, annotation and profiling framework. Bioinformatics 2016;32: 2520-3.\u003c/p\u003e\n\u003cp\u003e1. Kultima JR, Coelho LP, Forslund K, Huerta-Cepas J, Li SS, Driessen M, Voigt AY, Zeller G, Sunagawa S, Bork P. MOCAT2: a metagenomic assembly, annotation and profiling framework. Bioinformatics 2016;32: 2520-3.\u003c/p\u003e\n\u003cp\u003e1. Kultima JR, Coelho LP, Forslund K, Huerta-Cepas J, Li SS, Driessen M, Voigt AY, Zeller G, Sunagawa S, Bork P. MOCAT2: a metagenomic assembly, annotation and profiling framework. Bioinformatics 2016;32: 2520-3.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/8475354bf5df9db934396bc8.png"},{"id":106093939,"identity":"03dd7360-28f2-412e-96d1-c17d25c0b789","added_by":"auto","created_at":"2026-04-03 11:40:13","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1458856,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/02b84af6-100d-404f-ae77-9742d310a3b8.pdf"},{"id":96453636,"identity":"4f65f34b-9df0-40e6-98a5-fe7d23eedc3c","added_by":"auto","created_at":"2025-11-21 10:01:09","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":163240,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementTable.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/558325d50077551b96a7d45e.xlsx"},{"id":96452903,"identity":"1b07c9ea-b334-4f3e-b027-41374950f284","added_by":"auto","created_at":"2025-11-21 09:53:17","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1340841,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure S1. Overview of the micro-biomecomposition of healthy individuals, related to Figure 1.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. PCoA visualizing the beta-diversity of the 43 BALF samples from healthy volunteers grouped by gender, history of smoking, history of family-member cancer and ages.\u003c/p\u003e\n\u003cp\u003eB. Alpha-diversity of individuals grouped by gender, age, smoking history, and family cancer history. P-value was calculated using a two-sided Wilcoxon test.\u003c/p\u003e\n\u003cp\u003eC,D. Co-occurrence network of 155 species based on Spearman correlation coefficients among those species. Each node represents one species. Ecological relationships are represented by edges connecting two nodes. Grey lines indicate a positive correlation. In (C), Each node's size is proportional to the rank value; The background color indicates that 12 communications have been detected. In (D), the node color represent phyla; the node size is proportional to the relative abundance of the corresponding species.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAbbreviations: \u003c/strong\u003ePCoA, principal coordinate analysis.\u003c/p\u003e","description":"","filename":"FigureS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/4ae2d4256a468cb57f4c26c5.pdf"},{"id":96453139,"identity":"05c3e6c5-9ae6-4882-95a4-b9dc4c9bdbbd","added_by":"auto","created_at":"2025-11-21 09:58:25","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":514057,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure S2. WGCNA of metabolic composition in 43 healthy individuals, related to Figure 2.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. The dendrogram and modules of metabolic pathway detected by the WGCNA.\u003c/p\u003e\n\u003cp\u003eB. Pearson correlation analysis of modules and basic clinical information including age, gender, smoking history and family history of cancer.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAbbreviations: \u003c/strong\u003eWGCNA, weighted gene co-expression network analysis.\u003c/p\u003e","description":"","filename":"FigureS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/83900de5330fc88994c44200.pdf"},{"id":96377062,"identity":"cd0366b1-4b1d-4a0c-aab8-f07025f605b0","added_by":"auto","created_at":"2025-11-20 11:32:08","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":1527611,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure S3. Taxonomic annotation and phylogenetic tree of 56 MAGs, related to Figure 3.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA. Distribution of contig information for single-sample assembly per sample, including the number of contigs, length of the largest contig, and total length of all contigs. The dashed lines represent values of co-assembly.\u003c/p\u003e\n\u003cp\u003eB. The fraction of MAGs from the single-sample assembly and co-assembly.\u003c/p\u003e\n\u003cp\u003eC. Quality metrics across high-quality (n = 29) and medium-quality (n = 27) MAGs.\u003c/p\u003e\n\u003cp\u003eD. Heatmap showing the abundance of 29 high-quality MAGs. MAGs in the purple frame were from co-assembly. MAGs in red indicate novel species. MAGs in green indicate species ranked 18th based on relative abundance using Kranken2.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAbbreviations: MAGs, \u003c/strong\u003emetagenome-assembled genome.\u003c/p\u003e","description":"","filename":"FigureS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7937892/v1/02164086176bdaf77d9d4c29.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Deciphering the microbial landscape of lower respiratory tract in health and respiratory diseases","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe lung microbiota plays an important role in the human lung respiratory system[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Numerous studies have indicated that the diversity and composition of the respiratory microbiome in healthy individuals exhibit individual variability but are generally stable [\u003cspan additionalcitationids=\"CR3 CR4 CR5\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Using high-throughput sequencing, scientists can more accurately describe the composition of the lung microbiota in healthy individuals, including the abundance and functions of microorganisms [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Furthermore, the development of metagenomic binning technology has enabled the acquisition of nearly complete metagenome-assembled genomes (MAGs) on a large scale [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The microbiota composition in healthy human lungs has been partially characterized primarily through 16SrRNA amplicon sequencing [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. However, there have been limited assemblies of MAGs specifically targeting the healthy lung microbiome.\u003c/p\u003e\u003cp\u003eDue to clinical sample acquisition difficulty in healthy individuals and limited sequencing depth, the availability of sequenced genomes and functional information for most healthy individual lung microbes remains limited. In our previous studies, we showed pulmonary microbiome dysbiosis in lung cancer patients based on non-assembled metagenomic data[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. A large-scale and high-depth metagenomic sequencing scan is needed to perform to characterize microbiota compositions and functions in healthy human lungs. Here we address this problem by constructing co-network of bacterium based on their relative abundance, reconstructing draft prokaryotic genomes from 43 healthy individuals. Furthermore, we revealed the microbiota alterations in lung diseases focused on COVID-19 and non-small\u0026ensp;cell\u0026ensp;lung\u0026ensp;cancer (NSCLC) using microbial profiling of this healthy cohort. Our study will provide more comprehensive baseline data for research on human pulmonary microbiota, and better understand its mechanisms of action in both health and disease states.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eSample collection and processing\u003c/h2\u003e\u003cp\u003eAll details about the experimental samples are listed in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e, including 43 healthy individuals, 23 COVID-19 patients, and 28 NSCLC patients. Healthy individuals were recruited from Suining Central Hospital, Sichuan Province, in 2021. They underwent CT imaging, pulmonary function tests, electrocardiograms, and routine blood tests, all showing normal liver and kidney function.\u003c/p\u003e\u003cp\u003eCOVID-19 patients were diagnosed with SARS-CoV-2 infection through qPCR or antigen testing. NSCLC patients were histologically and endoscopically confirmed by examination from three pathologists. The samples of those patients were obtained from the West China Hospital of Sichuan University. All patients were followed up to confirm the final diagnosis.\u003c/p\u003e\u003cp\u003eIn this study, samples were collected via bronchoscopy for distal alveolar lavage. We used only bronchoscope working channel washes, which were done twice with a minimum volume of 15 mL of 0.9% saline solution. Immediately dispense the collected BALF into 1.8 mL sterile freezing tubes and store at -80\u0026deg;C until processing. The collection process was carried out following aseptic procedures to prevent contamination from environmental, human commensal, and miscellaneous bacteria.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eDNA extraction and quality control\u003c/h3\u003e\n\u003cp\u003eDNA extraction from BALF was performed using the DNA Microbiome kit (51704, Qiagen, USA) according to the manufacturer\u0026rsquo;s instructions. The host cells are differentially lysed, followed by enzymatic digestion, to efficiently remove host DNA. Subsequently, complete cell lysis is achieved through physical and chemical methods, facilitating the isolation of bacterial DNA.\u003c/p\u003e\n\u003ch3\u003eLibrary construction and sequencing\u003c/h3\u003e\n\u003cp\u003eUsing an NGS Nextera\u0026trade; DNA Flex Library Prep kit (Illumina), the DNA libraries were conducted. By performing an enzymatic reaction called segmentation, DNA is fragmented, followed by the addition of adapter sequences and PCR amplification. An Agilent 4200 Bioanalyzer (Agilent Technologies, Santa Clara, USA) was used to assess the quality of the DNA libraries. After a quality check, Illumina PE150 sequencing was conducted by pooling multiple libraries based on the required effective concentration and target data amount. Sequencing was performed on the Illumina NovaSeq 6000.\u003c/p\u003e\n\u003ch3\u003eData quality control\u003c/h3\u003e\n\u003cp\u003eFor the quality control of raw data obtained from the Illumina sequencing platform, mocat2 (v2.1.3) was used to filter out low-quality reads, and reads aligned to the GRCh38 Homo sapiens reference genome were removed [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. The remaining high-quality reads were further analyzed.\u003c/p\u003e\n\u003ch3\u003eTaxonomic classification\u003c/h3\u003e\n\u003cp\u003eTaxonomic classification was conducted with Kraken2 using the standard Kraken2 database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://benlangmead.github.io/aws-indexes/k2\u003c/span\u003e\u003cspan address=\"https://benlangmead.github.io/aws-indexes/k2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) containing archaea, bacteria, viruses, plasmids, human1, and UniVec_Core. Bracken [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] was then used to calculate the abundance of the identified species from the Kraken2 analysis.\u003c/p\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003eDiversity analysis\u003c/h2\u003e\u003cp\u003eBacterial diversity was assessed by calculating the Shannon index using the vegan package (v2.5.7) in R. The Wilcoxon test was conducted to determine the statistical significance of differences among groups of healthy individuals. Differences associated with a P-value of \u0026lt;\u0026thinsp;0.01 were considered significant.\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003eCo-occurrence network construction\u003c/h3\u003e\n\u003cp\u003eSpearman\u0026rsquo;s correlation coefficients of relationships across species with an absolute value\u0026thinsp;\u0026gt;\u0026thinsp;0.8 and statistically significant (P-value\u0026thinsp;\u0026lt;\u0026thinsp;0.01) were visualized using the R package igraph (v2.0.2) with a layout based on the Fruchterman-Reingold algorithm. We also calculated degree and page rank values to facilitate the identification of keystone species. Network module detection was performed using the greedy algorithm.\u003c/p\u003e\n\u003ch3\u003eMicrobiome-functional pathway analysis\u003c/h3\u003e\n\u003cp\u003eTo identify microbial pathways, the human removed reads were analysed with HUMAnN (v3.7) [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e] by using the databases of uniref 90[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e] and pathway MetaCyc[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. The \u0026ldquo;WGCNA\u0026rdquo; package[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e] was used to conduct WGCNA (weighted gene co-expression network analysis). For correlation analysis, spearman correlations were calculated using function corr.test in R.\u003c/p\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003eMetagenomic de novo assembly and binning\u003c/h2\u003e\u003cp\u003eThe Megahit module from metaWRAP-Binning was utilized to assemble clean data from the aforementioned processed Illumina sequencing data [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. Metagenomic binning was subsequently conducted using MaxBin2, metaBAT2, and CONCOCT [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. DAS Tool (v1.1.5) was used to aggregate multiple binning predictions into a new and enhanced bin set [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. CheckM2 was used to evaluate the quality of the assembled bins, which were screened based on the criteria of completeness (\u0026gt;\u0026thinsp;50%) and contamination (\u0026lt;\u0026thinsp;10%)[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. Then dRep (v3.0.0) was then used to remove duplicate bins with the parameter \"-sa 0.95\". After analysis, 56 MAGs were retained.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003eTaxonomic assignment and phylogenetic analysis\u003c/h2\u003e\u003cp\u003eMAGs were taxonomically classified using the Genome Taxonomy Database Toolkit (GTDB-Tk) (v2) with \"classify_wf\" function and default parameters[\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. The Average Nucleotide Identity (ANI) values between the five new MAGs and genomes from two reference cohorts were calculated using fastANI (v1.33). All phylogenetic trees of the 56 MAGs were constructed using iqtree2 (v2.2.3) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/iqtree/iqtree2\u003c/span\u003e\u003cspan address=\"https://github.com/iqtree/iqtree2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Finally, iTOL (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://itol.embl.de\u003c/span\u003e\u003cspan address=\"https://itol.embl.de\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was used to generate a phylogenetic tree of the 56 MAGs.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003eFunction annotation of MAGs\u003c/h2\u003e\u003cp\u003eHigh-quality MAGs were utilized for gene prediction by Prodigal (v2.6.3) software (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/hyattpd/Prodigal\u003c/span\u003e\u003cspan address=\"https://github.com/hyattpd/Prodigal\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. All complete genes were clustered at 90% protein identification using CD-HIT (v4.8.1) [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. The KEGG (Kyoto Encyclopedia of Genes and Genomes) annotation results were extracted using KofamKOALA (v1.3.0) software [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]. Additionally, key virulence factors were identified using the Virulence Factor Database (VFDB) [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e] via BLAST (v2.10.1+).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003eCalculation of the abundances of MAGs and genes\u003c/h2\u003e\u003cp\u003eThe quant_bins module from metaWRAP-Binning was used to quantify the abundance of each MAG in each sample. BWA MEM was used to align clean reads from each sample to gene catalogs after removing human contamination [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. The abundances were normalized to fragments per kilobase of gene sequence per million reads mapped (FPKM). The abundances of microbial taxa, KEGG pathways, and Virulence Factors (VFs) were calculated by combining the abundances of all the members within each category.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003eMicrobiome analysis of three groups\u003c/h2\u003e\u003cp\u003ePrincipal Coordinates Analysis (PCoA) based on the Bray\u0026ndash;Curtis distance matrix was utilized to visualize the compositional profiles of 212 species from samples belonging to three groups. The differences in microbiome composition between different phenotypes were calculated using permutational multivariate analysis of variance with distance matrices in the adonis2 function of the vegan package, with 999 permutations.\u003c/p\u003e\u003cp\u003eAssociations of specific microbial species with phenotypes were calculated using multivariate analysis by linear models (MaAsLin2) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://huttenhower.sph.harvard.edu/galaxy\u003c/span\u003e\u003cspan address=\"http://huttenhower.sph.harvard.edu/galaxy\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) with healthy and lung-diseased samples as references. The p-value of associations was computed by MaAsLin2, and the false discovery rate (FDR) was calculated using the Benjamini\u0026ndash;Hochberg correction. FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.01 was considered significant. Furthermore, the fold change was calculated using R (v4.3.2) with the Wilcoxon test.\u003c/p\u003e\u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003eDistribution of species composition in the healthy human population\u003c/h2\u003e\u003cp\u003eWe conducted meta-genomics profiling of bronchoalveolar lavage fluid (BALF) from 43 healthy individuals, including 11 males with smoking history (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea-b; Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). We obtained an average of 4,683,836,943 bases in each sample after quality control and removal of human-derived reads (Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). Using Kraken2 profiling, we identified 196 species with a relative abundance higher than 0.01 and present in more than 5% of individuals, accounting for 0.991 of the relative abundance per sample from healthy individuals (Table \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e-4). Those 196 species were assigned to eight phyla, with the dominant phyla being \u003cem\u003eFirmicutes\u003c/em\u003e, \u003cem\u003eProteobacteria\u003c/em\u003e, and \u003cem\u003eActinobacteria\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec). In Ibironke et al.\u0026rsquo;s study, the three phyla mentioned above were identified as dominant phyla in the human respiratory microbiome of five individuals based on full-length 16SrRNA sequencing[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. To identify microbial species may play a critical role in the maintaining of lung homeostasis, we examined our cohort for microbial taxa with high abundance and core microbes present in over 95% of individuals. We detected 15 high-abundance species as dominant species (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ed-e). Meanwhile, we observed 32 core species (\u0026gt;\u0026thinsp;5% of individuals) out of 196 species. 12 out of the 32 core microbes are also dominant species, indicating their critical roles in the lung ecosystem.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eWe further investigated the microbial component in our cohort through principal coordinate analysis (PCoA) and founded that there was no significantly separation based on gender, age, smoking history, or family cancer history (Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003ea). Furthermore, the alpha diversity of individuals grouped by the aforementioned clinical features was insignificant, indicating that differences in microbial communities between clinical groups were minor (Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eb).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTo identify keystone species and explore relationships among species. Spearman\u0026rsquo;s correlation coefficients were calculated, and the ecological relationships across 155 different microbes (a relative abundance\u0026thinsp;\u0026gt;\u0026thinsp;0.001 and present in more than 60% of individuals) were visualized using co-occurrence networks. There were 155 nodes and 1084 edges (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ef). The top five species with the highest degree and page rank values were \u003cem\u003eGranulicatella adiacens\u003c/em\u003e, \u003cem\u003eMogibacterium diversum\u003c/em\u003e, \u003cem\u003eStreptococcus australis\u003c/em\u003e, \u003cem\u003eMogibacterium pumilum\u003c/em\u003e, and \u003cem\u003eStreptococcus sp.116-D4\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ef). Two dominant species, \u003cem\u003eStaphylococcus aureus\u003c/em\u003e and \u003cem\u003eSalmonella enterica\u003c/em\u003e, have relatively low degree and page rank values (Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003ed). Additionally, 12 communities were detected (Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003ec). \u003cem\u003eFirmicutes\u003c/em\u003e, \u003cem\u003eProteobacteria\u003c/em\u003e, and \u003cem\u003eActinobacteria\u003c/em\u003e comprised the three largest communities. Module one was dominated by \u003cem\u003eProteobacteria\u003c/em\u003e, module two by \u003cem\u003eBacteroidetes\u003c/em\u003e and \u003cem\u003eFirmicutes\u003c/em\u003e, and module three by \u003cem\u003eFirmicutes\u003c/em\u003e and \u003cem\u003eActinobacteria\u003c/em\u003e. The phylum of the species was significantly associated with the corresponding module (P value\u0026thinsp;\u0026lt;\u0026thinsp;0.001, Fisher test), suggesting that species within the same phylum may perform similar roles in the healthy lung micro-ecology. The results indicate that the three major phyla (\u003cem\u003eFirmicutes\u003c/em\u003e, \u003cem\u003eProteobacteria\u003c/em\u003e, and \u003cem\u003eActinobacteria\u003c/em\u003e), based on abundance and co-network analysis played a crucial role in lung homeostasis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003eMetabolic profiling of microbe from the healthy human lung\u003c/h2\u003e\u003cp\u003eTo illustrate the metabolic composition of healthy individuals, we identified 374 pathways from the MetaCyc database. The top 15 relative abundances of metabolic terms are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003ea-b. These dominant metabolic terms include aerobic respiration I, aminoimidazole ribonucleotide biosynthesis, adenosine nucleotides de novo biosynthesis, guanosine ribonucleotides biosynthesis, and others. We conducted weighted gene co-expression network analysis (WGCNA) to explore the relationship between the 374 pathways and basic clinical information. Two modules were obtained (Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003ea). However, we did not find modules significantly related to the basic clinical information (Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003eb). Ethanolamine Utilization and Pyruvate Fermentation to Isobutanol was enriched in the female group compared to the male group (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). UMP biosynthesis II was positively correlated with age (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e2\u003c/span\u003ed). Our results suggest that the overall metabolic composition of microbes in BALF is not related to basic clinical information and indicates a stable state.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\u003ch2\u003eReconstruction of microbial genomes from the healthy human lung\u003c/h2\u003e\u003cp\u003eThe workflow of binning is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003ea. Firstly, assembly was performed based on single-sample and co-sample strategies after quality control on raw sequencing data and removing human contamination. The density of N50, the number of contigs, the length of the largest contigs, and the total length were shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003eb and Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003ea. The length of the largest contig and the total length of contigs from co-samples were larger than those from single samples. Microbial genomes representing individual bacterial species were then constructed from the assembled metagenomic sequencing data obtained from the 43 samples described above. We used the metaWRAP-Binning module to generate 844 bins from single-sample assemblies and 212 bins from mixed-sample assemblies. After dereplication, aggregation, and scoring filtering, we observed 121 bins. Furthermore, we obtained 84 metagenome-assembled genomes (MAGs) after quality assessment, meeting the criteria of \u0026gt;\u0026thinsp;50% completeness and \u0026lt;\u0026thinsp;10% contamination. These reconstructed microbial genomes underwent redundancy removal (average nucleotide identity ANI\u0026thinsp;\u0026gt;\u0026thinsp;0.95). A final set of 56 non-redundant MAGs were obtained. Particularly, 14 MAGs were from co-assembly (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003eb), indicating that co-assembly expands the number of detected MAGs. Among these 56 MAGs, 27 MAGs met the medium quality criteria (\u0026gt;\u0026thinsp;50% completeness and \u0026lt;\u0026thinsp;10% contamination), while 29 MAGs exhibited high quality (\u0026gt;\u0026thinsp;90% completeness and \u0026lt;\u0026thinsp;5% contamination). Furthermore, various indexes of high-quality MAGs were significantly higher than the medium-quality ones, including completeness, contamination, contig number, and the length of the largest contig length (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003ec).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eAll MAGs exhibited a comparatively high prevalence in the 43 metagenomes (Figure \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003ed), indicating that these MAGs were likely strains of core species. The 56 MAGs were subsequently classified into taxa using the Genome Taxonomy Database Toolkit (GTDB-Tk) (Table S5). Further analysis showed that 48 out of the 56 MAGs were identified at the species level (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003ed, those MAGs covered nine bacterial phyla. Most MAGs belonged to \u003cem\u003eActinobacteriota\u003c/em\u003e (18 MAGs), followed by \u003cem\u003eFirmicutes_A\u003c/em\u003e (12 MAGs), \u003cem\u003eFirmicutes\u003c/em\u003e (eight MAGs), \u003cem\u003eProteobacteria\u003c/em\u003e (three MAGs), \u003cem\u003eFusobacteriota\u003c/em\u003e (two MAGs), \u003cem\u003eBacteroidota\u003c/em\u003e (two MAGs), \u003cem\u003eSpirochaetota\u003c/em\u003e (one MAG), \u003cem\u003ePatescibacteria\u003c/em\u003e (one MAG), and \u003cem\u003eCampylobacterota\u003c/em\u003e (one MAG). Among those phyla, \u003cem\u003eActinobacteriota, Firmicutes, Proteobacteria, Fusobacteriota, Bacteroidota\u003c/em\u003e, and \u003cem\u003eSpirochaetota\u003c/em\u003e were also detected in Kraken2, which belonged to the 196 species set. Moreover, we successfully identified \u003cem\u003eCorynebacterium argentoratense\u003c/em\u003e and \u003cem\u003eHaemophilus_A parahaemolyticus\u003c/em\u003e, which had low mean relative abundance (\u0026lt;\u0026thinsp;1%).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003eNew species identified in healthy individual lung\u003c/h2\u003e\u003cp\u003eMeanwhile, we uncovered some fascinating results regarding the identification of new bacterial species. The other eight MAGs not identified at the species level were defined as potential novel species and assigned to three phyla: \u003cem\u003ePatescibacteria\u003c/em\u003e (three MAG), \u003cem\u003eFirmicutes_A\u003c/em\u003e (three MAG), and \u003cem\u003eFirmicutes\u003c/em\u003e (two MAG) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003ed). The phylogenetic tree of the 56 MAGs was constructed based on 120 conserved proteins. The taxonomy of these MAGs at the phylum level was consistent with the phylogenetic tree (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003ee, Table \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e). Particularly, three out of four MAGs in \u003cem\u003eProteobacteria\u003c/em\u003e were potential new species. Five out of the eight potential MAGs were highly qualified and had an ANI value less than 0.95. The five MAGs were further compared with bacterial genomes recently reported from two cohorts [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Among these bacterial species, ANI with these five MAGs was less than 0.95, indicating that they represent unknown species identified for the first time in this study.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003eFunctional characterizations of high-quality MAGs\u003c/h2\u003e\u003cp\u003eWe analyzed the gene functions of 29 high-quality MAGs to gain a better understanding of the functions of the lung microbiota. At KEGG level 1, there were six identified pathways, including metabolism (109 pathways in level 2), human diseases (32 pathways in level 2), organismal systems (27 pathways in level 2), environmental information processing (13 pathways in level 2), cellular processes (12 pathways in level 2), and genetic information processing (12 pathways in level 2). In total, we identified 205 pathways at level 2. Among these, 118 pathways (57.56%) were consistently present in all samples, indicating that these pathways represent core functions of the lung microbiota. The most abundant pathways were depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003ef. The largest number of pathways in level 2 were related to metabolism and genetic information processing, indicating their critical role in the lung micro-ecosystem. Particularly, the bacterial secretion system, protein export, and \u003cem\u003estaphylococcus aureus\u003c/em\u003e infection were also detected, indicating the cross-talk between bacteria and the host. The observation of \u003cem\u003estaphylococcus aureus\u003c/em\u003e infection is supported by the high relative abundance of \u003cem\u003estaphylococcus aureus\u003c/em\u003e detected using Kraken2.\u003c/p\u003e\u003cp\u003eThe prevalence of viral factors from 29 highly qualified MAGs was assessed based on the Virulence Factor Database (VFDB). A total of 575 virulence factors were identified. The top 20 high-abundance virulence factors are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e3\u003c/span\u003eg, including ABC transport-associated virulence genes.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec22\" class=\"Section2\"\u003e\u003ch2\u003eTaxonomic and functional profiles indicate minimal alterations in COVID-19 and NSCLC\u003c/h2\u003e\u003cp\u003eTo explore the differences in microbial profiling between healthy individuals and lung-diseased patients, we selected COVID-19 and NSCLC as representatives of infectious disease and cancer, respectively. We identified 212 species from 23 COVID-19 patients, 28 NSCLC patients, and 43 individuals in the healthy group after quality control and removal of human reads. Overall, the lung microbiome compositions, based on the beta-diversity metrics of the three groups, were minimally different (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). Moreover, COVID-19 samples were more scattered, while normal samples cluster closer together. The microbiome pairwise distances of NSCLC participants were more similar to those of the healthy group than those of individuals with COVID-19 (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e4\u003c/span\u003eb). Meanwhile, the top 15 species with high relative abundance in the healthy group changed more significantly in the COVID-19 group compared to NSCLC (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e4\u003c/span\u003ec).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTo further explore the microbiota changes between healthy and lung-diseased individuals, MaAsLin2 analysis was conducted based on species detected via Kraken2. We identified 47 high and 14 lower species in the comparison between the healthy group and the lung-diseased group, 23 high and 118 lower species in the COVID-19 versus healthy group comparison, and 44 high and 33 lower species in the NSCLC versus healthy group comparison (Table S6). Among the top 15 high-abundance bacteria in the healthy group, \u003cem\u003eEscherichia coli\u003c/em\u003e, \u003cem\u003eTropheryma whipplei\u003c/em\u003e, and \u003cem\u003eAeromonas hydrophila\u003c/em\u003e were significantly more abundant in the healthy group compared to the lung disease group. Seven out of 32 representative species in the healthy group showed enrichment compared to the lung disease group, emphasizing their potential role in maintaining lung homeostasis. The top 50 significantly different microbial species were displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e4\u003c/span\u003ed. Among those species, 20 were significantly enriched in the healthy group, while 30 were in the lung disease groups. Among those 50 species, \u003cem\u003eParvimonas micra\u003c/em\u003e also belonged to the 15 high-abundance species in the healthy group.\u003c/p\u003e\u003cp\u003eNext, we calculated differences in the relative abundance of bacteria among the three groups with a threshold of absolute value of fold change\u0026thinsp;\u0026gt;\u0026thinsp;1 and P-value\u0026thinsp;\u0026lt;\u0026thinsp;0.01 using the Wilcoxon test (Table S7 and S8; Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e4\u003c/span\u003ee-f). \u003cem\u003eRothia dentocariosa\u003c/em\u003e, \u003cem\u003eCorynebacterium kefirresidentii, Sphingobium yanoikuyae\u003c/em\u003e were enriched in the healthy group, while \u003cem\u003eSchaalia odontolytica\u003c/em\u003e was enriched in the NSCLC group. According to a case report, \u003cem\u003eRothia dentocariosa\u003c/em\u003e was cultured from BALF and caused opportunistic pulmonary infection [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. \u003cem\u003eCorynebacterium kefirresidentii\u003c/em\u003e was identified as a dominant member of the human skin microbiome through 16SrRNA sequencing [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. \u003cem\u003eSphingobium yanoikuyae\u003c/em\u003e can degrade carcinogenic products and may reduce in gastric cancer [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. \u003cem\u003eSchaalia odontolytica\u003c/em\u003e was linked to resistance to neoadjuvant chemoradiotherapy in rectal cancer [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The relative abundance of MAGs among the three groups was also assessed (Table S9 and S10; Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e4\u003c/span\u003ef). \u003cem\u003eActinomyces graevenitzii\u003c/em\u003e and \u003cem\u003eRothia mucilaginosa\u003c/em\u003e were significantly enriched in NSCLC. \u003cem\u003eActinomyces graevenitzii\u003c/em\u003e was isolated and cultured from the BALF of a patient with pulmonary actinomycosis [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study provides a comprehensive dataset of microbiota, an exhaustive catalog of MAGs, identifies potential new species in healthy lungs, and explores microbial alterations in lung diseases. Microbial profiling of 196 non-low-abundance species from 43 healthy individuals was conducted. The 15 dominant species consist of communal pathogens, opportunistic pathogens, and non-pathogenic commensals. For instance, \u003cem\u003eStaphylococcus aureus\u003c/em\u003e is detected in 30% of the population, mostly colonized in the nasal, throat, skin, and gastrointestinal tract [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], and has the potential to cause pneumonia. \u003cem\u003eKlebsiella pneumoniae\u003c/em\u003e has emerged as a major clinical health threat due to the rise of multidrug-resistant strains [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Interestingly, \u003cem\u003eRothia mucilaginosa\u003c/em\u003e has an anti-inflammatory effect induced by pathogens or lipopolysaccharides in chronic lung disease [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Proteins produced by \u003cem\u003eStreptococcus mitis\u003c/em\u003e and \u003cem\u003eStreptococcus oralis\u003c/em\u003e from viral pneumonia patients significantly increased virus replication in lung epithelial cells [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. And \u003cem\u003eStreptococcus mitis\u003c/em\u003e was linked to preserved lung function and favorable survival [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eTherefore, our results indicated that the dominant and keystone bacterial species consisted of both pathogens and non-pathogens, suggesting a complex microbial community in the lower respiratory tract of healthy lungs. It has been reported that pathogens, including \u003cem\u003eStreptococcus pneumoniae\u003c/em\u003e and \u003cem\u003eHaemophilus influenzae\u003c/em\u003e were found to be colonized in the airways of 20\u0026ndash;50% of healthy individuals [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. With an increasing number of bioinformatic methods being utilized to distinguish pathogens from background microorganisms, our findings may assist in identifying pathogens in the clinical diagnosis of pulmonary infections[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eWe also identified a significant amount of oral-associated microbiota in the lower respiratory tract (LTR) of healthy adult participant. In keystone species obtained through co-occurrence network analysis, \u003cem\u003eGranulicatella adiacens\u003c/em\u003e belongs to microbiota in the oral cavity, urogenital tract, and intestinal tract [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] and can cause periodontitis [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. \u003cem\u003eMogibacterium diversum\u003c/em\u003e was detected in human saliva samples [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. \u003cem\u003eSalmonella enterica\u003c/em\u003e has been isolated from stool samples of individuals with diarrhea in Grace[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Meanwhile, the dominant species \u003cem\u003eStreptococcus mitis, Streptococcus oralis\u003c/em\u003e were also oral involved. It have been reported that oral bacteria are likely the source of lung microbiota in healthy individuals [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. It was concluded that an increased presence of oral microbes entering the lungs was associated with reduced lung function and elevated levels of pro-inflammatory cytokines [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eOur study provided exhaustive genomics datasets for the lung microbiota. The co-assembly strategy increased the recovery ratio of MAGs, for instance, four of 29 high-quality MAGs were obtained using the co-assembly method. Furthermore, one out of five new species originated from that method, indicating that co-assembly enhanced the potential for reconstructing novel genomes. It was reported that co-assembly is an effective approach for reconstructing low-abundance MAGs [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. In this study, MAG6, MAG8, MAG9, and MAG12 were obtained through the co-assembly method, and their abundance was lower at the individual level compared to MAGs from the single-sample assembly. Moreover, in 29 high-quality MAGs, 15 were detected by Kraken2. Six were identified at the genus level, three were not identified at the genus level, and five new species were discovered. This indicates that assembly expanded the scope of microbiota identification. Meanwhile, \u003cem\u003eTropheryma whipplei\u003c/em\u003e (MAG1), \u003cem\u003eRothia mucilaginosa\u003c/em\u003e (MAG34), and \u003cem\u003eParvimonas micra\u003c/em\u003e (MAG56), which belonged to 29 MAGs, were detected among the top 15 high-abundance species according to Kraken2, highlighting their significant role in lung microecology.\u003c/p\u003e\u003cp\u003eWe observed that the microbiota of the diseased lung was altered compared to a healthy state, and the abundance of bacteria varied between the two different lung diseases. Interestingly, COVID-19 patients have a more distinct profile compared to NSCLS. For instance, individuals with COVID-19 had smaller Bray-Curtis distance values and experienced more microbial disturbance. Our results suggest that the microbial communities in infection-induced lung diseases may be more pronounced than in lung cancer, which requires further data for confirmation.\u003c/p\u003e\u003cp\u003eThere are several limitations in our study. First, the sample size was not large enough to elucidate the lung micro-ecology. Second, each sample needs a blank control experiment to eliminate environmental contamination. Third, some crucial species need to be validated through experiments. Overall, we present microbial abundance and MAGs, which will significantly enhance the capacity to conduct taxonomic grouping and metagenomic analyses for future studies on the pulmonary microbiota, as well as the associations between health and diseases. We also lay the groundwork for the future development of more personalized and precise management and therapy strategies for lung health.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by National Natural Science Foundation of China (Nos. 92159302, 32370628, 32170592); The Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences(2022-I2M-CoV19-006); State Key Laboratory Special Fund (No. 2060204); Chinese Academy of Medical Sciences Innovation Fund for Medical Sciences (No. 2023-12M-2-001); State Key Laboratory of Respiratory Health and Multimorbidity, State Key Laboratory Special Fund 2060204; the Science and Technology Project of Sichuan (Nos. 2022ZDZX0018, 2023NSFSC004, 2024NSFSC0402); Key R\u0026amp;D Support Plan of Chengdu Science and Technology Bureau (No. 2023-YF09-00039-SN) and 1.3.5 project for disciplines of excellence, West China Hospital, Sichuan University (No. ZYGD22009).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWML and ZFW conceived the study design. XW, YRY, YQZ and RJX collected clinical information about samples. YQL LLX and NNC CPL, THT and JZ performed genome sequencing. CC carried out bioinformatic analysis. SYW, HYW, DL, QLW, YC and XH interpreted the data. CC, YQL, HYW and SYW wrote the article. WML, HYW and DL provided clinical insights. All authors discussed the results and reviewed the article.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe are grateful to all participants who volunteered for this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData sharing statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe raw metagenomic sequencing\u0026nbsp;data generated in this study was\u0026nbsp; \u0026nbsp;available in the National Genomics Data Center (NGDC) Genome Sequence Archive (GSA) database (https://bigd.big.ac.cn/gsa-human/) under accession number HRA008205.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was conducted in accordance with the Declaration of Helsinki and approved by the ethics committees of all participating centers, reference number 2023.1175. Informed written consent was obtained from all patients at the participating institutions to permit the archiving of their biospecimens and their use in future studies. Basic and clinical information about the patients was collected during their hospital visits.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eNatalini, J.G., S. Singh, and L.N. Segal, \u003cem\u003eThe dynamic lung microbiome in health and disease.\u003c/em\u003e Nat Rev Microbiol, 2023. \u003cstrong\u003e21\u003c/strong\u003e(4): p. 222-235.\u003c/li\u003e\n \u003cli\u003eYuan, Y., et al., \u003cem\u003ePulmonary Actinomyces graevenitzii Infection: Case Report and Review of the Literature.\u003c/em\u003e Front Med (Lausanne), 2022. \u003cstrong\u003e9\u003c/strong\u003e: p. 916817.\u003c/li\u003e\n \u003cli\u003eCharlson, E.S., et al., \u003cem\u003eTopographical continuity of bacterial populations in the healthy human respiratory tract.\u003c/em\u003e Am J Respir Crit Care Med, 2011. \u003cstrong\u003e184\u003c/strong\u003e(8): p. 957-63.\u003c/li\u003e\n \u003cli\u003eStearns, J.C., et al., \u003cem\u003eCulture and molecular-based profiles show shifts in bacterial communities of the upper respiratory tract that occur with age.\u003c/em\u003e Isme j, 2015. \u003cstrong\u003e9\u003c/strong\u003e(5): p. 1268.\u003c/li\u003e\n \u003cli\u003eAhmed, B. and M.J. Cox, \u003cem\u003eComparison of the upper and lower airway microbiota in children with chronic lung diseases.\u003c/em\u003e 2018. \u003cstrong\u003e13\u003c/strong\u003e(8): p. e0201156.\u003c/li\u003e\n \u003cli\u003ePattaroni, C., et al., \u003cem\u003eEarly life inter-kingdom interactions shape the immunological environment of the airways.\u003c/em\u003e Microbiome, 2022. \u003cstrong\u003e10\u003c/strong\u003e(1): p. 34.\u003c/li\u003e\n \u003cli\u003e!!! INVALID CITATION !!!\u003c/li\u003e\n \u003cli\u003eNayfach, S., et al., \u003cem\u003eNew insights from uncultivated genomes of the global human gut microbiome.\u003c/em\u003e Nature, 2019. \u003cstrong\u003e568\u003c/strong\u003e(7753): p. 505-510.\u003c/li\u003e\n \u003cli\u003eMan, W.H., et al., \u003cem\u003eBacterial and viral respiratory tract microbiota and host characteristics in children with lower respiratory tract infections: a matched case-control study.\u003c/em\u003e Lancet Respir Med, 2019. \u003cstrong\u003e7\u003c/strong\u003e(5): p. 417-426.\u003c/li\u003e\n \u003cli\u003eHan, W., N. Wang, and M. Han, \u003cem\u003eIdentification of microbial markers associated with lung cancer based on multi-cohort 16\u0026thinsp;s rRNA analyses: A systematic review and meta-analysis.\u003c/em\u003e 2023. \u003cstrong\u003e12\u003c/strong\u003e(18): p. 19301-19319.\u003c/li\u003e\n \u003cli\u003eJin, J., et al., \u003cem\u003eDiminishing microbiome richness and distinction in the lower respiratory tract of lung cancer patients: A multiple comparative study design with independent validation.\u003c/em\u003e Lung Cancer, 2019. \u003cstrong\u003e136\u003c/strong\u003e: p. 129-135.\u003c/li\u003e\n \u003cli\u003eLi, Y., et al., \u003cem\u003eDysbiosis of lower respiratory tract microbiome are associated with proinflammatory states in non-small cell lung cancer patients.\u003c/em\u003e Thorac Cancer, 2024. \u003cstrong\u003e15\u003c/strong\u003e(2): p. 111-121.\u003c/li\u003e\n \u003cli\u003eIbironke, O., et al., \u003cem\u003eSpecies-level evaluation of the human respiratory microbiome.\u003c/em\u003e Gigascience, 2020. \u003cstrong\u003e9\u003c/strong\u003e(4).\u003c/li\u003e\n \u003cli\u003eZou, Y., et al., \u003cem\u003e1,520 reference genomes from cultivated human gut bacteria enable functional microbiome analyses.\u003c/em\u003e Nat Biotechnol, 2019. \u003cstrong\u003e37\u003c/strong\u003e(2): p. 179-185.\u003c/li\u003e\n \u003cli\u003eAlmeida, A., et al., \u003cem\u003eA unified catalog of 204,938 reference genomes from the human gut microbiome.\u003c/em\u003e Nat Biotechnol, 2021. \u003cstrong\u003e39\u003c/strong\u003e(1): p. 105-114.\u003c/li\u003e\n \u003cli\u003eSchiff, M.J. and M.H. Kaplan, \u003cem\u003eRothia dentocariosa pneumonia in an immunocompromised patient.\u003c/em\u003e Lung, 1987. \u003cstrong\u003e165\u003c/strong\u003e(5): p. 279-82.\u003c/li\u003e\n \u003cli\u003eSwaney, M.H., et al., \u003cem\u003eSweat and Sebum Preferences of the Human Skin Microbiota.\u003c/em\u003e 2023. \u003cstrong\u003e11\u003c/strong\u003e(1): p. e0418022.\u003c/li\u003e\n \u003cli\u003eHu, Y.L., et al., \u003cem\u003eThe Gastric Microbiome Is Perturbed in Advanced Gastric Adenocarcinoma Identified Through Shotgun Metagenomics.\u003c/em\u003e Front Cell Infect Microbiol, 2018. \u003cstrong\u003e8\u003c/strong\u003e: p. 433.\u003c/li\u003e\n \u003cli\u003eHuang, X., et al., \u003cem\u003eMetagenomic Analysis of Intratumoral Microbiome Linking to Response to Neoadjuvant Chemoradiotherapy in Rectal Cancer.\u003c/em\u003e Int J Radiat Oncol Biol Phys, 2023. \u003cstrong\u003e117\u003c/strong\u003e(5): p. 1255-1269.\u003c/li\u003e\n \u003cli\u003eDeinhardt-Emmer, S., et al., \u003cem\u003eVirulence patterns of Staphylococcus aureus strains from nasopharyngeal colonization.\u003c/em\u003e J Hosp Infect, 2018. \u003cstrong\u003e100\u003c/strong\u003e(3): p. 309-315.\u003c/li\u003e\n \u003cli\u003eHu Y, Feng Y, Zong Z. \u003cem\u003eGlobal distribution of Klebsiella pneumoniae producing extended-spectrum\u0026nbsp;\u003c/em\u003e\u003cem\u003e\u0026beta;\u003c/em\u003e\u003cem\u003e-lactamases in neonates\u003c/em\u003e. Precis Clin Med. 2024. \u003cstrong\u003e7\u003c/strong\u003e(4):pbae031.\u003c/li\u003e\n \u003cli\u003eRigauts, C., J. Aizawa, and S.L. Taylor, \u003cem\u003eR othia mucilaginosa is an anti-inflammatory bacterium in the respiratory tract of patients with chronic lung disease.\u003c/em\u003e 2022. \u003cstrong\u003e59\u003c/strong\u003e(5).\u003c/li\u003e\n \u003cli\u003eNishioka, K., et al., \u003cem\u003eProteins produced by Streptococcus species in the lower respiratory tract can modify antiviral responses against influenza virus in respiratory epithelial cells.\u003c/em\u003e Microbes Infect, 2021. \u003cstrong\u003e23\u003c/strong\u003e(1): p. 104764.\u003c/li\u003e\n \u003cli\u003eO\u0026apos;Dwyer, D.N., et al., \u003cem\u003eCommensal Oral Microbiota, Disease Severity and Mortality in Fibrotic Lung Disease.\u003c/em\u003e 2023.\u003c/li\u003e\n \u003cli\u003eCrawford, E.D., et al., Proc Natl Acad Sci U S A.\u003c/li\u003e\n \u003cli\u003eLangelier, C., et al., \u003cem\u003eMetagenomic Sequencing Detects Respiratory Pathogens in Hematopoietic Cellular Transplant Patients.\u003c/em\u003e Am J Respir Crit Care Med, 2018. \u003cstrong\u003e197\u003c/strong\u003e(4): p. 524-528.\u003c/li\u003e\n \u003cli\u003eRuoff, K.L., \u003cem\u003eNutritionally variant streptococci.\u003c/em\u003e Clin Microbiol Rev, 1991. \u003cstrong\u003e4\u003c/strong\u003e(2): p. 184-90.\u003c/li\u003e\n \u003cli\u003eBelstr\u0026oslash;m, D., et al., \u003cem\u003eDifferences in bacterial saliva profile between periodontitis patients and a control cohort.\u003c/em\u003e J Clin Periodontol, 2014. \u003cstrong\u003e41\u003c/strong\u003e(2): p. 104-12.\u003c/li\u003e\n \u003cli\u003eBaker, J.L., \u003cem\u003eUsing Nanopore Sequencing to Obtain Complete Bacterial Genomes from Saliva Samples.\u003c/em\u003e mSystems, 2022. \u003cstrong\u003e7\u003c/strong\u003e(5): p. e0049122.\u003c/li\u003e\n \u003cli\u003eMaraki, S., et al., \u003cem\u003eA 5-year study of the bacterial pathogens associated with acute diarrhoea on the island of Crete, Greece, and their resistance to antibiotics.\u003c/em\u003e Eur J Epidemiol, 2003. \u003cstrong\u003e18\u003c/strong\u003e(1): p. 85-90.\u003c/li\u003e\n \u003cli\u003eBassis, C.M., et al., \u003cem\u003eAnalysis of the upper respiratory tract microbiotas as the source of the lung and gastric microbiotas in healthy individuals.\u003c/em\u003e mBio, 2015. \u003cstrong\u003e6\u003c/strong\u003e(2): p. e00037.\u003c/li\u003e\n \u003cli\u003eZhang, J., et al., \u003cem\u003eDifferential Oral Microbial Input Determines Two Microbiota Pneumo-Types Associated with Health Status.\u003c/em\u003e 2022. \u003cstrong\u003e9\u003c/strong\u003e(32): p. e2203115.\u003c/li\u003e\n \u003cli\u003eStewart, R.D., et al., \u003cem\u003eAssembly of 913 microbial genomes from metagenomic sequencing of the cow rumen.\u003c/em\u003e Nat Commun, 2018. \u003cstrong\u003e9\u003c/strong\u003e(1): p. 870.\u003c/li\u003e\n \u003cli\u003eKultima, J.R., et al., \u003cem\u003eMOCAT2: a metagenomic assembly, annotation and profiling framework.\u003c/em\u003e Bioinformatics, 2016. \u003cstrong\u003e32\u003c/strong\u003e(16): p. 2520-3.\u003c/li\u003e\n \u003cli\u003eLu, J., et al., \u003cem\u003eBracken: Estimating species abundance in metagenomics data.\u003c/em\u003e PeerJ Computer Science, 2017. \u003cstrong\u003e3\u003c/strong\u003e: p. e104.\u003c/li\u003e\n \u003cli\u003eBeghini, F., et al., \u003cem\u003eIntegrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3.\u003c/em\u003e 2021. \u003cstrong\u003e10\u003c/strong\u003e.\u003c/li\u003e\n \u003cli\u003eSuzek, B.E., et al., \u003cem\u003eUniRef: comprehensive and non-redundant UniProt reference clusters.\u003c/em\u003e Bioinformatics, 2007. \u003cstrong\u003e23\u003c/strong\u003e(10): p. 1282-8.\u003c/li\u003e\n \u003cli\u003eCaspi, R., et al., \u003cem\u003eThe MetaCyc database of metabolic pathways and enzymes - a 2019 update.\u003c/em\u003e Nucleic Acids Res, 2020. \u003cstrong\u003e48\u003c/strong\u003e(D1): p. D445-d453.\u003c/li\u003e\n \u003cli\u003eLangfelder, P. and S. Horvath, \u003cem\u003eWGCNA: an R package for weighted correlation network analysis.\u003c/em\u003e BMC Bioinformatics, 2008. \u003cstrong\u003e9\u003c/strong\u003e: p. 559.\u003c/li\u003e\n \u003cli\u003eUritskiy, G.V., J. DiRuggiero, and J. Taylor, \u003cem\u003eMetaWRAP-a flexible pipeline for genome-resolved metagenomic data analysis.\u003c/em\u003e Microbiome, 2018. \u003cstrong\u003e6\u003c/strong\u003e(1): p. 158.\u003c/li\u003e\n \u003cli\u003eSieber, C.M.K., et al., \u003cem\u003eRecovery of genomes from metagenomes via a dereplication, aggregation and scoring strategy.\u003c/em\u003e Nat Microbiol, 2018. \u003cstrong\u003e3\u003c/strong\u003e(7): p. 836-843.\u003c/li\u003e\n \u003cli\u003eChklovski, A., et al., \u003cem\u003eCheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning.\u003c/em\u003e Nat Methods, 2023. \u003cstrong\u003e20\u003c/strong\u003e(8): p. 1203-1212.\u003c/li\u003e\n \u003cli\u003eChaumeil, P.A. and A.J. Mussig, \u003cem\u003eGTDB-Tk v2: memory friendly classification with the genome taxonomy database.\u003c/em\u003e 2022. \u003cstrong\u003e38\u003c/strong\u003e(23): p. 5315-5316.\u003c/li\u003e\n \u003cli\u003eHyatt, D., et al., \u003cem\u003eProdigal: prokaryotic gene recognition and translation initiation site identification.\u003c/em\u003e BMC Bioinformatics, 2010. \u003cstrong\u003e11\u003c/strong\u003e: p. 119.\u003c/li\u003e\n \u003cli\u003eLi, W. and A. Godzik, \u003cem\u003eCd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.\u003c/em\u003e Bioinformatics, 2006. \u003cstrong\u003e22\u003c/strong\u003e(13): p. 1658-9.\u003c/li\u003e\n \u003cli\u003eAramaki, T., et al., \u003cem\u003eKofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold.\u003c/em\u003e Bioinformatics, 2020. \u003cstrong\u003e36\u003c/strong\u003e(7): p. 2251-2252.\u003c/li\u003e\n \u003cli\u003eLiu, B., et al., \u003cem\u003eVFDB 2022: a general classification scheme for bacterial virulence factors.\u003c/em\u003e Nucleic Acids Res, 2022. \u003cstrong\u003e50\u003c/strong\u003e(D1): p. D912-d917.\u003c/li\u003e\n \u003cli\u003eLi, H. and R. Durbin, \u003cem\u003eFast and accurate short read alignment with Burrows-Wheeler transform.\u003c/em\u003e Bioinformatics, 2009. \u003cstrong\u003e25\u003c/strong\u003e(14): p. 1754-60.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"metagenome, lower respiratory tract, pulmonary microbiome, COVID-19, lung cancer","lastPublishedDoi":"10.21203/rs.3.rs-7937892/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7937892/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe microbiota plays an important role in maintaining lung homeostasis and the development of respiratory diseases. However, there have not been detailed characteristics of lung microbiome in health and respiratory diseases. Bronchoalveolar lavage fluid of 43 healthy individuals, 23 COVID patients, and 28 non-small cell lung cancerpatients was conducted to characterize microbial diversity and function by metagenomic sequencing. We identified 196 species in health lung after removing low abnudance species. The most abundant species were Staphylococcus aureus, Salmonella enterica, and Klebsiella pneumoniae. Keystone species were identified through co-network analysis, such as Granulicatella adiacens and Mogibacterium diversum. To further explore the microbial function in the healthy lung, we obtained 56 non-redundant metagenome-assembled genomes (MAGs) and subsequently classified them into nine phyla. Metabolism and genetic information processing exhibit high abundance, including bacterial secretion systems, protein export, and Staphylococcus aureus infection, indicating a close correlation between microbiota and the host. Furthermore, the microbiome pairwise distances of non-small cell lung cancer (NSCLC) participants were more similar to those of the healthy group than those of individuals with COVID-19. Rothia dentocariosa, Corynebacterium kefirresidentii, Sphingobium yanoikuyae were enriched in the healthy group, while Schaalia odontolytica, Actinomyces graevenitzii, and Rothia mucilaginosa was enriched in the NSCLC group. Overall, this study provides a comprehensive overview of the diversity and gene function of the lung microbiome, which will facilitate future studies of microbiota associated with respiratory diseases in humans.\u003c/p\u003e","manuscriptTitle":"Deciphering the microbial landscape of lower respiratory tract in health and respiratory diseases","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-20 11:32:03","doi":"10.21203/rs.3.rs-7937892/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"bea0bc7c-ea2d-40ba-b210-bd231b68ed99","owner":[],"postedDate":"November 20th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-02T16:40:23+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-20 11:32:03","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7937892","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7937892","identity":"rs-7937892","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.