Sequencing of 185 Streptococcus thermophilus and identification of fermentation biomarkers | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Sequencing of 185 Streptococcus thermophilus and identification of fermentation biomarkers Wenjun Liu, Linjie Wu, Jie Zhao, Weicheng Li, Yu Wang, Huijuan Zheng, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-61428/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Streptococcus ( S. ) thermophilus is an important dairy starter in the production of fermented dairy products has important significance, from natural fermentation in the past to industrial production today. While the genetic architecture underlying S. thermophilus traits and phenotypes is largely unknown. Here, we sequenced 185 S. thermophilus strains, which isolated from natural fermented dairy products of China and Mongolia and using comparative genomic and genome wide association study to provide novel point for genetic architecture underlying its traits and phenotypes. Genome analysis of S. thermophilus showed association of phylogeny with environmental and phenotypic features and revealed clades with high acid production potential or with substantial genome decay. A few S. thermophilus isolated from areas with high chloramphenicol emissions had a chloramphenicol-resistant gene CatB8 . Most importantly, we defined a growth score and identified a missense mutation G1118698T located at the gene Acn A that were both predictive of acidification capability of S. thermophilus . Our findings provide novel insight in S. thermophilus genetic traits, antibiotic resistant and predictive of acidification capability which both may had huge help in culture starter screening. General Microbiology Bacteriology Evolutionary Genetics Streptococcus thermophilus Whole genome sequencing Antibiotic resistance Growth score Acidification Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Streptococcus (S.) thermophilus is a predominant lactic acid bacterium (LAB) with rapid acidification capability and is a major dairy starter used in the milk fermentation and cheese production[ 1 ]. The starter strains of fermentation have vast economic value and produce a large amount fermented dairy products every year[ 2 ]. Traditional nomads have been making naturally fermented milk for thousands of years and these naturally fermented dairy products contain rich LAB resources[ 3 ]. However, previous researches mostly concentrate on genomics of industry strains of S. thermophilus. For example, one study showed that most of 47 industry S. thermophilus have similar genetic distance, indicated genome stability of industry strains[ 4 ]. The genetic architecture underlying natural S. thermophilus strains is largely unexplored. In this study, we isolated and sequenced 185 S. thermophilus strains from natural dairy products in traditional pasture of different regions in China and Mongolia, to uncover the genetic evolution of wild strains and to explore the relationship between genotype and fermentation characteristics. Phylogenetic tree of S. thermophilus revealed clades with high acid production potential or significant genome decay. We also found 13 S. thermophilus which have a chloramphenicol-resistant gene CatB8 isolated from areas with high chloramphenicol emissions. Besides, we defined a growth score and showed that the growth score could accurately predict acidification capability of S. thermophilus. With the growth score, we identified a missense mutation G1118698T by genome wide association study (GWAS) located at the gene Acn A that were significantly associated with the acidification. Our findings provided novel insights in S. thermophilus genetic traits and identified robust biomarkers screening of culture starter with high acidification capability. Results Phenotypic analysis and phylogenetic analysis of S. thermophilus From natural fermented dairy products in China and Mongolia, we isolated 185 S. thermophilus and found 61 strains having potential high acid production capability (H-Acid) using a preliminary fermentation experiment. Among the 185 S. thermophilus , 65 were from China and 120 were from Mongolia (Fig. 1 A). Whole genome sequencing (WGS) was performed for all 185 strains (mean coverage 374X, Table S1 ). After short read alignment to the reference genome CNRZ1066 of S. thermophilus , we identified 58,734 single nucleic acid polymorphisms (SNPs), 2,246 insertions and deletions (Indels, Table S2 ) and 450 copy number variations (CNVs, Table S3 ). As expected, ribosomal RNA (rRNA) and transfer RNA (tRNA) had the lowest mutation rates, followed by protein coding regions, pseudogenes and intergenic regions (Fig. 1 B). De novo assembly showed that the genome sizes of S. thermophilus ranged from 1.72 to 2.60 Mb and the numbers of genes ranged from 1,704 to 2,158 (Fig. 1 C-D). Pan-genome analysis of our 185 isolates as well as 32 S. thermophilus available from the NCBI[ 5 ] totally identified 7,629 genes including 827 core genes and 315 soft core genes (Fig. 1 E, Table S4-S6 ). The core and soft-core genes were mainly enriched in metabolism pathways like metabolic and biosynthesis of amino acids, which are essential for bacterial growth ( Fig. S1A ). Glycolysis/gluconeogenesis pathway and starch and sucrose metabolism pathways were enriched in shell genes (shared by 15%-95% samples). These two pathways are involved in carbohydrate utilization. Diversity in these pathways indicates that S. thermophilus strains might have undergone adaptive evolution in carbohydrate utilization pathways. Besides, the cloud genes (shared by less than 15% samples) were mainly enriched in quorum sensing and beta-lactam resistance pathways. 53% of cloud genes were homologous to genes in other Streptococcus strains and 8% to genes in Lactobacillales and Lactococcus , two bacterial species often found in natural fermented dairy products[ 6 ] ( Fig. S1B ). This indicated that a portion of cloud genes came from horizontal gene transfer (HGT). The phylogenetic tree of the 217 S. thermophilus revealed four large clades and was consistent with multilocus sequence typing (MLST) (Fig. 2 ) The within-clade average nucleotide identity (ANI) was significantly larger than the between-clade ANI and the principle component analysis also showed that the four clades are well-separated ( Fig. S2A-B ). Clade A, including strains from China, Mongolia and NCBI database, was closer to the roots of phylogenetic trees. In the other three clades, strains isolated from China and Mongolia also showed clear aggregation. Clade D was basically composed of strains isolated from yoghurts of Mongolia, the strains in Clade C were mainly from yoghurts of Mongolia and Xiniiang China. The clade B was mostly isolated from goat yoghurts in China and yoghurts in Mongolia. Genetic distance between isolates were significantly correlated with the geographical location of sampling sites ( Fig. S2C ). The dairy product type was also significantly correlated with the phylogenetic clades (Fisher’s exact test, p-value = 5⋅10 − 3 , Table S7 ). Interestingly, the H-Acid isolates and NCBI stains were significantly enriched in clade A and B (Fisher’s exact test, p-value = 5.1⋅10 − 7 , Fig. 2 ). Note that strains from NCBI were extensively used in industrial production and generally had high dairy fermentation capability. Furthermore, the cell-wall protease gene Prt S, a gene known to be associated with rapid growth and acidification rates at bacteria in milk[ 7 ], was significantly enriched in clade A (Fisher’s exact test, p-value = 1.3⋅10 − 28 , Table S8 ). These data implied that strains in clade A and B might have better fermentation potential than the strains in clade C and D. Genome decay and antibiotic resistance of S. thermophilus S. thermophilus in clade D had significantly fewer number of genes (Fig. 3 A), smaller genome sizes and more copy number losses than other clades, indicating that clade D might have undergone considerable genome decay (Fig. 3 B-C and Fig. S3A ). We found 131 genes that were prevalent in clade A-C (frequency > 0.5) but were significantly depleted in clade D (Fisher’s test, Benjamini–Hochberg adjusted p-value < 0.05, Table S9 ). These genes were significantly enriched in the pathways including quorum sensing, beta-lactam resistance, and ABC transporters ( Table S10 ). Further, clade D had significantly less quorum sensing genes compared with clade A and B (Fig. 3 D ) . Many of the depleted quorum sensing genes were Blp bacteriocins related genes. By comparing with the Blp protein family in Streptococcus pneumoniae ( Table S11 ), we found that blpB , blpM , blpH and blpR genes were significantly lost in strains from clade D comparing with other clades (Fig. 2 , Fisher’s test p-value < 0.001). BlpB protein, a transport accessory protein, is essential for the secretion of antimicrobial compounds. BlpM is a bacteriocin-encoding gene which directly influence the production of bacteriocins. Both BlpH and BlpR are members of two-component regulatory system, which allow bacteria to sense and respond to changes in different environment conditions. The knockout of blpB , blpH and blpR genes in S. thermophilus reduced production of bacteriocins compared with the wild type[ 8 ]. These suggest that strains in clade D may have lower production capacity of bacteriocins than strains in other clades. We found that 104 of the 131 depleted genes in D had homologs in 10 other streptococci species and these genes were also enriched in quorum sensing pathway, especially the bacteriocins cluster, followed by ABC transporters ( Table S12 ). This phenomenon was consistent with previous research, which found a striking level of genome decay in S. thermophilus compared with other streptococci[ 9 ]. Besides, 76 in 131 genes were identified in Streptococcus salivarius subspecies salivarius , implying that instead of being acquired by strains in clade A-C, the depleted genes in clade D were probably lost in strains from clade D in their adaptation to the milk niche. Due to the widespread misuse of antibiotics, the problems caused by bacterial resistance have received wide attention. By comparing pan-genome with Comprehensive Antibiotic Resistance Database[ 10 ], we found 77 genes related with resistance of 31 different antibiotics ( Table S13 ). Most antibiotic-related genes were associated with efflux (25 of 77, 32.47%) and target alteration (35 of 77, 45.45%). 16 strains in clade C had a gene homologous to the glycopeptides antibiotics resistance protein ARO3002945 ( Van H) [ 11 ]. These 16 strains with Van H were mostly isolated from Xinjiang Autonomous Region in China. Three antibiotics resistance-related genes, including two antibiotic efflux genes (ARO3000614 and ARO3004054) and one antibiotic target alteration gene ARO3004253 ( Van U) [ 12 ], are more likely being lost in clade D instead of being acquired by A-C. Besides, we also found that 13 strains in clade A had a gene homologous to the chloramphenicol resistance protein ARO3002680 ( CatB 8, Fig. 2 , Table S14 ). Most strains with CatB8 were isolated from Hongyuan prairie in Sichuan and Gannan prairie in Gansu, two high chloramphenicol emission provinces of China [ 13 ]. We identified nearby transposon sequences around CatB8 (~ 1 kb and ~ 5 kb) for 12 out of 13 strains with the CatB8 gene ( Table S15 ), indicating that the CatB8 gene was probably acquired by lateral gene transfer. CatB8 , chloramphenicol acetyltransferase, inactivates chloramphenicol by acetylation[ 14 ]. The acquisition of CatB8 might confer S. thermophilus resistance to chloramphenicol. In fact, we applied droplet digital PCR (ddPCR) for strains with CatB 8 and found a significantly higher expression of CatB8 in M17 cultures with 8 µg/ml chloramphenicol than cultures without chloramphenicol (Wilcoxon’s test, p-value = 0.0025, Fig. S3B , Table S16 ), indicating that the bacteria responded to the exposure of chloramphenicol by elevating the expression of CatB 8. These data suggested that the misuse of chloramphenicol might be closely related with the antibiotic resistant S. thermophilus and attention should be paid in the screening of potential starter to avoid the spread of resistance genes. Growth score and acidification of S. thermophilus Acidification is the most important characteristic of S. thermophilus as a starter in fermentation of dairy products. Rapid acidification can shorten fermentation time of yoghurt production and reduce the production costs. Screening of S. thermophilus strains with rapid acidification is one of most concerned problems in fermentation dairy enterprises. However, currently biomarkers for rapid acidification is still lacking. Theoretically, acidification is closely related with the growth rate of S. thermophilus. As most bacteria, S. thermophilus has a single circular genome. During replication, DNA sequences passed the replication fork should have two copies and those to be replicated should have only single copy. Thus, because millions of cells at different replication stages were used in WGS of S. thermophilus , a genomic region’s read depth should be negatively correlated with its distance to the replication origin and the strength of this correlation should reflect the growth rate of S. thermophilus . In our WGS data, the adjusted read depths (Methods) were indeed negatively correlated with the distance to the replication origin for all isolates. We defined a growth score as the negative value of Spearman’s correlation between them (Fig. 4 A-B, Fig S4A, Table S17 ). Clade A and B strains had significantly larger growth score than clade C and D ( Fig. S4B ). The growth score was significantly larger in H-Acid strains (Fig. 4 C) and in the strains with the PrtS gene (Fig. 4 D). These data implied that the growth score might provide an accurate marker for the acidification capability of S. thermophilus . Using the growth score, we performed Genome-Wide Association Study (GWAS) and stability selection to screen for genomic variations that might be related with the acidification capability of S. thermophilus (Methods). We found that 7 SNPs were significantly correlated with growth score by GWAS (Wilcoxon test, p-value < 10 − 5 or Bonferroni adjusted p-value < 0.05, Table S18 ). Among the 7 SNPs, the missense SNP A764991G located at the gene AsnC , which promotes the growth of the S. thermophilus in milk by regulating aspartic acid metabolism[ 15 , 16 ], had the highest selection probability and minimum p-value (1.4⋅10 − 8 ) in GWAS analysis. The missense SNP G1118698T located at the gene Acn A, which encodes aconitate hydratase A, also had high selection probability and was very significant in GWAS analysis (p-value = 4.5⋅10 − 6 , Fig. S4C-D ). AcnA involves in succinate and citrate production, contribute to acidification. 27 S. thermophilus with this mutation were all from clade B and D and were all in the non-H-Acid group ( Fig. S4E ). To confirm that the proposed growth score and the SNP G1118698T were associated with the acidification capability of S. thermophilus , we randomly selected 85 strains and performed fermentation experiments ( Table S19 ). We evaluated the acidification capability of S. thermophilus by the acidity and the acid production speed (Methods). The acidity and acid production speed were significantly higher in strains with the PrtS gene, and higher in clade A and B ( Fig. S4F-G ). The growth score was significantly positively correlated with acid production speed (Pearson’s correlation 0.4, p-value = 1.2⋅10 − 4 , Fig. 4 E) and acidity (Pearson’s correlation 0.35, p-value = 1.2⋅10 − 3 , Fig. 4 F). Similarly, the SNP G1118698T was significantly associated with acidity and the acidification speed (Fig. 4 G). These data suggested that the proposed growth score and the SNP G1118698T could serve as reliable biomarkers for screening S. thermophilus isolates with high acidification speed. Discussion We reported so far the largest WGS data of S. thermophilus isolated from natural fermentation dairy products. This large amount of data allowed us to systematically investigate the genomics landscape of S. thermophilus. We found that S. thermophilus had four large clades. Strains in clade A seemed to have high industry application potential, while strains in clade D might have undergone considerable genome decay through gene loss. We also identified novel biomarkers for the acidification capability of S. thermophilus. The data and novel discoveries in this paper provided valuable resources for understanding the evolution and genomics of S. thermophilus and for the industry application of S. thermophilus. Compared with many other streptococci species, S. thermophilus lives in a rather stable environment. Previous researches[ 17 ] discovered that many virulence-related genes were lost or became pseudogenes in S. thermophilus. Here, we revealed that in clade D strains of S. thermophilus had significantly smaller genomes than strains in other clades, indicating that the genome decay might be an ongoing process of S. thermophilus. This large genome decay was possibly due to the adaptation of S. thermophilus to their stable niche of milk. In fact, mathematical simulation[ 18 ] showed that stable environments often lead to smaller genomes than environments with greater variability. Many quorum-sensing genes, especially blp genes, were lost in clade D strains. By analyzing genomes of three S. thermophilus strains (LMG18311, CNRZ1066, and LMD-9), previous research[ 8 ] found that, although identified in all three strains, the blp gene clusters were only full functional in LMD-9. It was thus plausible that these blp genes conferred little or no survival advantages to S. thermophilus , and then the genes in this pathway gradually became inactive and eventually lost clade D S. thermophilus strains. The traditional method for evaluating the acidification capability of S. thermophilus strains is very labor-intensive, time-consuming and costly. With the advancement of sequencing technologies, WGS becomes very efficient and cost-effective. The growth score defined in this paper can be calculated only using WGS data and thus provided a very convenient and cost-effective surrogate to the traditional evaluation method. With this score, one can easily screen hundreds of S. thermophilus strains. In addition, since fast growth is often a desirable property for industrial bacteria, the growth score defined here might also serve as a robust criterion for evaluating other industrial bacteria and thus has a large application potential in industry. One disadvantage of this growth score is that it currently can only be used for monoculture bacteria. In industrial applications, multiple species of bacteria are often simultaneously used. WGS of the mixture of the different bacteria would cause read mapping ambiguity to different reference genomes and thus the growth score cannot be directly applied. However, we could only consider the genomic sequences that are unique to each species and generalize the growth score using these unique sequences for the mixed sequencing data. Materials And Methods Variant calling, assemble and annotation Genomic DNA was sequenced using an Illumina HiSeq 4000 platform (Illumina, San Diego, CA) generating 150-bp paired-end reads with an average insert size of 350 bps. All 185 S. thermophilus sequencing data were mapped to reference genome CNRZ1066 by BWA-mem[ 19 ] with default parameters. SNPs and Indels were called by GATK Unifiedgenotyper[ 20 ] and annotated by SnpEff[ 21 ]. SNPs having Indels within its 10 bp neighborhood were filtered. CNVs were called by CNV-BAC[ 22 ]. We performed de novo assembly using SOAPdenovo2[ 23 ] (k-mer = 71). The contigs were than annotated by Prokka[ 24 ]. Roary[ 25 ] was used for the pan-genome analysis. Core genes were defined as genes shared by all strains, soft core genes shared by at least 95% strains, shell genes shared by 15%-95% strains and cloud genes shared by less than15% strains. The origin of gene sequences in pan-genome were identified by comparing with nr database using blastp[ 26 ]. We used the species with highest bitscore and longest alignment length, identity > 40% and e-value < 10 − 6 as the final origin for each gene. Phylogenetic Analyses We used the Streptococcus salivarius CP013216 as an outgroup strain in the phylogenetic analysis. We aligned the outgroup strain genome, the 32 S. thermophilus genomes in the NCBI database, as well as the 185 assembled S. thermophilus sequences to the reference genome CNRZ1066 using the algorithm MumMer[27]. Neighbor-Joining tree was first generated using MEGA7[28] with default parameters. Then, we used ClonalFrameML[29] with NJ tree and alignment sequences to reconstruct the tree to remove influence of recombination. Among the genes that were prevalent in clades A-C (frequency > 0.5 in at least one of clades A-C) but less prevalent in clade D (frequency < 0.5 in clade D), we used Fisher’s exact test to identify genes significantly depleted genes in clade D. The genes with Benjamini–Hochberg adjusted p-value 1.5 were selected. This gave use 158 genes. Proteolysis and antibiotic resistance genes We compared annotated genes to the reference sequence of proteolysis genes from NCBI database using blastp[26]. We kept the alignments with e values less than 10 -5 and bit scores larger than 30. Antibiotic resistance genes were identified by comparing annotated genes with the sequences in the Comprehensive Antibiotic Resistance database[10]. Calculation of growth score and GWAS analysis We first normalized the read depth by considering local GC content and the mappability of short reads by BIC-seq2[30]. The adjusted read depth was calculated in 1000 bp bins as the ratio between the observed read count in the bin and the expected read count given by BIC-seq2. The replication origin of the reference CNRZ1066 was obtained from the DoirC database[31]. For each strain, we calculated the Spearman correlation between the bin’s adjusted read depth and its distance to the replication origin. For the GWAS analysis, we first performed a principle component analysis (PCA) based on SNPs and Indels with allele frequencies within (0.05, 0.95) and genes whose occurrence frequencies were in (0.05, 0.95). We then performed a linear regression using the growth score as the response variable and the first two PCA components as the covariates and calculated the residuals of the linear regression for each strain. This step was to remove potential confounding factors (such as hidden population structure) that might influence the growth score. Finally, we performed Wilcoxon’s rank test to identify nonsynonymous SNPs and Indels that were significantly correlated with the growth score residuals. For the stability selection, we first filtered the SNPs by controlling the false discovery rate less than 0.05. This gave us 690 SNPs. We then performed stability selection[32] using the lasso regression. Fermentation experiment In the preliminary acidification experiments, the S. thermophilus were inoculated into reconstituted skimmed milk. After 12h fermentation, titratable acidity was measured. Strains with curd time less than 12h and titratable acidity above 55 °T were defined as high acid production capability (H-Acid). The rest strains were defined as non-H-Acid group. Thus, we distinguish the 185 S. thermophilus strains into two groups preliminary. To test the acidification capability of S. thermophiles, strains from frozen stock were reactivated at 37℃ in M17 Broth (Oxoid) and subcultured twice at 24 h before use. Milk was prepared by adding 6% sucrose to 11.5 % reconstituted skimmed milk, which was then sterilized at 95 °C for 10 min and cooled to 42 °C before inoculation (about 6 log10 cfus mL -1 for each strain). Fermentation allowed to proceed at 42 °C until fermentation completed. The fermentation experiment was performed in triplicate. PH and titratable acidity (TA, °T) were measured in triplicate to evaluate fermentation progress. The pH was evaluated by pH meter (Mettler Toledo, Switzerland). Titratable acidity was measured using the method described in National Standards of the People’s Republic of China. Each sample (5.0 g) was mixed with 4.5 ml of distilled water and titrated with 0.1N NaOH in the presence of 0.5% phenolphthalein indicator to an end point of faint pink color. Minimum Inhibitory Concentration of Chloramphenicol We selected and reactivated twelve S. thermophilus isolates at 37℃ in M17 Broth (Oxoid) before use. A wide range of chloramphenicol concentration (spanning across a wide concentration range from 0.125 µg/mL to 64 µg/mL achieved by ten-fold dilution) were prepared before use. The minimum inhibitory concentration (MIC) was determined according to ISO Standard 10932:2010. Briefly, bacterial suspensions were diluted by 1000-fold (~3×10 5 cfu/mL) and tested against each chloramphenicol concentration. The MIC was recorded after incubating the bacterial cells for 48 h at 37 °C in strictly anaerobic conditions. Chloramphenicol resistance gene expression checked by droplet digital PCR The chloramphenicol resistance gene was quantified using QX100 droplet digital PCR (ddPCR, Bio-Rad), with the gene specific primer (Strep-F: 5’-AATGTTTAGCAATGACGGAAGCC-3’, Strep-R: 5’-TTCACCAATGTAAATCCCACCAC-3’). Quantification was performed using ddPCR as follows: initially, a final volume of 20 μL reaction mixture containing 2 μL cDNA, 10uL ddPCR Supermix for EvaGreen (Bio-Rad), 0.2 μL forward primer (20mM), 0.2 μL reverse primer (20mM) and 7.6 μL ddH2O were per-mixed; Each 20 μL reaction with 70 μL of droplet generation oil (Bio-Rad) was used to generate droplets; Droplets were generated by a droplet generator (Bio-Rad). The generated droplets with foil seal were then placed on a conventional PCR Thermocycler. After PCR, the PCR plate was loaded on the droplet reader (Bio-Rad), which automatically reads the droplets from each well of the plate. Analysis of the ddPCR data was performed with QuantaSoft analysis software (Bio-Rad) that accompanied the droplet reader. Availability of data and materials The data for the 185 isolated S. thermophilus has been deposited in the NCBI database under the BioProject ID: PRJNA594100. Supplementary tables see https://github.com/XiDsLab/ST185. Declarations Competing interests All authors declare that there were no conflicts of interest during completion of the current research. Author Contributions Z.S., H.Z., T.S. and R.X. conceived and designed the experiments. W.L., Y.W. and H.Z. performed the experiments. L.W., J.Z. and W.L. analyzed the data. R.X. supervised all data analysis. R.X., WL., L.W. and J.Z. drafted the manuscript. All authors read and approved the final manuscript. Acknowledgments We thank Professor Narisu and Yujun Cui for their suggestions. This research was supported by the National Natural Science Foundation of China (Grant no. 31622043, 3140066, 31771954, 11971039 and 71532001), China Agriculture Research System (Grant No. CARS-36). and National Key Basic Research Project of China (2016YFC0207705). References Ravyts, F., De Vuyst, L., and Leroy, F. (2012). Bacterial diversity and functionalities in food fermentations. Eng Life Sci 12, 356–367. Douillard, F.P., and de Vos, W.M. (2014). Functional genomics of lactic acid bacteria: from food to health. Microb Cell Fact 13 Suppl 1, S8. Song, Y., Sun, Z., Guo, C., Wu, Y., Liu, W., Yu, J., Menghe, B., Yang, R., and Zhang, H. (2016). Genetic diversity and population structure of Lactobacillus delbrueckii subspecies bulgaricus isolated from naturally fermented dairy foods. Sci Rep 6, 22704. Rasmussen, T.B., Danielsen, M., Valina, O., Garrigues, C., Johansen, E., and Pedersen, M.B. (2008). Streptococcus thermophilus core genome: comparative genome hybridization study of 47 strains. Appl Environ Microbiol 74, 4703–4710. Sayers, E.W., Agarwala, R., Bolton, E.E., Brister, J.R., Canese, K., Clark, K., Connor, R., Fiorini, N., Funk, K., Hefferon, T., et al. (2019). Database resources of the National Center for Biotechnology Information. Nucleic Acids Res 47, D23-D28. Macori, G., and Cotter, P.D. (2018). Novel insights into the microbiology of fermented dairy foods. Curr Opin Biotech 49, 172–178. Hols, P., Hancy, F., Fontaine, L., Grossiord, B., Prozzi, D., Leblond-Bourget, N., Decaris, B., Bolotin, A., Delorme, C., Ehrlich, S.D., et al. (2005). New insights in the molecular biology and physiology of Streptococcus thermophilus revealed by comparative genomics. Fems Microbiol Rev 29, 435–463. Fontaine, L., Boutry, C., Guedon, E., Guillot, A., Ibrahim, M., Grossiord, B., and Hols, P. (2007). Quorum-sensing regulation of the production of blp bacteriocins in Streptococcus thermophilus. J Bacteriol 189, 7195–7205. Bolotin, A., Quinquis, B., Renault, P., Sorokin, A., Ehrlich, S.D., Kulakauskas, S., Lapidus, A., Goltsman, E., Mazur, M., Pusch, G.D., et al. (2004). Complete sequence and comparative genome analysis of the dairy bacterium Streptococcus thermophilus. Nat Biotechnol 22, 1554–1558. Jia, B., Raphenya, A.R., Alcock, B., Waglechner, N., Guo, P., Tsang, K.K., Lago, B.A., Dave, B.M., Pereira, S., and Sharma, A.N. (2016). CARD 2017: expansion and model-centric curation of the comprehensive antibiotic resistance database. Nucleic Acids Res, gkw1004. Handwerger, S., Pucci, M.J., Volk, K.J., Liu, J.P., and Lee, M.S. (1992). The Cytoplasmic Peptidoglycan Precursor of Vancomycin-Resistant Enterococcus-Faecalis Terminates in Lactate. Journal of Bacteriology 174, 5982–5984. Depardieu, F., Bonora, M.G., Reynolds, P.E., and Courvalin, P. (2003). The vanG glycopeptide resistance operon from Enterococcus faecalis revisited. Mol Microbiol 50, 931–948. Zhang, Q.Q., Ying, G.G., Pan, C.G., Liu, Y.S., and Zhao, J.L. (2015). Comprehensive Evaluation of Antibiotics Emission and Fate in the River Basins of China: Source Analysis, Multimedia Modeling, and Linkage to Bacterial Resistance. Environ Sci Technol 49, 6772–6782. Schwarz, S., Kehrenberg, C., Doublet, B., and Cloeckaert, A. (2004). Molecular basis of bacterial resistance to chloramphenicol and florfenicol. Fems Microbiol Rev 28, 519–542. Kolling, R., Gielow, A., Seufert, W., Kucherer, C., and Messer, W. (1988). Asnc, a Multifunctional Regulator of Genes Located around the Replication Origin of Escherichia-Coli, Oric. Mol Gen Genet 212, 99–104. Arioli, S., Monnet, C., Guglielmetti, S., Parini, C., De Noni, I., Hogenboom, J., Halami, P.M., and Mora, D. (2007). Aspartate biosynthesis is essential for the growth of Streptococcus thermophilus in milk, and aspartate availability modulates the level of urease activity. Appl Environ Microb 73, 5789–5796. Bolotin, A., Quinquis, B., Renault, P., Sorokin, A., Ehrlich, S.D., Kulakauskas, S., Lapidus, A., Goltsman, E., Mazur, M., Pusch, G.D., et al. (2004). Complete sequence and comparative genome analysis of the dairy bacterium Streptococcus thermophilus. Nature Biotechnology 22, 1554–1558. Bentkowski, P., Van Oosterhout, C., and Mock, T. (2015). A Model of Genome Size Evolution for Prokaryotes in Stable and Fluctuating Environments. Genome Biol Evol 7, 2344–2351. Li, H., and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754–1760. McKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., Garimella, K., Altshuler, D., Gabriel, S., Daly, M., et al. (2010). The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 20, 1297–1303. Cingolani, P., Platts, A., Wang, L.L., Coon, M., Nguyen, T., Wang, L., Land, S.J., Lu, X., and Ruden, D.M. (2012). A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly 6, 80–92. Wu, L., Wang, H., Xia, Y., and Xi, R. (2020). CNV-BAC: Copy Number Variation Detection in Bacterial Circular Genome. Bioinformatics 36, 3890–3891. Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18. Seemann, T. (2014). Prokka: rapid prokaryotic genome annotation. Bioinformatics 30, 2068–2069. Page, A.J., Cummins, C.A., Hunt, M., Wong, V.K., Reuter, S., Holden, M.T., Fookes, M., Falush, D., Keane, J.A., and Parkhill, J. (2015). Roary: rapid large-scale prokaryote pan genome analysis. Bioinformatics 31, 3691–3693. Camacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., and Madden, T.L. (2009). BLAST+: architecture and applications. BMC Bioinformatics 10, 421. Delcher, A.L., Salzberg, S.L., and Phillippy, A.M. (2003). Using MUMmer to identify similar regions in large sequence sets. Current protocols in bioinformatics, 10.13. 11-10.13. 18. Kumar, S., Stecher, G., and Tamura, K. (2016). MEGA7: molecular evolutionary genetics analysis version 7.0 for bigger datasets. Mol Biol Evol 33, 1870–1874. Didelot, X., and Wilson, D.J. (2015). ClonalFrameML: efficient inference of recombination in whole bacterial genomes. PLoS computational biology 11, e1004041. Xi, R.B., Lee, S., Xia, Y.C., Kim, T.M., and Park, P.J. (2016). Copy number analysis of whole-genome data using BIC-seq2 and its application to detection of cancer susceptibility variants. Nucleic Acids Res 44, 6274–6286. Luo, H., and Gao, F. (2019). DoriC 10.0: an updated database of replication origins in prokaryotic genomes including chromosomes and plasmids. Nucleic Acids Res 47, D74-D77. Meinshausen, N., and Buhlmann, P. (2010). Stability selection. J R Stat Soc B 72, 417–473. Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryFigures.docx Supplemngt Figure-all FigureS4.pdf Growth score and acidification of S. thermophilus. FigureS1.pdf Pathway enrichment results and origin of genes in pan-genome. FigureS3.pdf Genome decay of clade D and ddPCR result for CatB8. FigureS2.pdf Genetic difference between 185 isolates. DatasetS1.xls dataset S1. Table S1. The sampling site, fermented dairy product, sequencing depth and preliminary fermentation experiment results for 185 isolates. Table S2. The location and annotation for SNPs and Indels. Table S3. The location and annotation for CNVs. Table S4. The sampling site and fermented dairy product for 32 S. thermophilus from NCBI database. Table S7. The natural fermented dairy products for 185 isolates in four clades. DatasetS2.xls dataset S2 Table S5. Genes for pan-genome analysis. Table S6. The origin of genes in pan-genome. DatasetS3.xls dataset S3: Table S8. Sequences of PrtS gene for 39 S. thermophilus. DatasetS4.xls dataset S4: Table S9. The annotation of 131 genes which are frequently lost in clade D. The sampling site and fermented dairy product for 32 S. thermophilus from NCBI database. Table S10. Pathway enrichment of 131 genes which are frequently lost in clade D. Table S11. Reference sequences for Blp proteins. Table S12. The comparison results of 131 absent genes in S. thermophilus with other 10 Streprococcus species. DatasetS5.xls dataset S5: Table S14. Sequences of CatB8 gene for 13 S. thermophilus. Table S15. Nearby transposon sequences around CatB8 in 13 samples. Table S16. Chloramphenicol resistance and expression of chloramphenicol resistance gene by ddPCR. DatasetS6.xls dataset S6: Table S17. The growth score for 185 S. thermophilus. Table S18. The annotation and p value of SNPs which significantly related with growth score in GWAS. Table S19. The fermentation experiment results for 85 S. thermophilus. Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-61428","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":1989087,"identity":"f7613268-83ae-4618-bbe6-16beaf4d9059","order_by":0,"name":"Wenjun Liu","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wenjun","middleName":"","lastName":"Liu","suffix":""},{"id":1989088,"identity":"701ebee5-693a-4cd3-aa7e-9b6bd12b975a","order_by":1,"name":"Linjie Wu","email":"","orcid":"","institution":"School of Mathematical Sciences and Center for Statistical Science, Peking University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Linjie","middleName":"","lastName":"Wu","suffix":""},{"id":1989089,"identity":"22a45a29-6d26-4703-9985-23810b88d5ab","order_by":2,"name":"Jie Zhao","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jie","middleName":"","lastName":"Zhao","suffix":""},{"id":1989090,"identity":"58d8466e-cbf3-4677-8760-f47314c05614","order_by":3,"name":"Weicheng Li","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Weicheng","middleName":"","lastName":"Li","suffix":""},{"id":1989091,"identity":"e5ed6d60-c9b3-4e67-97eb-160deb74fb82","order_by":4,"name":"Yu Wang","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yu","middleName":"","lastName":"Wang","suffix":""},{"id":1989092,"identity":"74bb3d4e-6d5e-4431-a1ca-b271a06898dd","order_by":5,"name":"Huijuan Zheng","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Huijuan","middleName":"","lastName":"Zheng","suffix":""},{"id":1989093,"identity":"b5fa0d29-f1a4-488b-a785-8848eb3d45c6","order_by":6,"name":"Tiansong Sun","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tiansong","middleName":"","lastName":"Sun","suffix":""},{"id":1989094,"identity":"b830f3a0-3ee2-4551-9983-c5d21a5a45c8","order_by":7,"name":"Heping Zhang","email":"","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Heping","middleName":"","lastName":"Zhang","suffix":""},{"id":1989095,"identity":"dbe566f0-7e78-4c7a-b556-57a5da5586f8","order_by":8,"name":"Ruibin Xi","email":"","orcid":"","institution":"School of Mathematical Sciences and Center for Statistical Science, Peking University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ruibin","middleName":"","lastName":"Xi","suffix":""},{"id":1989096,"identity":"95ac9ea2-e721-4946-988d-9182d537d837","order_by":9,"name":"Zhihong sun","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIie3PPQrCMBjG8bcE2iXYtVBor/CWgB9QPEul0ClKjyAIuniAjB6hR6gGnQTXDB0KgpOD4uLg4Bc4tnUTzH95l+cHCYBO96MRwPxx7HN5wrD/DQEWiDSJGxJ4kbZLTytjXDf3xVBe0rTwOvYiZiHmBCy5zqoIqlHEBB5YT5TLPceiBTRJVCVxOMYU5SBTywnjeCDg0HYl8QVH+SbSdLsojXEdAcWDyYvspqYLTQhuj4w8CENFSTDHJDbr/uLPOLvQm/RwtzXK6y3s25bcVD/skxO9r9ls/szOm291Op3uv7oDzkBND6Kx/OYAAAAASUVORK5CYII=","orcid":"","institution":"Inner Mongolia Agricultural University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Zhihong","middleName":"","lastName":"sun","suffix":""}],"badges":[],"createdAt":"2020-08-18 07:00:37","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-61428/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-61428/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":2253219,"identity":"6fd695c8-dd95-41be-b4c8-f82971755c66","added_by":"auto","created_at":"2020-09-04 17:20:11","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":93151,"visible":true,"origin":"","legend":"Descriptive statistics of 185 S. thermophilus. (A) Geographical distribution of 185 samples. (B) Boxplots of mutation rates for different genomic elements (SNPs and Indels). Conservative regions such as rRNA and tRNA showed smaller mutation rates. (C-D) The distribution of genome sizes and number of genes of the S. thermophilus strains (E) The number of genes for different gene type in pan-genome.\nNote: The designations employed and the presentation of the material on this map do not imply the expression of any opinion whatsoever on the part of Research Square concerning the legal status of any country, territory, city or area or of its authorities, or concerning the delimitation of its frontiers or boundaries. This map has been provided by the authors.\n","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/Fig1.png"},{"id":2253220,"identity":"165e4c0a-123b-4baf-af1a-96834f315298","added_by":"auto","created_at":"2020-09-04 17:20:12","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":374411,"visible":true,"origin":"","legend":"The phylogenetic tree as well as various genotypes and phenotypes of the S. thermophilus. The sampling locations, the dairy products from which the S. thermophilus was isolated and the pollution level of chloramphenicol are also shown. ","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/Fig2.png"},{"id":2253221,"identity":"9301a7d5-564b-48b5-aafd-90dd0216df4a","added_by":"auto","created_at":"2020-09-04 17:20:12","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":65487,"visible":true,"origin":"","legend":"Genome decay. (A-D) The boxplots of the number of genes, genome sizes, lengths of deletion regions and number of genes in quorum sensing for isolates from different clades, where p-values were calculated by Wilcoxon’s test.","description":"","filename":"Fig3final.jpg","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/Fig3final.jpg"},{"id":2253222,"identity":"9f1222a2-dd19-4b5b-b199-f249f11872b8","added_by":"auto","created_at":"2020-09-04 17:20:12","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":127381,"visible":true,"origin":"","legend":"Growth score and acidification. (A) Density plot of growth score (B) The smoothed adjusted read depth for all isolates (right) and the raw adjusted read depth for three example isolates (left). (C) Boxplots of growth scores of H-Acid and Non-H-Acid isolates. (D) Boxplot of growth scores in strains with or without the PrtS gene. (E-F) Scatter plot of acidity and speed of acid production versus the growth score in 85 strains. (G) Boxplots of acidity (left) and speed of acid production (right) in strains with or without the SNP G1118698T.","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/Fig4.png"},{"id":15668887,"identity":"51a94ffc-36b1-43b6-9bcc-d90b4ee0467e","added_by":"auto","created_at":"2021-11-18 13:50:36","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1150689,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/2045c3d9-2963-4b46-a861-77fc89ba7f64.pdf"},{"id":2253224,"identity":"22e7026b-81d6-47fd-8862-fbcdb8f41ba0","added_by":"auto","created_at":"2020-09-04 17:20:12","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":5533641,"visible":true,"origin":"","legend":"Supplemngt Figure-all","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/SupplementaryFigures.docx"},{"id":2253225,"identity":"9910b4de-3467-4277-a97c-142f658534be","added_by":"auto","created_at":"2020-09-04 17:20:13","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":822904,"visible":true,"origin":"","legend":"Growth score and acidification of S. thermophilus.","description":"","filename":"FigureS4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/FigureS4.pdf"},{"id":2253226,"identity":"7f99da35-3384-43a8-b70f-6545830d1ec9","added_by":"auto","created_at":"2020-09-04 17:20:13","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":199076,"visible":true,"origin":"","legend":"Pathway enrichment results and origin of genes in pan-genome.","description":"","filename":"FigureS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/FigureS1.pdf"},{"id":2253227,"identity":"ffbe8f39-01ad-4afe-9bb9-68cb5c52a54d","added_by":"auto","created_at":"2020-09-04 17:20:13","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":2144818,"visible":true,"origin":"","legend":"Genome decay of clade D and ddPCR result for CatB8.","description":"","filename":"FigureS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/FigureS3.pdf"},{"id":2253228,"identity":"81aee046-5ab7-455d-94bf-d8abe2dac1b2","added_by":"auto","created_at":"2020-09-04 17:20:14","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":2652827,"visible":true,"origin":"","legend":"Genetic difference between 185 isolates.","description":"","filename":"FigureS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/FigureS2.pdf"},{"id":2253229,"identity":"28e31903-9c78-4679-982b-9a2d3526c8be","added_by":"auto","created_at":"2020-09-04 17:20:14","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":7431680,"visible":true,"origin":"","legend":"dataset S1. Table S1. The sampling site, fermented dairy product, sequencing depth and preliminary fermentation experiment results for 185 isolates. Table S2. The location and annotation for SNPs and Indels. Table S3. The location and annotation for CNVs. Table S4. The sampling site and fermented dairy product for 32 S. thermophilus from NCBI database. Table S7. The natural fermented dairy products for 185 isolates in four clades.","description":"","filename":"DatasetS1.xls","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/DatasetS1.xls"},{"id":2253230,"identity":"b320a9b7-feea-4fb5-9ab1-0474a3fab506","added_by":"auto","created_at":"2020-09-04 17:20:15","extension":"xls","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":11735040,"visible":true,"origin":"","legend":"dataset S2 Table S5. Genes for pan-genome analysis. \nTable S6. The origin of genes in pan-genome.\n","description":"","filename":"DatasetS2.xls","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/DatasetS2.xls"},{"id":2253231,"identity":"f56b4b68-8ee6-454d-a106-5fecc1675917","added_by":"auto","created_at":"2020-09-04 17:20:15","extension":"xls","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":67072,"visible":true,"origin":"","legend":"dataset S3: Table S8. Sequences of PrtS gene for 39 S. thermophilus.","description":"","filename":"DatasetS3.xls","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/DatasetS3.xls"},{"id":2253232,"identity":"41eec524-6408-4bcb-ad55-9447afc56548","added_by":"auto","created_at":"2020-09-04 17:20:15","extension":"xls","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":93184,"visible":true,"origin":"","legend":"dataset S4: Table S9. The annotation of 131 genes which are frequently lost in clade D. The sampling site and fermented dairy product for 32 S. thermophilus from NCBI database.\nTable S10. Pathway enrichment of 131 genes which are frequently lost in clade D.\nTable S11. Reference sequences for Blp proteins.\nTable S12. The comparison results of 131 absent genes in S. thermophilus with other 10 Streprococcus species.","description":"","filename":"DatasetS4.xls","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/DatasetS4.xls"},{"id":2253233,"identity":"a9a09d38-1832-4f0a-b80e-62b8ffb5879f","added_by":"auto","created_at":"2020-09-04 17:20:15","extension":"xls","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":37376,"visible":true,"origin":"","legend":"dataset S5: Table S14. Sequences of CatB8 gene for 13 S. thermophilus.\nTable S15. Nearby transposon sequences around CatB8 in 13 samples.\nTable S16. Chloramphenicol resistance and expression of chloramphenicol resistance gene by ddPCR.\n","description":"","filename":"DatasetS5.xls","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/DatasetS5.xls"},{"id":2253234,"identity":"e0b7570a-9a75-46d6-81b7-502504d36afc","added_by":"auto","created_at":"2020-09-04 17:20:15","extension":"xls","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":46592,"visible":true,"origin":"","legend":"dataset S6: Table S17. The growth score for 185 S. thermophilus.\nTable S18. The annotation and p value of SNPs which significantly related with growth score in GWAS.\nTable S19. The fermentation experiment results for 85 S. thermophilus.\n","description":"","filename":"DatasetS6.xls","url":"https://assets-eu.researchsquare.com/files/rs-61428/v1/DatasetS6.xls"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"\u003cp\u003eSequencing of 185 \u003cem\u003eStreptococcus thermophilus\u003c/em\u003e and identification of fermentation biomarkers\u003c/p\u003e","fulltext":[{"header":"Introduction","content":" \u003cp\u003e \u003cem\u003eStreptococcus (S.) thermophilus\u003c/em\u003e is a predominant lactic acid bacterium (LAB) with rapid acidification capability and is a major dairy starter used in the milk fermentation and cheese production[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The starter strains of fermentation have vast economic value and produce a large amount fermented dairy products every year[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Traditional nomads have been making naturally fermented milk for thousands of years and these naturally fermented dairy products contain rich LAB resources[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eHowever, previous researches mostly concentrate on genomics of industry strains of \u003cem\u003eS. thermophilus.\u003c/em\u003e For example, one study showed that most of 47 industry \u003cem\u003eS. thermophilus\u003c/em\u003e have similar genetic distance, indicated genome stability of industry strains[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The genetic architecture underlying natural \u003cem\u003eS. thermophilus\u003c/em\u003e strains is largely unexplored.\u003c/p\u003e \u003cp\u003eIn this study, we isolated and sequenced 185\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e strains from natural dairy products in traditional pasture of different regions in China and Mongolia, to uncover the genetic evolution of wild strains and to explore the relationship between genotype and fermentation characteristics. Phylogenetic tree of \u003cem\u003eS. thermophilus\u003c/em\u003e revealed clades with high acid production potential or significant genome decay. We also found 13\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e which have a chloramphenicol-resistant gene \u003cem\u003eCatB8\u003c/em\u003e isolated from areas with high chloramphenicol emissions. Besides, we defined a growth score and showed that the growth score could accurately predict acidification capability of \u003cem\u003eS. thermophilus.\u003c/em\u003e With the growth score, we identified a missense mutation G1118698T by genome wide association study (GWAS) located at the gene \u003cem\u003eAcn\u003c/em\u003eA that were significantly associated with the acidification. Our findings provided novel insights in \u003cem\u003eS. thermophilus\u003c/em\u003e genetic traits and identified robust biomarkers screening of culture starter with high acidification capability.\u003c/p\u003e "},{"header":"Results","content":" \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePhenotypic analysis and phylogenetic analysis of S. thermophilus\u003c/h2\u003e \u003cp\u003eFrom natural fermented dairy products in China and Mongolia, we isolated 185\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e and found 61 strains having potential high acid production capability (H-Acid) using a preliminary fermentation experiment. Among the 185\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e, 65 were from China and 120 were from Mongolia (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). Whole genome sequencing (WGS) was performed for all 185 strains (mean coverage 374X, \u003cb\u003eTable S1\u003c/b\u003e). After short read alignment to the reference genome CNRZ1066 of \u003cem\u003eS. thermophilus\u003c/em\u003e, we identified 58,734 single nucleic acid polymorphisms (SNPs), 2,246 insertions and deletions (Indels, \u003cb\u003eTable S2\u003c/b\u003e) and 450 copy number variations (CNVs, \u003cb\u003eTable S3\u003c/b\u003e). As expected, ribosomal RNA (rRNA) and transfer RNA (tRNA) had the lowest mutation rates, followed by protein coding regions, pseudogenes and intergenic regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB). \u003cem\u003eDe novo\u003c/em\u003e assembly showed that the genome sizes of \u003cem\u003eS. thermophilus\u003c/em\u003e ranged from 1.72 to 2.60\u0026nbsp;Mb and the numbers of genes ranged from 1,704 to 2,158 (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC-D).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePan-genome analysis of our 185 isolates as well as 32\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e available from the NCBI[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] totally identified 7,629 genes including 827 core genes and 315 soft core genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE, \u003cb\u003eTable S4-S6\u003c/b\u003e). The core and soft-core genes were mainly enriched in metabolism pathways like metabolic and biosynthesis of amino acids, which are essential for bacterial growth (\u003cb\u003eFig. S1A\u003c/b\u003e). Glycolysis/gluconeogenesis pathway and starch and sucrose metabolism pathways were enriched in shell genes (shared by 15%-95% samples). These two pathways are involved in carbohydrate utilization. Diversity in these pathways indicates that \u003cem\u003eS. thermophilus\u003c/em\u003e strains might have undergone adaptive evolution in carbohydrate utilization pathways. Besides, the cloud genes (shared by less than 15% samples) were mainly enriched in quorum sensing and beta-lactam resistance pathways. 53% of cloud genes were homologous to genes in other \u003cem\u003eStreptococcus\u003c/em\u003e strains and 8% to genes in Lactobacillales and \u003cem\u003eLactococcus\u003c/em\u003e, two bacterial species often found in natural fermented dairy products[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] (\u003cb\u003eFig. S1B\u003c/b\u003e). This indicated that a portion of cloud genes came from horizontal gene transfer (HGT).\u003c/p\u003e \u003cp\u003eThe phylogenetic tree of the 217\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e revealed four large clades and was consistent with multilocus sequence typing (MLST) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) The within-clade average nucleotide identity (ANI) was significantly larger than the between-clade ANI and the principle component analysis also showed that the four clades are well-separated (\u003cb\u003eFig. S2A-B\u003c/b\u003e). Clade A, including strains from China, Mongolia and NCBI database, was closer to the roots of phylogenetic trees. In the other three clades, strains isolated from China and Mongolia also showed clear aggregation. Clade D was basically composed of strains isolated from yoghurts of Mongolia, the strains in Clade C were mainly from yoghurts of Mongolia and Xiniiang China. The clade B was mostly isolated from goat yoghurts in China and yoghurts in Mongolia. Genetic distance between isolates were significantly correlated with the geographical location of sampling sites (\u003cb\u003eFig. S2C\u003c/b\u003e). The dairy product type was also significantly correlated with the phylogenetic clades (Fisher\u0026rsquo;s exact test, p-value\u0026thinsp;=\u0026thinsp;5\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e, \u003cb\u003eTable S7\u003c/b\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eInterestingly, the H-Acid isolates and NCBI stains were significantly enriched in clade A and B (Fisher\u0026rsquo;s exact test, p-value\u0026thinsp;=\u0026thinsp;5.1\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Note that strains from NCBI were extensively used in industrial production and generally had high dairy fermentation capability. Furthermore, the cell-wall protease gene \u003cem\u003ePrt\u003c/em\u003eS, a gene known to be associated with rapid growth and acidification rates at bacteria in milk[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], was significantly enriched in clade A (Fisher\u0026rsquo;s exact test, p-value\u0026thinsp;=\u0026thinsp;1.3\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;28\u003c/sup\u003e, \u003cb\u003eTable S8\u003c/b\u003e). These data implied that strains in clade A and B might have better fermentation potential than the strains in clade C and D.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eGenome decay and antibiotic resistance of S. thermophilus\u003c/h2\u003e \u003cp\u003e \u003cem\u003eS. thermophilus\u003c/em\u003e in clade D had significantly fewer number of genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA), smaller genome sizes and more copy number losses than other clades, indicating that clade D might have undergone considerable genome decay (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB-C \u003cb\u003eand Fig. S3A\u003c/b\u003e). We found 131 genes that were prevalent in clade A-C (frequency\u0026thinsp;\u0026gt;\u0026thinsp;0.5) but were significantly depleted in clade D (Fisher\u0026rsquo;s test, Benjamini\u0026ndash;Hochberg adjusted p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05, \u003cb\u003eTable S9\u003c/b\u003e). These genes were significantly enriched in the pathways including quorum sensing, beta-lactam resistance, and ABC transporters (\u003cb\u003eTable S10\u003c/b\u003e). Further, clade D had significantly less quorum sensing genes compared with clade A and B (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD\u003cb\u003e)\u003c/b\u003e. Many of the depleted quorum sensing genes were \u003cem\u003eBlp\u003c/em\u003e bacteriocins related genes. By comparing with the Blp protein family in \u003cem\u003eStreptococcus pneumoniae\u003c/em\u003e (\u003cb\u003eTable S11\u003c/b\u003e), we found that \u003cem\u003eblpB\u003c/em\u003e, \u003cem\u003eblpM\u003c/em\u003e, \u003cem\u003eblpH\u003c/em\u003e and \u003cem\u003eblpR\u003c/em\u003e genes were significantly lost in strains from clade D comparing with other clades (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, Fisher\u0026rsquo;s test p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.001). BlpB protein, a transport accessory protein, is essential for the secretion of antimicrobial compounds. BlpM is a bacteriocin-encoding gene which directly influence the production of bacteriocins. Both BlpH and BlpR are members of two-component regulatory system, which allow bacteria to sense and respond to changes in different environment conditions. The knockout of \u003cem\u003eblpB\u003c/em\u003e, \u003cem\u003eblpH\u003c/em\u003e and \u003cem\u003eblpR\u003c/em\u003e genes in \u003cem\u003eS. thermophilus\u003c/em\u003e reduced production of bacteriocins compared with the wild type[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. These suggest that strains in clade D may have lower production capacity of bacteriocins than strains in other clades. We found that 104 of the 131 depleted genes in D had homologs in 10 other streptococci species and these genes were also enriched in quorum sensing pathway, especially the bacteriocins cluster, followed by ABC transporters (\u003cb\u003eTable S12\u003c/b\u003e). This phenomenon was consistent with previous research, which found a striking level of genome decay in \u003cem\u003eS. thermophilus\u003c/em\u003e compared with other streptococci[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Besides, 76 in 131 genes were identified in \u003cem\u003eStreptococcus salivarius\u003c/em\u003e subspecies \u003cem\u003esalivarius\u003c/em\u003e, implying that instead of being acquired by strains in clade A-C, the depleted genes in clade D were probably lost in strains from clade D in their adaptation to the milk niche.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eDue to the widespread misuse of antibiotics, the problems caused by bacterial resistance have received wide attention. By comparing pan-genome with Comprehensive Antibiotic Resistance Database[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], we found 77 genes related with resistance of 31 different antibiotics (\u003cb\u003eTable S13\u003c/b\u003e). Most antibiotic-related genes were associated with efflux (25 of 77, 32.47%) and target alteration (35 of 77, 45.45%). 16 strains in clade C had a gene homologous to the glycopeptides antibiotics resistance protein ARO3002945 (\u003cem\u003eVan\u003c/em\u003eH) [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. These 16 strains with \u003cem\u003eVan\u003c/em\u003eH were mostly isolated from Xinjiang Autonomous Region in China. Three antibiotics resistance-related genes, including two antibiotic efflux genes (ARO3000614 and ARO3004054) and one antibiotic target alteration gene ARO3004253 (\u003cem\u003eVan\u003c/em\u003eU) [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], are more likely being lost in clade D instead of being acquired by A-C. Besides, we also found that 13 strains in clade A had a gene homologous to the chloramphenicol resistance protein ARO3002680 (\u003cem\u003eCatB\u003c/em\u003e8, Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, \u003cb\u003eTable S14\u003c/b\u003e). Most strains with \u003cem\u003eCatB8\u003c/em\u003e were isolated from Hongyuan prairie in Sichuan and Gannan prairie in Gansu, two high chloramphenicol emission provinces of China [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. We identified nearby transposon sequences around \u003cem\u003eCatB8\u003c/em\u003e (~\u0026thinsp;1\u0026nbsp;kb and ~\u0026thinsp;5\u0026nbsp;kb) for 12 out of 13 strains with the \u003cem\u003eCatB8\u003c/em\u003e gene (\u003cb\u003eTable S15\u003c/b\u003e), indicating that the \u003cem\u003eCatB8\u003c/em\u003e gene was probably acquired by lateral gene transfer. \u003cem\u003eCatB8\u003c/em\u003e, chloramphenicol acetyltransferase, inactivates chloramphenicol by acetylation[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The acquisition of \u003cem\u003eCatB8\u003c/em\u003e might confer \u003cem\u003eS. thermophilus\u003c/em\u003e resistance to chloramphenicol. In fact, we applied droplet digital PCR (ddPCR) for strains with \u003cem\u003eCatB\u003c/em\u003e8 and found a significantly higher expression of \u003cem\u003eCatB8\u003c/em\u003e in M17 cultures with 8\u0026nbsp;\u0026micro;g/ml chloramphenicol than cultures without chloramphenicol (Wilcoxon\u0026rsquo;s test, p-value\u0026thinsp;=\u0026thinsp;0.0025, \u003cb\u003eFig. S3B\u003c/b\u003e, \u003cb\u003eTable S16\u003c/b\u003e), indicating that the bacteria responded to the exposure of chloramphenicol by elevating the expression of \u003cem\u003eCatB\u003c/em\u003e8. These data suggested that the misuse of chloramphenicol might be closely related with the antibiotic resistant \u003cem\u003eS. thermophilus\u003c/em\u003e and attention should be paid in the screening of potential starter to avoid the spread of resistance genes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eGrowth score and acidification of S. thermophilus\u003c/h2\u003e \u003cp\u003eAcidification is the most important characteristic of \u003cem\u003eS. thermophilus\u003c/em\u003e as a starter in fermentation of dairy products. Rapid acidification can shorten fermentation time of yoghurt production and reduce the production costs. Screening of \u003cem\u003eS. thermophilus\u003c/em\u003e strains with rapid acidification is one of most concerned problems in fermentation dairy enterprises. However, currently biomarkers for rapid acidification is still lacking. Theoretically, acidification is closely related with the growth rate of \u003cem\u003eS. thermophilus.\u003c/em\u003e As most bacteria, \u003cem\u003eS. thermophilus\u003c/em\u003e has a single circular genome. During replication, DNA sequences passed the replication fork should have two copies and those to be replicated should have only single copy. Thus, because millions of cells at different replication stages were used in WGS of \u003cem\u003eS. thermophilus\u003c/em\u003e, a genomic region\u0026rsquo;s read depth should be negatively correlated with its distance to the replication origin and the strength of this correlation should reflect the growth rate of \u003cem\u003eS. thermophilus\u003c/em\u003e. In our WGS data, the adjusted read depths (Methods) were indeed negatively correlated with the distance to the replication origin for all isolates. We defined a growth score as the negative value of Spearman\u0026rsquo;s correlation between them (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA-B, \u003cb\u003eFig S4A, Table S17\u003c/b\u003e). Clade A and B strains had significantly larger growth score than clade C and D (\u003cb\u003eFig. S4B\u003c/b\u003e). The growth score was significantly larger in H-Acid strains (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC) and in the strains with the \u003cem\u003ePrtS\u003c/em\u003e gene (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD). These data implied that the growth score might provide an accurate marker for the acidification capability of \u003cem\u003eS. thermophilus\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eUsing the growth score, we performed Genome-Wide Association Study (GWAS) and stability selection to screen for genomic variations that might be related with the acidification capability of \u003cem\u003eS. thermophilus\u003c/em\u003e (Methods). We found that 7 SNPs were significantly correlated with growth score by GWAS (Wilcoxon test, p-value\u0026thinsp;\u0026lt;\u0026thinsp;10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e or Bonferroni adjusted p-value \u0026lt; 0.05, \u003cb\u003eTable S18\u003c/b\u003e). Among the 7 SNPs, the missense SNP A764991G located at the gene \u003cem\u003eAsnC\u003c/em\u003e, which promotes the growth of the \u003cem\u003eS. thermophilus\u003c/em\u003e in milk by regulating aspartic acid metabolism[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], had the highest selection probability and minimum p-value (1.4\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e) in GWAS analysis. The missense SNP G1118698T located at the gene \u003cem\u003eAcn\u003c/em\u003eA, which encodes aconitate hydratase A, also had high selection probability and was very significant in GWAS analysis (p-value\u0026thinsp;=\u0026thinsp;4.5\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e, \u003cb\u003eFig. S4C-D\u003c/b\u003e). \u003cem\u003eAcnA\u003c/em\u003e involves in succinate and citrate production, contribute to acidification. 27\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e with this mutation were all from clade B and D and were all in the non-H-Acid group (\u003cb\u003eFig. S4E\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eTo confirm that the proposed growth score and the SNP G1118698T were associated with the acidification capability of \u003cem\u003eS. thermophilus\u003c/em\u003e, we randomly selected 85 strains and performed fermentation experiments (\u003cb\u003eTable S19\u003c/b\u003e). We evaluated the acidification capability of \u003cem\u003eS. thermophilus\u003c/em\u003e by the acidity and the acid production speed (Methods). The acidity and acid production speed were significantly higher in strains with the \u003cem\u003ePrtS\u003c/em\u003e gene, and higher in clade A and B (\u003cb\u003eFig. S4F-G\u003c/b\u003e). The growth score was significantly positively correlated with acid production speed (Pearson\u0026rsquo;s correlation 0.4, p-value\u0026thinsp;=\u0026thinsp;1.2\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;4\u003c/sup\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE) and acidity (Pearson\u0026rsquo;s correlation 0.35, p-value\u0026thinsp;=\u0026thinsp;1.2\u0026sdot;10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eF). Similarly, the SNP G1118698T was significantly associated with acidity and the acidification speed (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eG). These data suggested that the proposed growth score and the SNP G1118698T could serve as reliable biomarkers for screening \u003cem\u003eS. thermophilus\u003c/em\u003e isolates with high acidification speed.\u003c/p\u003e \u003c/div\u003e "},{"header":"Discussion","content":" \u003cp\u003eWe reported so far the largest WGS data of \u003cem\u003eS. thermophilus\u003c/em\u003e isolated from natural fermentation dairy products. This large amount of data allowed us to systematically investigate the genomics landscape of \u003cem\u003eS. thermophilus.\u003c/em\u003e We found that \u003cem\u003eS. thermophilus\u003c/em\u003e had four large clades. Strains in clade A seemed to have high industry application potential, while strains in clade D might have undergone considerable genome decay through gene loss. We also identified novel biomarkers for the acidification capability of \u003cem\u003eS. thermophilus.\u003c/em\u003e The data and novel discoveries in this paper provided valuable resources for understanding the evolution and genomics of \u003cem\u003eS. thermophilus\u003c/em\u003e and for the industry application of \u003cem\u003eS. thermophilus.\u003c/em\u003e\u003c/p\u003e \u003cp\u003eCompared with many other streptococci species, \u003cem\u003eS. thermophilus\u003c/em\u003e lives in a rather stable environment. Previous researches[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] discovered that many virulence-related genes were lost or became pseudogenes in \u003cem\u003eS. thermophilus.\u003c/em\u003e Here, we revealed that in clade D strains of \u003cem\u003eS. thermophilus\u003c/em\u003e had significantly smaller genomes than strains in other clades, indicating that the genome decay might be an ongoing process of \u003cem\u003eS. thermophilus.\u003c/em\u003e This large genome decay was possibly due to the adaptation of \u003cem\u003eS. thermophilus\u003c/em\u003e to their stable niche of milk. In fact, mathematical simulation[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] showed that stable environments often lead to smaller genomes than environments with greater variability. Many quorum-sensing genes, especially \u003cem\u003eblp\u003c/em\u003e genes, were lost in clade D strains. By analyzing genomes of three \u003cem\u003eS. thermophilus\u003c/em\u003e strains (LMG18311, CNRZ1066, and LMD-9), previous research[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] found that, although identified in all three strains, the \u003cem\u003eblp\u003c/em\u003e gene clusters were only full functional in LMD-9. It was thus plausible that these \u003cem\u003eblp\u003c/em\u003e genes conferred little or no survival advantages to \u003cem\u003eS. thermophilus\u003c/em\u003e, and then the genes in this pathway gradually became inactive and eventually lost clade D \u003cem\u003eS. thermophilus\u003c/em\u003e strains.\u003c/p\u003e \u003cp\u003eThe traditional method for evaluating the acidification capability of \u003cem\u003eS. thermophilus\u003c/em\u003e strains is very labor-intensive, time-consuming and costly. With the advancement of sequencing technologies, WGS becomes very efficient and cost-effective. The growth score defined in this paper can be calculated only using WGS data and thus provided a very convenient and cost-effective surrogate to the traditional evaluation method. With this score, one can easily screen hundreds of \u003cem\u003eS. thermophilus\u003c/em\u003e strains. In addition, since fast growth is often a desirable property for industrial bacteria, the growth score defined here might also serve as a robust criterion for evaluating other industrial bacteria and thus has a large application potential in industry. One disadvantage of this growth score is that it currently can only be used for monoculture bacteria. In industrial applications, multiple species of bacteria are often simultaneously used. WGS of the mixture of the different bacteria would cause read mapping ambiguity to different reference genomes and thus the growth score cannot be directly applied. However, we could only consider the genomic sequences that are unique to each species and generalize the growth score using these unique sequences for the mixed sequencing data.\u003c/p\u003e \u003cp\u003e "},{"header":"Materials And Methods","content":" \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eVariant calling, assemble and annotation\u003c/h2\u003e \u003cp\u003eGenomic DNA was sequenced using an Illumina HiSeq 4000 platform (Illumina, San Diego, CA) generating 150-bp paired-end reads with an average insert size of 350 bps. All 185\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e sequencing data were mapped to reference genome CNRZ1066 by BWA-mem[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] with default parameters. SNPs and Indels were called by GATK Unifiedgenotyper[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] and annotated by SnpEff[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. SNPs having Indels within its 10\u0026nbsp;bp neighborhood were filtered. CNVs were called by CNV-BAC[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. We performed \u003cem\u003ede novo\u003c/em\u003e assembly using SOAPdenovo2[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] (k-mer\u0026thinsp;=\u0026thinsp;71). The contigs were than annotated by Prokka[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Roary[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] was used for the pan-genome analysis. Core genes were defined as genes shared by all strains, soft core genes shared by at least 95% strains, shell genes shared by 15%-95% strains and cloud genes shared by less than15% strains. The origin of gene sequences in pan-genome were identified by comparing with nr database using blastp[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. We used the species with highest bitscore and longest alignment length, identity\u0026thinsp;\u0026gt;\u0026thinsp;40% and e-value\u0026thinsp;\u0026lt;\u0026thinsp;10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e as the final origin for each gene.\u003c/p\u003e \u003c/div\u003e \n\u003cp\u003e\u003cstrong\u003ePhylogenetic Analyses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe used the \u003cem\u003eStreptococcus salivarius\u003c/em\u003e CP013216 as an outgroup strain in the phylogenetic analysis. We aligned the outgroup strain genome, the 32 \u003cem\u003eS. thermophilus\u003c/em\u003e genomes in the NCBI database, as well as the 185 assembled \u003cem\u003eS. thermophilus\u003c/em\u003e sequences to the reference genome CNRZ1066 using the algorithm MumMer[27]. Neighbor-Joining tree was first generated using MEGA7[28] with default parameters. Then, we used ClonalFrameML[29] with NJ tree and alignment sequences to reconstruct the tree to remove influence of recombination. Among the genes that were prevalent in clades A-C (frequency \u0026gt; 0.5 in at least one of clades A-C) but less prevalent in clade D (frequency \u0026lt; 0.5 in clade D), we used Fisher\u0026rsquo;s exact test to identify genes significantly depleted genes in clade D. The genes with Benjamini\u0026ndash;Hochberg adjusted p-value \u0026lt; 0.05 and odds ratio \u0026gt; 1.5 were selected. This gave use 158 genes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProteolysis and antibiotic resistance genes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe compared annotated genes to the reference sequence of proteolysis genes from NCBI database using blastp[26]. We kept the alignments with e values less than 10\u003csup\u003e-5\u003c/sup\u003e and bit scores larger than 30. Antibiotic resistance genes were identified by comparing annotated genes with the sequences in the Comprehensive Antibiotic Resistance database[10].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCalculation of growth score and GWAS analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe first normalized the read depth by considering local GC content and the mappability of short reads by BIC-seq2[30]. The adjusted read depth was calculated in 1000 bp bins as the ratio between the observed read count in the bin and the expected read count given by BIC-seq2. The replication origin of the reference CNRZ1066 was obtained from the DoirC database[31]. For each strain, we calculated the Spearman correlation between the bin\u0026rsquo;s adjusted read depth and its distance to the replication origin. For the GWAS analysis, we first performed a principle component analysis (PCA) based on SNPs and Indels with allele frequencies within (0.05, 0.95) and genes whose occurrence frequencies were in (0.05, 0.95). We then performed a linear regression using the growth score as the response variable and the first two PCA components as the covariates and calculated the residuals of the linear regression for each strain. This step was to remove potential confounding factors (such as hidden population structure) that might influence the growth score. Finally, we performed Wilcoxon\u0026rsquo;s rank test to identify nonsynonymous SNPs and Indels that were significantly correlated with the growth score residuals. For the stability selection, we first filtered the SNPs by controlling the false discovery rate less than 0.05. This gave us 690 SNPs. We then performed stability selection[32] using the lasso regression.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFermentation experiment\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the preliminary acidification experiments, the \u003cem\u003eS. thermophilus\u003c/em\u003e were inoculated into reconstituted skimmed milk. After 12h fermentation, titratable acidity was measured. Strains with curd time less than 12h and titratable acidity above 55 \u0026deg;T were defined as high acid production capability (H-Acid). The rest strains were defined as non-H-Acid group. Thus, we distinguish the 185 \u003cem\u003eS. thermophilus\u003c/em\u003e strains into two groups preliminary. To test the acidification capability of \u003cem\u003eS. thermophiles,\u003c/em\u003e strains from frozen stock were reactivated at 37℃ in M17 Broth (Oxoid) and subcultured twice at 24 h before use. Milk was prepared by adding 6% sucrose to 11.5 % reconstituted skimmed milk, which was then sterilized at 95 \u0026deg;C for 10 min and cooled to 42 \u0026deg;C before inoculation (about 6 log10 cfus mL\u003csup\u003e-1\u003c/sup\u003e for each strain). Fermentation allowed to proceed at 42 \u0026deg;C until fermentation completed. The fermentation experiment was performed in triplicate. PH and titratable acidity (TA, \u0026deg;T) were measured in triplicate to evaluate fermentation progress. The pH was evaluated by pH meter (Mettler Toledo, Switzerland). Titratable acidity was measured using the method described in National Standards of the People\u0026rsquo;s Republic of China. Each sample (5.0 g) was mixed with 4.5 ml of distilled water and titrated with 0.1N NaOH in the presence of 0.5% phenolphthalein indicator to an end point of faint pink color.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMinimum Inhibitory Concentration of Chloramphenicol\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe selected and reactivated twelve \u003cem\u003eS. thermophilus \u003c/em\u003eisolates at 37℃ in M17 Broth (Oxoid) before use. A wide range of chloramphenicol concentration (spanning across a wide concentration range from 0.125 \u0026micro;g/mL to 64 \u0026micro;g/mL achieved by ten-fold dilution) were prepared before use. The minimum inhibitory concentration (MIC) was determined according to ISO Standard 10932:2010. Briefly, bacterial suspensions were diluted by 1000-fold (~3\u0026times;10\u003csup\u003e5\u003c/sup\u003ecfu/mL) and tested against each chloramphenicol concentration. The MIC was recorded after incubating the bacterial cells for 48 h at 37 \u0026deg;C in strictly anaerobic conditions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eChloramphenicol resistance gene expression checked by droplet digital PCR\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe chloramphenicol resistance gene was quantified using QX100 droplet digital PCR (ddPCR, Bio-Rad), with the gene specific primer (Strep-F: 5\u0026rsquo;-AATGTTTAGCAATGACGGAAGCC-3\u0026rsquo;, Strep-R: 5\u0026rsquo;-TTCACCAATGTAAATCCCACCAC-3\u0026rsquo;). Quantification was performed using ddPCR as follows: initially, a final volume of 20 \u0026mu;L reaction mixture containing 2 \u0026mu;L cDNA, 10uL ddPCR Supermix for EvaGreen (Bio-Rad), 0.2 \u0026mu;L forward primer (20mM), 0.2 \u0026mu;L reverse primer (20mM) and 7.6 \u0026mu;L ddH2O were per-mixed; Each 20 \u0026mu;L reaction with 70 \u0026mu;L of droplet generation oil (Bio-Rad) was used to generate droplets; Droplets were generated by a droplet generator (Bio-Rad). The generated droplets with foil seal were then placed on a conventional PCR Thermocycler. After PCR, the PCR plate was loaded on the droplet reader (Bio-Rad), which automatically reads the droplets from each well of the plate. Analysis of the ddPCR data was performed with QuantaSoft analysis software (Bio-Rad) that accompanied the droplet reader.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data for the 185 isolated \u003cem\u003eS. thermophilus \u003c/em\u003ehas been deposited in the NCBI database under the BioProject ID: PRJNA594100. Supplementary tables see https://github.com/XiDsLab/ST185.\u003c/p\u003e"},{"header":"Declarations","content":" \u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eAll authors declare that there were no conflicts of interest during completion of the current research.\u003c/p\u003e \u003c/p\u003e \u003ch2\u003eAuthor Contributions\u003c/h2\u003e \u003cp\u003eZ.S., H.Z., T.S. and R.X. conceived and designed the experiments. W.L., Y.W. and H.Z. performed the experiments. L.W., J.Z. and W.L. analyzed the data. R.X. supervised all data analysis. R.X., WL., L.W. and J.Z. drafted the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e \u003ch2\u003eAcknowledgments\u003c/h2\u003e \u003cp\u003eWe thank Professor Narisu and Yujun Cui for their suggestions. This research was supported by the National Natural Science Foundation of China (Grant no. 31622043, 3140066, 31771954, 11971039 and 71532001), China Agriculture Research System (Grant No. CARS-36). and National Key Basic Research Project of China (2016YFC0207705).\u003c/p\u003e "},{"header":"References","content":"\u003col\u003e\u003cli\u003e \u003cspan\u003eRavyts, F., De Vuyst, L., and Leroy, F. (2012). Bacterial diversity and functionalities in food fermentations. Eng Life Sci 12, 356\u0026ndash;367.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eDouillard, F.P., and de Vos, W.M. (2014). Functional genomics of lactic acid bacteria: from food to health. Microb Cell Fact 13 Suppl 1, S8.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eSong, Y., Sun, Z., Guo, C., Wu, Y., Liu, W., Yu, J., Menghe, B., Yang, R., and Zhang, H. (2016). Genetic diversity and population structure of Lactobacillus delbrueckii subspecies bulgaricus isolated from naturally fermented dairy foods. Sci Rep 6, 22704.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eRasmussen, T.B., Danielsen, M., Valina, O., Garrigues, C., Johansen, E., and Pedersen, M.B. (2008). Streptococcus thermophilus core genome: comparative genome hybridization study of 47 strains. Appl Environ Microbiol 74, 4703\u0026ndash;4710.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eSayers, E.W., Agarwala, R., Bolton, E.E., Brister, J.R., Canese, K., Clark, K., Connor, R., Fiorini, N., Funk, K., Hefferon, T., et al. (2019). Database resources of the National Center for Biotechnology Information. Nucleic Acids Res 47, D23-D28.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eMacori, G., and Cotter, P.D. (2018). Novel insights into the microbiology of fermented dairy foods. Curr Opin Biotech 49, 172\u0026ndash;178.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eHols, P., Hancy, F., Fontaine, L., Grossiord, B., Prozzi, D., Leblond-Bourget, N., Decaris, B., Bolotin, A., Delorme, C., Ehrlich, S.D., et al. (2005). New insights in the molecular biology and physiology of Streptococcus thermophilus revealed by comparative genomics. Fems Microbiol Rev 29, 435\u0026ndash;463.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eFontaine, L., Boutry, C., Guedon, E., Guillot, A., Ibrahim, M., Grossiord, B., and Hols, P. (2007). Quorum-sensing regulation of the production of blp bacteriocins in Streptococcus thermophilus. J Bacteriol 189, 7195\u0026ndash;7205.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eBolotin, A., Quinquis, B., Renault, P., Sorokin, A., Ehrlich, S.D., Kulakauskas, S., Lapidus, A., Goltsman, E., Mazur, M., Pusch, G.D., et al. (2004). Complete sequence and comparative genome analysis of the dairy bacterium Streptococcus thermophilus. Nat Biotechnol 22, 1554\u0026ndash;1558.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eJia, B., Raphenya, A.R., Alcock, B., Waglechner, N., Guo, P., Tsang, K.K., Lago, B.A., Dave, B.M., Pereira, S., and Sharma, A.N. (2016). CARD 2017: expansion and model-centric curation of the comprehensive antibiotic resistance database. Nucleic Acids Res, gkw1004.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eHandwerger, S., Pucci, M.J., Volk, K.J., Liu, J.P., and Lee, M.S. (1992). The Cytoplasmic Peptidoglycan Precursor of Vancomycin-Resistant Enterococcus-Faecalis Terminates in Lactate. Journal of Bacteriology 174, 5982\u0026ndash;5984.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eDepardieu, F., Bonora, M.G., Reynolds, P.E., and Courvalin, P. (2003). The vanG glycopeptide resistance operon from Enterococcus faecalis revisited. Mol Microbiol 50, 931\u0026ndash;948.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eZhang, Q.Q., Ying, G.G., Pan, C.G., Liu, Y.S., and Zhao, J.L. (2015). Comprehensive Evaluation of Antibiotics Emission and Fate in the River Basins of China: Source Analysis, Multimedia Modeling, and Linkage to Bacterial Resistance. Environ Sci Technol 49, 6772\u0026ndash;6782.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eSchwarz, S., Kehrenberg, C., Doublet, B., and Cloeckaert, A. (2004). Molecular basis of bacterial resistance to chloramphenicol and florfenicol. Fems Microbiol Rev 28, 519\u0026ndash;542.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eKolling, R., Gielow, A., Seufert, W., Kucherer, C., and Messer, W. (1988). Asnc, a Multifunctional Regulator of Genes Located around the Replication Origin of Escherichia-Coli, Oric. Mol Gen Genet 212, 99\u0026ndash;104.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eArioli, S., Monnet, C., Guglielmetti, S., Parini, C., De Noni, I., Hogenboom, J., Halami, P.M., and Mora, D. (2007). Aspartate biosynthesis is essential for the growth of Streptococcus thermophilus in milk, and aspartate availability modulates the level of urease activity. Appl Environ Microb 73, 5789\u0026ndash;5796.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eBolotin, A., Quinquis, B., Renault, P., Sorokin, A., Ehrlich, S.D., Kulakauskas, S., Lapidus, A., Goltsman, E., Mazur, M., Pusch, G.D., et al. (2004). Complete sequence and comparative genome analysis of the dairy bacterium Streptococcus thermophilus. Nature Biotechnology 22, 1554\u0026ndash;1558.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eBentkowski, P., Van Oosterhout, C., and Mock, T. (2015). A Model of Genome Size Evolution for Prokaryotes in Stable and Fluctuating Environments. Genome Biol Evol 7, 2344\u0026ndash;2351.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eLi, H., and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754\u0026ndash;1760.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eMcKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., Garimella, K., Altshuler, D., Gabriel, S., Daly, M., et al. (2010). The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 20, 1297\u0026ndash;1303.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eCingolani, P., Platts, A., Wang, L.L., Coon, M., Nguyen, T., Wang, L., Land, S.J., Lu, X., and Ruden, D.M. (2012). A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly 6, 80\u0026ndash;92.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eWu, L., Wang, H., Xia, Y., and Xi, R. (2020). CNV-BAC: Copy Number Variation Detection in Bacterial Circular Genome. Bioinformatics 36, 3890\u0026ndash;3891.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eLuo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eSeemann, T. (2014). Prokka: rapid prokaryotic genome annotation. Bioinformatics 30, 2068\u0026ndash;2069.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003ePage, A.J., Cummins, C.A., Hunt, M., Wong, V.K., Reuter, S., Holden, M.T., Fookes, M., Falush, D., Keane, J.A., and Parkhill, J. (2015). Roary: rapid large-scale prokaryote pan genome analysis. Bioinformatics 31, 3691\u0026ndash;3693.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eCamacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., and Madden, T.L. (2009). BLAST+: architecture and applications. BMC Bioinformatics 10, 421.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eDelcher, A.L., Salzberg, S.L., and Phillippy, A.M. (2003). Using MUMmer to identify similar regions in large sequence sets. Current protocols in bioinformatics, 10.13. 11-10.13. 18.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eKumar, S., Stecher, G., and Tamura, K. (2016). MEGA7: molecular evolutionary genetics analysis version 7.0 for bigger datasets. Mol Biol Evol 33, 1870\u0026ndash;1874.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eDidelot, X., and Wilson, D.J. (2015). ClonalFrameML: efficient inference of recombination in whole bacterial genomes. PLoS computational biology 11, e1004041.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eXi, R.B., Lee, S., Xia, Y.C., Kim, T.M., and Park, P.J. (2016). Copy number analysis of whole-genome data using BIC-seq2 and its application to detection of cancer susceptibility variants. Nucleic Acids Res 44, 6274\u0026ndash;6286.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eLuo, H., and Gao, F. (2019). DoriC 10.0: an updated database of replication origins in prokaryotic genomes including chromosomes and plasmids. Nucleic Acids Res 47, D74-D77.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e \u003cspan\u003eMeinshausen, N., and Buhlmann, P. (2010). Stability selection. J R Stat Soc B 72, 417\u0026ndash;473.\u003c/span\u003e \u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Streptococcus thermophilus, Whole genome sequencing, Antibiotic resistance, Growth score Acidification","lastPublishedDoi":"10.21203/rs.3.rs-61428/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-61428/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e \u003cem\u003eStreptococcus\u003c/em\u003e (\u003cem\u003eS.\u003c/em\u003e) \u003cem\u003ethermophilus\u003c/em\u003e is an important dairy starter in the production of fermented dairy products has important significance, from natural fermentation in the past to industrial production today. While the genetic architecture underlying \u003cem\u003eS. thermophilus\u003c/em\u003e traits and phenotypes is largely unknown. Here, we sequenced 185\u0026nbsp;\u003cem\u003eS. thermophilus\u003c/em\u003e strains, which isolated from natural fermented dairy products of China and Mongolia and using comparative genomic and genome wide association study to provide novel point for genetic architecture underlying its traits and phenotypes. Genome analysis of \u003cem\u003eS. thermophilus\u003c/em\u003e showed association of phylogeny with environmental and phenotypic features and revealed clades with high acid production potential or with substantial genome decay. A few \u003cem\u003eS. thermophilus\u003c/em\u003e isolated from areas with high chloramphenicol emissions had a chloramphenicol-resistant gene \u003cem\u003eCatB8\u003c/em\u003e. Most importantly, we defined a growth score and identified a missense mutation G1118698T located at the gene \u003cem\u003eAcn\u003c/em\u003eA that were both predictive of acidification capability of \u003cem\u003eS. thermophilus\u003c/em\u003e. Our findings provide novel insight in \u003cem\u003eS. thermophilus\u003c/em\u003e genetic traits, antibiotic resistant and predictive of acidification capability which both may had huge help in culture starter screening.\u003c/p\u003e","manuscriptTitle":"Sequencing of 185 Streptococcus thermophilus and identification of fermentation biomarkers","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-09-04 17:20:10","doi":"10.21203/rs.3.rs-61428/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"communications-biology","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commsbio","sideBox":"Learn more about [Communications Biology](http://www.nature.com/commsbio/)","snPcode":"","submissionUrl":"","title":"Communications Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"40d5fba7-45d6-45c6-b795-fe964ad3d72f","owner":[],"postedDate":"September 4th, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":438496,"name":"General Microbiology"},{"id":438497,"name":"Bacteriology"},{"id":438498,"name":"Evolutionary Genetics"}],"tags":[],"updatedAt":"2020-11-20T03:15:46+00:00","versionOfRecord":[],"versionCreatedAt":"2020-09-04 17:20:10","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-61428","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-61428","identity":"rs-61428","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.