Genome-wide exploration of biosynthetic gene clusters and their association to virulence in the entomopathogenic fungus Beauveria

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Species of the genus Beauveria are widely used as biological control agents due to their ability to infect and kill a broad range of arthropod pests. Despite their agricultural importance, the genomic diversity underlying virulence and secondary metabolism across the genus remains uncharacterized. In this study, we performed a comprehensive phylogenomic and biosynthetic gene cluster (BGC) analysis of 333 high-quality Beauveria genomes, including one newly sequenced Beauveria bassiana isolate (AS272) obtained from soil in southern Brazil. The genome of AS272 comprises 31.7 Mb with 9,778 predicted genes and high completeness (96.7% BUSCO). Comparative analysis revealed extensive conservation of orthogroups but indicated that the increased gene repertoire of AS272 is primarily associated with expansion of existing gene families, particularly transporters belonging to the major facilitator superfamily (MFS) and ATP-binding cassette (ABC) families. Phylogenomic reconstruction based on single-copy orthologs resolved major species clades within the genus and identified potential misidentified genomes in public databases. Genome mining using antiSMASH identified 14,602 BGCs across the dataset, with NRPS and PKS clusters dominating the biosynthetic landscape. Clustering with BiG-SCAPE revealed a mixture of highly conserved gene cluster families (GCFs) and numerous lineage-specific clusters, highlighting both evolutionary stability and diversification of secondary metabolism within Beauveria . Notably, B. bassiana lacked strictly conserved core GCFs across all isolates, instead exhibiting a heavy-tailed distribution of cluster frequencies consistent with substantial intraspecific genomic plasticity. Protein–protein interaction network analysis further indicated that several conserved BGCs are embedded within networks enriched for proteins associated with host–pathogen interactions. Together, these findings reveal a dynamic genomic architecture underlying secondary metabolism and pathogenic potential in Beauveria , providing a comparative framework for understanding virulence evolution and identifying candidate pathways for future functional and biotechnological studies.
Full text 129,904 characters · extracted from preprint-html · click to expand
Genome-wide exploration of biosynthetic gene clusters and their association to virulence in the entomopathogenic fungus Beauveria | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Genome-wide exploration of biosynthetic gene clusters and their association to virulence in the entomopathogenic fungus Beauveria Robson dos Santos Soares, Alexandra de Azevedo da Rocha, Matheus da Silva Camargo, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9203165/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract Species of the genus Beauveria are widely used as biological control agents due to their ability to infect and kill a broad range of arthropod pests. Despite their agricultural importance, the genomic diversity underlying virulence and secondary metabolism across the genus remains uncharacterized. In this study, we performed a comprehensive phylogenomic and biosynthetic gene cluster (BGC) analysis of 333 high-quality Beauveria genomes, including one newly sequenced Beauveria bassiana isolate (AS272) obtained from soil in southern Brazil. The genome of AS272 comprises 31.7 Mb with 9,778 predicted genes and high completeness (96.7% BUSCO). Comparative analysis revealed extensive conservation of orthogroups but indicated that the increased gene repertoire of AS272 is primarily associated with expansion of existing gene families, particularly transporters belonging to the major facilitator superfamily (MFS) and ATP-binding cassette (ABC) families. Phylogenomic reconstruction based on single-copy orthologs resolved major species clades within the genus and identified potential misidentified genomes in public databases. Genome mining using antiSMASH identified 14,602 BGCs across the dataset, with NRPS and PKS clusters dominating the biosynthetic landscape. Clustering with BiG-SCAPE revealed a mixture of highly conserved gene cluster families (GCFs) and numerous lineage-specific clusters, highlighting both evolutionary stability and diversification of secondary metabolism within Beauveria . Notably, B. bassiana lacked strictly conserved core GCFs across all isolates, instead exhibiting a heavy-tailed distribution of cluster frequencies consistent with substantial intraspecific genomic plasticity. Protein–protein interaction network analysis further indicated that several conserved BGCs are embedded within networks enriched for proteins associated with host–pathogen interactions. Together, these findings reveal a dynamic genomic architecture underlying secondary metabolism and pathogenic potential in Beauveria , providing a comparative framework for understanding virulence evolution and identifying candidate pathways for future functional and biotechnological studies. Beauveria Phylogenomics Biosynthetic gene clusters (BGCs) Virulence factors Secondary metabolites Comparative genomics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Beauveria is a genus of entomopathogenic fungi widely recognized for its use in the biological control of arthropod pests (de Sousa et al. 2025 ). Members of this genus infect their hosts using a common mechanism involving adhesion and germination of conidia on the insect cuticle, followed by penetration, internal colonization, and host death (Hong et al. 2024 ). Compared to conventional chemical insecticides, Beauveria species exhibit a narrower host range, lower environmental impact, and are considered safe for non-target organisms, including humans, animals, and plants (Mascarin and Jaronski 2016 ). These features support the development and commercialization of various Beauveria -based biopesticide formulations, which are applied in integrated pest management (IPM) strategies across different agricultural systems (Andreata et al. 2025 ). In Brazil, by the beginning of 2026, a total of 160 biocontrol products containing Beauveria sp. had been registered, highlighting its effectiveness in controlling insect pests (Ministério da Agricultura e Pecuária). The commercial importance of the genus Beauveria reflects its high pathogenic potential, which is supported by the diversity of molecules produced during the infection cycle, such as insecticidal toxins, immunosuppressive compounds, and antimicrobial metabolites (Oberti et al. 2025 ; Farag et al. 2025 ; Aly et al. 2025 ). Secondary metabolites are also produced during host infection (Chaithra et al. 2024 ), including oosporein, bassianolide, and beauvericin, which play pivotal roles in virulence by promoting host immune suppression, tissue colonization, and insect mortality (Wang et al. 2021 ). These compounds belong to distinct classes of secondary metabolites, including polyketides and nonribosomal peptides or depsipeptides, which are encoded by well-defined biosynthetic gene clusters (BGCs) comprising core synthase genes, such as polyketide synthases (PKSs) and nonribosomal peptide synthetases (NRPSs), along with tailoring enzymes, transporters, and regulatory elements (Keller 2019 ). The profile and production of these compounds vary among species and strains, highlighting the genomic plasticity of the genus and its ability to adapt to different hosts and environments (Solano-González et al. 2023 ). In other fungal genera, comparative genomic analyses have shown that gene annotation and the prediction of biosynthetic gene clusters (BGCs) associated with secondary metabolites are broadly distributed among species, representing a promising approach for investigating molecular mechanisms of virulence based on genomic data (Valero-Jiménez et al. 2016 ). Genomic plasticity contributes to diversification across species, yet core genes associated with pathogenicity, such as those involved in host adhesion, cuticle degradation, and stress tolerance, are commonly conserved among phylogenetically related species, highlighting their essential role in infective capacity and survival (Wang and Wang 2017 ). Despite advances in genome sequencing technologies and data availability, the lack of standardized curation and robust functional annotation in public databases continues to limit the effective use of genomic information in comparative studies. Rigorous curation of genomic datasets, including reannotation and functional validation of genes, is therefore essential to ensure the reliability of analyses aimed at elucidating metabolic diversity, virulence determinants, and evolutionary processes, particularly in entomopathogenic fungi (Wang and Wang 2017 ). Here, we conducted a comprehensive phylogenomic and biosynthetic gene cluster (BGC) analysis of 333 high-quality Beauveria genomes, including one newly sequenced by our group. Our approach integrates taxonomic profiling, BGC classification, and the functional association of secondary metabolite biosynthesis with virulence-related features in the newly sequenced strain. Together, our findings provide insights into the evolutionary conservation, diversification, and pathogenic potential encoded within the genomes of this agriculturally important fungal genus. Materials and Methods Genome sequencing, assembly and annotation B. bassiana AS272 strain was obtained from soil in the state of Rio Grande do Sul, Brazil. The isolation process used suspension of soil samples in sterile water, followed by plating on Oatmeal Agar-CTAB medium added of 50 mg/mL chloramphenicol. The pathogenicity of the isolates was confirmed by infection assays using the model arthropod Tenebrio molitor . The AS272 strain was maintained at 28°C on PDA medium, and identification was performed by sequencing the B locus intergenic region (Bloc) (accession number: PQ220121.1), primers – B51-F 5′-CGACCCGCCAACTACTTTGA-3′ and B31-R 5′-GTCTTCCAGTACCACTACGCC-3′. (Camargo, unpublished data). For DNA isolation, 3-day-old mycelia were used, following the CTAB method (Dar et al. 2025 ) with a bead beater (FastPrep®-24 5G - MP Biomedicals). Sequencing was performed in a paired-end mode (2 x 150 bp) using an Illumina NextSeq 2000 after library preparation by LacTAD ( https://www.lactad.unicamp.br/ ). Prior to assembly, raw sequencing reads were subjected to quality filtering using fastp v 0.23.4 (Chen et al. 2018 ) under default parameters. De novo genome assembly was subsequently performed using SPAdes v 4.2.0 (Bankevich et al. 2012 ) under default parameters. To enhance assembly contiguity, the resulting contigs were scaffolded with RagTag v 2.1.0 (Alonge et al. 2022 ), utilizing the Beauveria bassiana ARSEF 2860 reference genome (NCBI RefSeq: GCF_000280675.1) as a guide. The final genome sequence has been deposited in the NCBI database under the accession number GCA_052426205.1. The assembled and scaffolded genome was processed using Funannotate v 1.8.1 (Palmer and Stajich 2020 ) to perform FASTA header standardization, sequence sorting, and repeat masking. Gene prediction was subsequently conducted utilizing BRAKER3 v 3.0.6 (Gabriel et al. 2024 ), with the B. bassiana ARSEF 2860 reference proteome serving as the training dataset. Finally, the resulting GFF3 file was integrated back into the Funannotate pipeline to finalize the functional annotation. Secondary metabolite biosynthetic gene clusters (BGCs) were subsequently predicted from the BRAKER3 outputs (FASTA and GFF3 formats) utilizing antiSMASH v8.0.4 (Blin et al. 2025 ) configured in the fungal mode. Beauveria genome sequences A total of 333 assembled genomes, deposited as belonging to the genus Beauveria were retrieved from the NCBI database using NCBI datasets v 18.6.0 (O’Leary et al. 2024 ) and assessed for completeness using the BUSCO 5.8.2 tool (Benchmarking Universal Single-Copy Orthologs) (Manni et al. 2021 ) with the orthologous genes dataset of hypocreales (version 12). For phylogenomic analysis, all genomes sequences were used, while for functional prediction, only those with completeness scores above 92% were subsequently processed for gene prediction and functional annotation as described above using the Funannotate 1.8.1 pipeline (Palmer and Stajich 2020 ). The full list of accession codes and respective BUSCO results are provided in Table S1 . Phylogenomic analyses Phylogenomic analyses were performed using the BUSCO_phylogenomics pipeline (Waterhouse et al. 2018 ) ( https://github.com/jamiemcg/BUSCO_phylogenomics ), which generates phylogenomic matrices based on single-copy orthologous genes. The pipeline uses output files from BUSCO (Manni et al. 2021 ) to identify conserved single-copy genes, which are aligned with MAFFT (Katoh and Standley 2013 ). Nucleotide sequences of each orthologous group were aligned and concatenated into a supermatrix, with one partition considered per gene. Phylogenetic inference was conducted using IQ-TREE 2.1.4-beta (Minh et al. 2020 ), applying the TESTMERGE model to automatically merge partitions with compatible substitution models. Model selection was performed automatically by ModelFinder based on the Bayesian Information Criterion (BIC). Phylogenetic support was estimated using ultrafast bootstrap approximation (UFBoot) with 1,000 replicates. The final phylogenetic tree was constructed exclusively with nucleotide data and included all 333 genomes, with two Metarhizium species used as the outgroup. Secondary metabolite gene cluster analysis Gene cluster analysis was conducted using the local version of antiSMASH 8.0.4 (Blin et al. 2025 ). The classification of biosynthetic gene clusters (BGCs) into gene cluster families was carried out using BiG-SCAPE 2 (Biosynthetic Gene Similarity Clustering and Prospecting Engine) (Navarro-Muñoz et al. 2020 ). BGCs were grouped based on sequence similarity across a range of cutoff thresholds from 0.1 to 0.9. Clustering at a cutoff value of 0.6 provided an appropriate resolution for distinguishing BGC families. BGCs could be matched to previously characterized clusters by integrating the BiG-SCAPE algorithm with the Minimum Information about a Biosynthetic Gene cluster (MIBiG) database (Zdouc et al. 2025 ). Virulence Association Prediction To investigate the potential association between BGCs and virulence, the predicted protein-coding genes within selected BGCs were queried against the PHI-base 4.18 (Urban et al. 2025 ) database using a local BLASTP 2.12.0 (Camacho et al. 2009 ). Hits showing more than 50% identity and derived from fungal species were considered to be associated with virulence. Based on these results, a protein–protein interaction (PPI) network was constructed using the STRING 12.0 database (Szklarczyk et al. 2023 ) using the automatic annotation process in the STRING server (avilable at https://version-12-0.string-db.org/organism/STRG0A07FVB ). Material avaliability The codes used for shell processing, as well as for R analysis and figures plotting, are available as Supplementary Files 1 and 2, respectively. Intermediate and final processing files are available upon request. Results Characterization of B. bassiana AS272 genome sequence To characterize the genomic characteristics of the high-virulence B. bassiana isolate AS272, a comparative genomics approach was employed, using the NCBI reference strain B. bassiana ARSEF 2860 as a control. The assembly of isolate AS272 comprises 31,716,118 bp with a high level of completeness (96.7% complete BUSCOs). Braker3 prediction identified 9,778 protein-coding genes, leading to 10,937 predicted proteoforms, which is slightly higher compared to reference (10,364 predicted proteoforms). OrthoFinder analysis revealed a high degree of conservation, with 8,984 orthogroups shared between the two isolates, representing 99.5% of the total orthogroup repertoire. Notably, isolate AS272 exhibited a significantly higher percentage of genes assigned to orthogroups (98.1%; 10,730 genes) compared to the reference strain (95.7%; 9,921 genes). Furthermore, AS272 showed a marked reduction in unassigned "singleton" genes (207; 1.9%) relative to ARSEF 2860 (443; 4.3%), suggesting that the AS272 increased gene count is driven by the expansion of existing gene families (Fig. 1 A). A comparative analysis of gene family sizes identified significant quantitative differences between the two isolates across multiple orthogroups. While global orthogroup distribution remains relatively correlated (R = 0.26, p < 2.2e − 16), specific expansions could be observed in the ARSEF 2860 strain. Notably, families OG0000000 and OG0000001 showed the most substantial expansion in the reference strain ARSEF 2860 compared to AS272. Functional annotation via CDD search revealed that OG0000000 is predominantly composed of proteins containing the hAT family transposase domain (Dimer_Tnp_hAT, pfam05699) (Fig. 1 B), which are generally associated with genetic mobile elements, and often associated with genomic plasticity and structural variation. In contrast, OG0000001 is enriched with members of the MRP_assoc_pro superfamily (cl33195) and ABC transporter domains (cl38913), which are essential for the regulated efflux of metabolites and xenobiotics. Functional mapping via KEGG identified specific KEGG Orthologs (KO) with marked copy-number variation. The most significant expansion was observed in K08192 (an MFS transporter), which displayed a delta of 13 copies compared to the reference. This was followed by K22134 (MFS transporter), K08832 (serine-threonine kinase), and K05658 (ATP-Binding cassette). Conversely, a reduction was noted in K07375 (Tubulin beta), suggesting subtle structural or regulatory shifts relative to the reference strain (Fig. 1 C). Finally, isolate AS272 displays a robust increase of pathogenicity-related factors mapped against the PHI-base database using a stringent criteria (minimum of 50% identity over at least 50% of the length of the subject). Compared to ARSEF 2860, the strain AS272 exhibits a higher number of homologs linked to the "reduced virulence" and "loss of pathogenicity" categories, which points toward a potentially redundant virulence set of proteins (Fig. 1 D). These functional enrichments point toward specialized adaptations in nutrient transport and environmental sensing that likely support the infection phenotype of this soil-derived isolate. Phylogenomic analyses To establish the precise taxonomic identity and evolutionary context of the high-virulence isolate B. bassiana AS272, we performed a phylogenomic analysis using a maximum-likelihood approach. The phylogeny was reconstructed from a supermatrix of 247,446 bp derived from 204 genomic partitions, which represent 4,333 single-copy orthologs identified via BUSCO. The dataset included 335 sequences, comprising 333 Beauveria genomes (queried from NCBI as of March 2023) and two Metarhizium strains utilized as outgroups. The resulting fan phylogeny resolved the genus into several major clades (Fig. 2 ). To clarify evolutionary distances within dense terminal radiations, branch lengths were visualized using a logarithmic square-root transformation. The most expansive lineage is the B. bassiana clade, which contains the majority of investigated isolates. Highly supported sister clades, which are indicated by red internal branches (≥ 80% UFBoot), were clearly delineated for B. pseudobassiana , B. asiatica , B. varroae , and B. brongniartii . Additionally, a distinct clade was recovered pairing B. australis with B. medogensis . The tree remains rooted by a robust outgroup containing M. robertsii and M. anisopliae . Notably, B. bassiana AS272 (GCA_052426205.1) was positioned within a well-supported subclade composed exclusively of B. bassiana strains (Fig. 2 - arrow). This placement confirms its taxonomic identity and definitively separates it from sister species within the complex, such as B. pseudobassiana and B. brongniartii . Despite the overall robustness of the topology, several phylogenetic discrepancies were identified. B. bassiana MBC618 branched entirely outside the Beauveria genus, clustering with the Metarhizium outgroup with absolute support (100% UFBoot/100% SH-aLRT; Fig. 2 Arrow A). This significant incongruence suggests a potential misidentification or genomic sequence contamination for this specific isolate. Furthermore, B. rudraprayagi MTCC8017 was nested deeply within the B. bassiana clade, grouping directly with B. bassiana ARSEF 2597 (Fig. 2 - Arrow B). This finding supports the hypothesis that B. rudraprayagi may represent a variant of B. bassiana rather than a distinct phylogenetic species. Similarly, B. felina SYSU-MS7908 appeared phylogenetically distant from all other represented taxa, occupying an isolated basal position (Fig. 2 - Arrow C). Collectively, these discrepancies highlight the critical need for rigorous phylogenomic validation to address the incidence of misidentified genomes and cryptic diversity currently present in public repositories. Conservation of BGCs among Beauveria genus analyses With the taxonomic identity and evolutionary relationships of B. bassiana AS272 established, our focus shifted toward the functional genomic elements potentially associated with its virulence phenotype. Entomopathogenic fungi rely on a highly specialized set of secondary metabolites to facilitate host infection, evade insect immune responses, and compete in complex ecological niches. Therefore the specialized biosynthetic capacity of AS272 was inferred and compared to those observed in the broader evolutionary history of the genus. The Beauveria genomes that displayed BUSCO values above 92% here analyzed were first annotated using fungiSMASH. A total of 14,602 BGCs were predicted across the 333 genomes, revealing a high degree of biosynthetic diversity and inter-specific variation within the genus (Fig. 3 ). The total number of BGCs per genome ranged from approximately 35 to 55, with B. varroae and B. pseudobassiana exhibiting the most expansive biosynthetic repertoires, maintaining medians above 45 clusters per genome. Conversely, considering the content of BGCs, the genomes of B. asiatica, B. caledonica and B. rudraprayagi appeared more compact (Fig. 3 A). The distribution of Biosynthetic Gene Clusters (BGCs) by chemical class highlights a clear dominance of Non-Ribosomal Peptide Synthetases (NRPS) and Polyketide Synthases (PKS) across all species (Fig. 3 B- 3 G). NRPS BGCs represent the most abundant category identified (Fig. 3 C). Notably, B. varroae exhibited a significant expansion in this class, with a median of approximately 24 clusters, suggesting a specialized reliance on peptide-based metabolites compared to the genus average. PKS clusters, while slightly less abundant, were maintained across species, typically ranging between 5 and 20 per genome (Fig. 3 E). Interestingly, B. pseudobassiana displayed the highest intra-species variability in PKS content, indicating a high degree of genomic plasticity and potential for niche-specific chemical signatures within this lineage. In contrast to the high inter-specific variation observed in NRPS and PKS counts, the abundance of Terpene (Fig. 3 B) and PKS-NRP Hybrid (Fig. 3 D) BGCs was notably consistent across the genus. These classes exhibited a highly uniform distribution, with the majority of species harboring a stable repertoire of 6–9 terpene clusters and 2–6 hybrid clusters per genome. Conversely, Ribosomally synthesized and Post-translationally modified Peptides (RiPPs) BGCs (Fig. 3 G) were remarkably rare, often limited to one or two clusters or entirely absent, as seen in B. asiatica . A unique expansion in the "Others" category (Fig. 3 F) was also observed in B. pseudobassiana , hinting at the presence of unclassified or novel biosynthetic pathways that remain to be characterized (Fig. 3 ). The profile of distinct BGCs in B. bassiana AS272 follows the general distribution of BGCs in Beauveria genomes, totaling 39 predicted clusters. While its genomic repertoire aligns closely with the species median for most classes, the isolate exhibits specific counts of 19 NRPS, 8 PKS, 7 Terpene, and 4 PKS-NRP Hybrid clusters (red dots in Fig. 3 ). To further investigate the evolutionary relationships between the predicted BGCs, we performed a systematic clustering analysis using BiG-SCAPE. By applying a range of sequence similarity cutoffs (0.10 to 0.90), we evaluated the conservation of BGCs into Gene Cluster Families (GCFs) across the Beauveria genus. As expected, the number of identified GCFs decreased significantly as the clustering threshold became more stringent (higher cutoff), indicating that many clusters share common ancestral architectures (Fig. 4 A). This is reflected in the number of "singleton" BGCs—those unique to a single genome—which decreased sharply as the cutoff moved toward more inclusive similarity levels (Fig. 4 B). The mean size of the GCFs (number of BGCs per family) showed an exponential increase at higher cutoffs, particularly for the NRPS and PKS classes (Fig. 4 C). This suggests that a substantial portion of the Beauveria specialized metabolome consists of highly related biosynthetic pathways that likely produce similar core chemical scaffolds. At a selected cutoff of 0.60, which provides a balanced resolution between over-clustering and excessive fragmentation, the PKSI and Terpene classes exhibited the highest mean GCF sizes (Fig. 4 D). This indicates that these specific classes are characterized by large, multi-member families that are widely distributed across the genus. Finally, the proportion of singleton clusters varied markedly by chemical class (Fig. 4 E). While PKSI, Terpene, and NRPS clusters showed a low percentage of singletons (typically 25% singletons). This high degree of uniqueness in unclassified clusters highlights a reservoir of genomic novelty that may drive specific ecological adaptations or host-specialized interactions in individual Beauveria isolates, including the high-virulence AS272 strain. The distribution of Gene Cluster Families (GCFs) across the genus Beauveria was further evaluated. The analysis of GCF size, which is defined by the number of BGCs per family, revealed distinct evolutionary patterns among chemical classes (Fig. 5 ). The PKSI, Terpene, and NRPS classes exhibited the most expansive family sizes, with some GCFs containing over 100 individual BGCs. This indicates that a significant portion of the Beauveria specialized metabolome is composed of highly related biosynthetic pathways shared by a vast majority of the 333 genomes. The largest median GCF sizes were observed in the PKSI and Terpene families, suggesting the significant importance of these specific conserved pathways to the distinct lifestyles of Beauveria . In contrast, the PKS-NRP Hybrids, PKSother, RiPPs, and Others categories were characterized by significantly smaller GCFs, with medians frequently falling below five members. This suggests that these classes are more prone to diversification or are restricted to specific lineages. When filtering for GCFs containing known clusters from the MIBiG database, the mean GCF size dropped across all categories, although PKSI and NRPS retained the most robust presence. This drop highlights a substantial gap in the functional characterization of Beauveria secondary metabolism, as the majority of large, well-conserved families currently lack a corresponding characterized metabolite in public databases. To characterize the core biosynthetic potential of the genus Beauveria , a conservation analysis was performed on Gene Cluster Families (GCFs) containing at least 100 members. This threshold was selected because such a wide distribution across species and strains suggests these clusters encode for specialized metabolites that are fundamental to Beauveria biology. A total of 21 GCFs met this criterion, and their occupancy, defined as the percentage of species within a gene harboring the cluster, reveals distinct patterns of evolutionary stability and lineage-specific adaptation (Table S2 ). Several GCFs exhibited high occupancy across nearly the entire phylogeny, representing the stable biosynthetic framework of the genus. NRPS GCF 1963 and NRPS GCF 16320 emerged as two of the most ubiquitous pathways, with 100% occupancy across B. caledonica, B. medogensis , and B. rudraprayagi (GCF 1963) or B. asiatica, B. australis , and B. caledonica (GCF 16320). PKSI GCF 7963 was equally prominent, maintaining 100% occupancy in B. asiatica, B. australis, B. medogensis , and B. rudraprayagi. Terpene GCF 12284 demonstrated a unique distribution, present in 100% of B. australis, B. brongniartii, B. caledonica, B. medogensis , and B. rudraprayagi isolates, despite being nearly absent in B. bassiana (2%) and B. pseudobassiana (0%). In contrast to the ubiquitous core, several large GCFs served as definitive markers for specific clades. For B. bassiana and B. rudraprayagi , Terpene GCF 1989 showed high frequency in B. bassiana (85.9%) and 100% occupancy in B. rudraprayagi , but was absent in most other lineages. The sister species B. pseudobassiana and B. varroae share a unique biosynthetic profile, defined by 100% occupancy of NRPS GCF 16292, PKSI GCF 16319, and PKSother GCF 4094. Furthermore, NRPS GCF 15871 and NRPS GCF 19 showed nearly universal presence in these two clades (> 95%) while remaining restricted from the rest of the genus. The conservation analysis also highlighted GCFs with hits in the MIBiG database, providing insight into the chemical nature of these core clusters. Terpene GCF 1989 (linked to BGC0001839.1) and PKSI GCF 7963 (linked to BGC0001720.1) represent well-characterized pathways widely distributed across the genus. Notably, NRPS GCF 19 was associated with multiple MIBiG entries (BGC0000959.1, BGC0001136.1, and BGC0002035.1), suggesting it may represent a cluster family responsible for a known class of bioactive compounds like beauvericin or bassianolide. Within the B. bassiana complex, including the isolate AS272, occupancy for these 21 large GCFs typically fluctuated between 40% and 60%. This moderate frequency reflects the significant intra-specific genomic plasticity of the lineage. To characterize the species-specific biosynthetic landscape, we first attempted to identify a strictly conserved set of Gene Cluster Families (GCFs) within B. bassiana . Despite being the most abundant species in the dataset, comprising 199 of the 333 genomes analyzed, the results revealed that B. bassiana possesses zero strictly core GCFs when defined by 100% occupancy across all isolates. Given this observation, we evaluated the distribution of biosynthetic gene cluster families (GCFs) exclusively within the B. bassiana lineage using discrete power-law and log-normal modeling to investigate how these clusters are maintained across isolates. The analysis revealed a highly skewed distribution, with substantial variance in GCF occupancy among isolates (Fig. 6 ). However, model comparison using the Vuong likelihood ratio test did not support a significantly better fit of the power-law model over the log-normal alternative (LR = − 0.61, p = 0.543). Although a power-law fit yielded an estimated scaling exponent of α = 5.41 (xmin = 71), the absence of statistical support indicates that the observed pattern cannot be robustly classified as strictly scale-free. Instead, the distribution is consistently heavy-tailed, in which a small subset of GCFs is widely conserved, while most clusters occur at lower frequencies across isolates (Fig. 6 ). While no strictly core clusters were found, the modeling identified some GCFs, which represent high-frequency clusters that could form the functional backbone of the species' chemical repertoire without being universally present in every strain. The occupancy data for these hubs further illustrates this flexibility. The most prevalent family, Terpene GCF 1989, reached only 85.9% occupancy, followed by Terpene GCF 8171 (67.3%) and NRPS GCF 8144 (59.3%). Other major families such as PKSother 12262 (57.8%) and PKSI 8173 (56.8%) also displayed moderate to high frequency but remained absent in a significant portion of the population. Within the B. bassiana complex, including the high-virulence isolate AS272, occupancy for these 21 large GCFs typically fluctuated between 40% and 60%. This moderate frequency, compared to the strict 100% occupancy seen in smaller clades, reflects the significant intra-specific genomic plasticity of the lineage. This suggests that while B. bassiana maintains a core set of BGCs, it utilizes a more variable set of high-frequency clusters to potentially tailor its pathogenic profile to specific hosts or environmental niches. Association of B. bassiana AS272 BGCs with Virulence To investigate the functional context of conserved biosynthetic gene clusters and their potential association with pathogenicity, we performed a protein–protein interaction (PPI) network analysis. Interaction networks were constructed for the core genes of six highly conserved hub GCFs using the STRING database (version 12.0), based on the B. bassiana AS272 annotation. For this analysis, we selected GCFs with high prevalence across the genus (≥ 90 members), representing five biosynthetic classes: Terpene, NRPS, PKS–NRP hybrids, PKS-like, and T1PKS. Topological analysis indicated that these GCFs are embedded within highly interconnected interaction networks (Fig. 7 ). The resulting networks displayed substantial structural complexity, with an average node degree ranging from 15.23 (Terpene GCF 1989) to 43.75 (NRPS GCF 8144). A notable feature of these networks is their high clustering coefficients, ranging from 0.58 to 0.95, which exceeded expectations for comparable random networks. Correspondingly, clustering enrichment values ranged from 1.55 to 2.85, indicating a dense and modular organization in which biosynthetic core genes are connected with broader cellular interaction networks (Table S2 ). To assess potential associations with pathogenicity, network proteins were cross-referenced against the PHI-base database to identify homologs of experimentally characterized pathogen–host interaction factors. Across the six networks, the number of PHI-related nodes ranged from 8 to 18 (Fig. 7 ). Statistical enrichment analysis indicated that these associations are unlikely to arise by chance. In particular, the NRPS (GCF 8144) and PKS–NRP hybrid (GCF 12259) networks showed the strongest enrichment for known virulence-associated proteins ( FDR-adjusted p = 1.90 × 10⁻⁶). The PKSother network (GCF 12262) also exhibited significant enrichment (p = 6.40 × 10⁻⁵). Although the Terpene and T1PKS networks contained multiple PHI-base homologs (14 and 18 nodes, respectively), their enrichment did not reach the same level of statistical significance (FDR > 0.05). This pattern may reflect interactions with a broader or less well-characterized set of host-associated factors. Taken together, these results suggest that several conserved biosynthetic gene clusters in Beauveria , particularly NRPS and hybrid PKS–NRP systems, are embedded within interaction networks enriched for proteins associated with host interaction and pathogenicity (Table S2 ). Discussion Analysis of the B. bassiana isolate AS272 reveals a genomic architecture geared toward both metabolic versatility and pathogenic potential. While some studies emphasize the contribution of lineage-specific genes (singletons) to virulence, our results suggest that the aggressiveness of this isolate may instead be associated with an expansion of conserved gene families. The relatively low proportion of orphan genes compared to orthogroups indicates that the AS272 genome has likely evolved through gene duplication events, leading to increased functional redundancy and potential metabolic flexibility. A prominent feature of the genome is the expansion of transporters belonging to the Major Facilitator Superfamily (MFS) and ATP-binding cassette (ABC) superfamilies. These transporters are known to play key roles in toxin secretion during infection and may also contribute to protection against insect-derived defense metabolites and environmental stressors (Valero-Jiménez et al. 2016). In AS272, the increased copy number of genes such as K08192 suggests a potentially enhanced efflux capacity, which may facilitate survival and colonization within host environments (Solano-González et al. 2023). From a phylogenetic perspective, the phylogenomic reconstruction of 337 genomes revealed persistent misidentifications in public databases. The clustering of strain MBC618 within the Metarhizium genus, together with the isolated position of B. felina , highlights the limitations of traditional taxonomy based solely on morphology or single-locus markers such as the ITS region. Furthermore, the placement of B. rudraprayagi within the B. bassiana clade suggests that it may represent a variant or closely related lineage of B. bassiana , supporting previous studies that have emphasized the need for systematic revision within the Beauveria genus (Rehner et al. 2011). The use of single-copy orthologs identified through the BUSCO framework (Manni et al. 2021) provides a robust phylogenomic basis for species identification and for the accurate documentation of strains used in industrial or biotechnological applications. Another notable finding is the apparent absence of a strictly conserved core of biosynthetic gene clusters (BGCs) within B. bassiana . Instead, the distribution of these clusters follows a highly dynamic pattern, potentially enabling the fungus to maintain a flexible chemical arsenal and considerable genomic plasticity (Valero-Jiménez et al. 2016). While certain classes, such as terpene clusters, show relatively high conservation—suggesting roles in core biological processes—NRPS and PKS clusters appear to function primarily as sources of chemical innovation. Such intraspecific variation may explain the substantial differences in virulence observed among B. bassiana strains infecting the same host species, as individual strains may deploy distinct combinations of secondary metabolites derived from diverse biosynthetic hubs (Toopaang et al. 2025). Protein–protein interaction (PPI) network analysis further suggests that secondary metabolite biosynthesis is integrated within broader pathogenicity-related networks. Cross-referencing these networks with PHI-base identified several connections between core BGC proteins and previously described virulence-associated factors, indicating potential coordination between toxin biosynthesis and infection-related processes such as cuticle degradation. Although compounds such as beauvericin and oosporein are known to play important roles in suppressing host immune responses, most of the identified GCFs lacked direct counterparts when compared against the MIBiG database (Zdouc et al. 2025). The dynamic distribution of BGCs observed across B. bassiana isolates likely reflects the rapid evolutionary turnover typical of fungal secondary metabolite pathways. Biosynthetic clusters are frequently shaped by gene duplication, recombination, horizontal transfer, and lineage-specific gene loss, processes that collectively generate substantial chemical diversity within species. Such evolutionary plasticity allows entomopathogenic fungi to adapt to diverse ecological niches and host environments by modifying or expanding their repertoire of bioactive compounds. In the case of Beauveria , this flexible biosynthetic architecture may provide a selective advantage during host infection, where different insect species present distinct physiological defenses. Consequently, the diversification and differential retention of BGCs may represent an important evolutionary strategy enabling B. bassiana populations to maintain pathogenic versatility across a broad range of insect hosts. The absence of clear matches in MIBiG underscores the largely unexplored biosynthetic diversity within the Beauveria genus. This hidden repertoire of specialized metabolites may represent a substantial source of novel bioactive compounds with potential applications in biotechnology, agriculture, and pharmaceutical discovery, warranting further functional and chemical characterization. Statements and Declarations Funding Declaration This work was supported by grants from the Brazilian funding agencies Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq - 312994/2021-4 and 405015/2023-2), Fundação de Amparo à Pesquisa do Estado do Rio Grande do Sul (FAPERGS - 19/2551-0001708-1), and Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES). RSS and AAR was a recipient of a CAPES scholarships. HA and MSC were recipients of CNPq scholarships. Conflict of Interest: The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Ethical Approval: Not applicable. Contributions All authors contributed to the study conception and design. Robson dos Santos Soares, Matheus Camargo, Maria Eduarda Deluca João, and Charley Christian Staats performed material preparation and data collection. Robson dos Santos Soares, Alexandra Rocha, Henrique Antoniolli, and Charley Christian Staats performed the data analysis. Augusto Schrank and Charley Christian Staats were responsible for project administration and funding acquisition. The first draft of the manuscript was written by Robson dos Santos Soares , and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript. References Alonge M, Lebeigle L, Kirsche M, et al (2022) Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing. Genome Biol 23:258. https://doi.org/10.1186/s13059-022-02823-7 Aly HH, Meng Y, Wang D (2025) Comparative gene expression analysis of Beauveria bassiana against Spodoptera frugiperda. PeerJ 13:e19591. https://doi.org/10.7717/peerj.19591 Andreata MF de L, Mian S, Andrade G, et al (2025) The current increase and future perspectives of the microbial pesticides market in agriculture: the Brazilian example. Front Microbiol 16:1574269. https://doi.org/10.3389/fmicb.2025.1574269 Bankevich A, Nurk S, Antipov D, et al (2012) SPAdes: A new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol 19:455–477. https://doi.org/10.1089/cmb.2012.0021 Blin K, Shaw S, Vader L, et al (2025) antiSMASH 8.0: extended gene cluster detection capabilities and analyses of chemistry, enzymology, and regulation. Nucleic Acids Res 53:W32–W38. https://doi.org/10.1093/nar/gkaf334 Camacho C, Coulouris G, Avagyan V, et al (2009) BLAST+: architecture and applications. BMC Bioinformatics 10:421. https://doi.org/10.1186/1471-2105-10-421 Chaithra M, Prameeladevi T, Prasad L, et al (2024) Metabolomic profiling of virulent and non-virulent Beauveria bassiana strains: insights into the pathogenicity of Tetranychus truncatus. Arch Microbiol 206:311. https://doi.org/10.1007/s00203-024-04046-9 Chen S, Zhou Y, Chen Y, Gu J (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34:i884–i890. https://doi.org/10.1093/bioinformatics/bty560 Dar GJ, Nazir R, Wani SA, et al (2025) Optimizing a modified cetyltrimethylammonium bromide protocol for fungal DNA extraction: Insights from multilocus gene amplification. Open Life Sci 20:20221006. https://doi.org/10.1515/biol-2022-1006 de Sousa TK, Silva AT da, Soares FE de F (2025) Fungi-Based Bioproducts: A Review in the Context of One Health. Pathogens 14:. https://doi.org/10.3390/pathogens14050463 Farag PF, Elsisi AA, Elabd EW, et al (2025) Prediction of secreted uncharacterized protein structures from Beauveria bassiana ARSEF 2860 unravels novel toxins-like families. Sci Rep 15:17747. https://doi.org/10.1038/s41598-025-02618-3 Gabriel L, Brůna T, Hoff KJ, et al (2024) BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA. Genome Res 34:769–777. https://doi.org/10.1101/gr.278090.123 Hong S, Shang J, Sun Y, et al (2024) Fungal infection of insects: molecular insights and prospects. Trends Microbiol 32:302–316. https://doi.org/10.1016/j.tim.2023.09.005 Katoh K, Standley DM (2013) MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 30:772–780. https://doi.org/10.1093/molbev/mst010 Keller NP (2019) Fungal secondary metabolism: regulation, function and drug discovery. Nat Rev Microbiol 17:167–180. https://doi.org/10.1038/s41579-018-0121-1 Manni M, Berkeley MR, Seppey M, et al (2021) BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes. Mol Biol Evol 38:4647–4654. https://doi.org/10.1093/molbev/msab199 Mascarin GM, Jaronski ST (2016) The production and uses of Beauveria bassiana as a microbial insecticide. World J Microbiol Biotechnol 32:177. https://doi.org/10.1007/s11274-016-2131-3 Minh BQ, Schmidt HA, Chernomor O, et al (2020) IQ-TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Mol Biol Evol 37:1530–1534. https://doi.org/10.1093/molbev/msaa015 Ministério da Agricultura e Pecuária Agrofit - Sistema de Agrotóxicos Fitossanitários. https://agrofit.agricultura.gov.br/agrofit_cons/principal_agrofit_cons. Accessed 19 Feb 2026 Navarro-Muñoz JC, Selem-Mojica N, Mullowney MW, et al (2020) A computational framework to explore large-scale biosynthetic diversity. Nat Chem Biol 16:60–68. https://doi.org/10.1038/s41589-019-0400-9 O’Leary NA, Cox E, Holmes JB, et al (2024) Exploring and retrieving sequence and metadata for species across the tree of life with NCBI Datasets. Sci Data 11:732. https://doi.org/10.1038/s41597-024-03571-y Oberti H, Sessa L, Oliveira-Rizzo C, et al (2025) Novel genomic features in entomopathogenic fungus Beauveria bassiana ILB308: accessory genomic regions and putative virulence genes involved in the infection process of soybean pest Piezodorus guildinii. Pest Manag Sci 81:2323–2336. https://doi.org/10.1002/ps.8631 Palmer JM, Stajich J (2020) Funannotate v1.8.1: Eukaryotic genome annotation. Zenodo. https://doi.org/10.5281/zenodo.4054262 Rehner SA, Minnis AM, Sung G-H, et al (2011) Phylogeny and systematics of the anamorphic, entomopathogenic genus Beauveria . Mycologia 103:1055–1073. https://doi.org/10.3852/10-302 Solano-González S, Castro-Vásquez R, Molina-Bravo R (2023) Genomic Characterization and Functional Description of Beauveria bassiana Isolates from Latin America. J Fungi (Basel) 9:. https://doi.org/10.3390/jof9070711 Szklarczyk D, Kirsch R, Koutrouli M, et al (2023) The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res 51:D638–D646. https://doi.org/10.1093/nar/gkac1000 Toopaang W, Yoocha T, Naktang C, et al (2025) Transcriptomic insights into the interplay between polyketide biosynthesis and other secondary metabolite biosynthetic clusters and biological pathways in entomopathogen Beauveria bassiana . Front Microbiol 16:1583637. https://doi.org/10.3389/fmicb.2025.1583637 Urban M, Cuzick A, Seager J, et al (2025) PHI-base - the multi-species pathogen-host interaction database in 2025. Nucleic Acids Res 53:D826–D838. https://doi.org/10.1093/nar/gkae1084 Valero-Jiménez CA, Faino L, Spring In’t Veld D, et al (2016) Comparative genomics of Beauveria bassiana : uncovering signatures of virulence against mosquitoes. BMC Genomics 17:986. https://doi.org/10.1186/s12864-016-3339-1 Wang C, Wang S (2017) Insect pathogenic fungi: genomics, molecular interactions, and genetic improvements. Annu Rev Entomol 62:73–90. https://doi.org/10.1146/annurev-ento-031616-035509 Wang H, Peng H, Li W, et al (2021) The Toxins of Beauveria bassiana and the Strategies to Improve Their Virulence to Insects. Front Microbiol 12:705343. https://doi.org/10.3389/fmicb.2021.705343 Waterhouse RM, Seppey M, Simão FA, et al (2018) BUSCO applications from quality assessments to gene prediction and phylogenomics. Mol Biol Evol 35:543–548. https://doi.org/10.1093/molbev/msx319 Zdouc MM, Blin K, Louwen NLL, et al (2025) MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration. Nucleic Acids Res 53:D678–D690. https://doi.org/10.1093/nar/gkae1115 Additional Declarations No competing interests reported. Supplementary Files TableS1.csv Table S1: Detailed list of Beauveria and Metarhizium genomes used in this study, including NCBI accession numbers and assembly completeness metrics based on BUSCO analysis. TableS2.csv Table S2: Summary of Biosynthetic Gene Cluster Families (GCFs) identified across the Beauveria genus, including taxonomic distribution and similarity to known biosynthetic pathways (MIBiG). TableS3.csv Table S3: Topological network metrics and Pathogen-Host Interaction (PHI-base) gene enrichment analysis for selected Beauveria Gene Cluster Families. FileS1.sh FileS2.r Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 07 May, 2026 Reviewers agreed at journal 05 May, 2026 Reviewers agreed at journal 04 May, 2026 Reviewers invited by journal 04 May, 2026 Editor assigned by journal 27 Mar, 2026 Submission checks completed at journal 27 Mar, 2026 First submitted to journal 23 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9203165","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":636328899,"identity":"679337ad-c01f-4369-9ed7-b463f4768968","order_by":0,"name":"Robson dos Santos Soares","email":"","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":false,"prefix":"","firstName":"Robson","middleName":"dos Santos","lastName":"Soares","suffix":""},{"id":636328900,"identity":"262f62c6-d4eb-4f34-8690-967c80ec6f82","order_by":1,"name":"Alexandra de Azevedo da Rocha","email":"","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":false,"prefix":"","firstName":"Alexandra","middleName":"de Azevedo da","lastName":"Rocha","suffix":""},{"id":636328901,"identity":"55e4d116-bd7e-4567-9eb8-5868eedee7e7","order_by":2,"name":"Matheus da Silva Camargo","email":"","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":false,"prefix":"","firstName":"Matheus","middleName":"da Silva","lastName":"Camargo","suffix":""},{"id":636328902,"identity":"cb445c92-32de-42b0-bc22-dfb79b151ba4","order_by":3,"name":"Maria Eduarda Deluca João","email":"","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":false,"prefix":"","firstName":"Maria","middleName":"Eduarda Deluca","lastName":"João","suffix":""},{"id":636328903,"identity":"e25c7576-e081-42c5-95a2-aed321c3d2fd","order_by":4,"name":"Augusto Schrank","email":"","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":false,"prefix":"","firstName":"Augusto","middleName":"","lastName":"Schrank","suffix":""},{"id":636328904,"identity":"1c4ae9f9-0af1-4588-a479-67c677698998","order_by":5,"name":"Henrique da Rocha Moreira Antoniolli","email":"","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":false,"prefix":"","firstName":"Henrique","middleName":"da Rocha Moreira","lastName":"Antoniolli","suffix":""},{"id":636328905,"identity":"6f9cb769-ca0b-4ee6-bd47-111a528ad3e2","order_by":6,"name":"Charley Christian Staats","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvUlEQVRIiWNgGAWjYPACGwb2BiDF2ECMYjYgPsCQxsBzgEQth0nQIj+/+eHnD3/OJ/ZIH7/8gXHHPcJaDI6xGUscbLud2MOXUybBeKaYCC1sDAYSBxtuJ+7n4UljYGxLIMJhbeyffxz4cy6xh4cn+QNRWhiO8ZhJHGA7ANTCfkCCKC0Gx3LKLM62JRsDbWGTSDxDjMOaj2++UfHHThZoy+MPH3cQ4zAE4DFgIE0DAwP7AxI1jIJRMApGwUgBAMpMPBgWxVgCAAAAAElFTkSuQmCC","orcid":"","institution":"Federal University of Rio Grande do Sul","correspondingAuthor":true,"prefix":"","firstName":"Charley","middleName":"Christian","lastName":"Staats","suffix":""}],"badges":[],"createdAt":"2026-03-23 16:55:46","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9203165/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9203165/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109051496,"identity":"4c422885-0516-42a0-9c66-5718b539b4ee","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":252068,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparative genomic analysis of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eB. bassiana\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e AS272 with the reference strain ARSEF 2860. (A) \u003c/strong\u003eOrthology and Proteome Metrics. Bar chart comparing the total predicted genes, genes assigned to orthogroups, and unassigned genes for AS272 and ARSEF 2860. \u003cstrong\u003e(B) \u003c/strong\u003eOrthogroup Gene Family Expansion. Scatterplot comparing gene counts per orthogroup between the two strains. \u003cstrong\u003e(C) \u003c/strong\u003eFunctional Divergence of KEGG Orthologs (KO). Distribution of the most expanded and reduced functional categories in AS272. The horizontal axis represents the copy number delta (AS272 - ARSEF 2860).\u003cstrong\u003e (D) \u003c/strong\u003ePathogenicity-Related Factors (PHI-base). Distribution of genes mapped to the Pathogen-Host Interactions database.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/7f6afd53994c443bea1e685a.png"},{"id":109067871,"identity":"aa0f3bdb-3a9f-492a-8647-01446ea8bc8a","added_by":"auto","created_at":"2026-05-12 10:02:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":870203,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMaximum Likelihood radial phylogeny of the genus \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eBeauveria\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e. \u003c/strong\u003eMaximum likelihood radial phylogeny inferred from a concatenated supermatrix comprising 247,446 bp across 204 genomic partitions, representing 339 isolates (337 \u003cem\u003eBeauveria\u003c/em\u003e isolates and two \u003cem\u003eMetarhizium\u003c/em\u003eoutgroups). The phylogeny was rooted using \u003cem\u003eM.\u003c/em\u003e \u003cem\u003erobertsii\u003c/em\u003e ARSEF 23 and \u003cem\u003eM. anisopliae\u003c/em\u003eJEF-290. Internal branches with ultrafast bootstrap (UFB) support ≥ 80% are highlighted in red. Tip labels are color-coded according to species identification as indicated in the legend. To improve visualization of densely clustered clades, branch lengths were transformed using a logarithmic square-root scaling. The purple arrow marks the position of \u003cem\u003eB.\u003c/em\u003e \u003cem\u003ebassiana\u003c/em\u003e isolate AS272, while arrows labeled A, B, and C indicate incongruences in the tree.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/d11fb22f26f4a1339b839cb2.png"},{"id":109221667,"identity":"f210748d-dac7-482f-9ded-57bc02632dea","added_by":"auto","created_at":"2026-05-13 20:51:30","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":169137,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistribution and diversity of Biosynthetic Gene Clusters (BGCs) across the genus \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eBeauveria\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e.\u003c/strong\u003e Comparative analysis of BGC abundance across 333 genomes representing 10 taxonomic groups. The top panel illustrates the total BGC count per genome, followed by stratification by chemical class: Terpene, Non-ribosomal Peptide Synthetase (NRPS), Polyketide Synthase-Non-ribosomal Peptide Hybrids (PKS-NRP Hybrids), Polyketide Synthase (PKS), Others (unclassified clusters), and Ribosomally synthesized and Post-translationally modified Peptides (RiPPs). Box plots represent the interquartile range (IQR), with the central line indicating the median and whiskers extending to 1.5 × IQR. Individual genomic data points are overlaid as semi-transparent blue circles to illustrate intra-species variation. The red circle in each panel specifically highlights the BGC counts for the high-virulence isolate \u003cem\u003eB. bassiana\u003c/em\u003e AS27.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/bcc5eeaec7e6a8283cd8bbc4.png"},{"id":109067727,"identity":"3b86d5d8-ac58-4230-94e2-912aad9e96b2","added_by":"auto","created_at":"2026-05-12 10:00:21","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":223555,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparative analysis of biosynthetic potential and Gene Cluster Family (GCF) dynamics across genomic datasets.\u003c/strong\u003e (A) Total number of GCFs identified across a gradient of clustering stringency cutoffs (0.10 to 0.90). (B) Frequency of \"singleton\" clusters (BGCs that did not group into a family) across the same stringency gradient. (C) Mean GCF size (number of member clusters per family, log10 scale) as a function of clustering cutoff. (D) Comparison of mean GCF size at the selected cutoff (0.6) across biosynthetic classes, including Type I Polyketide Synthases (PKSI), Terpene, Non-ribosomal Peptide Synthetase (NRPS), and hybrids. (E) Percentage of singleton clusters at the selected cutoff (0.6) relative to the total number of BGCs within each biosynthetic category.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/038466799f14d1852658268a.png"},{"id":109051502,"identity":"1268ee28-05ea-4375-b1e7-59cb22673522","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":162199,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDiversity and functional characterization of Gene Cluster Families (GCFs) across the \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eBeauveria\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e genus.\u003c/strong\u003e (A) Distribution of GCF sizes across major chemical classes. The vertical axis (log10 scale) represents the number of BGCs grouped into each family at a 0.60 similarity cutoff. (B) Size distribution specifically for GCFs that contain hits from the MIBiG database, representing experimentally characterized biosynthetic pathways. In both panels, violin plots illustrate the density of families, while inner box plots indicate the median and interquartile range (IQR). Individual data points represent distinct GCFs.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/eb1ea26f2258f278bc68c1b8.png"},{"id":109051505,"identity":"93a7cf6e-1ed8-4847-950a-2d53c394070d","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":132469,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePower Law scaling and genomic fluidity of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eB. bassiana\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e Gene Cluster Families (GCFs).\u003c/strong\u003e Power Law scaling and genomic fluidity of B. bassiana Gene Cluster Families (GCFs). The log-log plot displays the Complementary Cumulative Distribution Function (CCDF), defined as P(X≥x), which represents the probability that a given GCF is shared by at least x number of isolates. The empirical data (open circles) were modeled using a Power Law distribution (solid red line) characterized by a scaling exponent α=5.41 and a lower threshold x\u003csub\u003emin\u003c/sub\u003e = 71.. A Log-Normal fit (dashed blue line) was overlaid for comparison. A Vuong likelihood ratio test (LR=−0.61,p=0.543) indicates that neither model provides a significantly better fit, suggesting the distribution of GCFs follows a heavy-tailed scaling behavior across the 199 \u003cem\u003eB. bassiana\u003c/em\u003e isolates.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/6caae8b5d3a6f1b118518a22.png"},{"id":109051504,"identity":"988912c9-3398-4d7f-8a25-072fb2bd3bb7","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":238698,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eProtein-Protein Interaction (PPI) Networks of Conserved Biosynthetic Gene Clusters (BGCs). \u003c/strong\u003eFunctional interaction landscapes for the six most prevalent Gene Cluster Families (GCFs) in the \u003cem\u003eBeauveria\u003c/em\u003e genus, centered on the \u003cem\u003eB. bassiana\u003c/em\u003e AS272 proteome. (A–F) Networks represent individual hub GCFs containing at least 90 members. Central hub nodes (circles with solid borders) represent biosynthetic enzymes coded by BGCs and cluster-defined proteins, while peripheral connector nodes represent predicted interacting partners. Node coloring indicates pathogenicity associations derived from PHI-base cross-referencing: red nodes are BGC-encoded virulence factors; green nodes are interacting partners with known roles in pathogen-host interactions; and white/orange nodes represent proteins with no characterized PHI-base homology. Line thickness (edges) corresponds to the STRING database confidence score for the interaction (\u0026gt;800).\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/6dd7105e65dcdbd90685f4d4.png"},{"id":109081432,"identity":"97b45c20-9b07-4a00-8c19-dfce27908776","added_by":"auto","created_at":"2026-05-12 12:18:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1841984,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/83e1c87c-f59e-42eb-b2ca-fc1a5623f4c5.pdf"},{"id":109051497,"identity":"f73de95e-f8b2-437a-b9c2-4080b43d504f","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"csv","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":19307,"visible":true,"origin":"","legend":"\u003cp\u003eTable S1: Detailed list of \u003cem\u003eBeauveria\u003c/em\u003e and \u003cem\u003eMetarhizium\u003c/em\u003egenomes used in this study, including NCBI accession numbers and assembly completeness metrics based on BUSCO analysis.\u003c/p\u003e","description":"","filename":"TableS1.csv","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/7730274cb51e2afb07a1b4ef.csv"},{"id":109081190,"identity":"3cfc4e06-bd3e-46fa-94a7-ed28f22ac2f4","added_by":"auto","created_at":"2026-05-12 12:05:18","extension":"csv","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":732,"visible":true,"origin":"","legend":"\u003cp\u003eTable S2: Summary of Biosynthetic Gene Cluster Families (GCFs) identified across the Beauveria genus, including taxonomic distribution and similarity to known biosynthetic pathways (MIBiG).\u003c/p\u003e","description":"","filename":"TableS2.csv","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/1caaee267ea6c2ff20431794.csv"},{"id":109067723,"identity":"f02e8898-42f5-4223-9c45-0a82b1aa274c","added_by":"auto","created_at":"2026-05-12 10:00:17","extension":"csv","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":614,"visible":true,"origin":"","legend":"\u003cp\u003eTable S3: Topological network metrics and Pathogen-Host Interaction (PHI-base) gene enrichment analysis for selected Beauveria Gene Cluster Families.\u003c/p\u003e","description":"","filename":"TableS3.csv","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/4a3ee01dabddce7f754e4054.csv"},{"id":109051500,"identity":"b52c6b8a-7752-4be4-b760-8f66a1655da4","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"sh","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":8304,"visible":true,"origin":"","legend":"","description":"","filename":"FileS1.sh","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/d491cf3fe9405f19e1049155.sh"},{"id":109051501,"identity":"345141c5-5220-48c2-a0bd-b77315689086","added_by":"auto","created_at":"2026-05-12 06:52:46","extension":"r","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":39064,"visible":true,"origin":"","legend":"","description":"","filename":"FileS2.r","url":"https://assets-eu.researchsquare.com/files/rs-9203165/v1/f3dacea4cd9f9d10692bbd1c.r"}],"financialInterests":"No competing interests reported.","formattedTitle":"Genome-wide exploration of biosynthetic gene clusters and their association to virulence in the entomopathogenic fungus Beauveria","fulltext":[{"header":"Introduction","content":"\u003cp\u003e \u003cem\u003eBeauveria\u003c/em\u003e is a genus of entomopathogenic fungi widely recognized for its use in the biological control of arthropod pests (de Sousa et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Members of this genus infect their hosts using a common mechanism involving adhesion and germination of conidia on the insect cuticle, followed by penetration, internal colonization, and host death (Hong et al. \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). Compared to conventional chemical insecticides, \u003cem\u003eBeauveria\u003c/em\u003e species exhibit a narrower host range, lower environmental impact, and are considered safe for non-target organisms, including humans, animals, and plants (Mascarin and Jaronski \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). These features support the development and commercialization of various \u003cem\u003eBeauveria\u003c/em\u003e-based biopesticide formulations, which are applied in integrated pest management (IPM) strategies across different agricultural systems (Andreata et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). In Brazil, by the beginning of 2026, a total of 160 biocontrol products containing \u003cem\u003eBeauveria\u003c/em\u003e sp. had been registered, highlighting its effectiveness in controlling insect pests (Minist\u0026eacute;rio da Agricultura e Pecu\u0026aacute;ria).\u003c/p\u003e \u003cp\u003eThe commercial importance of the genus \u003cem\u003eBeauveria\u003c/em\u003e reflects its high pathogenic potential, which is supported by the diversity of molecules produced during the infection cycle, such as insecticidal toxins, immunosuppressive compounds, and antimicrobial metabolites (Oberti et al. \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2025\u003c/span\u003e; Farag et al. \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2025\u003c/span\u003e; Aly et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Secondary metabolites are also produced during host infection (Chaithra et al. \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), including oosporein, bassianolide, and beauvericin, which play pivotal roles in virulence by promoting host immune suppression, tissue colonization, and insect mortality (Wang et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). These compounds belong to distinct classes of secondary metabolites, including polyketides and nonribosomal peptides or depsipeptides, which are encoded by well-defined biosynthetic gene clusters (BGCs) comprising core synthase genes, such as polyketide synthases (PKSs) and nonribosomal peptide synthetases (NRPSs), along with tailoring enzymes, transporters, and regulatory elements (Keller \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The profile and production of these compounds vary among species and strains, highlighting the genomic plasticity of the genus and its ability to adapt to different hosts and environments (Solano-Gonz\u0026aacute;lez et al. \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). In other fungal genera, comparative genomic analyses have shown that gene annotation and the prediction of biosynthetic gene clusters (BGCs) associated with secondary metabolites are broadly distributed among species, representing a promising approach for investigating molecular mechanisms of virulence based on genomic data (Valero-Jim\u0026eacute;nez et al. \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). Genomic plasticity contributes to diversification across species, yet core genes associated with pathogenicity, such as those involved in host adhesion, cuticle degradation, and stress tolerance, are commonly conserved among phylogenetically related species, highlighting their essential role in infective capacity and survival (Wang and Wang \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2017\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eDespite advances in genome sequencing technologies and data availability, the lack of standardized curation and robust functional annotation in public databases continues to limit the effective use of genomic information in comparative studies. Rigorous curation of genomic datasets, including reannotation and functional validation of genes, is therefore essential to ensure the reliability of analyses aimed at elucidating metabolic diversity, virulence determinants, and evolutionary processes, particularly in entomopathogenic fungi (Wang and Wang \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Here, we conducted a comprehensive phylogenomic and biosynthetic gene cluster (BGC) analysis of 333 high-quality \u003cem\u003eBeauveria\u003c/em\u003e genomes, including one newly sequenced by our group. Our approach integrates taxonomic profiling, BGC classification, and the functional association of secondary metabolite biosynthesis with virulence-related features in the newly sequenced strain. Together, our findings provide insights into the evolutionary conservation, diversification, and pathogenic potential encoded within the genomes of this agriculturally important fungal genus.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eGenome sequencing, assembly and annotation\u003c/h2\u003e \u003cp\u003e \u003cem\u003eB. bassiana\u003c/em\u003e AS272 strain was obtained from soil in the state of Rio Grande do Sul, Brazil. The isolation process used suspension of soil samples in sterile water, followed by plating on Oatmeal Agar-CTAB medium added of 50 mg/mL chloramphenicol. The pathogenicity of the isolates was confirmed by infection assays using the model arthropod \u003cem\u003eTenebrio molitor\u003c/em\u003e. The AS272 strain was maintained at 28\u0026deg;C on PDA medium, and identification was performed by sequencing the B locus intergenic region (Bloc) (accession number: PQ220121.1), primers \u0026ndash; B51-F 5\u0026prime;-CGACCCGCCAACTACTTTGA-3\u0026prime; and B31-R 5\u0026prime;-GTCTTCCAGTACCACTACGCC-3\u0026prime;. (Camargo, unpublished data). For DNA isolation, 3-day-old mycelia were used, following the CTAB method (Dar et al. \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2025\u003c/span\u003e) with a bead beater (FastPrep\u0026reg;-24 5G - MP Biomedicals). Sequencing was performed in a paired-end mode (2 x 150 bp) using an Illumina NextSeq 2000 after library preparation by LacTAD (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.lactad.unicamp.br/\u003c/span\u003e\u003cspan address=\"https://www.lactad.unicamp.br/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003ePrior to assembly, raw sequencing reads were subjected to quality filtering using fastp v 0.23.4 (Chen et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) under default parameters. \u003cem\u003eDe novo\u003c/em\u003e genome assembly was subsequently performed using SPAdes v 4.2.0 (Bankevich et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2012\u003c/span\u003e) under default parameters. To enhance assembly contiguity, the resulting contigs were scaffolded with RagTag v 2.1.0 (Alonge et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), utilizing the \u003cem\u003eBeauveria bassiana\u003c/em\u003e ARSEF 2860 reference genome (NCBI RefSeq: GCF_000280675.1) as a guide. The final genome sequence has been deposited in the NCBI database under the accession number GCA_052426205.1. The assembled and scaffolded genome was processed using Funannotate v 1.8.1 (Palmer and Stajich \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) to perform FASTA header standardization, sequence sorting, and repeat masking. Gene prediction was subsequently conducted utilizing BRAKER3 v 3.0.6 (Gabriel et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2024\u003c/span\u003e), with the \u003cem\u003eB. bassiana\u003c/em\u003e ARSEF 2860 reference proteome serving as the training dataset. Finally, the resulting GFF3 file was integrated back into the Funannotate pipeline to finalize the functional annotation. Secondary metabolite biosynthetic gene clusters (BGCs) were subsequently predicted from the BRAKER3 outputs (FASTA and GFF3 formats) utilizing antiSMASH v8.0.4 (Blin et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2025\u003c/span\u003e) configured in the fungal mode.\u003c/p\u003e \u003cp\u003e \u003cb\u003eBeauveria\u003c/b\u003e \u003cb\u003egenome sequences\u003c/b\u003e\u003c/p\u003e \u003cp\u003eA total of 333 assembled genomes, deposited as belonging to the genus \u003cem\u003eBeauveria\u003c/em\u003e were retrieved from the NCBI database using NCBI datasets v 18.6.0 (O\u0026rsquo;Leary et al. \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) and assessed for completeness using the BUSCO 5.8.2 tool (Benchmarking Universal Single-Copy Orthologs) (Manni et al. \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) with the orthologous genes dataset of hypocreales (version 12). For phylogenomic analysis, all genomes sequences were used, while for functional prediction, only those with completeness scores above 92% were subsequently processed for gene prediction and functional annotation as described above using the Funannotate 1.8.1 pipeline (Palmer and Stajich \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The full list of accession codes and respective BUSCO results are provided in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePhylogenomic analyses \u003c/h3\u003e\n\u003cp\u003ePhylogenomic analyses were performed using the BUSCO_phylogenomics pipeline (Waterhouse et al. \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/jamiemcg/BUSCO_phylogenomics\u003c/span\u003e\u003cspan address=\"https://github.com/jamiemcg/BUSCO_phylogenomics\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), which generates phylogenomic matrices based on single-copy orthologous genes. The pipeline uses output files from BUSCO (Manni et al. \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) to identify conserved single-copy genes, which are aligned with MAFFT (Katoh and Standley \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2013\u003c/span\u003e). Nucleotide sequences of each orthologous group were aligned and concatenated into a supermatrix, with one partition considered per gene. Phylogenetic inference was conducted using IQ-TREE 2.1.4-beta (Minh et al. \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), applying the TESTMERGE model to automatically merge partitions with compatible substitution models. Model selection was performed automatically by ModelFinder based on the Bayesian Information Criterion (BIC). Phylogenetic support was estimated using ultrafast bootstrap approximation (UFBoot) with 1,000 replicates. The final phylogenetic tree was constructed exclusively with nucleotide data and included all 333 genomes, with two \u003cem\u003eMetarhizium\u003c/em\u003e species used as the outgroup.\u003c/p\u003e\n\u003ch3\u003eSecondary metabolite gene cluster analysis\u003c/h3\u003e\n\u003cp\u003eGene cluster analysis was conducted using the local version of antiSMASH 8.0.4 (Blin et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). The classification of biosynthetic gene clusters (BGCs) into gene cluster families was carried out using BiG-SCAPE 2 (Biosynthetic Gene Similarity Clustering and Prospecting Engine) (Navarro-Mu\u0026ntilde;oz et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). BGCs were grouped based on sequence similarity across a range of cutoff thresholds from 0.1 to 0.9. Clustering at a cutoff value of 0.6 provided an appropriate resolution for distinguishing BGC families. BGCs could be matched to previously characterized clusters by integrating the BiG-SCAPE algorithm with the Minimum Information about a Biosynthetic Gene cluster (MIBiG) database (Zdouc et al. \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eVirulence Association Prediction\u003c/h3\u003e\n\u003cp\u003eTo investigate the potential association between BGCs and virulence, the predicted protein-coding genes within selected BGCs were queried against the PHI-base 4.18 (Urban et al. \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2025\u003c/span\u003e) database using a local BLASTP 2.12.0 (Camacho et al. \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). Hits showing more than 50% identity and derived from fungal species were considered to be associated with virulence. Based on these results, a protein\u0026ndash;protein interaction (PPI) network was constructed using the STRING 12.0 database (Szklarczyk et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) using the automatic annotation process in the STRING server (avilable at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://version-12-0.string-db.org/organism/STRG0A07FVB\u003c/span\u003e\u003cspan address=\"https://version-12-0.string-db.org/organism/STRG0A07FVB\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eMaterial avaliability\u003c/h3\u003e\n\u003cp\u003eThe codes used for shell processing, as well as for R analysis and figures plotting, are available as Supplementary Files 1 and 2, respectively. Intermediate and final processing files are available upon request.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eCharacterization of \u003cem\u003eB. bassiana\u003c/em\u003e AS272 genome sequence\u003c/p\u003e \u003cp\u003eTo characterize the genomic characteristics of the high-virulence \u003cem\u003eB. bassiana\u003c/em\u003e isolate AS272, a comparative genomics approach was employed, using the NCBI reference strain \u003cem\u003eB. bassiana\u003c/em\u003e ARSEF 2860 as a control. The assembly of isolate AS272 comprises 31,716,118 bp with a high level of completeness (96.7% complete BUSCOs). Braker3 prediction identified 9,778 protein-coding genes, leading to 10,937 predicted proteoforms, which is slightly higher compared to reference (10,364 predicted proteoforms). OrthoFinder analysis revealed a high degree of conservation, with 8,984 orthogroups shared between the two isolates, representing 99.5% of the total orthogroup repertoire. Notably, isolate AS272 exhibited a significantly higher percentage of genes assigned to orthogroups (98.1%; 10,730 genes) compared to the reference strain (95.7%; 9,921 genes). Furthermore, AS272 showed a marked reduction in unassigned \"singleton\" genes (207; 1.9%) relative to ARSEF 2860 (443; 4.3%), suggesting that the AS272 increased gene count is driven by the expansion of existing gene families (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA comparative analysis of gene family sizes identified significant quantitative differences between the two isolates across multiple orthogroups. While global orthogroup distribution remains relatively correlated (R\u0026thinsp;=\u0026thinsp;0.26, p\u0026thinsp;\u0026lt;\u0026thinsp;2.2e\u0026thinsp;\u0026minus;\u0026thinsp;16), specific expansions could be observed in the ARSEF 2860 strain. Notably, families OG0000000 and OG0000001 showed the most substantial expansion in the reference strain ARSEF 2860 compared to AS272. Functional annotation via CDD search revealed that OG0000000 is predominantly composed of proteins containing the hAT family transposase domain (Dimer_Tnp_hAT, pfam05699) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB), which are generally associated with genetic mobile elements, and often associated with genomic plasticity and structural variation. In contrast, OG0000001 is enriched with members of the MRP_assoc_pro superfamily (cl33195) and ABC transporter domains (cl38913), which are essential for the regulated efflux of metabolites and xenobiotics.\u003c/p\u003e \u003cp\u003eFunctional mapping via KEGG identified specific KEGG Orthologs (KO) with marked copy-number variation. The most significant expansion was observed in K08192 (an MFS transporter), which displayed a delta of 13 copies compared to the reference. This was followed by K22134 (MFS transporter), K08832 (serine-threonine kinase), and K05658 (ATP-Binding cassette). Conversely, a reduction was noted in K07375 (Tubulin beta), suggesting subtle structural or regulatory shifts relative to the reference strain (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC). Finally, isolate AS272 displays a robust increase of pathogenicity-related factors mapped against the PHI-base database using a stringent criteria (minimum of 50% identity over at least 50% of the length of the subject). Compared to ARSEF 2860, the strain AS272 exhibits a higher number of homologs linked to the \"reduced virulence\" and \"loss of pathogenicity\" categories, which points toward a potentially redundant virulence set of proteins (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD). These functional enrichments point toward specialized adaptations in nutrient transport and environmental sensing that likely support the infection phenotype of this soil-derived isolate.\u003c/p\u003e\n\u003ch3\u003ePhylogenomic analyses\u003c/h3\u003e\n\u003cp\u003eTo establish the precise taxonomic identity and evolutionary context of the high-virulence isolate \u003cem\u003eB. bassiana\u003c/em\u003e AS272, we performed a phylogenomic analysis using a maximum-likelihood approach. The phylogeny was reconstructed from a supermatrix of 247,446 bp derived from 204 genomic partitions, which represent 4,333 single-copy orthologs identified via BUSCO. The dataset included 335 sequences, comprising 333 \u003cem\u003eBeauveria\u003c/em\u003e genomes (queried from NCBI as of March 2023) and two \u003cem\u003eMetarhizium\u003c/em\u003e strains utilized as outgroups.\u003c/p\u003e \u003cp\u003eThe resulting fan phylogeny resolved the genus into several major clades (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). To clarify evolutionary distances within dense terminal radiations, branch lengths were visualized using a logarithmic square-root transformation. The most expansive lineage is the \u003cem\u003eB. bassiana\u003c/em\u003e clade, which contains the majority of investigated isolates. Highly supported sister clades, which are indicated by red internal branches (\u0026ge;\u0026thinsp;80% UFBoot), were clearly delineated for \u003cem\u003eB. pseudobassiana\u003c/em\u003e, \u003cem\u003eB. asiatica\u003c/em\u003e, \u003cem\u003eB. varroae\u003c/em\u003e, and \u003cem\u003eB. brongniartii\u003c/em\u003e. Additionally, a distinct clade was recovered pairing \u003cem\u003eB. australis\u003c/em\u003e with \u003cem\u003eB. medogensis\u003c/em\u003e. The tree remains rooted by a robust outgroup containing \u003cem\u003eM. robertsii\u003c/em\u003e and \u003cem\u003eM. anisopliae\u003c/em\u003e. Notably, \u003cem\u003eB. bassiana\u003c/em\u003e AS272 (GCA_052426205.1) was positioned within a well-supported subclade composed exclusively of \u003cem\u003eB. bassiana\u003c/em\u003e strains (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e - arrow). This placement confirms its taxonomic identity and definitively separates it from sister species within the complex, such as \u003cem\u003eB. pseudobassiana\u003c/em\u003e and \u003cem\u003eB. brongniartii\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eDespite the overall robustness of the topology, several phylogenetic discrepancies were identified. \u003cem\u003eB. bassiana\u003c/em\u003e MBC618 branched entirely outside the \u003cem\u003eBeauveria\u003c/em\u003e genus, clustering with the \u003cem\u003eMetarhizium\u003c/em\u003e outgroup with absolute support (100% UFBoot/100% SH-aLRT; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e Arrow A). This significant incongruence suggests a potential misidentification or genomic sequence contamination for this specific isolate. Furthermore, \u003cem\u003eB. rudraprayagi\u003c/em\u003e MTCC8017 was nested deeply within the \u003cem\u003eB. bassiana\u003c/em\u003e clade, grouping directly with \u003cem\u003eB. bassiana\u003c/em\u003e ARSEF 2597 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e - Arrow B). This finding supports the hypothesis that \u003cem\u003eB. rudraprayagi\u003c/em\u003e may represent a variant of \u003cem\u003eB. bassiana\u003c/em\u003e rather than a distinct phylogenetic species. Similarly, \u003cem\u003eB. felina\u003c/em\u003e SYSU-MS7908 appeared phylogenetically distant from all other represented taxa, occupying an isolated basal position (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e - Arrow C). Collectively, these discrepancies highlight the critical need for rigorous phylogenomic validation to address the incidence of misidentified genomes and cryptic diversity currently present in public repositories.\u003c/p\u003e \u003cp\u003e \u003cb\u003eConservation of BGCs among\u003c/b\u003e \u003cb\u003eBeauveria\u003c/b\u003e \u003cb\u003egenus analyses\u003c/b\u003e\u003c/p\u003e \u003cp\u003eWith the taxonomic identity and evolutionary relationships of \u003cem\u003eB. bassiana\u003c/em\u003e AS272 established, our focus shifted toward the functional genomic elements potentially associated with its virulence phenotype. Entomopathogenic fungi rely on a highly specialized set of secondary metabolites to facilitate host infection, evade insect immune responses, and compete in complex ecological niches. Therefore the specialized biosynthetic capacity of AS272 was inferred and compared to those observed in the broader evolutionary history of the genus. The \u003cem\u003eBeauveria\u003c/em\u003e genomes that displayed BUSCO values above 92% here analyzed were first annotated using fungiSMASH. A total of 14,602 BGCs were predicted across the 333 genomes, revealing a high degree of biosynthetic diversity and inter-specific variation within the genus (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The total number of BGCs per genome ranged from approximately 35 to 55, with \u003cem\u003eB. varroae\u003c/em\u003e and \u003cem\u003eB. pseudobassiana\u003c/em\u003e exhibiting the most expansive biosynthetic repertoires, maintaining medians above 45 clusters per genome. Conversely, considering the content of BGCs, the genomes of \u003cem\u003eB. asiatica, B. caledonica\u003c/em\u003e and \u003cem\u003eB. rudraprayagi\u003c/em\u003e appeared more compact (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe distribution of Biosynthetic Gene Clusters (BGCs) by chemical class highlights a clear dominance of Non-Ribosomal Peptide Synthetases (NRPS) and Polyketide Synthases (PKS) across all species (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB-\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eG). NRPS BGCs represent the most abundant category identified (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). Notably, \u003cem\u003eB. varroae\u003c/em\u003e exhibited a significant expansion in this class, with a median of approximately 24 clusters, suggesting a specialized reliance on peptide-based metabolites compared to the genus average. PKS clusters, while slightly less abundant, were maintained across species, typically ranging between 5 and 20 per genome (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). Interestingly, \u003cem\u003eB. pseudobassiana\u003c/em\u003e displayed the highest intra-species variability in PKS content, indicating a high degree of genomic plasticity and potential for niche-specific chemical signatures within this lineage.\u003c/p\u003e \u003cp\u003eIn contrast to the high inter-specific variation observed in NRPS and PKS counts, the abundance of Terpene (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB) and PKS-NRP Hybrid (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD) BGCs was notably consistent across the genus. These classes exhibited a highly uniform distribution, with the majority of species harboring a stable repertoire of 6\u0026ndash;9 terpene clusters and 2\u0026ndash;6 hybrid clusters per genome. Conversely, Ribosomally synthesized and Post-translationally modified Peptides (RiPPs) BGCs (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eG) were remarkably rare, often limited to one or two clusters or entirely absent, as seen in \u003cem\u003eB. asiatica\u003c/em\u003e. A unique expansion in the \"Others\" category (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eF) was also observed in \u003cem\u003eB. pseudobassiana\u003c/em\u003e, hinting at the presence of unclassified or novel biosynthetic pathways that remain to be characterized (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The profile of distinct BGCs in \u003cem\u003eB. bassiana\u003c/em\u003e AS272 follows the general distribution of BGCs in \u003cem\u003eBeauveria\u003c/em\u003e genomes, totaling 39 predicted clusters. While its genomic repertoire aligns closely with the species median for most classes, the isolate exhibits specific counts of 19 NRPS, 8 PKS, 7 Terpene, and 4 PKS-NRP Hybrid clusters (red dots in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo further investigate the evolutionary relationships between the predicted BGCs, we performed a systematic clustering analysis using BiG-SCAPE. By applying a range of sequence similarity cutoffs (0.10 to 0.90), we evaluated the conservation of BGCs into Gene Cluster Families (GCFs) across the \u003cem\u003eBeauveria\u003c/em\u003e genus. As expected, the number of identified GCFs decreased significantly as the clustering threshold became more stringent (higher cutoff), indicating that many clusters share common ancestral architectures (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). This is reflected in the number of \"singleton\" BGCs\u0026mdash;those unique to a single genome\u0026mdash;which decreased sharply as the cutoff moved toward more inclusive similarity levels (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe mean size of the GCFs (number of BGCs per family) showed an exponential increase at higher cutoffs, particularly for the NRPS and PKS classes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). This suggests that a substantial portion of the \u003cem\u003eBeauveria\u003c/em\u003e specialized metabolome consists of highly related biosynthetic pathways that likely produce similar core chemical scaffolds. At a selected cutoff of 0.60, which provides a balanced resolution between over-clustering and excessive fragmentation, the PKSI and Terpene classes exhibited the highest mean GCF sizes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD). This indicates that these specific classes are characterized by large, multi-member families that are widely distributed across the genus. Finally, the proportion of singleton clusters varied markedly by chemical class (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE). While PKSI, Terpene, and NRPS clusters showed a low percentage of singletons (typically\u0026thinsp;\u0026lt;\u0026thinsp;5%), suggesting they belong to well-established, genus-wide families, the \"Others\" and RiPPs categories were dominated by unique, strain-specific clusters (\u0026gt;\u0026thinsp;25% singletons). This high degree of uniqueness in unclassified clusters highlights a reservoir of genomic novelty that may drive specific ecological adaptations or host-specialized interactions in individual \u003cem\u003eBeauveria\u003c/em\u003e isolates, including the high-virulence AS272 strain.\u003c/p\u003e \u003cp\u003eThe distribution of Gene Cluster Families (GCFs) across the genus \u003cem\u003eBeauveria\u003c/em\u003e was further evaluated. The analysis of GCF size, which is defined by the number of BGCs per family, revealed distinct evolutionary patterns among chemical classes (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). The PKSI, Terpene, and NRPS classes exhibited the most expansive family sizes, with some GCFs containing over 100 individual BGCs. This indicates that a significant portion of the \u003cem\u003eBeauveria\u003c/em\u003e specialized metabolome is composed of highly related biosynthetic pathways shared by a vast majority of the 333 genomes. The largest median GCF sizes were observed in the PKSI and Terpene families, suggesting the significant importance of these specific conserved pathways to the distinct lifestyles of \u003cem\u003eBeauveria\u003c/em\u003e. In contrast, the PKS-NRP Hybrids, PKSother, RiPPs, and Others categories were characterized by significantly smaller GCFs, with medians frequently falling below five members. This suggests that these classes are more prone to diversification or are restricted to specific lineages. When filtering for GCFs containing known clusters from the MIBiG database, the mean GCF size dropped across all categories, although PKSI and NRPS retained the most robust presence. This drop highlights a substantial gap in the functional characterization of \u003cem\u003eBeauveria\u003c/em\u003e secondary metabolism, as the majority of large, well-conserved families currently lack a corresponding characterized metabolite in public databases.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo characterize the core biosynthetic potential of the genus \u003cem\u003eBeauveria\u003c/em\u003e, a conservation analysis was performed on Gene Cluster Families (GCFs) containing at least 100 members. This threshold was selected because such a wide distribution across species and strains suggests these clusters encode for specialized metabolites that are fundamental to \u003cem\u003eBeauveria\u003c/em\u003e biology. A total of 21 GCFs met this criterion, and their occupancy, defined as the percentage of species within a gene harboring the cluster, reveals distinct patterns of evolutionary stability and lineage-specific adaptation (Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). Several GCFs exhibited high occupancy across nearly the entire phylogeny, representing the stable biosynthetic framework of the genus. NRPS GCF 1963 and NRPS GCF 16320 emerged as two of the most ubiquitous pathways, with 100% occupancy across \u003cem\u003eB. caledonica, B. medogensis\u003c/em\u003e, and \u003cem\u003eB. rudraprayagi\u003c/em\u003e (GCF 1963) or \u003cem\u003eB. asiatica, B. australis\u003c/em\u003e, and \u003cem\u003eB. caledonica\u003c/em\u003e (GCF 16320). PKSI GCF 7963 was equally prominent, maintaining 100% occupancy in \u003cem\u003eB. asiatica, B. australis, B. medogensis\u003c/em\u003e, and \u003cem\u003eB. rudraprayagi.\u003c/em\u003e Terpene GCF 12284 demonstrated a unique distribution, present in 100% of \u003cem\u003eB. australis, B. brongniartii, B. caledonica, B. medogensis\u003c/em\u003e, and \u003cem\u003eB. rudraprayagi\u003c/em\u003e isolates, despite being nearly absent in \u003cem\u003eB. bassiana\u003c/em\u003e (2%) and \u003cem\u003eB. pseudobassiana\u003c/em\u003e (0%).\u003c/p\u003e \u003cp\u003eIn contrast to the ubiquitous core, several large GCFs served as definitive markers for specific clades. For \u003cem\u003eB. bassiana\u003c/em\u003e and \u003cem\u003eB. rudraprayagi\u003c/em\u003e, Terpene GCF 1989 showed high frequency in \u003cem\u003eB. bassiana\u003c/em\u003e (85.9%) and 100% occupancy in \u003cem\u003eB. rudraprayagi\u003c/em\u003e, but was absent in most other lineages. The sister species \u003cem\u003eB. pseudobassiana\u003c/em\u003e and \u003cem\u003eB. varroae\u003c/em\u003e share a unique biosynthetic profile, defined by 100% occupancy of NRPS GCF 16292, PKSI GCF 16319, and PKSother GCF 4094. Furthermore, NRPS GCF 15871 and NRPS GCF 19 showed nearly universal presence in these two clades (\u0026gt;\u0026thinsp;95%) while remaining restricted from the rest of the genus.\u003c/p\u003e \u003cp\u003eThe conservation analysis also highlighted GCFs with hits in the MIBiG database, providing insight into the chemical nature of these core clusters. Terpene GCF 1989 (linked to BGC0001839.1) and PKSI GCF 7963 (linked to BGC0001720.1) represent well-characterized pathways widely distributed across the genus. Notably, NRPS GCF 19 was associated with multiple MIBiG entries (BGC0000959.1, BGC0001136.1, and BGC0002035.1), suggesting it may represent a cluster family responsible for a known class of bioactive compounds like beauvericin or bassianolide.\u003c/p\u003e \u003cp\u003eWithin the \u003cem\u003eB. bassiana\u003c/em\u003e complex, including the isolate AS272, occupancy for these 21 large GCFs typically fluctuated between 40% and 60%. This moderate frequency reflects the significant intra-specific genomic plasticity of the lineage. To characterize the species-specific biosynthetic landscape, we first attempted to identify a strictly conserved set of Gene Cluster Families (GCFs) within \u003cem\u003eB. bassiana\u003c/em\u003e. Despite being the most abundant species in the dataset, comprising 199 of the 333 genomes analyzed, the results revealed that \u003cem\u003eB. bassiana\u003c/em\u003e possesses zero strictly core GCFs when defined by 100% occupancy across all isolates.\u003c/p\u003e \u003cp\u003eGiven this observation, we evaluated the distribution of biosynthetic gene cluster families (GCFs) exclusively within the \u003cem\u003eB. bassiana\u003c/em\u003e lineage using discrete power-law and log-normal modeling to investigate how these clusters are maintained across isolates. The analysis revealed a highly skewed distribution, with substantial variance in GCF occupancy among isolates (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). However, model comparison using the Vuong likelihood ratio test did not support a significantly better fit of the power-law model over the log-normal alternative (LR\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.61, p\u0026thinsp;=\u0026thinsp;0.543). Although a power-law fit yielded an estimated scaling exponent of α\u0026thinsp;=\u0026thinsp;5.41 (xmin\u0026thinsp;=\u0026thinsp;71), the absence of statistical support indicates that the observed pattern cannot be robustly classified as strictly scale-free. Instead, the distribution is consistently heavy-tailed, in which a small subset of GCFs is widely conserved, while most clusters occur at lower frequencies across isolates (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWhile no strictly core clusters were found, the modeling identified some GCFs, which represent high-frequency clusters that could form the functional backbone of the species' chemical repertoire without being universally present in every strain. The occupancy data for these hubs further illustrates this flexibility. The most prevalent family, Terpene GCF 1989, reached only 85.9% occupancy, followed by Terpene GCF 8171 (67.3%) and NRPS GCF 8144 (59.3%). Other major families such as PKSother 12262 (57.8%) and PKSI 8173 (56.8%) also displayed moderate to high frequency but remained absent in a significant portion of the population. Within the \u003cem\u003eB. bassiana\u003c/em\u003e complex, including the high-virulence isolate AS272, occupancy for these 21 large GCFs typically fluctuated between 40% and 60%. This moderate frequency, compared to the strict 100% occupancy seen in smaller clades, reflects the significant intra-specific genomic plasticity of the lineage. This suggests that while \u003cem\u003eB. bassiana\u003c/em\u003e maintains a core set of BGCs, it utilizes a more variable set of high-frequency clusters to potentially tailor its pathogenic profile to specific hosts or environmental niches.\u003c/p\u003e \u003cp\u003e \u003cb\u003eAssociation of\u003c/b\u003e \u003cb\u003eB. bassiana\u003c/b\u003e \u003cb\u003eAS272 BGCs with Virulence\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTo investigate the functional context of conserved biosynthetic gene clusters and their potential association with pathogenicity, we performed a protein\u0026ndash;protein interaction (PPI) network analysis. Interaction networks were constructed for the core genes of six highly conserved hub GCFs using the STRING database (version 12.0), based on the \u003cem\u003eB. bassiana\u003c/em\u003e AS272 annotation. For this analysis, we selected GCFs with high prevalence across the genus (\u0026ge;\u0026thinsp;90 members), representing five biosynthetic classes: Terpene, NRPS, PKS\u0026ndash;NRP hybrids, PKS-like, and T1PKS.\u003c/p\u003e \u003cp\u003eTopological analysis indicated that these GCFs are embedded within highly interconnected interaction networks (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e). The resulting networks displayed substantial structural complexity, with an average node degree ranging from 15.23 (Terpene GCF 1989) to 43.75 (NRPS GCF 8144). A notable feature of these networks is their high clustering coefficients, ranging from 0.58 to 0.95, which exceeded expectations for comparable random networks. Correspondingly, clustering enrichment values ranged from 1.55 to 2.85, indicating a dense and modular organization in which biosynthetic core genes are connected with broader cellular interaction networks (Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo assess potential associations with pathogenicity, network proteins were cross-referenced against the PHI-base database to identify homologs of experimentally characterized pathogen\u0026ndash;host interaction factors. Across the six networks, the number of PHI-related nodes ranged from 8 to 18 (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e). Statistical enrichment analysis indicated that these associations are unlikely to arise by chance. In particular, the NRPS (GCF 8144) and PKS\u0026ndash;NRP hybrid (GCF 12259) networks showed the strongest enrichment for known virulence-associated proteins ( FDR-adjusted p\u0026thinsp;=\u0026thinsp;1.90 \u0026times; 10⁻⁶). The PKSother network (GCF 12262) also exhibited significant enrichment (p\u0026thinsp;=\u0026thinsp;6.40 \u0026times; 10⁻⁵). Although the Terpene and T1PKS networks contained multiple PHI-base homologs (14 and 18 nodes, respectively), their enrichment did not reach the same level of statistical significance (FDR\u0026thinsp;\u0026gt;\u0026thinsp;0.05). This pattern may reflect interactions with a broader or less well-characterized set of host-associated factors. Taken together, these results suggest that several conserved biosynthetic gene clusters in \u003cem\u003eBeauveria\u003c/em\u003e, particularly NRPS and hybrid PKS\u0026ndash;NRP systems, are embedded within interaction networks enriched for proteins associated with host interaction and pathogenicity (Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eAnalysis of the \u003cem\u003eB. bassiana\u003c/em\u003e isolate AS272 reveals a genomic architecture geared toward both metabolic versatility and pathogenic potential. While some studies emphasize the contribution of lineage-specific genes (singletons) to virulence, our results suggest that the aggressiveness of this isolate may instead be associated with an expansion of conserved gene families. The relatively low proportion of orphan genes compared to orthogroups indicates that the AS272 genome has likely evolved through gene duplication events, leading to increased functional redundancy and potential metabolic flexibility.\u003c/p\u003e\n\u003cp\u003eA prominent feature of the genome is the expansion of transporters belonging to the Major Facilitator Superfamily (MFS) and ATP-binding cassette (ABC) superfamilies. These transporters are known to play key roles in toxin secretion during infection and may also contribute to protection against insect-derived defense metabolites and environmental stressors (Valero-Jim\u0026eacute;nez et al. 2016). In AS272, the increased copy number of genes such as K08192 suggests a potentially enhanced efflux capacity, which may facilitate survival and colonization within host environments (Solano-Gonz\u0026aacute;lez et al. 2023). \u003c/p\u003e\n\u003cp\u003eFrom a phylogenetic perspective, the phylogenomic reconstruction of 337 genomes revealed persistent misidentifications in public databases. The clustering of strain MBC618 within the \u003cem\u003eMetarhizium\u003c/em\u003e genus, together with the isolated position of \u003cem\u003eB. felina\u003c/em\u003e, highlights the limitations of traditional taxonomy based solely on morphology or single-locus markers such as the ITS region. Furthermore, the placement of \u003cem\u003eB. rudraprayagi\u003c/em\u003e within the \u003cem\u003eB. bassiana\u003c/em\u003e clade suggests that it may represent a variant or closely related lineage of \u003cem\u003eB. bassiana\u003c/em\u003e, supporting previous studies that have emphasized the need for systematic revision within the \u003cem\u003eBeauveria\u003c/em\u003e genus (Rehner et al. 2011). The use of single-copy orthologs identified through the BUSCO framework (Manni et al. 2021) provides a robust phylogenomic basis for species identification and for the accurate documentation of strains used in industrial or biotechnological applications.\u003c/p\u003e\n\u003cp\u003eAnother notable finding is the apparent absence of a strictly conserved core of biosynthetic gene clusters (BGCs) within \u003cem\u003eB. bassiana\u003c/em\u003e. Instead, the distribution of these clusters follows a highly dynamic pattern, potentially enabling the fungus to maintain a flexible chemical arsenal and considerable genomic plasticity (Valero-Jim\u0026eacute;nez et al. 2016). While certain classes, such as terpene clusters, show relatively high conservation\u0026mdash;suggesting roles in core biological processes\u0026mdash;NRPS and PKS clusters appear to function primarily as sources of chemical innovation. Such intraspecific variation may explain the substantial differences in virulence observed among \u003cem\u003eB. bassiana\u003c/em\u003e strains infecting the same host species, as individual strains may deploy distinct combinations of secondary metabolites derived from diverse biosynthetic hubs (Toopaang et al. 2025). \u003c/p\u003e\n\u003cp\u003eProtein\u0026ndash;protein interaction (PPI) network analysis further suggests that secondary metabolite biosynthesis is integrated within broader pathogenicity-related networks. Cross-referencing these networks with PHI-base identified several connections between core BGC proteins and previously described virulence-associated factors, indicating potential coordination between toxin biosynthesis and infection-related processes such as cuticle degradation. Although compounds such as beauvericin and oosporein are known to play important roles in suppressing host immune responses, most of the identified GCFs lacked direct counterparts when compared against the MIBiG database (Zdouc et al. 2025).\u003c/p\u003e\n\u003cp\u003eThe dynamic distribution of BGCs observed across \u003cem\u003eB. bassiana\u003c/em\u003e isolates likely reflects the rapid evolutionary turnover typical of fungal secondary metabolite pathways. Biosynthetic clusters are frequently shaped by gene duplication, recombination, horizontal transfer, and lineage-specific gene loss, processes that collectively generate substantial chemical diversity within species. Such evolutionary plasticity allows entomopathogenic fungi to adapt to diverse ecological niches and host environments by modifying or expanding their repertoire of bioactive compounds. In the case of \u003cem\u003eBeauveria\u003c/em\u003e, this flexible biosynthetic architecture may provide a selective advantage during host infection, where different insect species present distinct physiological defenses. Consequently, the diversification and differential retention of BGCs may represent an important evolutionary strategy enabling \u003cem\u003eB. bassiana\u003c/em\u003e populations to maintain pathogenic versatility across a broad range of insect hosts.\u003c/p\u003e\n\u003cp\u003eThe absence of clear matches in MIBiG underscores the largely unexplored biosynthetic diversity within the \u003cem\u003eBeauveria \u003c/em\u003egenus. This hidden repertoire of specialized metabolites may represent a substantial source of novel bioactive compounds with potential applications in biotechnology, agriculture, and pharmaceutical discovery, warranting further functional and chemical characterization.\u003c/p\u003e\n"},{"header":"Statements and Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding Declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by grants from the Brazilian funding agencies Conselho Nacional de Desenvolvimento Cient\u0026iacute;fico e Tecnol\u0026oacute;gico (CNPq - 312994/2021-4 and 405015/2023-2), Funda\u0026ccedil;\u0026atilde;o de Amparo \u0026agrave; Pesquisa do Estado do Rio Grande do Sul (FAPERGS - 19/2551-0001708-1), and Coordena\u0026ccedil;\u0026atilde;o de Aperfei\u0026ccedil;oamento de Pessoal de N\u0026iacute;vel Superior (CAPES). RSS and AAR was a recipient of a CAPES scholarships. HA and MSC were recipients of CNPq scholarships.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest:\u0026nbsp;\u003c/strong\u003eThe authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Approval:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eContributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAll authors\u003c/strong\u003e contributed to the study conception and design. \u003cstrong\u003eRobson dos Santos Soares, Matheus Camargo, Maria Eduarda Deluca Jo\u0026atilde;o, and Charley Christian Staats\u003c/strong\u003e performed material preparation and data collection. \u003cstrong\u003eRobson dos Santos Soares, Alexandra Rocha, Henrique Antoniolli, and Charley Christian Staats\u003c/strong\u003e performed the data analysis. \u003cstrong\u003eAugusto Schrank and Charley Christian Staats\u003c/strong\u003e were responsible for project administration and funding acquisition. The first draft of the manuscript was written by \u003cstrong\u003eRobson dos Santos Soares\u003c/strong\u003e, and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAlonge M, Lebeigle L, Kirsche M, et al (2022) Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing. Genome Biol 23:258. https://doi.org/10.1186/s13059-022-02823-7\u003c/li\u003e\n\u003cli\u003eAly HH, Meng Y, Wang D (2025) Comparative gene expression analysis of \u003cem\u003eBeauveria bassiana\u003c/em\u003e against Spodoptera frugiperda. PeerJ 13:e19591. https://doi.org/10.7717/peerj.19591\u003c/li\u003e\n\u003cli\u003eAndreata MF de L, Mian S, Andrade G, et al (2025) The current increase and future perspectives of the microbial pesticides market in agriculture: the Brazilian example. Front Microbiol 16:1574269. https://doi.org/10.3389/fmicb.2025.1574269\u003c/li\u003e\n\u003cli\u003eBankevich A, Nurk S, Antipov D, et al (2012) SPAdes: A new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol 19:455\u0026ndash;477. https://doi.org/10.1089/cmb.2012.0021\u003c/li\u003e\n\u003cli\u003eBlin K, Shaw S, Vader L, et al (2025) antiSMASH 8.0: extended gene cluster detection capabilities and analyses of chemistry, enzymology, and regulation. Nucleic Acids Res 53:W32\u0026ndash;W38. https://doi.org/10.1093/nar/gkaf334\u003c/li\u003e\n\u003cli\u003eCamacho C, Coulouris G, Avagyan V, et al (2009) BLAST+: architecture and applications. BMC Bioinformatics 10:421. https://doi.org/10.1186/1471-2105-10-421\u003c/li\u003e\n\u003cli\u003eChaithra M, Prameeladevi T, Prasad L, et al (2024) Metabolomic profiling of virulent and non-virulent \u003cem\u003eBeauveria bassiana\u003c/em\u003e strains: insights into the pathogenicity of Tetranychus truncatus. Arch Microbiol 206:311. https://doi.org/10.1007/s00203-024-04046-9\u003c/li\u003e\n\u003cli\u003eChen S, Zhou Y, Chen Y, Gu J (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34:i884\u0026ndash;i890. https://doi.org/10.1093/bioinformatics/bty560\u003c/li\u003e\n\u003cli\u003eDar GJ, Nazir R, Wani SA, et al (2025) Optimizing a modified cetyltrimethylammonium bromide protocol for fungal DNA extraction: Insights from multilocus gene amplification. Open Life Sci 20:20221006. https://doi.org/10.1515/biol-2022-1006\u003c/li\u003e\n\u003cli\u003ede Sousa TK, Silva AT da, Soares FE de F (2025) Fungi-Based Bioproducts: A Review in the Context of One Health. Pathogens 14:. https://doi.org/10.3390/pathogens14050463\u003c/li\u003e\n\u003cli\u003eFarag PF, Elsisi AA, Elabd EW, et al (2025) Prediction of secreted uncharacterized protein structures from \u003cem\u003eBeauveria bassiana\u003c/em\u003e ARSEF 2860 unravels novel toxins-like families. Sci Rep 15:17747. https://doi.org/10.1038/s41598-025-02618-3\u003c/li\u003e\n\u003cli\u003eGabriel L, Brůna T, Hoff KJ, et al (2024) BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA. Genome Res 34:769\u0026ndash;777. https://doi.org/10.1101/gr.278090.123\u003c/li\u003e\n\u003cli\u003eHong S, Shang J, Sun Y, et al (2024) Fungal infection of insects: molecular insights and prospects. Trends Microbiol 32:302\u0026ndash;316. https://doi.org/10.1016/j.tim.2023.09.005\u003c/li\u003e\n\u003cli\u003eKatoh K, Standley DM (2013) MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol 30:772\u0026ndash;780. https://doi.org/10.1093/molbev/mst010\u003c/li\u003e\n\u003cli\u003eKeller NP (2019) Fungal secondary metabolism: regulation, function and drug discovery. Nat Rev Microbiol 17:167\u0026ndash;180. https://doi.org/10.1038/s41579-018-0121-1\u003c/li\u003e\n\u003cli\u003eManni M, Berkeley MR, Seppey M, et al (2021) BUSCO Update: Novel and Streamlined Workflows along with Broader and Deeper Phylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes. Mol Biol Evol 38:4647\u0026ndash;4654. https://doi.org/10.1093/molbev/msab199\u003c/li\u003e\n\u003cli\u003eMascarin GM, Jaronski ST (2016) The production and uses of \u003cem\u003eBeauveria bassiana\u003c/em\u003e as a microbial insecticide. World J Microbiol Biotechnol 32:177. https://doi.org/10.1007/s11274-016-2131-3\u003c/li\u003e\n\u003cli\u003eMinh BQ, Schmidt HA, Chernomor O, et al (2020) IQ-TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Mol Biol Evol 37:1530\u0026ndash;1534. https://doi.org/10.1093/molbev/msaa015\u003c/li\u003e\n\u003cli\u003eMinist\u0026eacute;rio da Agricultura e Pecu\u0026aacute;ria Agrofit - Sistema de Agrot\u0026oacute;xicos Fitossanit\u0026aacute;rios. https://agrofit.agricultura.gov.br/agrofit_cons/principal_agrofit_cons. Accessed 19 Feb 2026\u003c/li\u003e\n\u003cli\u003eNavarro-Mu\u0026ntilde;oz JC, Selem-Mojica N, Mullowney MW, et al (2020) A computational framework to explore large-scale biosynthetic diversity. Nat Chem Biol 16:60\u0026ndash;68. https://doi.org/10.1038/s41589-019-0400-9\u003c/li\u003e\n\u003cli\u003eO\u0026rsquo;Leary NA, Cox E, Holmes JB, et al (2024) Exploring and retrieving sequence and metadata for species across the tree of life with NCBI Datasets. Sci Data 11:732. https://doi.org/10.1038/s41597-024-03571-y\u003c/li\u003e\n\u003cli\u003eOberti H, Sessa L, Oliveira-Rizzo C, et al (2025) Novel genomic features in entomopathogenic fungus \u003cem\u003eBeauveria bassiana \u003c/em\u003eILB308: accessory genomic regions and putative virulence genes involved in the infection process of soybean pest Piezodorus guildinii. Pest Manag Sci 81:2323\u0026ndash;2336. https://doi.org/10.1002/ps.8631\u003c/li\u003e\n\u003cli\u003ePalmer JM, Stajich J (2020) Funannotate v1.8.1: Eukaryotic genome annotation. Zenodo. https://doi.org/10.5281/zenodo.4054262\u003c/li\u003e\n\u003cli\u003eRehner SA, Minnis AM, Sung G-H, et al (2011) Phylogeny and systematics of the anamorphic, entomopathogenic genus \u003cem\u003eBeauveria\u003c/em\u003e. Mycologia 103:1055\u0026ndash;1073. https://doi.org/10.3852/10-302\u003c/li\u003e\n\u003cli\u003eSolano-Gonz\u0026aacute;lez S, Castro-V\u0026aacute;squez R, Molina-Bravo R (2023) Genomic Characterization and Functional Description of \u003cem\u003eBeauveria bassiana\u003c/em\u003e Isolates from Latin America. J Fungi (Basel) 9:. https://doi.org/10.3390/jof9070711\u003c/li\u003e\n\u003cli\u003eSzklarczyk D, Kirsch R, Koutrouli M, et al (2023) The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res 51:D638\u0026ndash;D646. https://doi.org/10.1093/nar/gkac1000\u003c/li\u003e\n\u003cli\u003eToopaang W, Yoocha T, Naktang C, et al (2025) Transcriptomic insights into the interplay between polyketide biosynthesis and other secondary metabolite biosynthetic clusters and biological pathways in entomopathogen \u003cem\u003eBeauveria bassiana\u003c/em\u003e. Front Microbiol 16:1583637. https://doi.org/10.3389/fmicb.2025.1583637\u003c/li\u003e\n\u003cli\u003eUrban M, Cuzick A, Seager J, et al (2025) PHI-base - the multi-species pathogen-host interaction database in 2025. Nucleic Acids Res 53:D826\u0026ndash;D838. https://doi.org/10.1093/nar/gkae1084\u003c/li\u003e\n\u003cli\u003eValero-Jim\u0026eacute;nez CA, Faino L, Spring In\u0026rsquo;t Veld D, et al (2016) Comparative genomics of \u003cem\u003eBeauveria bassiana\u003c/em\u003e: uncovering signatures of virulence against mosquitoes. BMC Genomics 17:986. https://doi.org/10.1186/s12864-016-3339-1\u003c/li\u003e\n\u003cli\u003eWang C, Wang S (2017) Insect pathogenic fungi: genomics, molecular interactions, and genetic improvements. Annu Rev Entomol 62:73\u0026ndash;90. https://doi.org/10.1146/annurev-ento-031616-035509\u003c/li\u003e\n\u003cli\u003eWang H, Peng H, Li W, et al (2021) The Toxins of \u003cem\u003eBeauveria bassiana\u003c/em\u003e and the Strategies to Improve Their Virulence to Insects. Front Microbiol 12:705343. https://doi.org/10.3389/fmicb.2021.705343\u003c/li\u003e\n\u003cli\u003eWaterhouse RM, Seppey M, Sim\u0026atilde;o FA, et al (2018) BUSCO applications from quality assessments to gene prediction and phylogenomics. Mol Biol Evol 35:543\u0026ndash;548. https://doi.org/10.1093/molbev/msx319\u003c/li\u003e\n\u003cli\u003eZdouc MM, Blin K, Louwen NLL, et al (2025) MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration. Nucleic Acids Res 53:D678\u0026ndash;D690. https://doi.org/10.1093/nar/gkae1115\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"functional-and-integrative-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"fige","sideBox":"Learn more about [Functional \u0026 Integrative Genomics](http://link.springer.com/journal/10142)","snPcode":"10142","submissionUrl":"https://submission.nature.com/new-submission/10142/3","title":"Functional \u0026 Integrative Genomics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Beauveria, Phylogenomics, Biosynthetic gene clusters (BGCs), Virulence factors, Secondary metabolites, Comparative genomics","lastPublishedDoi":"10.21203/rs.3.rs-9203165/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9203165/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSpecies of the genus \u003cem\u003eBeauveria\u003c/em\u003e are widely used as biological control agents due to their ability to infect and kill a broad range of arthropod pests. Despite their agricultural importance, the genomic diversity underlying virulence and secondary metabolism across the genus remains uncharacterized. In this study, we performed a comprehensive phylogenomic and biosynthetic gene cluster (BGC) analysis of 333 high-quality \u003cem\u003eBeauveria\u003c/em\u003e genomes, including one newly sequenced \u003cem\u003eBeauveria bassiana\u003c/em\u003e isolate (AS272) obtained from soil in southern Brazil. The genome of AS272 comprises 31.7 Mb with 9,778 predicted genes and high completeness (96.7% BUSCO). Comparative analysis revealed extensive conservation of orthogroups but indicated that the increased gene repertoire of AS272 is primarily associated with expansion of existing gene families, particularly transporters belonging to the major facilitator superfamily (MFS) and ATP-binding cassette (ABC) families. Phylogenomic reconstruction based on single-copy orthologs resolved major species clades within the genus and identified potential misidentified genomes in public databases. Genome mining using antiSMASH identified 14,602 BGCs across the dataset, with NRPS and PKS clusters dominating the biosynthetic landscape. Clustering with BiG-SCAPE revealed a mixture of highly conserved gene cluster families (GCFs) and numerous lineage-specific clusters, highlighting both evolutionary stability and diversification of secondary metabolism within \u003cem\u003eBeauveria\u003c/em\u003e. Notably, \u003cem\u003eB. bassiana\u003c/em\u003e lacked strictly conserved core GCFs across all isolates, instead exhibiting a heavy-tailed distribution of cluster frequencies consistent with substantial intraspecific genomic plasticity. Protein\u0026ndash;protein interaction network analysis further indicated that several conserved BGCs are embedded within networks enriched for proteins associated with host\u0026ndash;pathogen interactions. Together, these findings reveal a dynamic genomic architecture underlying secondary metabolism and pathogenic potential in \u003cem\u003eBeauveria\u003c/em\u003e, providing a comparative framework for understanding virulence evolution and identifying candidate pathways for future functional and biotechnological studies.\u003c/p\u003e","manuscriptTitle":"Genome-wide exploration of biosynthetic gene clusters and their association to virulence in the entomopathogenic fungus Beauveria","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-12 06:52:41","doi":"10.21203/rs.3.rs-9203165/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"292846645665406846618701857307551363390","date":"2026-05-07T15:33:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"242449554132320753586093085027248417917","date":"2026-05-05T06:47:16+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"62653822817705920692429440922566127220","date":"2026-05-04T11:13:09+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-05-04T08:02:47+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-27T12:32:45+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-27T12:32:25+00:00","index":"","fulltext":""},{"type":"submitted","content":"Functional \u0026 Integrative Genomics","date":"2026-03-23T16:46:48+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"functional-and-integrative-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"fige","sideBox":"Learn more about [Functional \u0026 Integrative Genomics](http://link.springer.com/journal/10142)","snPcode":"10142","submissionUrl":"https://submission.nature.com/new-submission/10142/3","title":"Functional \u0026 Integrative Genomics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"c3da5ba5-b54f-42e2-af33-7fb0d1736196","owner":[],"postedDate":"May 12th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"292846645665406846618701857307551363390","date":"2026-05-07T15:33:01+00:00","index":26,"fulltext":""},{"type":"reviewerAgreed","content":"242449554132320753586093085027248417917","date":"2026-05-05T06:47:16+00:00","index":23,"fulltext":""},{"type":"reviewerAgreed","content":"62653822817705920692429440922566127220","date":"2026-05-04T11:13:09+00:00","index":22,"fulltext":""},{"type":"reviewersInvited","content":"10","date":"2026-05-04T08:02:47+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-12T06:52:41+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-12 06:52:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9203165","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9203165","identity":"rs-9203165","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00