Aspergillus flavus pangenome (AflaPan) uncovers novel aflatoxin and secondary metabolite associated gene clusters | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Aspergillus flavus pangenome (AflaPan) uncovers novel aflatoxin and secondary metabolite associated gene clusters Sunil S. Gangurde, Walid Korani, Prasad Bajaj, Hui Wang, Jake C. Fountain, and 10 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3958535/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 30 Apr, 2024 Read the published version in BMC Plant Biology → Version 1 posted 9 You are reading this latest preprint version Abstract Background Aspergillus flavus is an important agricultural and food safety threat due to its production of carcinogenic aflatoxins. It has high level of genetic diversity that is adapted to various environments. Recently, we reported two reference genomes of A. flavus isolates, AF13 ( MAT1-2 and highly aflatoxigenic isolate) and NRRL3357 ( MAT1-1 and moderate aflatoxin producer). Where, an insertion of 310 kb in AF13 included an aflatoxin producing gene bZIP transcription factor, named atfC . Observations of significant genomic variants between these isolates of contrasting phenotypes prompted an investigation into variation among other agricultural isolates of A. flavus with the goal of discovering novel genes potentially associated with aflatoxin production regulation. Present study was designed with three main objectives: (1) collection of large number of A. flavus isolates from diverse sources including maize plants and field soils; (2) whole genome sequencing of collected isolates and development of a pangenome; and (3) pangenome-wide association study (Pan-GWAS) to identify novel secondary metabolite cluster genes. Results Pangenome analysis of 346 A. flavus isolates identified a total of 17,855 unique orthologous gene clusters, with mere 41% (7,315) core genes and 59% (10,540) accessory genes indicating accumulation of high genomic diversity during domestication. 5,994 orthologous gene clusters in accessory genome not annotated in either the A. flavus AF13 or NRRL3357 reference genomes. Pan-genome wide association analysis of the genomic variations identified 391 significant associated pan-genes associated with aflatoxin production. Interestingly, most of the significantly associated pan-genes (94%; 369 associations) belonged to accessory genome indicating that genome expansion has resulted in the incorporation of new genes associated with aflatoxin and other secondary metabolites. Conclusion In summary, this study provides complete pangenome framework for the species of Aspergillus flavus along with associated genes for pathogen survival and aflatoxin production. The large accessory genome indicated large genome diversity in the species A. flavus , however AflaPan is a closed pangenome represents optimum diversity of species A. flavus . Most importantly, the newly identified aflatoxin producing gene clusters will be a new source for seeking aflatoxin mitigation strategies and needs new attention in research. Aspergillus flavus aflatoxin core genome accessory genome genomic diversity pan-GWAS Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Aspergillus flavus is an opportunistic saprophyte, infects varieties of crops such as maize and peanuts both pre- and postharvest [ 1 ]. These fungi also produce secondary metabolites including aflatoxins which contaminate the grains and are toxic and carcinogenic leading to inducing acute and chronic health issues in both humans and animals [ 2 ]. In addition to aflatoxins, A. flavus also produces cyclopiazonic acid (CPA), which harms liver, kidneys, and gastrointestinal tract [ 3 ]. Because of its harmful impact on human health, the aflatoxin is considered as “hidden and slow poison” and the research community is trying its best to managing through this menace through a combination of improved genetics and post-harvest (farm to industry) management on one side while understanding the reason of compulsion for pathogen to produce toxin on the other side [ 4 ]. Aflatoxin contamination is an annual issue in the southern United States and throughout the world, but occasional outbreaks happen in the “cooler areas” like mid-western Corn Belt of the United States [ 5 ]. Aspergillus flavus belongs to Aspergillus section Flavi , which splits into eight clades and currently contains 33 species. Of these, two species only produce aflatoxin B1 and B2 ( A. pseudotamarii and A. togoensis ), and 14 species are able to produce aflatoxins B1, B2, G1 and G2 [ 6 ]. Large differences in gene contents can occur among genomes of A. flavus isolates, with only a portion of genes being universal, or core, to all genomes. For instance, A. flavus and A. oryzae possess core identical gene models and sequences, however A. flavus is a very devastating fungal pathogen produces aflatoxin, while A. oryzae is a beneficial fungus used widely for fermentation in food industries [ 7 – 8 ]. Interestingly, comparative genomic analysis of these two Aspergilli clearly shows that A. oryzae is a domesticated ecotype of wild A. flavus [ 9 ]. Therefore, a single reference genome could not represent complete diversity of a species [ 10 ]. A recent pan-metabolome analysis of 94 A. flavus isolates support this with identification of 7821 biosynthetic gene clusters (BGCs) including 25% population specific BGCs and 92 unique BGCs [ 11 ]. A. flavus isolate NRRL 3357 was used as a reference genome for multiple omics studies [ 12 ]. Recently, the NRRL 3357 assembly was updated and re-annotated for eight chromosomes [ 13 ]. At the same time, [ 14 ] reported two new A. flavus reference genomes comparatively, which revealed a large insertion potentially contributing to stress tolerance and aflatoxin production in isolate AF13, a MAT1-2 mating type in comparison with NRRL 3557, a MAT1-1 mating type. This study between AF13 and NRRL 3357 confirmed that AF13 has an insertion of 310 kb with a bZIP transcription factor named atfC [ 14 ], which shares homology with A. flavus bZIP transcription factor atfA [ 15 ] and atfB [ 16 ]. These transcription factors have been shown to regulate the production of aflatoxin and its precursors in response to oxidative stress and to coordinate oxidative stress responsive genes including catalase [ 17 ]. If using only the NRRL 3357 reference genome, the presence of this insertion the atfC gene would not have been recognized due to reference bias. Therefore, use of single reference genome, either NRRL 3357 or AF13, hinders the potential of omics studies to fully describe and investigate complex pathways such as aflatoxin or other secondary metabolite biosynthesis in A. flavus . It is essential to study the comprehensive genetic diversity of A. flavus using a large number of isolates from diverse geographical regions and ecosystems [ 9 ]. Pangenomics has emerged as an effective approach to explain the widespread diversity present in a species, and super-pangenome can be used to represent the diversity present in an entire genus [ 10 ]. A pangenome represents all the genes present across a species, a pan-genes can be subdivided into “core-genome” present in all the individuals in a species and “accessory genes” not present across all the isolates which is also called as “dispensable genome” [ 18 ]. For instance, three versions of A. fumigatus pangenome have been developed, emphasizing identification of azole drug resistance genes in A. fumigatus [ 19 – 21 ]. Therefore, a pangenomic approach is necessary to fully capture the genetic diversity present within the A. flavus species, particularly in agricultural environments, and can be used to perform in depth studies of a plethora of secondary metabolite gene clusters and their regulation. To accomplish this, present study was designed with three main objectives: (1) collection of large number of A. flavus isolates from diverse sources including maize plants and field soils; (2) whole genome sequencing of collected isolates and development of a pangenome; and (3) pangenome-wide association study (Pan-GWAS) to identify novel secondary metabolite cluster genes. This new A. flavus pangenome, here named as AflaPan, can serve as a new reference for comparative genomics studies by the Aspergillus research community with greater representation of the genetic diversity of this species than any single isolate reference can provide. Material and Methods Aspergillus flavus isolates A total of 225 isolates (98 from infected corns; 127 from soils) were newly collected and sequenced in this study. 98 isolates were associated with corn fields from infected plant parts of corn at harvest in Mississippi Delta, including 11 isolates from corn leaves, 38 from corn silks, 30 from corn tassels, and 19 from dust air-spora of corn combine harvester (Figure 1a). Among the remaining 127 isolates from different field soils with diverse cropping systems in southern Georgia in the fall harvesting season, there were 10 isolated from corn and peanut rotation fields of Tift and Turner County, 26 continuous peanut fields in Irwin and Turner County, 11 from corn and cotton rotation fields of Turner County, 17 from corn and peanut rotation fields of Irwin County, 8 from corn, cotton and peanut rotation fields in Berrien and Tift County, 16 from continuous sunflower fields of Tift County, 2 from continuous soybean fields of Tift County, and 37 continuous sorghum field of Tift County (Figure 1b). For A. flavus isolation, protocol of [22] was followed with modification. Briefly, 10 g of soil sample was added to a bottle containing 99 mL of phosphate buffer solution. The bottles were placed on a shaker for 30 minutes, and afterwards, the mixtures were serially diluted with sterile phosphate buffer. In triplicate, 100 µL of each serial dilution was spread on 100 mm × 15 mm petri dishes containing Modified Dichloran Rose Bengal (MDRB) agar [23]. The inoculated plates were incubated in the dark for 3-5 days at 37°C. After incubation, up to 20 A. flavus colonies for each sample were randomly selected from the MDRB plates and transferred to individual 60 mm × 15 mm Petri dishes containing potato dextrose agar amended with 0.3% β-cyclodextrin. The inoculated plates were incubated in the dark at 28°C for 5-7 days and characterized for sclerotia and aflatoxin production. In addition to 225 newly collected isolates, raw DNA sequence data for 121 isolates was also downloaded from previously reported genome assemblies from National Centre for Biotechnology Information Sequence Read Archive (NCBI-SRA). Finally, a total of 346 A. flavus isolates (225 new isolates and 121 from NCBI-SRA) representing 11 states of the United States were used for development of A. flavus pangenome (Figure 1c). A detailed list of isolates with metadata, geographical site, mating types, and sclerotia types is provided in Supplementary Table 1. DNA extraction and whole genome sequencing For short read sequencing for each isolate collection, a normal CTAB DNA isolation protocol was used [14]. Each isolate was cultured in yeast extract sucrose (YES) medium with 2% yeast extract, 1% sucrose for five days at 30 ̊ in the dark. Mycelial mats from each culture were collected and ground in a chilled mortar and pestle with liquid nitrogen. The ground mycelia (1-2 g) were then combined with 15 mL of CTAB extraction buffer (0.1M Tris pH 8.0, 1.4M NaCl, 20mM EDTA, 2% (w/v) CTAB, 4% (w/v) polyvinylpyrrolidone (PVP-40), and 0.5% (v/v) b-mercaptoethanol), mixed by inversion, and incubated in a water bath at 65 ̊C for 45 min with occasional inversion. The lysate was then combined with 15 mL of chloroform:isoamyl alcohol (24:1), mixed by inversion, and centrifuged at 8,000 · g for 15 min at 4 ̊. The upper phase was then transferred to a new 50 mL centrifuge tube. The chloroform separation was then performed a second time, and the upper phase was then combined with one volume of cold isopropanol for DNA precipitation. The DNA was then pelleted by centrifuging at 8,000 · g for 15 min at 4 ̊, and washed with 70% ethanol. The pellets were then dried and suspended in 100 mL TE buffer (10mM Tris pH 8.0, 1 mM EDTA pH 8.0). RNase-A was then added to a final concentration of 5 mg/mL and the samples were incubated at 37 ̊ for 1 hr. The obtained DNA was then stored at -20 ̊ until used. Isolated DNA was quantified with both a Nanodrop ND-1000 spectrophotometer (ThermoFisher, Waltham, MA, USA) and a Qubit 3.0 fluorometer (ThermoFisher) and checked using gel electrophoresis. Isolated DNA (500 ng) from each sample was then shipped to the Novogene Corp. (Sacamento, CA) for quality checking, sequencing, and initial data filtering [24]. Libraries (350bp insert size) were generated using NEBNext® DNA Library Prep Kit (New England Biolabs, Ipswich, MA) following manufacturer’s instructions. In brief, initially dA-tailing followed by adapter, further the fragments (of 300-350 bp) were PCR enriched with P5 and indexed by P7 oligos. The quality of sequencing libraries was again checked on Qubit® 3.0 flurometer to determine the concentration of each library. Agilent® 2100 bioanalyzer was used to assess the insert size by using 1ng/ul DNA from sequencing libraries. Finally, the quantitative real time PCR (qRT-PCR) was performed to detect the effective concentration of each library. The libraries with appropriate insert size and with >2mM concentration were qualified for high throughput sequencing on Illumina HiSeq 4000 platform. Pair-end sequencing data was generated for each isolate with the read length PE150 at each end. Sequencing data for these 225 isolates of A. flavus used in this study has been submitted to NCBI with bioproject ID: PRJNA915632. Genome assembly and annotations The raw sequencing data quality analysis and trimming of adapter sequences for each isolate was carried out using Trimmomatic v0.40 [25]. Trimmed reads were assembled using SPAdes v.3.15.4 [26], with AF13 genome assembly as a reference using mismatch and short indel correction with the Burrow-Wheeler Aligner, BWA-MEM v.0.7.12 [27]. Assemblies were improved using Pilon v.1.22 [28]. The assembly summary statistics were calculated using QUAST v5.0.2 [29]. Functional annotation and gene prediction were performed using Funannotate pipeline v.1.8.14 [30]. In brief, the A. flavus assemblies were initially prepared for gene prediction and functional annotation. The small and repetitive contings (shorter than 500bp) in the assembly were removed by using the function funannotate clean which implements minimap2 with “leave one out” methodology, where unaligned shortest contigs to N50 of assembly were filtered out with parameters percent coverage overlap (--cov=95) and percent identity of overlap (--pident=95). All assemblies were soft-masked using function funannotate mask with default method tantan RepeatMasker (v.4.0.8) for repetitive elements using Dfam and RepBase repeat libraries [31]. Masked assemblies were used for gene prediction using funannotate predict with the help of Evidence Modeler algorithm. Curated protein models from closely related species Aspergillus oryzae were provided as evidence for gene prediction. Species confirmation, identification of mating types, and types of sclerotia Sclerotia are survival structures produced by A. flavus to survive in extreme environmental conditions [32]. Wild-type strains of A. flavus are known to produce sclerotia in culture after 5 days of incubation at ideal growth temperatures or a short period of cold storage. A. flavus can be categorized into three groups based on the type of sclerotia production: (1) no sclerotia, (2) small type (S) sclerotia (400 µm. Large (L) type sclerotia are very common, not uniform, and different shapes. Notably, S- sclerotia producing A. flavus isolates are observed to be almost always positive for aflatoxin production, while L- sclerotia producers were observed to have little to no effect on the presence or concentration of aflatoxins [22]. Furthermore, S-type sclerotia strains, on average, produce more aflatoxin than the L-type sclerotia strains. In this study, to determine the sclerotia size of each A. flavus isolate, the samples were viewed under a microscope, and the microscope reticule was used to measure the size of the sclerotia. For sclerotia producing isolates, 100 sclerotia were measured to calculate average sclerotia size, and the average sclerotia size was used to assign the type of sclerotia production for A. flavus . We also confirmed the sclerotia types in individual isolates using the cypA and norB sequences. A. flavus and A. parasiticus are closely related species, but A. flavus produces only aflatoxins B1 and B2, while, A. parasiticus can produce B1, B2, G1 and G2. A. flavus genome is missing the portions of genes ( cypA and norB ) upstream to the polyketide synthase gene [33]. Therefore, we used A. parasiticus specific cypA-norB aflatoxin pathway gene cluster marker genes (NCBI accession no. AY371490) to detect the A. parasiticus isolates [34]. Based on blast results, a total of 225 isolates were confirmed as A. flavus, while a total of 39 isolates collected from soil were identified as A. parasiticus isolates. We collected a total of 346 A. flavus isolate genomes (225 from this study and 121 from NCBI-SRA) for construction of a pan-genome of A. flavus . We identified the large sclerotia (L) types using the norB-cypA region genomic sequence from strain NRRL3357 (NCBI accession no. AY566564), while, the short sclerotia (S) types were identified using locus of aflatoxin biosynthetic gene cluster from isolate AF70 (NCBI accession no. AY510453) [33]. In Aspergilli , A. flavus and A. parasiticus are heterothallic species with two mating type loci MAT1-1 and MAT1-2 located on chromosome 6 associated with sexual reproduction [35]. Mating types were identified among all isolates using the reported marker sequences. Mating types MAT1-1 identified using complete cds sequence of mating-type protein MAT alpha 1 (NCBI accession no. EU357934), while the MAT1-2 types were identified using the complete cds sequence of a putative DNA lyase gene also called as mating-type HMG (High Mobility Group)-box protein MAT1-2 (NCBI accession no. EU357937) [36]. Read alignment and variant calling AF13 reference genome (v1.0, Gene Bank assembly accession number GCA_014117485.1) was used for mapping and variant calling. The genome files for AF13 reference genome were downloaded from NCBI from bio-project ID PRJNA606266 [14]. Sequence reads for 346 A. flavus isolate were aligned to AF13 reference genome using Burrows-Wheeler Aligner (BWA v.0.7.17) [27]. The alignment files were converted into BAM files using Samtools v.1.10 [37]. Picard tools v.2.18.3 (http://broadinstitute.github.io/picard) used to mark the PCR duplicates using MarkDuplicate. HaplotypeCaller implemented in GATK (v.4.1.0) was used to call the SNPs using AF13 as a reference genome [38]. VariantFiltration followed by SelectVariants workflow was used to filter variants using parameters [-window-size = 10, -QualBy-Dept <2.0, -MapQual <40.0, -Qscore <100, -MapQualityRankSum 3.0, -FisherStrandBias >60.0, -ReadPosRankSum <-8.0.]. Whole-genome phylogeny estimation SNV (Single Nucleotide Variants) based phylogenetic tree was constructed by filtering the variants at minor allele frequency 0.25 and the loci with zero coverage in any isolate. Variants from each isolate were concatenated and used as input in RAxML v.8.2.12 to construct maximum likelihood (ML) tree [39]. Annotations for origin, source, toxicity, sclerotia types, mating types, genome size (Mb), and number of scaffolds for each isolate were plotted on phylogenetic tree using iTOL (Interactive Tree of Life) [40]. The final tree was visualized in iTOL by customizing the parameters. Genetic clusters among 346 A. flavus isolates identified using discriminant analysis of principle components [41]. The PCA plot was visualized using ggplot2 in R version 4.2.2 [42]. Population structure and diversity analysis Population structure of 346 A. flavus isolates was studied by identifying optimum number of groups ( k ) corresponding to the lowest Bayesian information criterion (BIC). The STRUCTURE v.2.3.4 software was used to study the population structure and admixture [43]. Diversity parameters such as nucleotide diversity ( π ), fixation index ( F ST ), and Tajima’s D diversity index [44] were analyzed for each genetic cluster to understand genetic diversity and distinctness of genetic clusters. VCFtools v.0.1.13 was used to calculate the genetic diversity parameters (nucleotide diversity ( π ), fixation index ( F ST ), and Tajima’s D) in a non-overlapping window of 10000bp [45]. Nucleotide diversity ( π ) is used as measure of variability of SNPs nucleotides within individual genetic clusters. Nucleotide diversity ( π ) calculated in VCFtools using parameters --site-pi --window-pi 10000. Tajima’s D diversity index allows us to see whether a population is neutrally evolving or whether it is under selection. Tajima’s D was calculated in VCFtools with parameters --TajimaD 10000. Genome-wide fixation index ( Fst ) is a measure of population differentiation. Fst ranges from 0 to 1, where is Fst =0 means two populations are genetically similar. While, Fst=1 means the populations are significantly genetically dissimilar to each other. Fst was calculated using VCFtools with parameters --weir-fst-pop --fst-window-size 10000. All diversity statistics was visualized using ggplot2 in R. The linkage disequilibrium decay for each population was calculated using VCFtools in 10000Kb window. The linkage disequilibrium ( r 2 ) was plotted against the distance between the two loci in the equilibrium (Kb). Development of a pangenome, AflaPan, and annotation of orthogroups The protein sequences from 346 isolates obtained from funannotate were used to identify the orthologous gene clusters using OrthoFinder [46]. Protein sequences from the AF13 reference genome were also included to identify the functions of orthologous gene clusters in AflaPan. If a gene sequence from AF13 genome is grouped in an orthologous cluster then the whole cluster is annotated as the annotation of the grouped sequence from AF13 reference genome. However, some orthologous gene clusters were not grouped with single sequence of AF13 genome. We called these orthologous gene clusters as a novel gene clusters when compared with AF13 reference genome. A customized bash script was used to extract the most extended protein sequence for every orthologous group to functionally annotate the orthogroups. Blast+ was used to find the highest similar proteins against SwissProt (https://www.expasy.org/) [47] and FunigDB (https://fungidb.org/) [48]. The sequences remain functionally annotated with annotations from AF13 reference genome and FungiDB were further annotated using InterProScan (https://github.com/ebi-pf-team/interproscan) [49]. We used the criteria E-value cutoff of 1 × 10 5 , percent identity > 70%, minimum query coverage > 50% and minimum subject coverage > 50% [19]. The novel orthogroups were included in the pangenome only if they are present in at least 5% of the isolates in the study, to avoid false gene prediction. The presence/absence matrix of orthologous groups used to describe the framework of AflaPan genome. Further the presence/absence matrix was used for Pan-GWAS analysis, to check the utility of the AflaPan genome. Estimation of aflatoxin content (ppb) Up to 20 A. flavus colonies for each isolate were randomly selected from the Modified Dichloran Rose Bengal (MDRB) plates and transferred to individual 60 mm × 15 mm petri dishes containing potato dextrose agar amended with 0.3% β-cyclodextrin. The inoculated plates were incubated in the dark at 28°C for 7 days. After incubation, evaluated for aflatoxigenicity using three cultural methods using yellow pigment test, UV fluorescence test, and ammonia vapor test [22]. The percentage of aflatoxin-producing colonies for each sample was calculated based on the qualitative analysis of the three cultural methods. To quantitate the aflatoxin production by A. flavus isolates, aflatoxins were extracted from agar plugs of 7-day old cultures and analyzed by Enzyme Linked Immuno-Sorbent Assay (ELISA). First, agar plugs of fresh cultures were extracted at 1:5 ratio with methanol-water (70:30, v / v ). The extracts were filtered through 0.25 µm pore-size nylon syringe filters. The filtrates were stored at 4°C until use. For ELISA, Veratox® for Aflatoxin kits (Neogen Inc., Lansing, MI, USA) used to determine the total aflatoxin concentration. If the results exceeded the quantitative parameters of the protocol, the filtrates were serially diluted and re-analyzed. The final aflatoxin content (ppb) was calculated using the dilution factor determined by number dilutions. A Neogen Stat-Fax 303 Plus microwell reader was used to measure the absorbance of the ELISA reaction and to calculate the final concentration of aflatoxins in solution [22]. Pan-GWAS analysis to identify the aflatoxin production associated gene clusters The presence/absence variance matrix (PAV) for 17,855 orthologous gene clusters of AflaPan and phenotyping data generated for aflatoxin production was used for pan-genome wide association analysis. The presence/absence matrix of 1/0, where ‘0’ represents absence of orthogroups, and ‘1’ means presence of orthogroups in the isolates. The objective behind Pan-GWAS analysis was to identify the association of accessory genes with regulation of aflatoxin production in A. flavus . Pan-GWAS analysis was perform using mixed-linear-model (MLM) implemented in TASSEL v.5.0 [50]. MLM calculates both fixed and random effects in the population. For MLM, a kinship (K) was calculated using PAV matrix and used jointly with population structure (Q). The Q+K approach helps to increase the power of genome-wide association analysis [51]. A Bonferroni threshold was calculated at P -value 0.05, Bonferroni threshold=0.05/17855 (Orthogroups in AflaPan) =2.8 × 10 -6 . The orthogroups showing P value less than or equal to Bonferroni threshold of 2.8 × 10 -6 were called as significantly associated orthogroups [52]. GWAS results were visualized using Manhattan plots and Quintile-Quintile (QQ) plots generated in the R package ‘CMplot’ [53]. Availability of data and materials Sequencing statistics, assembly statistics, SNP statistics, and Pan-GWAS analysis are provided in the attached supplementary files. The newly developed genome assemblies for 225 isolates used in this study are available at https://zenodo.org/deposit/7615243. Raw sequencing data and metadata for each isolate is available through National Center for Biotechnology Information (NCBI) - Sequence Read Archive (SRA) with Bioproject ID PRJNA915632. Newly sequenced 225 fungal isolate cultures are available upon request by contacting the corresponding author. The 121 isolates from public data can be requested from corresponding authors of respective articles [24, 11, 54]. Results De novo genome assemblies of A. flavus isolates A total of 225 isolates were newly collected and sequenced as part of this study (Figure 1a, 1b). In addition, we downloaded raw sequencing data for 121 previously sequenced isolates from NCBI sequence read archive. Among these, 7 isolates were sequenced by Fountain et al. (2020b) [24], 94 sequenced for studying pan-metabolomics [11], and 20 draft genomes from [54]. Overall, a total of 346 isolate genomes were included in this study, representing the diversity from 11 states (California, Arizona, Oklahoma, Texas, Louisiana, Mississippi, Georgia, Florida, North Carolina, Pennsylvania and Iowa) of the United States and soils of crops including peanut, sunflower, soybean, sorghum, corn, sesame and cotton. Of these 346 isolates, 177 were aflatoxigenic and 169 were non-aflatoxigenic. For sclerotia, there were 128 isolates with large sclerotia (L-strains) and 218 isolates showing small sclerotia (S-strains). Finally, for mating types, 159 showed mating type MAT1-1 and 187 isolates showed mating type MAT1-2 ( Supplementary Table 1 ). De novo genome assemblies were developed for 346 A. flavus isolates using paired end Illumina sequencing reads. The mean number of contigs and scaffolds in the assemblies for soil isolates were 597 and 570, with average of 726 and 697 contigs and scaffolds, respectively. Average length of contigs and scaffolds of soil isolates was 89 kb and 98 kb, respectively. The assemblies of isolates collected from corn dust had a total of 111 and 131, contigs and scaffolds, respectively, the lowest number of the newly sequenced isolate genomes. In contrast, the assemblies of isolates collected from corn leaf tissues had the highest number 151 and 167 of contigs and scaffolds, respectively. The assembly mean contig N50 and mean scaffold N50 were 403,194 bp and 523,939 bp. Mean genome size of the assemblies of corn isolates and soil isolates was 36.9 Mb and 37.6 Mb, respectively. It is observed that soil genomes are slightly longer than corn isolates, with maximum 43.2 Mb genome size in soil isolates and maximum 38.5 Mb in corn isolates. Among the corn isolates there was no significant difference in number of contings and scaffolds collected from different corn parts such as silk, leaf, tassel and dust, and no significant difference in the genome sizes of corn isolates was observed ( Supplementary Table 2 ). To perform population genomics studies, we retrieved ~1.02 million single nucleotide variants (SNVs) using AF13 as reference, approximately 28 SNVs per kilo bases (kb) among 346 A. flavus isolates. Large number of SNVs across the isolates indicated prominent genetic diversity in the species of A. flavus . On an average of 1,021,529 SNVs per isolate (ranged from 795,241- 1,027,945 SNVs) were used for diversity analysis. An average of 4,871 SNVs showed heterozygous calls, and 1,016,658 SNVs were homozygous per isolate. Only homozygous SNV calls were used for population genomics studies ( Supplementary Figure 1; Supplementary Table 3 ). Diversity in the population of 346 A. flavus isolates A total of four genetic clusters ( K =4) were identified in the collection of 346 A. flavus isolates using population structure analysis. The largest genetic cluster (P1) included 179 isolates, with majority of isolates from soil. Second largest cluster (P2) included 158 isolates grouped from soil as well as corn. Third and fourth clusters were minor with merely 5 and 4 isolates, respectively. Highly toxigenic isolates collected from the state of Louisiana were grouped in clusters 3 and 4 (Figure 1d). Four principle components identified using discriminant analysis of principle components further confirmed the presence of four genetic clusters (Figure 1e). Linkage-disequilibrium (LD) decay for each genetic cluster was calculated and plotted r 2 values against distance between two SNVs (Kb) using ggplot2. Rapid LD decay was observed for genetic cluster P1 (10 Kb), followed by genetic cluster P3 (12 Kb), cluster P4 (18 Kb) and cluster P2 (25 Kb). The population growth results in reduction of LD the largest genetic cluster (P1) showed least LD decay (Figure 1f). Genome-wide nucleotide diversity ( π ), fixation index ( F ST ) and Tajima’s D diversity was analyzed individually for each cluster to understand the distinctness of genetic clusters. Genome-wide nucleotide diversity ( π ) for each cluster was visualized using circos plot, showed significant number of hot spots with high nucleotide diversity in each cluster (Figure 1g). Highest average nucleotide diversity ( π =3.2 × 10 -3 ) was observed for cluster P1, followed by cluster P4 ( π =2.6 × 10 -3 ), and cluster P3 ( π =2.0 × 10 -3 ). The lowest nucleotide diversity was observed in genetic cluster P2 ( π =8.4 × 10 -4 ) (Figure 1h). Further, the chromosome wise nucleotide diversity across 346 isolates was represented using boxplots. Chromosome 8 showed highest nucleotide diversity (mean π =5.0 × 10 -3 , with range 1.9 × 10 -2 – 3.8 × 10 -7 ), followed by chromosomes 5, 6 and 7 (mean π =3.7 × 10 -3 , with range 1.85 × 10 -2 – 5.5 × 10 -7 ), chromosomes 3 and 4 (mean π =3.5 × 10 -3 , with range 1.7 × 10 -2 – 5.4 × 10 -7 ), and chromosomes 1 and 2 (mean π =3.2× 10 -3 , with range 1.9 × 10 -2 – 2.5 × 10 -5 ) ( Supplementary figure 2 ). The population differentiation was investigated by calculating fixation index ( F ST ) between the pairs of genetic clusters. Fixation index ( F ST ) value can range from 0 to 1. Where, F ST =0 indicates the two clusters are completely sharing each other, whereas F ST =1 means there is no sharing or complete differentiation between two clusters. In this study, higher F ST values (>0.6) of the cluster P2 with clusters P1, P3 and P4 indicated that P2 cluster is most differentiated cluster than other three clusters. For instance, fixation index ( F ST ) of 0.822 was recorded between clusters P2 and P4, followed by 0.814 between P2 and P3, and 0.619 between P2 and P1. Lowest fixation index ( F ST =0.346) was recorded between the clusters P3 and P4, followed by P1 and P3 ( F ST =0.346), P1 and P4 ( F ST =0.460) (Figure 1h). Genome-wide Tajima’s D diversity index was calculated for each cluster to study the abundance of rare alleles in each cluster. Tajima’s D diversity helps to determine whether the cluster is evolving neutrally, or under selection pressure. Tajima’s D=0, means the cluster is evolving neutrally without any evidence of selection, Tajima’s D0, means scarcity of rare alleles in cluster or population contraction. Tajima’s D for cluster P1 was -0.13, indicating that cluster P1 is evolving neutrally with partial selection sweep. Interestingly, Tajima’s D for cluster P2 was -1.77, indicated abundance of rare alleles in P2. The Tajima’s D for cluster P3 (0.81) and P4 (0.81) was greater than zero indicating that these two populations lack rare alleles and therefore the clusters are resulted in the population contraction (Figure 1i). Chromosome wise Tajima’s D index was higher for chromosome 6 and 8 (Supplementary Figure 3). Whole-genome phylogeny and characteristics of genetic clusters Whole genome phylogeny was constructed and confirmed the presence of 4 clusters within population of 346 A. flavus isolates. Morphological characteristics of each cluster were investigated for their origin (geographical site), source (soil or corn), mating types, sclerotia types, toxigenicity, genome size (Mb), and number scaffolds in each assembly (Figure 2). In the major cluster P1, the majority of isolates were from soil (136) and 43 were from corn. In cluster P1, a total of 126 isolates had small sclerotia (S-types), and 53 had large sclerotia types. In case of mating types, cluster P1 included 94 isolates with mating type MAT1-1 , whereas 85 were MAT1-2 . Interestingly, in cluster P1, a total of 111 isolates were aflatoxigenic and 68 were non-aflatoxigenic. Second major cluster P2 included a total of 103 isolates from soil and 55 from corn. In this cluster 58 isolates were aflatoxigenic and 100 isolates were non-aflatoxigenic. There were 74 isolates with large sclerotia (L) types and 84 with small sclerotia (S) types. In case of mating types, 61 isolates showed ( MAT1-1 ) types and 97 showed ( MAT1-2 ) types. In cluster P1, majority of isolates were aflatoxigenic, whereas in P2 majority of isolates were non-aflatoxigenic. Cluster P3 and P4 were minor genetic clusters with 5 and 4 isolates, respectively. In genetic cluster P3 all isolates were from collected from soil with aflatoxigenic, small sclerotia types, and four with MAT1-1 and one MAT1-2 mating types. In cluster P4, three isolates were aflatoxigenic and a one isolate was non-aflatoxigenic from the state of Iowa. All aflatoxigenic isolates in cluster P3 and P4 were from the states of Louisiana and Oklahoma (Figure 2). 310 kb insertion is rare and identified only in soil samples Recently a highly toxigenic A. flavus strain AF13 was sequenced, and a 310 kb insertion was identified in AF13 genome, but was absent in isolate NRRL3357. This 310 kb insertion includes a bZIP transcription factor named atfC [14]. The presence of 310 kb region was analyzed across 346 isolates, there were only 5 isolates with this 310 kb insertion. Interestingly, all these 5 isolates are from collected from soil samples in the state of Louisiana, and all are toxigenic isolates grouped in cluster P3 and P4 (Figure 2). However, there were about 57% shared identity with this 310 kb insertion among the other 341 isolates. Therefore, 310 kb insertion is rare in A. flavus species and the isolates with this insertion are highly toxigenic [14] ( Supplementary Table 1 ). AflaPan contains 7,315 core and 10,540 accessory genes Genome assemblies 346 isolates were used to develop pan-genome framework for A. flavus , called AflaPan genome . A pan-genome represents total genes in the species, including core genes shared by all the individuals of the species and accessory genes that are not present in all the individuals of the species. In this study, we collected 346 genomes of A. flavus and identified 17,855 unique orthologous gene clusters in this pangenome, called AflaPan genome. In AflaPan, of the 17,855, there were 7,315 orthologous gene clusters as core genome (41.1% of the pangenome) present in all 346 genomes, 3,887 as softcore orthogroups in >95% genomes (21.7% of the pangenome), 3,132 as shell genes in 5-95% genomes (17.5% of the pangenome), and 3,521 as cloud genes present in less than 5% of the genomes (19.7% of the pangenome) (Figure 3a, 3b). We also report here 5,994 novel gene clusters that were not annotated in the AF13 reference genome [14]. Of the new gene clusters, 451 were identified in soft-core (12% of total soft-core genes), 2,069 (66% of total shell genes) in shell and 3,474 (99% of total cloud genes) in cloud (Figure 3c). The gene clusters in this AflaPan genome were compared with the gene models in current reference genomes AF13 and NRRL3357 of A. flavus . Reference genomes AF13 and NRRL3357 contains of 13,188 and 13,487 annotated gene models, respectively. However, in AflaPan, a total of 17,855 unique orthologous gene clusters were annotated. The AflaPan genome includes more diverse genome segments that were missing in current single reference genome assemblies (Figure 3d). Significant differences between the lengths of protein sequences were observed in core and accessory genomes. The average protein length in core genome was 579 amino acids (AAs), with a range of 68 to 6885 AAs. While the average protein lengths in soft-core genome was 463 AAs, with range of 60 to 7781 AAs. Average protein length in shell genome was 361 AAs, with the range of 50 to 3917 AAs. In clouds and singletons, the average protein length was 228 AAs and 233 AAs, respectively. Therefore, the conservative proteins were longer in length. Overall, the average protein length was decreased from core-softcore-shell-clouds and singletons ( Figure 3e; Supplementary Table 4; Supplementary Table 5 ). Furthermore, this Aflapan was closed and captured comprehensive genetic diversity of the intraspecies of A. flavus . Any addition of A. flavus genomes will not substantially increase the number of pan-genes or core genes of A. flavus pan-genome (Figure 3f). The number of pan-genes and core genes will be not substantially increased beyond this large collection of 346 A. flavus genome in this AflaPan (Figure 3g). Toxin producing gene clusters in AflaPan The orthologous gene clusters in AflaPan were annotated using blasts in InterProScan, SWissProt, and FungiDB databases. A total of 17,038 orthogroups were annotated using FungiDB database. There were 12,691 orthologous gene clusters annotated using SwissProt and 15,740 annotated using InterProScan. Therefore, there were a total of 12,500 orthogroups annotated using all three databases and 399 orthologous gene clusters were not annotated using either of the three databases ( Supplementary Table 5 ). Based on the annotations a total of 89 toxins producing orthologous gene clusters were identified in core genome. Toxins producing gene clusters (TPCs) in core genome include 3-hydroxyacyl-CoA dehydrogenase-like protein, 6-hydroxy-D-nicotine oxidase, acyl-CoA dehydrogenase, satratoxin biosynthesis SC3 cluster protein, astacin-like metalloprotease toxin, aflatoxin B1 (AFB1) aldehyde reductase, viriditoxin biosynthesis cluster protein D ( VdtD ), latrotoxin producing gene cluster ( LTX ), AM-toxin biosynthesis protein or cytochrome P450 monooxygenase AMT3, AAL-toxin biosynthesis cluster protein, cercosporin toxin biosynthesis cluster, viridicatumtoxin synthesis proteins ( vrtL and vrtT ), aflatoxin biosynthesis protein O, and transcription activator AMTR1 (AM toxin). Cytochrome P450 monooxygenase is potent mycotoxins producing gene clusters, and in AflaPan, we identified ALT8 , aflU , aflV , and AKT7 toxin producing cytochrome P450 monooxygenases. TPCs in core genome indicated that these clusters are widely present in the species A. flavus and annotated in reference genomes AF13 and NRRL3357. A total of 43 TPCs identified in soft-core genome including 5 new gene clusters integrated from soil isolates of Louisiana and Oklahoma. New TPCs in soft-core genome included toxin subunit YenA2, taurine hydroxylase-like protein SAT17, and cytochrome P450 monooxygenase AKT7. Similarly, in shell genome, a total of 45 TPCs were annotated. Interestingly, a total of 23 new gene clusters were identified in shell genome linked with toxin production. New TPCs in shell genome include MFS gliotoxin efflux transporter ( gliA ), called gliotoxin biosynthesis protein A, alpha-latrocrustotoxin-Lt1a, killer toxin subunits alpha/beta, double-stranded DNA deaminase toxin A, and acyl-CoA dehydrogenase AFT10-1. Particularly, in shell genome there were 33 new TPCs integrated from soil isolates and 3 were from corn samples. In cloud genome, a total of 36 TPCs were identified, and all were new or non-AF13 TPCs. Among singletons we identified 8 new TPCs and all were originated from soil samples of Georgia and Mississippi. The TPCs in singletons included C3 and PZP-like alpha-2-macroglobulin domain-containing protein, ADP-ribosylating toxin CARDS (Community-Acquired Respiratory Distress Syndrome toxin), norsolorinic acid reductase ( aflE ), FAD-dependent monooxygenase, aflatoxin cluster transcriptional coactivator ( aflS ), cytochrome P450 monooxygenase ( alfG ), and abhydrolase domain-containing protein AFT2. Here we observed that the extended genome portions of soil included new toxin producing genes ( Supplementary Table 5 ). In order to check the association of new orthologous gene clusters with toxin production and secondary metabolite production, AflaPan-genome wide association analysis (AflaPan-GWAS) was conducted using phenotypes of total aflatoxins produced. Secondary metabolite producing gene clusters A. flavus also produces a variety of secondary metabolites which may be useful for pharmaceutical industries. In current AF13 and NRRL3357 reference genomes of A. flavus [14] there were 80 and 78 secondary metabolite producing gene clusters. In this AflaPan genome, there were a total of 160 secondary metabolite producing gene clusters annotated for 10 important secondary metabolites, including aflatrem (14), aflavarin (3), aflatoxin (63), asparasone (6), cyclopiazonic acid (25), ditryptophenaline (3), kojic acid (3), leporins (4), piperazines (19), and ustiloxin (20). Aflatrem is tremorgenic mycotoxin, the only class of mycotoxins impacting central nervous system [55]. In AflaPan, a total of 14 orthogroups were annotated as aflatrem biosynthetic gene clusters. Among them, 8 orthogroups identified in core, and 5 in soft-core and 1 in shell genome. It indicated that aflatrem producing gene clusters are highly conserved in A. flavus . Aflavarin is also an important sclerotial metabolite exhibiting insecticidal activity [55]. In AflaPan, a total of 3 aflavarin producing orthologous gene clusters were annotated, and there were two in core genome and one in cloud genome. Aflatoxin is a potent mycotoxin produced by A. flavus classified as polyketide synthase (PKS) backbone genes. Based on annotation, there were a total of 63 aflatoxin producing secondary metabolite gene clusters identified in this AflaPan genome. Among them, 20 were identified in core, 10 in soft-core, 21 in shell genome, 9 in cloud, and 3 as singletons. Aflatoxin producing gene clusters are uniformly distributed in core and accessory genomes of AflaPan. Moreover, Pan-GWAS analysis was conducted using phenotyping data of aflatoxin produced by 225 isolates, which uncovered 391 orthogroups associated with toxin production, including these 63 aflatoxin producing secondary metabolite gene clusters. Asparasone is another sclerotial metabolite of structural class PKS [56]. In AflaPan, a total of 6 asparasone producing orthologous gene clusters were identified. Among them, there were one in each core and shell, 2 in soft core, and 2 in cloud genome of AflaPan. Cyclopiazonic acid (CPA) is an indole-tetramic acid mycotoxin, which belongs to structural group dimethyl tryptophan synthases (DMATs) or hybrid PKS (polyketide synthases)-NRPSs (non-ribosomal peptide synthetases) produced by a number of species of Aspergillus [56]. A total of 25 CPA producing orthogroups were identified in AflaPan. Among them, 7 identified in core, 4 in soft core, 6 in shell, and 8 in cloud genomes of AflaPan. Ditriptophenaline (DTP) produces diketopiperazine alkaloid in A. flavus . There were three DTP producing orthogroups identified in cloud genome of AflaPan. Kojic acid (5-hydroxy-2-hydroxymethyl-1, 4-pyrone) belongs to structural group oxydoreductases, a skin lightning agent widely used in cosmetics industries. Moreover, it also has antibacterial activity and therefore used as pharmaceuticals, pesticides and insecticides [57]. In AflaPan, three orthologous gene clusters were identified associated with Kojic acid biosynthesis. Two of them were in core and one was identified in cloud genome of AflaPan. Leporin A is an antiinsectan N-methoxy-2-pyridone sclerotial metabolite isolated from the sclerotia of A. leporis . Leporin belongs to structural group hybrid PKS-NRPS. In AflaPan, four orthogroups were identified linked with leporin biosynthesis. Among them, three were in soft-core genome and one was identified as a singleton in soil isolate GA6_18. Piperazines are class of saturated N-heterocycles, acting as central nervous system stimulants widely used in medicinal chemistry [58]. In AflaPan, a total of 19 orthogroups identified as piperazines biosynthesis gene clusters. Among them, there were 5 identified in core, 2 in soft-core, 2 in shell, 9 in clouds, and one as singleton identified in soil isolate GA6_18. Interestingly, in AflaPan, there were 10 new orthogroups identified associated with Piperazines biosynthesis. Ustiloxin B is a mycotoxin with cyclic peptide belonging to structural group ribosomally synthesized and post-translationally modified peptides (RiPPs). In AflaPan, there were 20 orthogroups annotated as Ustiloxin B biosynthetic clusters. Among them, there were 9 identified in core genome, 5 in soft-core genome, 3 in shell genome, and 3 in cloud genome of AflaPan ( Supplementary Table 6 ). Pan-GWAS uncovered gene clusters associated with aflatoxin biosynthesis pathways Pan-GWAS analysis using PAV matrix of 17,855 orthogroups and the newly collected 225 isolates of A. flavus uncovered new and existing genes associated with aflatoxin production. Pan-GWAS analysis identified a total of 391 orthogroups significantly associated with aflatoxin production ( Figure 4 ). Of these 391, a total of 91 orthogroups were not annotated either with NCBI blast or InterProScan. There were 369 orthogroups (94.4%) from shell genome. Important orthologous gene clusters with significant association to aflatoxin production were listed in Table 2. Surprisingly, of the 369 shell genes, there were 256 (69.4%) orthogroups which were not annotated and may be absent in the AF13 reference assembly. The new orthogroups significantly associated with aflatoxin production were potential intermediates in various aflatoxin biosynthetic pathways in A. flavus . There were 15 orthogroups from soft-core and seven from cloud genome significantly associated with aflatoxin production, including aspergilol biosynthesis cluster protein, aspyridones biosynthesis protein E ( apdE ), and imizoquin biosynthesis cluster protein D ( imqD ) ( Supplementary Table 7 ). Among highly associated aflatoxin producing gene clusters there were cytochrome P450 monooxygenase ( apdE , acrF , aflV , AKT7 , aneD ), ABC multidrug transporters ( atrB , atrD , MDR2 ), aflatoxin biosynthesis regulatory protein R ( aflR ), aflatoxin biosynthesis protein S ( aflS ), demethylsterigmatocystin 6-O-methyltransferase ( aflO ), aflatoxin biosynthesis protein T ( aflT ), fatty acid synthase alpha subunit ( aflA ), beta-cyclopiazonate dehydrogenase, and fumagillin biosynthesis polyketide synthase. aflaoxin producing orthologous gene clusters homologous to aflatoxin producing genes in other fungi were also showed strong association with aflatoxin production in this AflaPan, for instance, terrein biosynthesis cluster protein ( terG ) from A. terreus , ABC multidrug transporter ( atrB ) from A. nidulans , Azaphilone biosynthesis cluster protein ( azaJ ) from A. niger , double-stranded DNA deaminase aflatoxin A from Burkholderia cenocepacia , glutathione S-transferase omega from Saccharomyces cerevisiae , and AM-toxin biosynthesis regulator 1 ( AMTR-1 ) from Alternaria alternata ( Supplementary Table 7 ). Discussion Aflatoxin is detrimental to human health specially affecting liver and gall bladder in addition to problems of growth stunting and impaired development. This aflatoxin producing pathogen, Aspergillus flavus, has acquired large geographical diversity and broad host range over the time and sustained under varied climatic conditions. The aflatoxin contamination has been global food safety concern for humans and animals, many a times even leading to death. Isolated efforts have been made though for identifying the genomic regions and candidate genes controlling different aflatoxin resistance mechanisms [4], however, less efforts are being made to in depth understanding adoption of isolates specially under changing climate conditions. Nevertheless, the recently sequenced two reference-level genomes for two A. flavus isolates (isolate AF13 of mating type MAT1-2 with high producer of aflatoxin and isolate NRRL3357 of mating type MAT1-1 with a moderate producer) facilitated identification of 310 kb region being the major aflatoxin producing gene cluster [14]. As of now not much understanding has been developed on the extent of diversity among pathogens from different geographies and crops. Therefore, this comprehensive research with large number of sequenced genomes has been conducted to solve the puzzle of evolution of isolates, intensity of infection to plants and varied level of aflatoxin production under different climatic conditions. We collected and sequenced new set of 225 A. flavus isolates and also included 121 publicly available sequences for genome analysis. This is the maiden effort to develop pangenome of A. flavus , called AflaPan, and perform other genomic analysis using 346 genome sequences. The AflaPan had 17,855 unique orthologous gene clusters including 7,315 core genes (41%) and 10,540 accessory genes indicating significant genome diversity between A. flavus isolates. It is interesting to note that there were 5,994 orthologous gene clusters that had been not annotated in A. flavus AF13 and NRRL3357 reference genomes [14]. The pan-GWAS study using AflaPan identified 391 significant associated orthologous gene clusters for aflatoxin production, indicating the usefulness of AflaPan. It’s worthwhile noting that AflaPan represents a large diversity of 346 A. flavus genomes from 11 states of United States with aflatoxin contamination issues in various crops (Fig. 1). Although no such study has been reported for A. flavus , however, recently two pangenomes have been reported for related species A. fumigatus with 300 [19] and 260 [21] isolates. It is important to note that these two pangenomes are closed in nature similar to AflaPan indicating sufficient sample size for capturing complete species genome diversity. The accessory genome is a variable portion of pangenome and higher accessory gene content indicates higher diversity among the individuals. AflaPan had large proportion of accessory genome (58.9%; 10,540 genes) ( Figure 3a ) as compared to A. fumigatus (41% accessory genome; [19] and Candida albicans (mere 9% accessory genome; [18]. The high accessory genome content in A. flavus indicates the faster pace of genome evolution which might be its strength for its successful adaptation to different climatic conditions, geographies and host plants. This finding also signals warning to research community for much faster devising and implementing solutions to tackle the aflatoxin contamination problem. Pangenomes can be classified as open pangenomes or closed pangenomes [59]. If the number of accessory and core genes of a pangenome are stable and not increasing after adding an optimum number of genomes, then it is referred to as a closed pangenome. In contrast, if the number of accessory and core genes of a pangenome are substantially increasing upon addition of individual genomes then it is referred as an open pangenome. An open pangenome can be substantially improved by sequencing more individuals resulting in more accessory and core genes being included in the existing pangenome [60]. In AflaPan, the number of accessory and core genes were not substantially increased after adding 300 A. flavus genomes. Therefore, AflaPan is a closed pangenome, and effectively captured the overall genomic diversity of the A. flavus species. In contrast, the pangenome of the wheat fungal pathogen Zymoseptoria tritici was developed with only 19 isolates which can be improved further by including more isolates [61]. Similarly, recently another open pangenome of an onion bacterial pathogen Pantoea ananatis was developed using 81 P. ananatis genomes which showed an increase of 50 new genes with addition of each isolate [62]. While exploring possibilities for genome expansion or horizontal gene transfer [63] in A. flavus genomes, the isolates from soil samples had larger genomes in comparison with isolates collected from infected corn tissues. The expanded soil isolates’ genomes included novel aflatoxin producing gene clusters identified in the Pan-GWAS analysis using AflaPan. One important objective behind developing AflaPan was to conduct Pan-GWAS study and to identify unreported aflatoxin producing gene clusters. In this study, the newly sequenced isolates of 225 were used for phenotyping aflatoxin production and Pan-GWAS analysis to identify the gene clusters associated with aflatoxin production. Pan-GWAS analysis for aflatoxin production, using a presence/absence matrix based on 17,855 pan-genes across 225 isolates, identified 391 pan-genes significantly associated with aflatoxin production. Among them, 369 pan-genes (94.4%) were located in the shell genome of AflaPan. Our findings suggest that the shell genome of AflaPan is enriched with aflatoxin producing gene clusters. Shell genome includes pan-genes which are present in 95% of the total individuals ( Supplementary Table 7 ). Interestingly, a total of 256 new pan-genes which are not annotated in present reference genomes NRRL3357 and AF13 were significantly associated with aflatoxin production. These 256 pan-genes could be important targets for developing aflatoxin mitigation strategies using gene silencing or genome editing approaches. Earlier, Pan-GWAS has been conducted in several fungi as well as bacterial pathogens to identify causal genes. For instance, Pan-GWAS in Pantoea ananatis pangenome identified 14 strongly associated genes in HiVir/PASVIL clusters responsible for pathogenicity in onion [61]. Recently, human pathogen A. fumigatus has developed resistance against fungicide triazoles. Pan-genome wide association analysis using A. fumigatus pangenome of 300 isolates identified 12 genes associated with triazole resistance [19]. Another, Pan-GWAS using A. fumigatus pangenome of 218 isolates identified azole drug resistant genes in ergosterol biosynthetic pathway [20]. Two genes in A. flavus , aflS and aflR , mainly regulate the aflatoxin biosynthesis pathway [64], with aflR encoding a DNA-binding binuclear zinc cluster (Zn(II) 2 Cys 6 ) transcription factor required for expression of a majority of structural genes in aflatoxin biosynthesis pathway [65]. Here, Pan-GWAS analysis using AflaPan identified significant association with aflR with -log10( p ) = 7.6. aflS has also been shown to be essential for aflatoxin biosynthesis and is required for activation of aflR [66]. Disruption of aflS in A. parasiticus resulted in lower expression of some AFB1 producing genes, such as aflC , aflD , aflM and aflP and reduced aflatoxin levels [67]. It is also observed that aflS is strongly associated with aflatoxin production with -log10( p ) =7.64. In addition, there are several aflatoxin biosynthesis genes including aflV , aflT , aflA and aflO which showed strong associations with aflatoxin production according to the Pan-GWAS analysis ( Table 2 ). aflV showed to be associated with aflatoxin production with -log10( p ) =6.3. aflV encodes cytochrome P450 monoxygenase ( cypX ) has been shown to catalyze the reaction from averufin (AVR) to hydroxyversicolorone (HVN) and HVN further involves in aflatoxin biosynthesis [68]. It is reported that aflT is not directly linked with aflatoxin biosynthesis, though it is present in aflatoxin producing gene cluster [69]. However, Pan-GWAS analysis indicated that aflT is significantly associated with aflatoxin production with -log10( p ) =7.5. The fas (fatty acid synthases) genes, namely aflA also referred as ‘ fas-2 ’ and aflB referred as ‘ fas-1 ’ produces α and β sub-units, respectively, both subunits transform the hexanoate units into a polyketide structure in aflatoxin biosynthesis pathway [70]. In this study, aflA ( fas-2 ) showed strong association with aflatoxins with -log10( p ) =7.6. The aflO ( omtB ) homolog in A. nidulans required for conversion of demethylsterigmatocystin (DMST) to sterigmatocystin (ST) in A. nidulans [71]. Pan-GWAS analysis showed significant association with previously reported aflatoxin producing gene clusters [72] such as O-methyltransferase A ( OmtA/AflP ) and versicolorin dehydrogenase /ketoreductase ( Ver1/AflM ) with –log10( p ) value of 7.4 and 6.8, respectively. OmtA/AflP and Ver1/AflM had been targeted for host-induced gene silencing (HIGS) in peanuts resulting in reduced aflatoxin contamination [73]. Moreover, Ver1/AflM was also targeted for host induced gene silencing in maize [74]. A. flavus is also a reservoir for important secondary metabolites [75]. In this study, we investigated 10 secondary metabolites and their associated genes, investigating the importance of AflaPan genome. Interestingly, in AflaPan, there were 165 pan-genes annotated to associate with production of 10 important secondary metabolites ( Supplementary Table 6 ). Recently, pan-secondary metabolome based on 94 A. flavus isolates identified 7,821 biosynthetic gene clusters (BGCs) responsible for production of variety of metabolites [11]. The secondary metabolite producing gene clusters for these 10 metabolites were identified in core as well as accessory genomes of this AflaPan. Kojic acid (KA) has commercial use in industries including cosmetics, especially in skin care product to prevent the exposure to UV radiations [56]. In AflaPan, three pan-genes were identified as Kojic acid biosynthesis genes. In addition, the pan-genes also associated with the production of aflavarin, aflatoxin, asparasone, cyclopiazonic acid, ditryptophenaline, leporins, piperazines, and ustiloxin. In summary, this study provides complete pangenome framework for the species of Aspergillus flavus along with associated genes for pathogen survival and aflatoxin production. This pangenome can be used for performing further studies on target traits as well as developing required diagnostic genotyping panel for isolate identification. Furthermore, the information generated through this study can also be used for identifying individuals from different diversity clusters for developing individual genome assemblies leading development of pangenome reference. Most importantly, the newly identified aflatoxin producing gene clusters will be a new source for seeking aflatoxin mitigation strategies and needs new attention in research. Declarations COMPETING INTEREST The authors declare that they have no competing interests ETHICS APPROVAL AND CONSENT TO PARTICIPATE The plant samples and soil samples were collected from the fields with permission from field owners. CONSENT FOR PUBLICATION All authors approved the manuscript for publication Availability of data and material Analyzed data on sequencing statistics, assembly statistics, SNP statistics and Pan-GWAS analysis are provided in the attached supplementary files. The newly developed genome assemblies for 225 isolates used in this study are available at https://zenodo.org/deposit/7615243. Associated raw sequencing data for each isolate are available through National Center for Biotechnology Information (NCBI) - Sequence Read Archive (SRA) with Bioproject ID PRJNA915632. Fungal isolates are available upon request by contacting corresponding author. The 121 isolates from public data can be requested from corresponding authors. Competing interests The authors declare that they have no competing interests. Funding This work is partially supported by the U.S. Department of Agriculture Agricultural Research Service (USDA-ARS). Georgia Agricultural Commodity Commission for Corn, National Corn Growers Association, Aflatoxin Mitigation Center of Excellence (AMCOE), Georgia Peanut Commission, National Peanut Board, The Peanut Research Foundation and MARS-Wrigley. References Fountain JC, Scully BT, Ni X, Kemerait RC, Lee RD, Chen ZY, Guo B. Environmental influences on maize- Aspergillus flavus interactions and aflatoxin production. Front Micro. 2014;5:40. Liu Y, Wu F. Global burden of aflatoxin-induced hepatocellular carcinoma: a risk assessment. Environ Health Perspect. 2019;118:818–24. Burdock GA, Flamm WG. Safety assessment of the mycotoxin cyclopiazonic acid. Int J Toxicol. 2000;19:195–218. Pandey MK, Kumar R, Pandey AK, Soni P, Gangurde SS, Sudini HK, et al. Mitigating aflatoxin contamination in groundnut through a combination of genetic resistance and post-harvest management practices. Toxins. 2019;11:315. Guo BZ, Yu J, Holbrook CC, Cleveland TE, Nierman WC, Scully BT. Strategies in prevention of preharvest aflatoxin contamination in peanuts: Aflatoxin biosynthesis, genetics and genomics. Peanut Sci. 2009;36:11–20. Frisvad JC, Hubka V, Ezekiel CN, Hong SB, Nováková A, Chen AJ, et al. Taxonomy of Aspergillus section Flavi and their production of aflatoxins, ochratoxins and other mycotoxins. Stud Mycol. 2019;93:1–63. Payne GA, Nierman WC, Wortman JR, Pritchard BL, Brown D, Dean RA, et al. Whole genome comparison of Aspergillus flavus and A. oryzae . Med Mycol. 2006;44:9–S11. Chang PK, Ehrlich KC. What does genetic diversity of Aspergillus flavus tell us about Aspergillus oryzae ? Int J Food Microbiol. 2010;138:189–99. Gibbons JG, Salichos L, Slot JC, Rinker DC, McGary KL, King JG, et al. The evolutionary imprint of domestication on genome variation and function of the filamentous fungus Aspergillus oryzae . Curr Biol. 2012;22:1403–9. Khan AW, Garg V, Roorkiwal M, Golicz AA, Edwards D, Varshney RK. Super-Pangenome by integrating the wild side of a species for accelerated crop improvement. Trends Plant Sci. 2020;25:148–58. Drott MT, Rush TA, Satterlee TR, Giannone RJ, Abraham PE, Greco C, et al. Microevolution in the pansecondary metabolome of Aspergillus flavus and its potential macroevolutionary implications for filamentous fungi. Proc Natl Acad Sci USA. 2021;118:e2021683118. Nierman WC, Yu J, Fedorova-Abrams ND, Losada L, Cleveland TE. Genome sequence of Aspergillus flavus NRRL 3357, a strain that causes aflatoxin contamination of food and feed. Genome Announc. 2015;3:e00168–15. Skerker JM, Pianalto KM, Mondo SJ, Yang K, Arkin AP, Keller NP, et al. Chromosome assembled and annotated genome sequence of Aspergillus flavus NRRL 3357. G3. Genes Genomes Genet. 2021;11:213. Fountain JC, Clevenger JP, Nadon B, Youngblood RC, Korani W, Chang PK, et al. Two new Aspergillus flavus reference genomes reveal a large insertion potentially contributing to isolate stress tolerance and aflatoxin production. Genes Genomes Genet. 2020a;G3:10: 3515–31. Balázs A, Pócsi I, Hamari Z, Leiter É, Emri T, Miskei M, et al. AtfA bZIP-type transcription factor regulates oxidative and osmotic stress responses in Aspergillus nidulans . Mol Genet Genomics. 2010;283:289–303. Wee JS, Hong Y, Roze L, Day D, Chanda A, et al. The fungal bZIP transcription factor AtfB controls virulence-associated processes in Aspergillus parasiticus . Toxins. 2017;9:287. Fountain JC, Bajaj P, Nayak SN, Yang L, Pandey MK, et al. Responses of Aspergillus flavus to oxidative stress are related to fungal development regulator, antioxidant enzyme, and secondary metabolite biosynthetic gene expression. Front Microbiol. 2016;7:2048. McCarthy C, Fitzpatrick DA. Pan-genome analyses of model fungal species. Microb Genom. 2019;5:e000243. Barber AE, Sae-Ong T, Kang K, Seelbinder B, Li J, Walther G, et al. Aspergillus fumigatus pan-genome analysis identifies genetic variants associated with human infection. Nat Microbiol. 2021;6:1526–36. Rhodes J, Abdolrasouli A, Dunne K, Sewell TR, Zhang Y, Ballard E, et al. Population genomics confirms acquisition of drug-resistant Aspergillus fumigatus infection by humans from the environment. Nat Microbiol. 2022;7:663–74. Lofgren LA, Ross BS, Cramer RA, Stajich JE. The pan-genome of Aspergillus fumigatus provides a high-resolution view of its population structure revealing high levels of lineage-specific diversity driven by recombination. PLoS Biol. 2022;20:e3001890. Abbas HK, Weaver MA, Zablotowiez RM, Horn BW, Shier WT. Relationship between aflatoxin production and sclerotia formation among isolates of Aspergillus section Flavi from the Mississippi Delta. Eur J Plant Path. 2005;112:283–7. Horn BW, Dorner JW. Soil populations of Aspergillus species from section Flavi along a transect through peanut-growing regions of the United States. Mycologia. 1998;90:767–76. Fountain JC, Clevenger JP, Nadon B, Wang H, Abbas HK, et al. Draft genome sequences of one Aspergillus parasiticus isolate and nine Aspergillus flavus isolates with varying stress tolerance and aflatoxin production. Microbiol Resour Announc. 2020b;9:e00478–20. Bolger AM, Lohse M, Usadel B, Trimmomatic. A flexible trimmer for Illumina Sequence Data. Bioinformatics. 2014;30:2114–20. Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, et al. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 2012;19:455–77. Li H. 2013. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv preprint arXiv:1303.3997. Walker BJ, Abeel T, Shea T, Priest M, Abouelliel A, et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS ONE. 2014;9:e112963. Mikheenko A, Prjibelski A, Saveliev V, Antipov D, Gurevich A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics. 2018;34:142–50. Palmer JM, Stajich J. Funannotate v1. 8.1: Eukaryotic genome annotation. Zenodo 2020; https://doi.org/10.5281/zenodo , 4054262. Bao W, Kojima KK, Kohany O. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mob Dna. 2015;6:1–6. Wicklow DT. Sporogenic germination of sclerotia of Aspergillus flavus and Aspergillus parasiticus . Trans Br Mycol Soc. 1984;82:621–4. Ehrlich KC, Chang PK, Yu J, Cotty PJ. Aflatoxin biosynthesis cluster gene cypA is required for G- aflatoxin formation. Appl Environ Microbiol. 2004;70:6518–24. Chang PK, Cary JW, Bhatnagar D, Cleveland TE, Bennett JW, et al. Cloning of the Aspergillus parasiticus apa-2 gene associated with the regulation of aflatoxin biosynthesis. Appl Environ Microbiol. 1993;59:3273–9. Moore GG, Elliott JL, Singh R, Horn BW, Dorner JW, Stone EA, et al. Sexuality generates diversity in the aflatoxin gene cluster: evidence on a global scale. PLoS Pathog. 2013;9:e1003574. Ramirez-Prado JH, Moore GG, Horn BW, Carbone I. Characterization and population analysis of the mating-type genes in Aspergillus flavus and Aspergillus parasiticus . Fungal Genet Biol. 2008;45:1292–9. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. Genome Project Data Processing Subgroup, the sequence alignment/map format and SAMtools. Bioinformatics. 2009;25:2078–9. Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, et al. From FastQ data to high-confidence variant calls: the genome analysis toolkit best practices pipeline. Curr Protoc Bioinf. 2013;43:11–10. Stamatakis A, RAxML-VI-HPC. Maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22:2688–90. Letunic I, Bork P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 2021;49:293–6. Jombart T, Devillard S, Balloux F. Discriminant analysis of principal components: a new method for the analysis of genetically structured populations. BMC Genet. 2010;11:94. Wickham H. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York. 2016; ISBN 978-3-319-24277-4. Falush D, Stephens M, Pritchard JK. Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies. Genetics. 2003;164:1567–87. Tajima F. Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics. 1989;123:585–95. Danecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA. The variant call format and VCFtools. Bioinformatics. 2011;27:2156–8. Emms DM, Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019;20:238. Gasteiger E, Gattiker A, Hoogland C, Ivanyi I, Appel RD, Bairoch A, ExPASy. The proteomics server for in-depth protein knowledge and analysis. Nucleic Acids Res. 2003;31:3784–8. Basenko EY, Pulman JA, Shanmugasundram A, Harb OS, Crouch K, Starns D. FungiDB: an integrated bioinformatic resource for fungi and oomycetes. J Fungi. 2018;20:39. Jones P, Binns D, Chang HY, Fraser M, Li W, McAnulla C, McWilliam H, et al. InterProScan 5: genome-scale protein function classification. Bioinformatics. 2014;30:1236–40. Bradbury PJ, Zhang Z, Kroon DE, Casstevens TM, Ramdoss Y, Buckler ES. TASSEL: software for association mapping of complex traits in diverse samples. Bioinformatics. 2007;23:2633–5. Henderson CR. Best linear unbiased estimation and prediction under a selection model. Biometrics. 1975;31:423–47. Gangurde SS, Wang H, Yaduru S, Pandey MK, Fountain JC, Chu Y, et al. Nested-association mapping (NAM)‐based genetic dissection uncovers candidate genes for seed and pod weights in peanut ( Arachis hypogaea ). Plant Biotechnol J. 2020;18:1457–71. Yin L. R package CMplot v 3.1.3 2016; https://github.com/YinLiLin/CMplot . Gebru ST, Mammel MK, Gangiredla J, Tartera C, Cary JW, Moore GG, et al. Draft genome sequences of 20 Aspergillus flavus isolates from corn kernels and cornfield soils in Louisiana. Microbiol Resour Announc. 2020;9:e00826–20. TePaske MR, Gloer JB, Wicklow DT, Dowd PF. Aflavarin and β-Aflatrem: new anti-insectan metabolites from the sclerotia of Aspergillus flavus . J Nat Prod. 1992;55:1080–6. Vinokurova NG, Ivanushkina NE, Khmel’Nitskaya II, Arinbasarov MU. Synthesis of α-cyclopiazonic acid by fungi of the genus Aspergillus . Appl Biochem Microbiol. 2007;43:435–8. Rosfarizan M, Mohd SM, Nurashikin S, Madihah MS, Arbakariya BA. Kojic acid: Applications and development of fermentation process for production. Biotechnol Mol Biol Rev. 2010;5:24–37. Chandrika TN, Shrestha SK, Ngo HX, Tsodikov OV, Howard KC, Garneau-Tsodikova S. Alkylated piperazines and piperazine-azole hybrids as antifungal agents. J Med Chem. 2018;61:158–73. Richard GF, Cham. https://doi.org/10.1007/978-3-030-38281-0_12 . Bosi E, Fani R, Fondi M. Defining orthologs and pangenome size metrics. Bacterial Pangenomics: Methods Mol Biol. 2015;1231:191–202. Badet T, Oggenfuss U, Abraham L, McDonald BA, Croll D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen Zymoseptoria tritici . BMC Biol. 2020;18:1–18. Agarwal G, Gitaitis RD, Dutta B. Pan-genome of novel Pantoea stewartii subsp. indologenes reveals genes involved in onion pathogenicity and evidence of lateral gene transfer. Microorganisms. 2021;9:1761. Kelkar YD, Ochman H. Causes and consequences of genome expansion in fungi. Genome Biol Evol. 2012;4:13–23. Georgianna DR, Payne GA. Genetic regulation of aflatoxin biosynthesis: from gene to genome. Fungal Genet Biol. 2009;46:113–25. Price MS, Yu J, Nierman WC, Kim HS, Pritchard B, Jacobus CA, et al. The aflatoxin pathway regulator AflR induces gene transcription inside and outside of the aflatoxin biosynthetic cluster. FEMS Microbiol Lett. 2006;255:275–9. Chang PK. Lack of interaction between AFLR and AFLJ contributes to nonaflatoxigenicity of Aspergillus sojae . J Biotechnol. 2004a;107:245–53. Chang PK. The Aspergillus parasiticus protein AFLJ interacts with the aflatoxin pathway-specific regulator AFLR. Mol Genet Genomics. 2003;268:711–9. Wen Y, Hatabayashi H, Arai H, Kitamoto HK, Yabe K. Function of the cypX and moxY Genes in Aflatoxin Biosynthesis in Aspergillus parasiticus . Appl Environ Microbiol. 2005;71:3192–8. Chang PK, Yu J, Yu JH. aflT , a MFS transporter-encoding gene located in the aflatoxin gene cluster, does not have a significant role in aflatoxin secretion. Fungal Genet Biol 2004b; 41:911–920. Roze LV, Hong SY, Linz JE. Aflatoxin Biosynthesis: Current Frontiers. Annu Rev Food Sci Technol. 2013;4:293–311. Kelkar HS, Keller NP, Adams TH. Aspergillus nidulans stcP encodes an O-methyltransferase that is required for sterigmatocystin biosynthesis. Appl Environ Microbiol. 1996;62:4296–8. Yu J. Current understanding on aflatoxin biosynthesis and future perspective in reducing aflatoxin contamination. Toxins. 2012;4:1024–57. Sharma KK, Pothana A, Prasad K, Shah D, Kaur J, Bhatnagar D, et al. Peanuts that keep aflatoxin at bay: a threshold that matters. Plant Biotechnol J. 2018;16:1024–33. Raruang Y, Omolehin O, Hu D, Wei Q, Han ZQ, Rajasekaran K. Host induced gene silencing targeting Aspergillus flavus aflM reduced aflatoxin contamination in transgenic maize under field conditions. Front Microbiol. 2020;11:754. Cary JW, Gilbert MK, Lebar MD, Majumdar R, Calvo AM. Aspergillus flavus secondary metabolites: more than just aflatoxins. Food Saf. 2018;6:7–32. Tables Table 1 Summary of newly sequenced genome assemblies from corn and soil samples used for development of an Aspergillus flavus pangenome (AflaPan). Feature Soil isolates Corn isolates Minimum Maximum Average Minimum Maximum Average Scaffolds 149 7,949 696.8 111.0 1230 269.7 Contigs 170 8,104 726.3 146.0 1,436.0 295.7 Genome size (Mb) 36.6 43.2 37.6 36.5 38.5 37.0 Average scaffold size (Kb) 5.4 248.6 98.0 27.8 274.1 161.5 Average contig size (Kb) 5.3 217.9 87.8 26.8 253.1 144.4 Contig N50 (Kb) 11.1 1,439.6 307.4 346.3 1,054.5 603.7 Scaffold N50 (Kb) 11.1 2,316.2 406.2 378.6 1,367.5 775.6 Table 2 Pan-GWAS identified orthogroups associated with aflatoxin (ppb) production in Aspergillus flavus Orthogroup a -log10(p) b Genome c Length d Annotation e OG0012151 7.64 Shell 444 Aflatoxin biosynthesis protein R ( aflR ) OG0012150 7.64 Shell 438 Aflatoxin biosynthesis protein S ( aflS ) OG0011833 7.06 Shell 458 Fusaristatin A biosynthesis cluster protein OG0000052 9.97 Soft-core 764 Aspyridones biosynthesis protein F OG0012617 6.33 Shell Novel 493 Neosartiricin biosynthesis protein R OG0012284 6.54 Shell 317 Cyclochlorotine biosynthesis protein O OG0012070 7.16 Shell 534 Cytochrome P450 monooxygenase; Acurin A biosynthesis cluster protein ( acrF ) OG0012018 6.27 Shell 508 Cytochrome P450 monooxygenase aflV; Aflatoxin biosynthesis protein ( aflV ) OG0012566 7.12 Shell 470 Cytochrome P450 monooxygenase; AK-toxin biosynthesis protein 7 ( AKT7 ) OG0000168 6.55 Shell Novel 521 Cytochrome P450 monooxygenase aneD; Aculenes biosynthesis cluster protein D ( aneD ) OG0010996 16.78 Soft-core 215 Cytochrome P450 monooxygenase apdE; Aspyridones biosynthesis protein E ( apdE ) OG0013427 22.40 Shell Novel 341 Dehydrogenase azaJ; Azaphilone biosynthesis cluster protein azaJ ( azaJ ) OG0011609 7.22 Shell Novel 345 Dehydrogenase RED2; T-toxin biosynthesis protein (RED2) OG0012391 7.34 Shell 116 Demethylsterigmatocystin 6-O-methyltransferase; Aflatoxin biosynthesis protein O ( aflO ) OG0012144 7.47 Shell 548 Efflux pump aflT; Aflatoxin biosynthesis protein T ( aflT ) OG0013752 6.10 Shell Novel 261 Efflux pump atB; Terreic acid biosynthesis cluster protein B ( atB ) OG0012168 6.89 Shell Novel 541 Efflux pump mlcE; Compactin biosynthesis protein E ( mlcE ) OG0011998 6.55 Shell 651 Efflux pump radE; Radicicol biosynthesis cluster protein ( radE ) OG0012554 7.92 Shell Novel 1,102 Efflux pump terG; Terrein biosynthesis cluster protein ( terG ) OG0012658 6.81 Shell Novel 508 Efflux pump terG; Terrein biosynthesis cluster protein ( terG ) OG0012155 9.12 Shell Novel 208 Enterobactin synthase component B; Enterobactin biosynthesis bifunctional protein ( EntB ) OG0012119 7.64 Shell 1,737 Fatty acid synthase alpha subunit aflA; Aflatoxin biosynthesis protein A ( aflA ) OG0012283 6.46 Shell Novel 636 Fumagillin dodecapentaenoate synthase; Fumagillin biosynthesis polyketide synthase OG0014250 10.44 Shell Novel 220 Methyltransferase pytC; Pyranterreones biosynthesis cluster protein C ( pytC ) OG0012525 6.62 Shell 330 MFS gliotoxin efflux transporter gliA; Gliotoxin biosynthesis protein A ( gliA ) OG0012623 6.44 Shell 497 MFS transporter cpaT; Cyclopiazonic acid biosynthesis cluster protein T ( cpaT ) OG0013082 7.78 Shell Novel 317 Notoamide biosynthesis activator; Notoamide biosynthesis cluster protein L ( NotL ) OG0011985 6.13 Shell 643 Dehydrogenase patE; Patulin biosynthesis cluster protein E ( patE ) OG0000472 6.52 Soft-core 720 Peptide transporter imqD; Imizoquin biosynthesis cluster protein D ( imqD ) OG0012742 7.29 Shell Novel 707 Prenyltransferase phnF; Phenalenone biosynthesis cluster protein F ( phnF ) OG0012565 7.81 Shell 251 Probable tetra hydroxynaphthalene reductase; Conidial pigment biosynthesis cluster OG0012428 6.26 Shell 286 Aspercryptin biosynthesis cluster protein L ( atnL ) OG0012853 17.19 Shell 726 Cyclochlorotine biosynthesis protein T ( cctT ) OG0012547 6.95 Shell 1,060 Melleolides biosynthesis cluster protein OG0012333 6.85 Shell Novel 599 Fusicoccin A biosynthetic gene clusters protein OG0012630 6.81 Shell Novel 451 Satratoxin biosynthesis SC3 cluster protein 17 ( SAT17 ) a Orthogroups of AflaPan significantly associated with toxin production b -log10 transformation of p values identified in Pan-GWAS analysis c Genome (Core, soft-core, shell or cloud) of the orthogroups. Where, the new genes are denoted by suffix novel. For instance, shell novel means, the new orthogroups in the shell genome. Whereas, ‘Shell’ means the orthogroups already existing in previous reference genomes. d Protein sequence length of orthogroup e Annotation of orthogroups Additional Declarations No competing interests reported. Supplementary Files PanGWASSupplementaryTables.xlsx Supplementary Material Supplementary Table 1.Metadata of the 346 Aspergillus flavus isolates used in development of AflaPan genome. Supplementary Table 2.Statistics of genome assemblies of 346 Aspergillus flavus isolates Supplementary Table 3. Statistics of single nucleotide polymorphism (SNP) calls across 346 isolates of A. flavus relative to the AF13 reference assembly. Supplementary Table 4. Protein sequences of 17,855 orthogroups in AflaPan genome Supplementary Table 5.Annotation of 17,855 orthogroups in AflaPan genome Supplementary Table 6. Secondary metabolite producing pan-genes in AflaPan for 10 selected metabolites Supplementary Table 7.Pan-genome wide association analysis identified orthogroups significantly associated with aflatoxin production in culture. SupplementaryFigures.docx Cite Share Download PDF Status: Published Journal Publication published 30 Apr, 2024 Read the published version in BMC Plant Biology → Version 1 posted Editorial decision: Revision requested 12 Mar, 2024 Reviews received at journal 12 Mar, 2024 Reviews received at journal 26 Feb, 2024 Reviewers agreed at journal 23 Feb, 2024 Reviewers agreed at journal 23 Feb, 2024 Reviewers invited by journal 22 Feb, 2024 Editor assigned by journal 20 Feb, 2024 Submission checks completed at journal 20 Feb, 2024 First submitted to journal 15 Feb, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3958535","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":274012581,"identity":"0036b253-91d0-432c-9fe9-fc0084bb828c","order_by":0,"name":"Sunil S. Gangurde","email":"","orcid":"","institution":"University of Georgia","correspondingAuthor":false,"prefix":"","firstName":"Sunil","middleName":"S.","lastName":"Gangurde","suffix":""},{"id":274012582,"identity":"454f8c2f-434d-491a-bde9-748f1676928c","order_by":1,"name":"Walid Korani","email":"","orcid":"","institution":"HudsonAlpha Institute for Biotechnology","correspondingAuthor":false,"prefix":"","firstName":"Walid","middleName":"","lastName":"Korani","suffix":""},{"id":274012583,"identity":"2e2ad805-0d01-479d-ad40-1a52f0d72aef","order_by":2,"name":"Prasad Bajaj","email":"","orcid":"","institution":"International Crop Research Institute for the Semi-Arid Tropics (ICRISAT)","correspondingAuthor":false,"prefix":"","firstName":"Prasad","middleName":"","lastName":"Bajaj","suffix":""},{"id":274012584,"identity":"f6970e4b-6859-4d7b-a05d-a1d68ca502df","order_by":3,"name":"Hui Wang","email":"","orcid":"","institution":"University of Georgia","correspondingAuthor":false,"prefix":"","firstName":"Hui","middleName":"","lastName":"Wang","suffix":""},{"id":274012585,"identity":"87dffc8e-0724-4f75-9da6-b5d588574fc8","order_by":4,"name":"Jake C. Fountain","email":"","orcid":"","institution":"University of Georgia","correspondingAuthor":false,"prefix":"","firstName":"Jake","middleName":"C.","lastName":"Fountain","suffix":""},{"id":274012586,"identity":"7ad9f11a-672a-4b0f-9b42-97f5c6f46820","order_by":5,"name":"Gaurav Agarwal","email":"","orcid":"","institution":"Michigan State University","correspondingAuthor":false,"prefix":"","firstName":"Gaurav","middleName":"","lastName":"Agarwal","suffix":""},{"id":274012587,"identity":"6b521ffb-1521-469e-8197-e5648c529a22","order_by":6,"name":"Manish K. Pandey","email":"","orcid":"","institution":"International Crop Research Institute for the Semi-Arid Tropics (ICRISAT)","correspondingAuthor":false,"prefix":"","firstName":"Manish","middleName":"K.","lastName":"Pandey","suffix":""},{"id":274012588,"identity":"1093e77a-72b9-4488-838a-a3303671adf0","order_by":7,"name":"Hamed K. Abbas","email":"","orcid":"","institution":"USDA-ARS","correspondingAuthor":false,"prefix":"","firstName":"Hamed","middleName":"K.","lastName":"Abbas","suffix":""},{"id":274012589,"identity":"6a1b0880-762c-48fb-96a1-05143dc39d67","order_by":8,"name":"Perng-Kuang Chang","email":"","orcid":"","institution":"USDA-ARS","correspondingAuthor":false,"prefix":"","firstName":"Perng-Kuang","middleName":"","lastName":"Chang","suffix":""},{"id":274012590,"identity":"0fbed524-317c-4d58-9728-1277efa49c50","order_by":9,"name":"C. Corley Holbrook","email":"","orcid":"","institution":"USDA-ARS","correspondingAuthor":false,"prefix":"","firstName":"C.","middleName":"Corley","lastName":"Holbrook","suffix":""},{"id":274012591,"identity":"c5f30362-a8ee-410c-9bcc-b53bd936210e","order_by":10,"name":"Robert C. Kemerait","email":"","orcid":"","institution":"University of Georgia","correspondingAuthor":false,"prefix":"","firstName":"Robert","middleName":"C.","lastName":"Kemerait","suffix":""},{"id":274012592,"identity":"146603b1-7895-4615-add6-6c5a0eefa4af","order_by":11,"name":"Rajeev K. Varshney","email":"","orcid":"","institution":"Murdoch University","correspondingAuthor":false,"prefix":"","firstName":"Rajeev","middleName":"K.","lastName":"Varshney","suffix":""},{"id":274012593,"identity":"43e92b76-cebe-4355-a269-d249736dcebc","order_by":12,"name":"Bhabesh Dutta","email":"","orcid":"","institution":"University of Georgia","correspondingAuthor":false,"prefix":"","firstName":"Bhabesh","middleName":"","lastName":"Dutta","suffix":""},{"id":274012594,"identity":"4d2c984b-c383-4a1d-afea-f30f2f5f42f9","order_by":13,"name":"Josh P. Clevenger","email":"","orcid":"","institution":"HudsonAlpha Institute for Biotechnology","correspondingAuthor":false,"prefix":"","firstName":"Josh","middleName":"P.","lastName":"Clevenger","suffix":""},{"id":274012595,"identity":"2bd9f2c6-7ec8-4bd5-be10-de44248132ab","order_by":14,"name":"Baozhu Guo","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIiWNgGAWjYBACeyBmBmJ5IHWAgYGNCC2GDRAtQJotgTgtBgcgWoDKeQyItKX98OHPhW0MCQbHe75J/iizY5B37zHAq8WeJy3BeCZIy5mz26R5ziUzGJ45g1+LYUOOQTJvG0Pyhhu526QZ25gZDGekJeD3y/k3BoeBWhI33Mh5JvmzrZ4ILTdyDJuhWtgkeNsOM8hLJB/A77AZz5KZec5JGM48c8zYmufccR4DnsP4tdjzJx/+zFNmI893vPnhzR9l1XLy7Y0NeLVAgQScxWOA3w5sQJ4oO0bBKBgFo2AkAQDVTUW1if8bQgAAAABJRU5ErkJggg==","orcid":"","institution":"USDA-ARS","correspondingAuthor":true,"prefix":"","firstName":"Baozhu","middleName":"","lastName":"Guo","suffix":""}],"badges":[],"createdAt":"2024-02-15 11:46:24","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3958535/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3958535/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12870-024-04950-8","type":"published","date":"2024-05-01T00:46:08+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":51474097,"identity":"ee3602bf-320c-4461-b123-69f00bba2d14","added_by":"auto","created_at":"2024-02-22 08:55:54","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":7481243,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIsolate collection and diversity analysis. \u003c/strong\u003eIn this study, among 346 isolates, a total of 225 isolates were newly collected and sequenced, and 121 isolates were downloaded from NCBI-SRA. Of the 225 newly sequenced isolates, a) 98 isolates from various infected plant parts (Tassel, leaf, dust and silk) of corn, b) 127 isolates from soil samples from fields of various cropping systems of either rotation or continuous cropping of crops such as cotton, peanut, corn, sunflower, soybean, sorghum from three counties namely Tifton, Turner, Irwin, Berrien. c) A total of 346 isolates represents a total of 10 states of United States with representatives aflatoxigenic and non-aflatoxigenic isolates from each state. d) Population structure analysis identified four groups. Each color represents one group. Four groups are labelled as P1, P2, P3, and P4 respectively. e) Discriminant analysis of principle components confirmed four groups in the \u003cem\u003eAspergillus flavus\u003c/em\u003e population. f) Linkage disequilibrium (LD) decay estimates in each population. Pairwise linkage disequilibrium (r\u003csup\u003e2\u003c/sup\u003e) was calculated for each SNP in each population individually, in a window of 10Kb using plink. Four lines represents LD of four genetic populations. g) Circos plot illustrates nucleotide diversity (\u003cem\u003eπ\u003c/em\u003e) in each population. From outside to inside the tracks, chr represents 8 chromosomes of \u003cem\u003eA. flavus\u003c/em\u003e, tracks from P1 to P4 represents nucleotide diversity present in four genetic populations of \u003cem\u003eA. flavus\u003c/em\u003e. h) Means of nucleotide diversity (\u003cem\u003eπ\u003c/em\u003e) and \u003cem\u003eFst\u003c/em\u003e values calculated between the groups. The grey color links between the groups shows the \u003cem\u003eFst\u003c/em\u003e values between the groups. While the values inside the circles indicates the mean nucleotide diversity (\u003cem\u003eπ\u003c/em\u003e) present in that particular group. e) Box plots illustrates distribution of Tajima’D measures calculated in 10kb window for the groups P1, P2, P3 and P4.\u003c/p\u003e","description":"","filename":"Figure1diversity.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/b1472878da18fbdea631ace0.jpg"},{"id":51474104,"identity":"d2aae2a8-4b4b-499d-978f-001ae316bf60","added_by":"auto","created_at":"2024-02-22 08:55:55","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":10838919,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWhole genome phylogeny of soil and corn \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eA. flavus\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e isolates. \u003c/strong\u003eA whole genome phylogeny constructed using ~1.02 million SNPs called on whole genome sequencing data on 346 isolates. The tracks from inside to outside shows information of isolates on name of isolate, origin, source, sclerotia types, mating types, toxicity, genome size and number scaffolds in each genome. 1) Isolates- represents the name of each isolates, 2) Origin- illustrates the geographic site of each isolate from 10 states of United States, 3) Source- shows the information on whether it is isolated from corn (tassel, leaf, silk, dust) and soil, 4) Sclerotia- shows information on sclerotia types short sclerotia (S-types) and long sclerotia (L-types), 5) Mating types- explains the diversity of isolates based on two mating types namely, \u003cem\u003eMAT1-1\u003c/em\u003e and \u003cem\u003eMAT1-2\u003c/em\u003e, 6) Toxicity of isolates- illustrates the diversity based on aflatoxigenic (toxin producers) and non-aflatoxigenic (not produce toxins), 7) Genome size (Mb)- the bar plots shows the genome size of each isolate in million base pairs (Mb), 8) Scaffolds- number of scaffolds in genome assembly of each isolate.\u003c/p\u003e","description":"","filename":"Figure2dendrogram.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/fa2f7fa3ae142f928b7b2f2e.jpg"},{"id":51474098,"identity":"0b801678-968c-45b1-96a3-e8b39358302e","added_by":"auto","created_at":"2024-02-22 08:55:54","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":7057249,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of the pangenome of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eAspergillus flavus\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e (AflaPan). \u003c/strong\u003ea) Total orthologous gene clusters and proportions of core, soft-core, shell, clouds and singletons genes in AflaPan genome. b) A bi-color plot illustrates presence/absence matrix of 17,855 orthologous gene clusters identified across 346 \u003cem\u003eA. flavus\u003c/em\u003e genomes. Vertical axis represents number of genomes, while horizontal axis represents number of orthologous gene clusters, categorized into core, soft-core, shell and cloud. Where, core orthologous gene clusters are present in all 346 genomes, softcore (orthogroups present in \u0026gt;95% of the isolates), shell (orthogroups present in 5–95% of the isolates) and cloud (orthogroups present in less than 5% of the isolates) genomes. c) Stacked bar graph shows proportion of novel gene clusters when compared with current reference genome AF13. d) Comparison of gene modules of AflaPan genome and current reference genomes of \u003cem\u003eA. flavus, \u003c/em\u003eAF13 and NRRL3357. In AF13 and NRRL3357 there are a total 13,188 and 13,487 genes annotated respectively, in AflaPan we annotated 17,855 orthogroups, e) Comparison of length of protein sequences in core and accessory genomes. There is significant difference between the length of protein sequences in core and accessory genomes. The length of protein sequences of core genome is significantly longer than accessory genome. f) Orthologous gene clusters in accessory and core genome increases with increasing number of genomes. Blue and red color represents pan- accessory and core genes respectively. The black dotted line represents the range. g) New gene clusters were decreased with increasing the number of genomes in AflaPan genome. After adding 300 genomes the there was no significant increase in the new genes in AflaPan.\u003c/p\u003e","description":"","filename":"Figure3aflapan.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/6835ce983d50ef4651a52998.jpg"},{"id":51474102,"identity":"717a51bb-4049-434c-9393-2a50678e6ca4","added_by":"auto","created_at":"2024-02-22 08:55:55","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1994912,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePan-GWAS analysis for isolate aflatoxin production in culture. \u003c/strong\u003ea) Manhattan plot for aflatoxins (ppb), b) QQ plot for total toxins (ppb). On X-axis represents the portions of soft-core, shell, clouds and singletons. Y-axis represents the \u003cem\u003ep\u003c/em\u003e values transformed as –log10(\u003cem\u003ep\u003c/em\u003e). The novel genes when compared with AF13 reference genome in the accessory genomes were visualized in separate groups with suffix ‘Novel’. For instance, soft-core_Novel represents the new genes in soft-core genome, Shell_Novel represents new genes in shell genome, Cloud_Novel represents new genes in cloud genome. Single_novel represents novel singletons. The Bonferroni threshold is marked at –log10(\u003cem\u003ep\u003c/em\u003e) = 2.8 × 10\u003csup\u003e-6\u003c/sup\u003e with dotted red line.\u0026nbsp;\u003c/p\u003e","description":"","filename":"Figure4pangwas.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/1e22e5e41c641bc228f7e497.jpg"},{"id":55697911,"identity":"38f7ee44-930e-4ab1-b022-3e6d83990d19","added_by":"auto","created_at":"2024-05-02 02:25:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1784797,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/cb5355b2-7c22-49ee-ad4d-9d5fd0b5977b.pdf"},{"id":51474105,"identity":"d289d19b-726a-4e2c-bc69-a505d4cd2457","added_by":"auto","created_at":"2024-02-22 08:55:56","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":6747560,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary Material\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 1.\u003c/strong\u003eMetadata of the 346 \u003cem\u003eAspergillus flavus\u003c/em\u003eisolates used in development of AflaPan genome.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 2.\u003c/strong\u003eStatistics of genome assemblies of 346 \u003cem\u003eAspergillus flavus\u003c/em\u003e isolates\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 3.\u003c/strong\u003e Statistics of single nucleotide polymorphism (SNP) calls across 346 isolates of \u003cem\u003eA. flavus \u003c/em\u003erelative to the AF13 reference assembly.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 4. \u003c/strong\u003eProtein sequences of 17,855 orthogroups in AflaPan genome\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 5.\u003c/strong\u003eAnnotation of 17,855 orthogroups in AflaPan genome\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 6. \u003c/strong\u003eSecondary metabolite producing pan-genes in AflaPan for 10 selected metabolites\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Table 7.\u003c/strong\u003ePan-genome wide association analysis identified orthogroups significantly associated with aflatoxin production in culture.\u003c/p\u003e","description":"","filename":"PanGWASSupplementaryTables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/4985d49fb10d38f5a0f506c3.xlsx"},{"id":51474108,"identity":"5f5895a0-38c6-4419-a9cd-ed990756eccf","added_by":"auto","created_at":"2024-02-22 08:55:56","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":5429642,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-3958535/v1/805e06dc241b24efd494a897.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Aspergillus flavus pangenome (AflaPan) uncovers novel aflatoxin and secondary metabolite associated gene clusters","fulltext":[{"header":"Introduction","content":"\u003cp\u003e \u003cem\u003eAspergillus flavus\u003c/em\u003e is an opportunistic saprophyte, infects varieties of crops such as maize and peanuts both pre- and postharvest [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. These fungi also produce secondary metabolites including aflatoxins which contaminate the grains and are toxic and carcinogenic leading to inducing acute and chronic health issues in both humans and animals [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. In addition to aflatoxins, \u003cem\u003eA. flavus\u003c/em\u003e also produces cyclopiazonic acid (CPA), which harms liver, kidneys, and gastrointestinal tract [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Because of its harmful impact on human health, the aflatoxin is considered as \u0026ldquo;hidden and slow poison\u0026rdquo; and the research community is trying its best to managing through this menace through a combination of improved genetics and post-harvest (farm to industry) management on one side while understanding the reason of compulsion for pathogen to produce toxin on the other side [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Aflatoxin contamination is an annual issue in the southern United States and throughout the world, but occasional outbreaks happen in the \u0026ldquo;cooler areas\u0026rdquo; like mid-western Corn Belt of the United States [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cem\u003eAspergillus flavus\u003c/em\u003e belongs to \u003cem\u003eAspergillus\u003c/em\u003e section \u003cem\u003eFlavi\u003c/em\u003e, which splits into eight clades and currently contains 33 species. Of these, two species only produce aflatoxin B1 and B2 (\u003cem\u003eA. pseudotamarii\u003c/em\u003e and \u003cem\u003eA. togoensis\u003c/em\u003e), and 14 species are able to produce aflatoxins B1, B2, G1 and G2 [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Large differences in gene contents can occur among genomes of \u003cem\u003eA. flavus\u003c/em\u003e isolates, with only a portion of genes being universal, or core, to all genomes. For instance, \u003cem\u003eA. flavus\u003c/em\u003e and \u003cem\u003eA. oryzae\u003c/em\u003e possess core identical gene models and sequences, however \u003cem\u003eA. flavus\u003c/em\u003e is a very devastating fungal pathogen produces aflatoxin, while \u003cem\u003eA. oryzae\u003c/em\u003e is a beneficial fungus used widely for fermentation in food industries [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Interestingly, comparative genomic analysis of these two Aspergilli clearly shows that \u003cem\u003eA. oryzae\u003c/em\u003e is a domesticated ecotype of wild \u003cem\u003eA. flavus\u003c/em\u003e [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Therefore, a single reference genome could not represent complete diversity of a species [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. A recent pan-metabolome analysis of 94 \u003cem\u003eA. flavus\u003c/em\u003e isolates support this with identification of 7821 biosynthetic gene clusters (BGCs) including 25% population specific BGCs and 92 unique BGCs [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cem\u003eA. flavus\u003c/em\u003e isolate NRRL 3357 was used as a reference genome for multiple omics studies [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Recently, the NRRL 3357 assembly was updated and re-annotated for eight chromosomes [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. At the same time, [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] reported two new \u003cem\u003eA. flavus\u003c/em\u003e reference genomes comparatively, which revealed a large insertion potentially contributing to stress tolerance and aflatoxin production in isolate AF13, a \u003cem\u003eMAT1-2\u003c/em\u003e mating type in comparison with NRRL 3557, a \u003cem\u003eMAT1-1\u003c/em\u003e mating type. This study between AF13 and NRRL 3357 confirmed that AF13 has an insertion of 310 kb with a bZIP transcription factor named \u003cem\u003eatfC\u003c/em\u003e [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], which shares homology with \u003cem\u003eA. flavus\u003c/em\u003e bZIP transcription factor \u003cem\u003eatfA\u003c/em\u003e [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] and \u003cem\u003eatfB\u003c/em\u003e [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. These transcription factors have been shown to regulate the production of aflatoxin and its precursors in response to oxidative stress and to coordinate oxidative stress responsive genes including catalase [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. If using only the NRRL 3357 reference genome, the presence of this insertion the \u003cem\u003eatfC\u003c/em\u003e gene would not have been recognized due to reference bias.\u003c/p\u003e \u003cp\u003eTherefore, use of single reference genome, either NRRL 3357 or AF13, hinders the potential of omics studies to fully describe and investigate complex pathways such as aflatoxin or other secondary metabolite biosynthesis in \u003cem\u003eA. flavus\u003c/em\u003e. It is essential to study the comprehensive genetic diversity of \u003cem\u003eA. flavus\u003c/em\u003e using a large number of isolates from diverse geographical regions and ecosystems [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Pangenomics has emerged as an effective approach to explain the widespread diversity present in a species, and super-pangenome can be used to represent the diversity present in an entire genus [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. A pangenome represents all the genes present across a species, a pan-genes can be subdivided into \u0026ldquo;core-genome\u0026rdquo; present in all the individuals in a species and \u0026ldquo;accessory genes\u0026rdquo; not present across all the isolates which is also called as \u0026ldquo;dispensable genome\u0026rdquo; [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. For instance, three versions of \u003cem\u003eA. fumigatus\u003c/em\u003e pangenome have been developed, emphasizing identification of azole drug resistance genes in \u003cem\u003eA. fumigatus\u003c/em\u003e [\u003cspan additionalcitationids=\"CR20\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTherefore, a pangenomic approach is necessary to fully capture the genetic diversity present within the \u003cem\u003eA. flavus\u003c/em\u003e species, particularly in agricultural environments, and can be used to perform in depth studies of a plethora of secondary metabolite gene clusters and their regulation. To accomplish this, present study was designed with three main objectives: (1) collection of large number of \u003cem\u003eA. flavus\u003c/em\u003e isolates from diverse sources including maize plants and field soils; (2) whole genome sequencing of collected isolates and development of a pangenome; and (3) pangenome-wide association study (Pan-GWAS) to identify novel secondary metabolite cluster genes. This new \u003cem\u003eA. flavus\u003c/em\u003e pangenome, here named as AflaPan, can serve as a new reference for comparative genomics studies by the Aspergillus research community with greater representation of the genetic diversity of this species than any single isolate reference can provide.\u003c/p\u003e"},{"header":"Material and Methods","content":"\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAspergillus flavus\u003c/em\u003e\u003c/strong\u003e\u003cstrong\u003e isolates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of 225 isolates (98 from infected corns; 127 from soils) were newly collected and sequenced in this study. 98 isolates were associated with corn fields from infected plant parts of corn at harvest in Mississippi Delta, including 11 isolates from corn leaves, 38 from corn silks, 30 from corn tassels, and 19 from dust air-spora of corn combine harvester (Figure 1a). Among the remaining 127 isolates from different field soils with diverse cropping systems in southern Georgia in the fall harvesting season, there were 10 isolated from corn and peanut rotation fields of Tift and Turner County, 26 continuous peanut fields in Irwin and Turner County, 11 from corn and cotton rotation fields of Turner County, 17 from corn and peanut rotation fields of Irwin County, 8 from corn, cotton and peanut rotation fields in Berrien and Tift County, 16 from continuous sunflower fields of Tift County, 2 from continuous soybean fields of Tift County, and 37 continuous sorghum field of Tift County (Figure 1b).\u003c/p\u003e\n\u003cp\u003eFor \u003cem\u003eA. flavus \u003c/em\u003eisolation, protocol of [22] was followed with modification. Briefly, 10 g of soil sample was added to a bottle containing 99 mL of phosphate buffer solution. The bottles were placed on a shaker for 30 minutes, and afterwards, the mixtures were serially diluted with sterile phosphate buffer. In triplicate, 100 \u0026micro;L of each serial dilution was spread on 100 mm \u0026times; 15 mm petri dishes containing Modified Dichloran Rose Bengal (MDRB) agar [23]. The inoculated plates were incubated in the dark for 3-5 days at 37\u0026deg;C. After incubation, up to 20 \u003cem\u003eA. flavus\u003c/em\u003e colonies for each sample were randomly selected from the MDRB plates and transferred to individual 60 mm \u0026times; 15 mm Petri dishes containing potato dextrose agar amended with 0.3% \u0026beta;-cyclodextrin. The inoculated plates were incubated in the dark at 28\u0026deg;C for 5-7 days and characterized for sclerotia and aflatoxin production.\u003c/p\u003e\n\u003cp\u003eIn addition to 225 newly collected isolates, raw DNA sequence data for 121 isolates was also downloaded from previously reported genome assemblies from National Centre for Biotechnology Information Sequence Read Archive (NCBI-SRA). Finally, a total of 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates (225 new isolates and 121 from NCBI-SRA) representing 11 states of the United States were used for development of \u003cem\u003eA. flavus \u003c/em\u003epangenome (Figure 1c). A detailed list of isolates with metadata, geographical site, mating types, and sclerotia types is provided in Supplementary Table 1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDNA extraction and whole genome sequencing \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor short read sequencing for each isolate collection, a normal CTAB DNA isolation protocol was used [14]. Each isolate was cultured in yeast extract sucrose (YES) medium with 2% yeast extract, 1% sucrose for five days at 30 ̊ in the dark. Mycelial mats from each culture were collected and ground in a chilled mortar and pestle with liquid nitrogen. The ground mycelia (1-2 g) were then combined with 15 mL of CTAB extraction buffer (0.1M Tris pH 8.0, 1.4M NaCl, 20mM EDTA, 2% (w/v) CTAB, 4% (w/v) polyvinylpyrrolidone (PVP-40), and 0.5% (v/v) b-mercaptoethanol), mixed by inversion, and incubated in a water bath at 65 ̊C for 45 min with occasional inversion. The lysate was then combined with 15 mL of chloroform:isoamyl alcohol (24:1), mixed by inversion, and centrifuged at 8,000 \u0026middot; g for 15 min at 4 ̊. The upper phase was then transferred to a new 50 mL centrifuge tube. The chloroform separation was then performed a second time, and the upper phase was then combined with one volume of cold isopropanol for DNA precipitation. The DNA was then pelleted by centrifuging at 8,000 \u0026middot; g for 15 min at 4 ̊, and washed with 70% ethanol. The pellets were then dried and suspended in 100 mL TE buffer (10mM Tris pH 8.0, 1 mM EDTA pH 8.0). RNase-A was then added to a final concentration of 5 mg/mL and the samples were incubated at 37 ̊ for 1 hr. The obtained DNA was then stored at -20 ̊ until used. Isolated DNA was quantified with both a Nanodrop ND-1000 spectrophotometer (ThermoFisher, Waltham, MA, USA) and a Qubit 3.0 fluorometer (ThermoFisher) and checked using gel electrophoresis.\u003c/p\u003e\n\u003cp\u003eIsolated DNA (500 ng) from each sample was then shipped to the Novogene Corp. (Sacamento, CA) for quality checking, sequencing, and initial data filtering [24]. Libraries (350bp insert size) were generated using NEBNext\u0026reg; DNA Library Prep Kit (New England Biolabs, Ipswich, MA) following manufacturer\u0026rsquo;s instructions. In brief, initially dA-tailing followed by adapter, further the fragments (of 300-350 bp) were PCR enriched with P5 and indexed by P7 oligos. The quality of sequencing libraries was again checked on Qubit\u0026reg; 3.0 flurometer to determine the concentration of each library. Agilent\u0026reg; 2100 bioanalyzer was used to assess the insert size by using 1ng/ul DNA from sequencing libraries. Finally, the quantitative real time PCR (qRT-PCR) was performed to detect the effective concentration of each library. The libraries with appropriate insert size and with \u0026gt;2mM concentration were qualified for high throughput sequencing on Illumina HiSeq 4000 platform. Pair-end sequencing data was generated for each isolate with the read length PE150 at each end. Sequencing data for these 225 isolates of \u003cem\u003eA. flavus\u003c/em\u003e used in this study has been submitted to NCBI with bioproject ID: PRJNA915632.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGenome assembly and annotations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe raw sequencing data quality analysis and trimming of adapter sequences for each isolate was carried out using Trimmomatic v0.40 [25]. Trimmed reads were assembled using SPAdes v.3.15.4 [26], with AF13 genome assembly as a reference using mismatch and short indel correction with the Burrow-Wheeler Aligner, BWA-MEM v.0.7.12 [27]. Assemblies were improved using Pilon v.1.22 [28]. The assembly summary statistics were calculated using QUAST v5.0.2 [29]. Functional annotation and gene prediction were performed using Funannotate pipeline v.1.8.14 [30]. In brief, the \u003cem\u003eA. flavus\u003c/em\u003e assemblies were initially prepared for gene prediction and functional annotation. The small and repetitive contings (shorter than 500bp) in the assembly were removed by using the function funannotate clean which implements minimap2 with \u0026ldquo;leave one out\u0026rdquo; methodology, where unaligned shortest contigs to N50 of assembly were filtered out with parameters percent coverage overlap (--cov=95) and percent identity of overlap (--pident=95). All assemblies were soft-masked using function funannotate mask with default method tantan RepeatMasker (v.4.0.8) for repetitive elements using Dfam and RepBase repeat libraries [31]. Masked assemblies were used for gene prediction using funannotate predict with the help of Evidence Modeler algorithm. Curated protein models from closely related species \u003cem\u003eAspergillus oryzae\u003c/em\u003e were provided as evidence for gene prediction.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSpecies confirmation, identification of mating types, and types of sclerotia\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSclerotia are survival structures produced by \u003cem\u003eA. flavus\u003c/em\u003e to survive in extreme environmental conditions [32]. Wild-type strains of \u003cem\u003eA. flavus \u003c/em\u003eare known to produce sclerotia in culture after 5 days of incubation at ideal growth temperatures or a short period of cold storage. \u003cem\u003eA. flavus\u003c/em\u003e can be categorized into three groups based on the type of sclerotia production: (1) no sclerotia, (2) small type (S) sclerotia (\u0026lt; 400 \u0026micro;m size, uniform, and rare to find), and (3) large type (L) sclerotia (\u0026gt;400 \u0026micro;m. Large (L) type sclerotia are very common, not uniform, and different shapes. Notably, S- sclerotia producing \u003cem\u003eA. flavus \u003c/em\u003eisolates are observed to be almost always positive for aflatoxin production, while L- sclerotia producers were observed to have little to no effect on the presence or concentration of aflatoxins [22].\u003c/p\u003e\n\u003cp\u003eFurthermore, S-type sclerotia strains, on average, produce more aflatoxin than the L-type sclerotia strains. In this study, to determine the sclerotia size of each \u003cem\u003eA. flavus \u003c/em\u003eisolate, the samples were viewed under a microscope, and the microscope reticule was used to measure the size of the sclerotia. For sclerotia producing isolates, 100 sclerotia were measured to calculate average sclerotia size, and the average sclerotia size was used to assign the type of sclerotia production for \u003cem\u003eA. flavus\u003c/em\u003e. We also confirmed the sclerotia types in individual isolates using the \u003cem\u003ecypA \u003c/em\u003eand \u003cem\u003enorB\u003c/em\u003e sequences. \u003cem\u003eA. flavus\u003c/em\u003e and \u003cem\u003eA. parasiticus\u003c/em\u003e are closely related species, but \u003cem\u003eA. flavus\u003c/em\u003e produces only aflatoxins B1 and B2, while, \u003cem\u003eA. parasiticus\u003c/em\u003e can produce B1, B2, G1 and G2. \u003cem\u003eA. flavus\u003c/em\u003e genome is missing the portions of genes (\u003cem\u003ecypA \u003c/em\u003eand \u003cem\u003enorB\u003c/em\u003e) upstream to the polyketide synthase gene [33]. Therefore, we used \u003cem\u003eA. parasiticus\u003c/em\u003e specific \u003cem\u003ecypA-norB\u003c/em\u003e aflatoxin pathway gene cluster marker genes (NCBI accession no. AY371490) to detect the \u003cem\u003eA. parasiticus\u003c/em\u003e isolates [34]. Based on blast results, a total of 225 isolates were confirmed as \u003cem\u003eA. flavus,\u003c/em\u003e while a total of 39 isolates collected from soil were identified as \u003cem\u003eA. parasiticus\u003c/em\u003e isolates. We collected a total of 346 \u003cem\u003eA. flavus\u003c/em\u003e isolate genomes (225 from this study and 121 from NCBI-SRA) for construction of a pan-genome of \u003cem\u003eA. flavus\u003c/em\u003e. We identified the large sclerotia (L) types using the \u003cem\u003enorB-cypA\u003c/em\u003e region genomic sequence from strain NRRL3357 (NCBI accession no. AY566564), while, the short sclerotia (S) types were identified using locus of aflatoxin biosynthetic gene cluster from isolate AF70 (NCBI accession no. AY510453) [33].\u003c/p\u003e\n\u003cp\u003eIn \u003cem\u003eAspergilli\u003c/em\u003e, \u003cem\u003eA. flavus\u003c/em\u003e and \u003cem\u003eA. parasiticus\u003c/em\u003e are heterothallic species with two mating type loci \u003cem\u003eMAT1-1\u003c/em\u003e and \u003cem\u003eMAT1-2\u003c/em\u003e located on chromosome 6 associated with sexual reproduction [35]. Mating types were identified among all isolates using the reported marker sequences. Mating types \u003cem\u003eMAT1-1\u003c/em\u003e identified using complete cds sequence of mating-type protein MAT alpha 1 (NCBI accession no. EU357934), while the \u003cem\u003eMAT1-2\u003c/em\u003e types were identified using the complete cds sequence of a putative DNA lyase gene also called as mating-type HMG (High Mobility Group)-box protein \u003cem\u003eMAT1-2\u003c/em\u003e (NCBI accession no. EU357937) [36].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRead alignment and variant calling\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAF13 reference genome (v1.0, Gene Bank assembly accession number GCA_014117485.1) was used for mapping and variant calling. The genome files for AF13 reference genome were downloaded from NCBI from bio-project ID PRJNA606266 [14]. Sequence reads for 346 \u003cem\u003eA. flavus\u003c/em\u003e isolate were aligned to AF13 reference genome using Burrows-Wheeler Aligner (BWA v.0.7.17) [27]. The alignment files were converted into BAM files using Samtools v.1.10 [37]. Picard tools v.2.18.3 (http://broadinstitute.github.io/picard) used to mark the PCR duplicates using MarkDuplicate. HaplotypeCaller implemented in GATK (v.4.1.0) was used to call the SNPs using AF13 as a reference genome [38]. VariantFiltration followed by SelectVariants workflow was used to filter variants using parameters [-window-size = 10, -QualBy-Dept \u0026lt;2.0, -MapQual \u0026lt;40.0, -Qscore \u0026lt;100, -MapQualityRankSum \u0026lt;-12.5, -StrandOddsRatio \u0026gt; 3.0, -FisherStrandBias \u0026gt;60.0, -ReadPosRankSum \u0026lt;-8.0.].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWhole-genome phylogeny estimation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSNV (Single Nucleotide Variants) based phylogenetic tree was constructed by filtering the variants at minor allele frequency 0.25 and the loci with zero coverage in any isolate. Variants from each isolate were concatenated and used as input in RAxML v.8.2.12 to construct maximum likelihood (ML) tree [39]. Annotations for origin, source, toxicity, sclerotia types, mating types, genome size (Mb), and number of scaffolds for each isolate were plotted on phylogenetic tree using iTOL (Interactive Tree of Life) [40]. The final tree was visualized in iTOL by customizing the parameters. Genetic clusters among 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates identified using discriminant analysis of principle components [41]. The PCA plot was visualized using ggplot2 in R version 4.2.2 [42].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePopulation structure and diversity analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePopulation structure of 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates was studied by identifying optimum number of groups (\u003cem\u003ek\u003c/em\u003e) corresponding to the lowest Bayesian information criterion (BIC). The STRUCTURE v.2.3.4 software was used to study the population structure and admixture [43]. Diversity parameters such as nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e), fixation index (\u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e), and Tajima\u0026rsquo;s D diversity index [44] were analyzed for each genetic cluster to understand genetic diversity and distinctness of genetic clusters. VCFtools v.0.1.13 was used to calculate the genetic diversity parameters (nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e), fixation index (\u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e), and Tajima\u0026rsquo;s D) in a non-overlapping window of 10000bp [45]. Nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e) is used as measure of variability of SNPs nucleotides within individual genetic clusters. Nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e) calculated in VCFtools using parameters --site-pi --window-pi 10000. Tajima\u0026rsquo;s D diversity index allows us to see whether a population is neutrally evolving or whether it is under selection. Tajima\u0026rsquo;s D was calculated in VCFtools with parameters --TajimaD 10000. Genome-wide fixation index (\u003cem\u003eFst\u003c/em\u003e) is a measure of population differentiation. Fst ranges from 0 to 1, where is \u003cem\u003eFst\u003c/em\u003e=0 means two populations are genetically similar. While, Fst=1 means the populations are significantly genetically dissimilar to each other. Fst was calculated using VCFtools with parameters --weir-fst-pop --fst-window-size 10000. All diversity statistics was visualized using ggplot2 in R. The linkage disequilibrium decay for each population was calculated using VCFtools in 10000Kb window. The linkage disequilibrium (\u003cem\u003er\u003csup\u003e2\u003c/sup\u003e\u003c/em\u003e) was plotted against the distance between the two loci in the equilibrium (Kb).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDevelopment of a pangenome, AflaPan, and annotation of orthogroups\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe protein sequences from 346 isolates obtained from funannotate were used to identify the orthologous gene clusters using OrthoFinder [46]. Protein sequences from the AF13 reference genome were also included to identify the functions of orthologous gene clusters in AflaPan. If a gene sequence from AF13 genome is grouped in an orthologous cluster then the whole cluster is annotated as the annotation of the grouped sequence from AF13 reference genome. However, some orthologous gene clusters were not grouped with single sequence of AF13 genome. We called these orthologous gene clusters as a novel gene clusters when compared with AF13 reference genome. A customized bash script was used to extract the most extended protein sequence for every orthologous group to functionally annotate the orthogroups. Blast+ was used to find the highest similar proteins against SwissProt (https://www.expasy.org/) [47] and FunigDB (https://fungidb.org/) [48]. The sequences remain functionally annotated with annotations from AF13 reference genome and FungiDB were further annotated using InterProScan (https://github.com/ebi-pf-team/interproscan) [49]. We used the criteria E-value cutoff of 1 \u0026times; 10\u003csup\u003e5\u003c/sup\u003e, percent identity \u0026gt; 70%, minimum query coverage \u0026gt; 50% and minimum subject coverage \u0026gt; 50% [19]. The novel orthogroups were included in the pangenome only if they are present in at least 5% of the isolates in the study, to avoid false gene prediction. The presence/absence matrix of orthologous groups used to describe the framework of AflaPan genome. Further the presence/absence matrix was used for Pan-GWAS analysis, to check the utility of the AflaPan genome.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEstimation of aflatoxin content (ppb) \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUp to 20 \u003cem\u003eA. flavus\u003c/em\u003e colonies for each isolate were randomly selected from the Modified Dichloran Rose Bengal (MDRB) plates and transferred to individual 60 mm \u0026times; 15 mm petri dishes containing potato dextrose agar amended with 0.3% \u0026beta;-cyclodextrin. The inoculated plates were incubated in the dark at 28\u0026deg;C for 7 days. After incubation, evaluated for aflatoxigenicity using three cultural methods using yellow pigment test, UV fluorescence test, and ammonia vapor test [22]. The percentage of aflatoxin-producing colonies for each sample was calculated based on the qualitative analysis of the three cultural methods.\u003c/p\u003e\n\u003cp\u003eTo quantitate the aflatoxin production by \u003cem\u003eA. flavus \u003c/em\u003eisolates, aflatoxins were extracted from agar plugs of 7-day old cultures and analyzed by Enzyme Linked Immuno-Sorbent Assay (ELISA). First, agar plugs of fresh cultures were extracted at 1:5 ratio with methanol-water (70:30, \u003cem\u003ev\u003c/em\u003e/\u003cem\u003ev\u003c/em\u003e). The extracts were filtered through 0.25 \u0026micro;m pore-size nylon syringe filters. The filtrates were stored at 4\u0026deg;C until use. For ELISA, Veratox\u0026reg; for Aflatoxin kits (Neogen Inc., Lansing, MI, USA) used to determine the total aflatoxin concentration. If the results exceeded the quantitative parameters of the protocol, the filtrates were serially diluted and re-analyzed. The final aflatoxin content (ppb) was calculated using the dilution factor determined by number dilutions. A Neogen Stat-Fax 303 Plus microwell reader was used to measure the absorbance of the ELISA reaction and to calculate the final concentration of aflatoxins in solution [22].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePan-GWAS analysis to identify the aflatoxin production associated gene clusters\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe presence/absence variance matrix (PAV) for 17,855 orthologous gene clusters of AflaPan and phenotyping data generated for aflatoxin production was used for pan-genome wide association analysis. The presence/absence matrix of 1/0, where \u0026lsquo;0\u0026rsquo; represents absence of orthogroups, and \u0026lsquo;1\u0026rsquo; means presence of orthogroups in the isolates. The objective behind Pan-GWAS analysis was to identify the association of accessory genes with regulation of aflatoxin production in \u003cem\u003eA. flavus\u003c/em\u003e. Pan-GWAS analysis was perform using mixed-linear-model (MLM) implemented in TASSEL v.5.0 [50]. MLM calculates both fixed and random effects in the population. For MLM, a kinship (K) was calculated using PAV matrix and used jointly with population structure (Q). The Q+K approach helps to increase the power of genome-wide association analysis [51]. A Bonferroni threshold was calculated at \u003cem\u003eP\u003c/em\u003e-value 0.05, Bonferroni threshold=0.05/17855 (Orthogroups in AflaPan) =2.8 \u0026times; 10\u003csup\u003e-6\u003c/sup\u003e. The orthogroups showing P value less than or equal to Bonferroni threshold of 2.8 \u0026times; 10\u003csup\u003e-6\u003c/sup\u003e were called as significantly associated orthogroups [52]. GWAS results were visualized using Manhattan plots and Quintile-Quintile (QQ) plots generated in the R package \u0026lsquo;CMplot\u0026rsquo; [53].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSequencing statistics, assembly statistics, SNP statistics, and Pan-GWAS analysis are provided in the attached supplementary files. The newly developed genome assemblies for 225 isolates used in this study are available at https://zenodo.org/deposit/7615243. Raw sequencing data and metadata for each isolate is available through National Center for Biotechnology Information (NCBI) - Sequence Read Archive (SRA) with Bioproject ID PRJNA915632. Newly sequenced 225 fungal isolate cultures are available upon request by contacting the corresponding author. The 121 isolates from public data can be requested from corresponding authors of respective articles [24, 11, 54].\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eDe novo genome assemblies of \u003c/strong\u003e\u003cstrong\u003e\u003cem\u003eA. flavus\u003c/em\u003e\u003c/strong\u003e\u003cstrong\u003e isolates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of 225 isolates were newly collected and sequenced as part of this study (Figure 1a, 1b). In addition, we downloaded raw sequencing data for 121 previously sequenced isolates from NCBI sequence read archive. Among these, 7 isolates were sequenced by Fountain et al. (2020b) [24], 94 sequenced for studying pan-metabolomics [11], and 20 draft genomes from [54]. Overall, a total of 346 isolate genomes were included in this study, representing the diversity from 11 states (California, Arizona, Oklahoma, Texas, Louisiana, Mississippi, Georgia, Florida, North Carolina, Pennsylvania and Iowa) of the United States and soils of crops including peanut, sunflower, soybean, sorghum, corn, sesame and cotton. Of these 346 isolates, 177 were aflatoxigenic and 169 were non-aflatoxigenic. For sclerotia, there were 128 isolates with large sclerotia (L-strains) and 218 isolates showing small sclerotia (S-strains). Finally, for mating types, 159 showed mating type \u003cem\u003eMAT1-1\u003c/em\u003e and 187 isolates showed mating type \u003cem\u003eMAT1-2\u003c/em\u003e (\u003cstrong\u003eSupplementary Table 1\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eDe novo genome assemblies were developed for 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates using paired end Illumina sequencing reads. The mean number of contigs and scaffolds in the assemblies for soil isolates were 597 and 570, with average of 726 and 697 contigs and scaffolds, respectively. Average length of contigs and scaffolds of soil isolates was 89 kb and 98 kb, respectively. The assemblies of isolates collected from corn dust had a total of 111 and 131, contigs and scaffolds, respectively, the lowest number of the newly sequenced isolate genomes. In contrast, the assemblies of isolates collected from corn leaf tissues had the highest number 151 and 167 of contigs and scaffolds, respectively. The assembly mean contig N50 and mean scaffold N50 were 403,194 bp and 523,939 bp. Mean genome size of the assemblies of corn isolates and soil isolates was 36.9 Mb and 37.6 Mb, respectively. It is observed that soil genomes are slightly longer than corn isolates, with maximum 43.2 Mb genome size in soil isolates and maximum 38.5 Mb in corn isolates. Among the corn isolates there was no significant difference in number of contings and scaffolds collected from different corn parts such as silk, leaf, tassel and dust, and no significant difference in the genome sizes of corn isolates was observed (\u003cstrong\u003eSupplementary Table 2\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eTo perform population genomics studies, we retrieved ~1.02 million single nucleotide variants (SNVs) using AF13 as reference, approximately 28 SNVs per kilo bases (kb) among 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates. Large number of SNVs across the isolates indicated prominent genetic diversity in the species of \u003cem\u003eA. flavus\u003c/em\u003e. On an average of 1,021,529 SNVs per isolate (ranged from 795,241- 1,027,945 SNVs) were used for diversity analysis. An average of 4,871 SNVs showed heterozygous calls, and 1,016,658 SNVs were homozygous per isolate. Only homozygous SNV calls were used for population genomics studies (\u003cstrong\u003eSupplementary Figure 1;\u003c/strong\u003e\u003cstrong\u003eSupplementary Table 3\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDiversity in the population of 346 \u003c/strong\u003e\u003cstrong\u003e\u003cem\u003eA. flavus\u003c/em\u003e\u003c/strong\u003e\u003cstrong\u003e isolates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA total of four genetic clusters (\u003cem\u003eK\u003c/em\u003e=4) were identified in the collection of 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates using population structure analysis. The largest genetic cluster (P1) included 179 isolates, with majority of isolates from soil. Second largest cluster (P2) included 158 isolates grouped from soil as well as corn. Third and fourth clusters were minor with merely 5 and 4 isolates, respectively. Highly toxigenic isolates collected from the state of Louisiana were grouped in clusters 3 and 4 (Figure 1d). Four principle components identified using discriminant analysis of principle components further confirmed the presence of four genetic clusters (Figure 1e). Linkage-disequilibrium (LD) decay for each genetic cluster was calculated and plotted \u003cem\u003er\u003csup\u003e2\u003c/sup\u003e\u003c/em\u003e values against distance between two SNVs (Kb) using ggplot2. Rapid LD decay was observed for genetic cluster P1 (10 Kb), followed by genetic cluster P3 (12 Kb), cluster P4 (18 Kb) and cluster P2 (25 Kb). The population growth results in reduction of LD the largest genetic cluster (P1) showed least LD decay (Figure 1f).\u003c/p\u003e\n\u003cp\u003eGenome-wide nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e), fixation index (\u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e) and Tajima\u0026rsquo;s D diversity was analyzed individually for each cluster to understand the distinctness of genetic clusters. Genome-wide nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e) for each cluster was visualized using circos plot, showed significant number of hot spots with high nucleotide diversity in each cluster (Figure 1g). Highest average nucleotide diversity (\u003cem\u003e\u0026pi;\u003c/em\u003e=3.2 \u0026times; 10\u003csup\u003e-3\u003c/sup\u003e) was observed for cluster P1, followed by cluster P4 (\u003cem\u003e\u0026pi;\u003c/em\u003e=2.6 \u0026times; 10\u003csup\u003e-3\u003c/sup\u003e), and cluster P3 (\u003cem\u003e\u0026pi;\u003c/em\u003e=2.0 \u0026times; 10\u003csup\u003e-3\u003c/sup\u003e). The lowest nucleotide diversity was observed in genetic cluster P2 (\u003cem\u003e\u0026pi;\u003c/em\u003e=8.4 \u0026times; 10\u003csup\u003e-4\u003c/sup\u003e) (Figure 1h). Further, the chromosome wise nucleotide diversity across 346 isolates was represented using boxplots. Chromosome 8 showed highest nucleotide diversity (mean \u003cem\u003e\u0026pi;\u003c/em\u003e=5.0 \u0026times; 10\u003csup\u003e-3\u003c/sup\u003e, with range 1.9 \u0026times; 10\u003csup\u003e-2 \u003c/sup\u003e\u0026ndash; 3.8 \u0026times; 10\u003csup\u003e-7\u003c/sup\u003e), followed by chromosomes 5, 6 and 7 (mean \u003cem\u003e\u0026pi;\u003c/em\u003e=3.7 \u0026times; 10\u003csup\u003e-3\u003c/sup\u003e, with range 1.85 \u0026times; 10\u003csup\u003e-2 \u003c/sup\u003e\u0026ndash; 5.5 \u0026times; 10\u003csup\u003e-7\u003c/sup\u003e), chromosomes 3 and 4 (mean \u003cem\u003e\u0026pi;\u003c/em\u003e=3.5 \u0026times; 10\u003csup\u003e-3\u003c/sup\u003e, with range 1.7 \u0026times; 10\u003csup\u003e-2 \u003c/sup\u003e\u0026ndash; 5.4 \u0026times; 10\u003csup\u003e-7\u003c/sup\u003e), and chromosomes 1 and 2 (mean \u003cem\u003e\u0026pi;\u003c/em\u003e=3.2\u0026times; 10\u003csup\u003e-3\u003c/sup\u003e, with range 1.9 \u0026times; 10\u003csup\u003e-2 \u003c/sup\u003e\u0026ndash; 2.5 \u0026times; 10\u003csup\u003e-5\u003c/sup\u003e) (\u003cstrong\u003eSupplementary figure 2\u003c/strong\u003e). The population differentiation was investigated by calculating fixation index (\u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e) between the pairs of genetic clusters. Fixation index (\u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e) value can range from 0 to 1. Where, \u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e =0 indicates the two clusters are completely sharing each other, whereas \u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e =1 means there is no sharing or complete differentiation between two clusters. In this study, higher \u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e values (\u0026gt;0.6) of the cluster P2 with clusters P1, P3 and P4 indicated that P2 cluster is most differentiated cluster than other three clusters. For instance, fixation index (\u003cem\u003eF\u003csub\u003eST\u003c/sub\u003e\u003c/em\u003e) of 0.822 was recorded between clusters P2 and P4, followed by 0.814 between P2 and P3, and 0.619 between P2 and P1. Lowest fixation index (\u003cem\u003eF\u003csub\u003eST \u003c/sub\u003e\u003c/em\u003e=0.346) was recorded between the clusters P3 and P4, followed by P1 and P3 (\u003cem\u003eF\u003csub\u003eST \u003c/sub\u003e\u003c/em\u003e=0.346), P1 and P4 (\u003cem\u003eF\u003csub\u003eST \u003c/sub\u003e\u003c/em\u003e=0.460) (Figure 1h).\u003c/p\u003e\n\u003cp\u003eGenome-wide Tajima\u0026rsquo;s D diversity index was calculated for each cluster to study the abundance of rare alleles in each cluster. Tajima\u0026rsquo;s D diversity helps to determine whether the cluster is evolving neutrally, or under selection pressure. Tajima\u0026rsquo;s D=0, means the cluster is evolving neutrally without any evidence of selection, Tajima\u0026rsquo;s D\u0026lt;0, means excess of rare alleles in the clusters or population expansion, Tajima\u0026rsquo;s D\u0026gt;0, means scarcity of rare alleles in cluster or population contraction. Tajima\u0026rsquo;s D for cluster P1 was -0.13, indicating that cluster P1 is evolving neutrally with partial selection sweep. Interestingly, Tajima\u0026rsquo;s D for cluster P2 was -1.77, indicated abundance of rare alleles in P2. The Tajima\u0026rsquo;s D for cluster P3 (0.81) and P4 (0.81) was greater than zero indicating that these two populations lack rare alleles and therefore the clusters are resulted in the population contraction (Figure 1i). Chromosome wise Tajima\u0026rsquo;s D index was higher for chromosome 6 and 8 (Supplementary Figure 3).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWhole-genome phylogeny and characteristics of genetic clusters\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWhole genome phylogeny was constructed and confirmed the presence of 4 clusters within population of 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates. Morphological characteristics of each cluster were investigated for their origin (geographical site), source (soil or corn), mating types, sclerotia types, toxigenicity, genome size (Mb), and number scaffolds in each assembly (Figure 2). In the major cluster P1, the majority of isolates were from soil (136) and 43 were from corn. In cluster P1, a total of 126 isolates had small sclerotia (S-types), and 53 had large sclerotia types. In case of mating types, cluster P1 included 94 isolates with mating type \u003cem\u003eMAT1-1\u003c/em\u003e, whereas 85 were \u003cem\u003eMAT1-2\u003c/em\u003e. Interestingly, in cluster P1, a total of 111 isolates were aflatoxigenic and 68 were non-aflatoxigenic. Second major cluster P2 included a total of 103 isolates from soil and 55 from corn. In this cluster 58 isolates were aflatoxigenic and 100 isolates were non-aflatoxigenic. There were 74 isolates with large sclerotia (L) types and 84 with small sclerotia (S) types. In case of mating types, 61 isolates showed (\u003cem\u003eMAT1-1\u003c/em\u003e) types and 97 showed (\u003cem\u003eMAT1-2\u003c/em\u003e) types. In cluster P1, majority of isolates were aflatoxigenic, whereas in P2 majority of isolates were non-aflatoxigenic. Cluster P3 and P4 were minor genetic clusters with 5 and 4 isolates, respectively. In genetic cluster P3 all isolates were from collected from soil with aflatoxigenic, small sclerotia types, and four with \u003cem\u003eMAT1-1\u003c/em\u003e and one \u003cem\u003eMAT1-2\u003c/em\u003e mating types. In cluster P4, three isolates were aflatoxigenic and a one isolate was non-aflatoxigenic from the state of Iowa. All aflatoxigenic isolates in cluster P3 and P4 were from the states of Louisiana and Oklahoma (Figure 2).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e310 kb insertion is rare and identified only in soil samples\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRecently a highly toxigenic \u003cem\u003eA. flavus\u003c/em\u003e strain AF13 was sequenced, and a 310 kb insertion was identified in AF13 genome, but was absent in isolate NRRL3357. This 310 kb insertion includes a bZIP transcription factor named \u003cem\u003eatfC\u003c/em\u003e [14].\u003c/p\u003e\n\u003cp\u003eThe presence of 310 kb region was analyzed across 346 isolates, there were only 5 isolates with this 310 kb insertion. Interestingly, all these 5 isolates are from collected from soil samples in the state of Louisiana, and all are toxigenic isolates grouped in cluster P3 and P4 (Figure 2). However, there were about 57% shared identity with this 310 kb insertion among the other 341 isolates. Therefore, 310 kb insertion is rare in \u003cem\u003eA. flavus\u003c/em\u003e species and the isolates with this insertion are highly toxigenic [14] (\u003cstrong\u003eSupplementary Table 1\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAflaPan contains 7,315 core and 10,540 accessory genes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGenome assemblies 346 isolates were used to develop pan-genome framework for \u003cem\u003eA. flavus\u003c/em\u003e, called AflaPan genome\u003cem\u003e. \u003c/em\u003eA pan-genome represents total genes in the species, including core genes shared by all the individuals of the species and accessory genes that are not present in all the individuals of the species. In this study, we collected 346 genomes of \u003cem\u003eA. flavus\u003c/em\u003e and identified 17,855 unique orthologous gene clusters in this pangenome, called AflaPan genome. In AflaPan, of the 17,855, there were 7,315 orthologous gene clusters as core genome (41.1% of the pangenome) present in all 346 genomes, 3,887 as softcore orthogroups in \u0026gt;95% genomes (21.7% of the pangenome), 3,132 as shell genes in 5-95% genomes (17.5% of the pangenome), and 3,521 as cloud genes present in less than 5% of the genomes (19.7% of the pangenome) (Figure 3a, 3b). We also report here 5,994 novel gene clusters that were not annotated in the AF13 reference genome [14]. Of the new gene clusters, 451 were identified in soft-core (12% of total soft-core genes), 2,069 (66% of total shell genes) in shell and 3,474 (99% of total cloud genes) in cloud (Figure 3c).\u003c/p\u003e\n\u003cp\u003eThe gene clusters in this AflaPan genome were compared with the gene models in current reference genomes AF13 and NRRL3357 of \u003cem\u003eA. flavus\u003c/em\u003e. Reference genomes AF13 and NRRL3357 contains of 13,188 and 13,487 annotated gene models, respectively. However, in AflaPan, a total of 17,855 unique orthologous gene clusters were annotated. The AflaPan genome includes more diverse genome segments that were missing in current single reference genome assemblies (Figure 3d).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSignificant differences between the lengths of protein sequences were observed in core and accessory genomes. The average protein length in core genome was 579 amino acids (AAs), with a range of 68 to 6885 AAs. While the average protein lengths in soft-core genome was 463 AAs, with range of 60 to 7781 AAs. Average protein length in shell genome was 361 AAs, with the range of 50 to 3917 AAs. In clouds and singletons, the average protein length was 228 AAs and 233 AAs, respectively. Therefore, the conservative proteins were longer in length. Overall, the average protein length was decreased from core-softcore-shell-clouds and singletons (\u003cstrong\u003eFigure 3e; Supplementary Table 4; Supplementary Table 5\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eFurthermore, this Aflapan was closed and captured comprehensive genetic diversity of the intraspecies of \u003cem\u003eA. flavus\u003c/em\u003e. Any addition of \u003cem\u003eA. flavus\u003c/em\u003e genomes will not substantially increase the number of pan-genes or core genes of \u003cem\u003eA. flavus\u003c/em\u003e pan-genome (Figure 3f). The number of pan-genes and core genes will be not substantially increased beyond this large collection of 346 \u003cem\u003eA. flavus\u003c/em\u003e genome in this AflaPan (Figure 3g).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eToxin producing gene clusters in AflaPan \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe orthologous gene clusters in AflaPan were annotated using blasts in InterProScan, SWissProt, and FungiDB databases. A total of 17,038 orthogroups were annotated using FungiDB database. There were 12,691 orthologous gene clusters annotated using SwissProt and 15,740 annotated using InterProScan. Therefore, there were a total of 12,500 orthogroups annotated using all three databases and 399 orthologous gene clusters were not annotated using either of the three databases (\u003cstrong\u003eSupplementary Table 5\u003c/strong\u003e). Based on the annotations a total of 89 toxins producing orthologous gene clusters were identified in core genome. Toxins producing gene clusters (TPCs) in core genome include 3-hydroxyacyl-CoA dehydrogenase-like protein, 6-hydroxy-D-nicotine oxidase, acyl-CoA dehydrogenase, satratoxin biosynthesis SC3 cluster protein, astacin-like metalloprotease toxin, aflatoxin B1 (AFB1) aldehyde reductase, viriditoxin biosynthesis cluster protein D (\u003cem\u003eVdtD\u003c/em\u003e), latrotoxin producing gene cluster (\u003cem\u003eLTX\u003c/em\u003e), AM-toxin biosynthesis protein or cytochrome P450 monooxygenase AMT3, AAL-toxin biosynthesis cluster protein, cercosporin toxin biosynthesis cluster, viridicatumtoxin synthesis proteins (\u003cem\u003evrtL\u003c/em\u003e and \u003cem\u003evrtT\u003c/em\u003e), aflatoxin biosynthesis protein O, and transcription activator AMTR1 (AM toxin). Cytochrome P450 monooxygenase is potent mycotoxins producing gene clusters, and in AflaPan, we identified \u003cem\u003eALT8\u003c/em\u003e, \u003cem\u003eaflU\u003c/em\u003e, \u003cem\u003eaflV\u003c/em\u003e, and \u003cem\u003eAKT7\u003c/em\u003e toxin producing cytochrome P450 monooxygenases. TPCs in core genome indicated that these clusters are widely present in the species \u003cem\u003eA. flavus\u003c/em\u003e and annotated in reference genomes AF13 and NRRL3357.\u003c/p\u003e\n\u003cp\u003eA total of 43 TPCs identified in soft-core genome including 5 new gene clusters integrated from soil isolates of Louisiana and Oklahoma. New TPCs in soft-core genome included toxin subunit YenA2, taurine hydroxylase-like protein SAT17, and cytochrome P450 monooxygenase AKT7. Similarly, in shell genome, a total of 45 TPCs were annotated. Interestingly, a total of 23 new gene clusters were identified in shell genome linked with toxin production. New TPCs in shell genome include MFS gliotoxin efflux transporter (\u003cem\u003egliA\u003c/em\u003e), called gliotoxin biosynthesis protein A, alpha-latrocrustotoxin-Lt1a, killer toxin subunits alpha/beta, double-stranded DNA deaminase toxin A, and acyl-CoA dehydrogenase AFT10-1. Particularly, in shell genome there were 33 new TPCs integrated from soil isolates and 3 were from corn samples. In cloud genome, a total of 36 TPCs were identified, and all were new or non-AF13 TPCs. Among singletons we identified 8 new TPCs and all were originated from soil samples of Georgia and Mississippi. The TPCs in singletons included C3 and PZP-like alpha-2-macroglobulin domain-containing protein, ADP-ribosylating toxin CARDS (Community-Acquired Respiratory Distress Syndrome toxin), norsolorinic acid reductase (\u003cem\u003eaflE\u003c/em\u003e), FAD-dependent monooxygenase, aflatoxin cluster transcriptional coactivator (\u003cem\u003eaflS\u003c/em\u003e), cytochrome P450 monooxygenase (\u003cem\u003ealfG\u003c/em\u003e), and abhydrolase domain-containing protein AFT2. Here we observed that the extended genome portions of soil included new toxin producing genes (\u003cstrong\u003eSupplementary Table 5\u003c/strong\u003e). In order to check the association of new orthologous gene clusters with toxin production and secondary metabolite production, AflaPan-genome wide association analysis (AflaPan-GWAS) was conducted using phenotypes of total aflatoxins produced.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSecondary metabolite producing gene clusters\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eA. flavus\u003c/em\u003e also produces a variety of secondary metabolites which may be useful for pharmaceutical industries. In current AF13 and NRRL3357 reference genomes of \u003cem\u003eA. flavus\u003c/em\u003e [14] there were 80 and 78 secondary metabolite producing gene clusters. In this AflaPan genome, there were a total of 160 secondary metabolite producing gene clusters annotated for 10 important secondary metabolites, including aflatrem (14), aflavarin (3), aflatoxin (63), asparasone (6), cyclopiazonic acid (25), ditryptophenaline (3), kojic acid (3), leporins (4), piperazines (19), and ustiloxin (20).\u003c/p\u003e\n\u003cp\u003eAflatrem is tremorgenic mycotoxin, the only class of mycotoxins impacting central nervous system [55]. In AflaPan, a total of 14 orthogroups were annotated as aflatrem biosynthetic gene clusters. Among them, 8 orthogroups identified in core, and 5 in soft-core and 1 in shell genome. It indicated that aflatrem producing gene clusters are highly conserved in \u003cem\u003eA. flavus\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003eAflavarin is also an important sclerotial metabolite exhibiting insecticidal activity [55]. In AflaPan, a total of 3 aflavarin producing orthologous gene clusters were annotated, and there were two in core genome and one in cloud genome.\u003c/p\u003e\n\u003cp\u003eAflatoxin is a potent mycotoxin produced by \u003cem\u003eA. flavus\u003c/em\u003e classified as polyketide synthase (PKS) backbone genes. Based on annotation, there were a total of 63 aflatoxin producing secondary metabolite gene clusters identified in this AflaPan genome. Among them, 20 were identified in core, 10 in soft-core, 21 in shell genome, 9 in cloud, and 3 as singletons. Aflatoxin producing gene clusters are uniformly distributed in core and accessory genomes of AflaPan. Moreover, Pan-GWAS analysis was conducted using phenotyping data of aflatoxin produced by 225 isolates, which uncovered 391 orthogroups associated with toxin production, including these 63 aflatoxin producing secondary metabolite gene clusters.\u003c/p\u003e\n\u003cp\u003eAsparasone is another sclerotial metabolite of structural class PKS [56]. In AflaPan, a total of 6 asparasone producing orthologous gene clusters were identified. Among them, there were one in each core and shell, 2 in soft core, and 2 in cloud genome of AflaPan. Cyclopiazonic acid (CPA) is an indole-tetramic acid mycotoxin, which belongs to structural group dimethyl tryptophan synthases (DMATs) or hybrid PKS (polyketide synthases)-NRPSs (non-ribosomal peptide synthetases) produced by a number of species of \u003cem\u003eAspergillus\u003c/em\u003e [56]. A total of 25 CPA producing orthogroups were identified in AflaPan. Among them, 7 identified in core, 4 in soft core, 6 in shell, and 8 in cloud genomes of AflaPan. Ditriptophenaline (DTP) produces diketopiperazine alkaloid in \u003cem\u003eA. flavus\u003c/em\u003e. There were three DTP producing orthogroups identified in cloud genome of AflaPan.\u003c/p\u003e\n\u003cp\u003eKojic acid (5-hydroxy-2-hydroxymethyl-1, 4-pyrone) belongs to structural group oxydoreductases, a skin lightning agent widely used in cosmetics industries. Moreover, it also has antibacterial activity and therefore used as pharmaceuticals, pesticides and insecticides [57]. In AflaPan, three orthologous gene clusters were identified associated with Kojic acid biosynthesis. Two of them were in core and one was identified in cloud genome of AflaPan. Leporin A is an antiinsectan N-methoxy-2-pyridone sclerotial metabolite isolated from the sclerotia of \u003cem\u003eA. leporis\u003c/em\u003e. Leporin belongs to structural group hybrid PKS-NRPS. In AflaPan, four orthogroups were identified linked with leporin biosynthesis. Among them, three were in soft-core genome and one was identified as a singleton in soil isolate GA6_18.\u003c/p\u003e\n\u003cp\u003ePiperazines are class of saturated N-heterocycles, acting as central nervous system stimulants widely used in medicinal chemistry [58]. In AflaPan, a total of 19 orthogroups identified as piperazines biosynthesis gene clusters. Among them, there were 5 identified in core, 2 in soft-core, 2 in shell, 9 in clouds, and one as singleton identified in soil isolate GA6_18. Interestingly, in AflaPan, there were 10 new orthogroups identified associated with Piperazines biosynthesis. Ustiloxin B is a mycotoxin with cyclic peptide belonging to structural group ribosomally synthesized and post-translationally modified peptides (RiPPs). In AflaPan, there were 20 orthogroups annotated as Ustiloxin B biosynthetic clusters. Among them, there were 9 identified in core genome, 5 in soft-core genome, 3 in shell genome, and 3 in cloud genome of AflaPan (\u003cstrong\u003eSupplementary Table 6\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePan-GWAS uncovered gene clusters associated with aflatoxin biosynthesis pathways \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePan-GWAS analysis using PAV matrix of 17,855 orthogroups and the newly collected 225 isolates of \u003cem\u003eA. flavus\u003c/em\u003e uncovered new and existing genes associated with aflatoxin production. Pan-GWAS analysis identified a total of 391 orthogroups significantly associated with aflatoxin production (\u003cstrong\u003eFigure 4\u003c/strong\u003e). Of these 391, a total of 91 orthogroups were not annotated either with NCBI blast or InterProScan. There were 369 orthogroups (94.4%) from shell genome. Important orthologous gene clusters with significant association to aflatoxin production were listed in Table 2. Surprisingly, of the 369 shell genes, there were 256 (69.4%) orthogroups which were not annotated and may be absent in the AF13 reference assembly. The new orthogroups significantly associated with aflatoxin production were potential intermediates in various aflatoxin biosynthetic pathways in \u003cem\u003eA. flavus\u003c/em\u003e. There were 15 orthogroups from soft-core and seven from cloud genome significantly associated with aflatoxin production, including aspergilol biosynthesis cluster protein, aspyridones biosynthesis protein E (\u003cem\u003eapdE\u003c/em\u003e), and imizoquin biosynthesis cluster protein D (\u003cem\u003eimqD\u003c/em\u003e) (\u003cstrong\u003eSupplementary Table 7\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003eAmong highly associated aflatoxin producing gene clusters there were cytochrome P450 monooxygenase (\u003cem\u003eapdE\u003c/em\u003e,\u003cem\u003e acrF\u003c/em\u003e, \u003cem\u003eaflV\u003c/em\u003e, \u003cem\u003eAKT7\u003c/em\u003e, \u003cem\u003eaneD\u003c/em\u003e), ABC multidrug transporters (\u003cem\u003eatrB\u003c/em\u003e, \u003cem\u003eatrD\u003c/em\u003e, \u003cem\u003eMDR2\u003c/em\u003e), aflatoxin biosynthesis regulatory protein R (\u003cem\u003eaflR\u003c/em\u003e), aflatoxin biosynthesis protein S (\u003cem\u003eaflS\u003c/em\u003e), demethylsterigmatocystin 6-O-methyltransferase (\u003cem\u003eaflO\u003c/em\u003e), aflatoxin biosynthesis protein T (\u003cem\u003eaflT\u003c/em\u003e), fatty acid synthase alpha subunit (\u003cem\u003eaflA\u003c/em\u003e), beta-cyclopiazonate dehydrogenase, and fumagillin biosynthesis polyketide synthase. aflaoxin producing orthologous gene clusters homologous to aflatoxin producing genes in other fungi were also showed strong association with aflatoxin production in this AflaPan, for instance, terrein biosynthesis cluster protein (\u003cem\u003eterG\u003c/em\u003e) from \u003cem\u003eA. terreus\u003c/em\u003e, ABC multidrug transporter (\u003cem\u003eatrB\u003c/em\u003e) from \u003cem\u003eA. nidulans\u003c/em\u003e, Azaphilone biosynthesis cluster protein (\u003cem\u003eazaJ\u003c/em\u003e) from \u003cem\u003eA. niger\u003c/em\u003e, double-stranded DNA deaminase aflatoxin A from \u003cem\u003eBurkholderia cenocepacia\u003c/em\u003e, glutathione S-transferase omega from \u003cem\u003eSaccharomyces cerevisiae\u003c/em\u003e, and AM-toxin biosynthesis regulator 1 (\u003cem\u003eAMTR-1\u003c/em\u003e) from \u003cem\u003eAlternaria alternata\u003c/em\u003e (\u003cstrong\u003eSupplementary Table 7\u003c/strong\u003e).\u003c/p\u003e\n"},{"header":"Discussion","content":"\n\u003cp\u003eAflatoxin is detrimental to human health specially affecting liver and gall bladder in addition to problems of growth stunting and impaired development. This aflatoxin producing pathogen, \u003cem\u003eAspergillus flavus, \u003c/em\u003ehas acquired large geographical diversity and broad host range over the time and sustained under varied climatic conditions. The aflatoxin contamination has been global food safety concern for humans and animals, many a times even leading to death. Isolated efforts have been made though for identifying the genomic regions and candidate genes controlling different aflatoxin resistance mechanisms [4], however, less efforts are being made to in depth understanding adoption of isolates specially under changing climate conditions. Nevertheless, the recently sequenced two reference-level genomes for two \u003cem\u003eA. flavus\u003c/em\u003e isolates (isolate AF13 of mating type\u003cem\u003e MAT1-2\u003c/em\u003e with high producer of aflatoxin and isolate NRRL3357 of mating type \u003cem\u003eMAT1-1\u003c/em\u003e with a moderate producer) facilitated identification of 310 kb region being the major aflatoxin producing gene cluster [14]. As of now not much understanding has been developed on the extent of diversity among pathogens from different geographies and crops. Therefore, this comprehensive research with large number of sequenced genomes has been conducted to solve the puzzle of evolution of isolates, intensity of infection to plants and varied level of aflatoxin production under different climatic conditions. We collected and sequenced new set of 225 \u003cem\u003eA. flavus\u003c/em\u003e isolates and also included 121 publicly available sequences for genome analysis. This is the maiden effort to develop pangenome of \u003cem\u003eA. flavus\u003c/em\u003e, called AflaPan, and perform other genomic analysis using 346 genome sequences. The AflaPan had 17,855 unique orthologous gene clusters including 7,315 core genes (41%) and 10,540 accessory genes indicating significant genome diversity between \u003cem\u003eA. flavus\u003c/em\u003e isolates. It is interesting to note that there were 5,994 orthologous gene clusters that had been not annotated in \u003cem\u003eA. flavus\u003c/em\u003e AF13 and NRRL3357 reference genomes [14]. The pan-GWAS study using AflaPan identified 391 significant associated orthologous gene clusters for aflatoxin production, indicating the usefulness of AflaPan. It\u0026rsquo;s worthwhile noting that AflaPan represents a large diversity of 346 \u003cem\u003eA. flavus\u003c/em\u003e genomes from 11 states of United States with aflatoxin contamination issues in various crops (Fig. 1). Although no such study has been reported for \u003cem\u003eA. flavus\u003c/em\u003e, however, recently two pangenomes have been reported for related species \u003cem\u003eA. fumigatus\u003c/em\u003e with 300 [19] and 260 [21] isolates. It is important to note that these two pangenomes are closed in nature similar to AflaPan indicating sufficient sample size for capturing complete species genome diversity.\u003c/p\u003e\n\u003cp\u003eThe accessory genome is a variable portion of pangenome and higher accessory gene content indicates higher diversity among the individuals. AflaPan had large proportion of accessory genome (58.9%; 10,540 genes) (\u003cstrong\u003eFigure 3a\u003c/strong\u003e) as compared to \u003cem\u003eA. fumigatus\u003c/em\u003e (41% accessory genome; [19] and \u003cem\u003eCandida albicans\u003c/em\u003e (mere 9% accessory genome; [18]. The high accessory genome content in \u003cem\u003eA. flavus\u003c/em\u003e indicates the faster pace of genome evolution which might be its strength for its successful adaptation to different climatic conditions, geographies and host plants. This finding also signals warning to research community for much faster devising and implementing solutions to tackle the aflatoxin contamination problem.\u003c/p\u003e\n\u003cp\u003ePangenomes can be classified as open pangenomes or closed pangenomes [59]. If the number of accessory and core genes of a pangenome are stable and not increasing after adding an optimum number of genomes, then it is referred to as a closed pangenome. In contrast, if the number of accessory and core genes of a pangenome are substantially increasing upon addition of individual genomes then it is referred as an open pangenome. An open pangenome can be substantially improved by sequencing more individuals resulting in more accessory and core genes being included in the existing pangenome [60]. In AflaPan, the number of accessory and core genes were not substantially increased after adding 300 \u003cem\u003eA. flavus\u003c/em\u003e genomes. Therefore, AflaPan is a closed pangenome, and effectively captured the overall genomic diversity of the \u003cem\u003eA. flavus\u003c/em\u003e species. In contrast, the pangenome of the wheat fungal pathogen \u003cem\u003eZymoseptoria tritici\u003c/em\u003e was developed with only 19 isolates which can be improved further by including more isolates [61]. Similarly, recently another open pangenome of an onion bacterial pathogen \u003cem\u003ePantoea ananatis\u003c/em\u003e was developed using 81 \u003cem\u003eP. ananatis\u003c/em\u003e genomes which showed an increase of 50 new genes with addition of each isolate [62]. While exploring possibilities for genome expansion or horizontal gene transfer [63] in \u003cem\u003eA. flavus\u003c/em\u003e genomes, the isolates from soil samples had larger genomes in comparison with isolates collected from infected corn tissues. The expanded soil isolates\u0026rsquo; genomes included novel aflatoxin producing gene clusters identified in the Pan-GWAS analysis using AflaPan.\u003c/p\u003e\n\u003cp\u003eOne important objective behind developing AflaPan was to conduct Pan-GWAS study and to identify unreported aflatoxin producing gene clusters. In this study, the newly sequenced isolates of 225 were used for phenotyping aflatoxin production and Pan-GWAS analysis to identify the gene clusters associated with aflatoxin production. Pan-GWAS analysis for aflatoxin production, using a presence/absence matrix based on 17,855 pan-genes across 225 isolates, identified 391 pan-genes significantly associated with aflatoxin production. Among them, 369 pan-genes (94.4%) were located in the shell genome of AflaPan. Our findings suggest that the shell genome of AflaPan is enriched with aflatoxin producing gene clusters. Shell genome includes pan-genes which are present in 95% of the total individuals (\u003cstrong\u003eSupplementary Table 7\u003c/strong\u003e). Interestingly, a total of 256 new pan-genes which are not annotated in present reference genomes NRRL3357 and AF13 were significantly associated with aflatoxin production. These 256 pan-genes could be important targets for developing aflatoxin mitigation strategies using gene silencing or genome editing approaches. Earlier, Pan-GWAS has been conducted in several fungi as well as bacterial pathogens to identify causal genes. For instance, Pan-GWAS in \u003cem\u003ePantoea ananatis\u003c/em\u003e pangenome identified 14 strongly associated genes in HiVir/PASVIL clusters responsible for pathogenicity in onion [61]. Recently, human pathogen \u003cem\u003eA. fumigatus\u003c/em\u003e has developed resistance against fungicide triazoles. Pan-genome wide association analysis using \u003cem\u003eA. fumigatus\u003c/em\u003e pangenome of 300 isolates identified 12 genes associated with triazole resistance [19]. Another, Pan-GWAS using \u003cem\u003eA. fumigatus\u003c/em\u003e pangenome of 218 isolates identified azole drug resistant genes in ergosterol biosynthetic pathway [20].\u003c/p\u003e\n\u003cp\u003eTwo genes in \u003cem\u003eA. flavus\u003c/em\u003e, \u003cem\u003eaflS\u003c/em\u003e and \u003cem\u003eaflR\u003c/em\u003e, mainly regulate the aflatoxin biosynthesis pathway [64], with \u003cem\u003eaflR \u003c/em\u003eencoding a DNA-binding binuclear zinc cluster (Zn(II)\u003csub\u003e2\u003c/sub\u003eCys\u003csub\u003e6\u003c/sub\u003e) transcription factor required for expression of a majority of structural genes in aflatoxin biosynthesis pathway [65]. Here, Pan-GWAS analysis using AflaPan identified significant association with \u003cem\u003eaflR\u003c/em\u003e with -log10(\u003cem\u003ep\u003c/em\u003e) = 7.6. \u003cem\u003eaflS\u003c/em\u003e has also been shown to be essential for aflatoxin biosynthesis and is required for activation of \u003cem\u003eaflR\u003c/em\u003e [66]. Disruption of \u003cem\u003eaflS\u003c/em\u003e in \u003cem\u003eA. parasiticus\u003c/em\u003e resulted in lower expression of some AFB1 producing genes, such as \u003cem\u003eaflC\u003c/em\u003e, \u003cem\u003eaflD\u003c/em\u003e, \u003cem\u003eaflM\u003c/em\u003e and \u003cem\u003eaflP\u003c/em\u003e and reduced aflatoxin levels [67]. It is also observed that \u003cem\u003eaflS\u003c/em\u003e is strongly associated with aflatoxin production with -log10(\u003cem\u003ep\u003c/em\u003e) =7.64. In addition, there are several aflatoxin biosynthesis genes including \u003cem\u003eaflV\u003c/em\u003e, \u003cem\u003eaflT\u003c/em\u003e, \u003cem\u003eaflA\u003c/em\u003e and \u003cem\u003eaflO\u003c/em\u003e which showed strong associations with aflatoxin production according to the Pan-GWAS analysis (\u003cstrong\u003eTable 2\u003c/strong\u003e). \u003cem\u003eaflV\u003c/em\u003e showed to be associated with aflatoxin production with -log10(\u003cem\u003ep\u003c/em\u003e) =6.3. \u003cem\u003eaflV\u003c/em\u003e encodes cytochrome P450 monoxygenase (\u003cem\u003ecypX\u003c/em\u003e) has been shown to catalyze the reaction from averufin (AVR) to hydroxyversicolorone (HVN) and HVN further involves in aflatoxin biosynthesis [68]. It is reported that \u003cem\u003eaflT\u003c/em\u003e is not directly linked with aflatoxin biosynthesis, though it is present in aflatoxin producing gene cluster [69]. However, Pan-GWAS analysis indicated that \u003cem\u003eaflT\u003c/em\u003e is significantly associated with aflatoxin production with -log10(\u003cem\u003ep\u003c/em\u003e) =7.5. The \u003cem\u003efas \u003c/em\u003e(fatty acid synthases) genes, namely \u003cem\u003eaflA\u003c/em\u003e also referred as \u0026lsquo;\u003cem\u003efas-2\u003c/em\u003e\u0026rsquo; and \u003cem\u003eaflB\u003c/em\u003e referred as \u0026lsquo;\u003cem\u003efas-1\u003c/em\u003e\u0026rsquo; produces \u0026alpha; and \u0026beta; sub-units, respectively, both subunits transform the hexanoate units into a polyketide structure in aflatoxin biosynthesis pathway [70]. In this study, \u003cem\u003eaflA\u003c/em\u003e (\u003cem\u003efas-2\u003c/em\u003e) showed strong association with aflatoxins with -log10(\u003cem\u003ep\u003c/em\u003e) =7.6. The \u003cem\u003eaflO\u003c/em\u003e (\u003cem\u003eomtB\u003c/em\u003e) homolog in \u003cem\u003eA. nidulans\u003c/em\u003e required for conversion of demethylsterigmatocystin (DMST) to sterigmatocystin (ST) in \u003cem\u003eA. nidulans\u003c/em\u003e [71].\u003c/p\u003e\n\u003cp\u003ePan-GWAS analysis showed significant association with previously reported aflatoxin producing gene clusters [72] such as O-methyltransferase A (\u003cem\u003eOmtA/AflP\u003c/em\u003e) and versicolorin dehydrogenase /ketoreductase (\u003cem\u003eVer1/AflM\u003c/em\u003e) with \u0026ndash;log10(\u003cem\u003ep\u003c/em\u003e) value of 7.4 and 6.8, respectively. \u003cem\u003eOmtA/AflP\u003c/em\u003e and \u003cem\u003eVer1/AflM\u003c/em\u003e had been targeted for host-induced gene silencing (HIGS) in peanuts resulting in reduced aflatoxin contamination [73]. Moreover, \u003cem\u003eVer1/AflM\u003c/em\u003e was also targeted for host induced gene silencing in maize [74].\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eA. flavus\u003c/em\u003e is also a reservoir for important secondary metabolites [75]. In this study, we investigated 10 secondary metabolites and their associated genes, investigating the importance of AflaPan genome. Interestingly, in AflaPan, there were 165 pan-genes annotated to associate with production of 10 important secondary metabolites (\u003cstrong\u003eSupplementary Table 6\u003c/strong\u003e). Recently, pan-secondary metabolome based on 94 \u003cem\u003eA. flavus\u003c/em\u003e isolates identified 7,821 biosynthetic gene clusters (BGCs) responsible for production of variety of metabolites [11]. The secondary metabolite producing gene clusters for these 10 metabolites were identified in core as well as accessory genomes of this AflaPan. Kojic acid (KA) has commercial use in industries including cosmetics, especially in skin care product to prevent the exposure to UV radiations [56]. In AflaPan, three pan-genes were identified as Kojic acid biosynthesis genes. In addition, the pan-genes also associated with the production of aflavarin, aflatoxin, asparasone, cyclopiazonic acid, ditryptophenaline, leporins, piperazines, and ustiloxin.\u003c/p\u003e\n\u003cp\u003eIn summary, this study provides complete pangenome framework for the species of \u003cem\u003eAspergillus flavus \u003c/em\u003ealong with associated genes for pathogen survival and aflatoxin production. This pangenome can be used for performing further studies on target traits as well as developing required diagnostic genotyping panel for isolate identification. Furthermore, the information generated through this study can also be used for identifying individuals from different diversity clusters for developing individual genome assemblies leading development of pangenome reference. Most importantly, the newly identified aflatoxin producing gene clusters will be a new source for seeking aflatoxin mitigation strategies and needs new attention in research.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCOMPETING INTEREST\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eETHICS APPROVAL AND CONSENT TO PARTICIPATE\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe plant samples and soil samples were collected from the fields with permission from field owners.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCONSENT FOR PUBLICATION\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors approved the manuscript for publication\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and material\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAnalyzed data on sequencing statistics, assembly statistics, SNP statistics and Pan-GWAS analysis are provided in the attached supplementary files. The newly developed genome assemblies for 225 isolates used in this study are available at https://zenodo.org/deposit/7615243. Associated raw sequencing data for each isolate are available through National Center for Biotechnology Information (NCBI) - Sequence Read Archive (SRA) with Bioproject ID PRJNA915632. Fungal isolates are available upon request by contacting corresponding author. The 121 isolates from public data can be requested from corresponding authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work is partially supported by the U.S. Department of Agriculture Agricultural Research Service (USDA-ARS). Georgia Agricultural Commodity Commission for Corn, National Corn Growers Association, Aflatoxin Mitigation Center of Excellence (AMCOE), Georgia Peanut Commission, National Peanut Board, The Peanut Research Foundation and MARS-Wrigley.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eFountain JC, Scully BT, Ni X, Kemerait RC, Lee RD, Chen ZY, Guo B. Environmental influences on maize-\u003cem\u003eAspergillus flavus\u003c/em\u003e interactions and aflatoxin production. Front Micro. 2014;5:40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Y, Wu F. Global burden of aflatoxin-induced hepatocellular carcinoma: a risk assessment. Environ Health Perspect. 2019;118:818\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBurdock GA, Flamm WG. Safety assessment of the mycotoxin cyclopiazonic acid. Int J Toxicol. 2000;19:195\u0026ndash;218.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePandey MK, Kumar R, Pandey AK, Soni P, Gangurde SS, Sudini HK, et al. Mitigating aflatoxin contamination in groundnut through a combination of genetic resistance and post-harvest management practices. Toxins. 2019;11:315.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo BZ, Yu J, Holbrook CC, Cleveland TE, Nierman WC, Scully BT. Strategies in prevention of preharvest aflatoxin contamination in peanuts: Aflatoxin biosynthesis, genetics and genomics. Peanut Sci. 2009;36:11\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrisvad JC, Hubka V, Ezekiel CN, Hong SB, Nov\u0026aacute;kov\u0026aacute; A, Chen AJ, et al. Taxonomy of \u003cem\u003eAspergillus\u003c/em\u003e section \u003cem\u003eFlavi\u003c/em\u003e and their production of aflatoxins, ochratoxins and other mycotoxins. Stud Mycol. 2019;93:1\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePayne GA, Nierman WC, Wortman JR, Pritchard BL, Brown D, Dean RA, et al. Whole genome comparison of \u003cem\u003eAspergillus flavus\u003c/em\u003e and \u003cem\u003eA. oryzae\u003c/em\u003e. Med Mycol. 2006;44:9\u0026ndash;S11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang PK, Ehrlich KC. What does genetic diversity of \u003cem\u003eAspergillus flavus\u003c/em\u003e tell us about \u003cem\u003eAspergillus oryzae\u003c/em\u003e? Int J Food Microbiol. 2010;138:189\u0026ndash;99.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGibbons JG, Salichos L, Slot JC, Rinker DC, McGary KL, King JG, et al. The evolutionary imprint of domestication on genome variation and function of the filamentous fungus \u003cem\u003eAspergillus oryzae\u003c/em\u003e. Curr Biol. 2012;22:1403\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan AW, Garg V, Roorkiwal M, Golicz AA, Edwards D, Varshney RK. Super-Pangenome by integrating the wild side of a species for accelerated crop improvement. Trends Plant Sci. 2020;25:148\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDrott MT, Rush TA, Satterlee TR, Giannone RJ, Abraham PE, Greco C, et al. Microevolution in the pansecondary metabolome of \u003cem\u003eAspergillus flavus\u003c/em\u003e and its potential macroevolutionary implications for filamentous fungi. Proc Natl Acad Sci USA. 2021;118:e2021683118.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNierman WC, Yu J, Fedorova-Abrams ND, Losada L, Cleveland TE. Genome sequence of \u003cem\u003eAspergillus flavus\u003c/em\u003e NRRL 3357, a strain that causes aflatoxin contamination of food and feed. Genome Announc. 2015;3:e00168\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSkerker JM, Pianalto KM, Mondo SJ, Yang K, Arkin AP, Keller NP, et al. Chromosome assembled and annotated genome sequence of \u003cem\u003eAspergillus flavus\u003c/em\u003e NRRL 3357. G3. Genes Genomes Genet. 2021;11:213.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFountain JC, Clevenger JP, Nadon B, Youngblood RC, Korani W, Chang PK, et al. Two new \u003cem\u003eAspergillus flavus\u003c/em\u003e reference genomes reveal a large insertion potentially contributing to isolate stress tolerance and aflatoxin production. Genes Genomes Genet. 2020a;G3:10: 3515\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBal\u0026aacute;zs A, P\u0026oacute;csi I, Hamari Z, Leiter \u0026Eacute;, Emri T, Miskei M, et al. AtfA bZIP-type transcription factor regulates oxidative and osmotic stress responses in \u003cem\u003eAspergillus nidulans\u003c/em\u003e. Mol Genet Genomics. 2010;283:289\u0026ndash;303.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWee JS, Hong Y, Roze L, Day D, Chanda A, et al. The fungal bZIP transcription factor \u003cem\u003eAtfB\u003c/em\u003e controls virulence-associated processes in \u003cem\u003eAspergillus parasiticus\u003c/em\u003e. Toxins. 2017;9:287.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFountain JC, Bajaj P, Nayak SN, Yang L, Pandey MK, et al. Responses of \u003cem\u003eAspergillus flavus\u003c/em\u003e to oxidative stress are related to fungal development regulator, antioxidant enzyme, and secondary metabolite biosynthetic gene expression. Front Microbiol. 2016;7:2048.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcCarthy C, Fitzpatrick DA. Pan-genome analyses of model fungal species. Microb Genom. 2019;5:e000243.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarber AE, Sae-Ong T, Kang K, Seelbinder B, Li J, Walther G, et al. \u003cem\u003eAspergillus fumigatus\u003c/em\u003e pan-genome analysis identifies genetic variants associated with human infection. Nat Microbiol. 2021;6:1526\u0026ndash;36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRhodes J, Abdolrasouli A, Dunne K, Sewell TR, Zhang Y, Ballard E, et al. Population genomics confirms acquisition of drug-resistant Aspergillus fumigatus infection by humans from the environment. Nat Microbiol. 2022;7:663\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLofgren LA, Ross BS, Cramer RA, Stajich JE. The pan-genome of \u003cem\u003eAspergillus fumigatus\u003c/em\u003e provides a high-resolution view of its population structure revealing high levels of lineage-specific diversity driven by recombination. PLoS Biol. 2022;20:e3001890.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbbas HK, Weaver MA, Zablotowiez RM, Horn BW, Shier WT. Relationship between aflatoxin production and sclerotia formation among isolates of \u003cem\u003eAspergillus\u003c/em\u003e section \u003cem\u003eFlavi\u003c/em\u003e from the Mississippi Delta. Eur J Plant Path. 2005;112:283\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHorn BW, Dorner JW. Soil populations of \u003cem\u003eAspergillus\u003c/em\u003e species from section \u003cem\u003eFlavi\u003c/em\u003e along a transect through peanut-growing regions of the United States. Mycologia. 1998;90:767\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFountain JC, Clevenger JP, Nadon B, Wang H, Abbas HK, et al. Draft genome sequences of one \u003cem\u003eAspergillus parasiticus\u003c/em\u003e isolate and nine \u003cem\u003eAspergillus flavus\u003c/em\u003e isolates with varying stress tolerance and aflatoxin production. Microbiol Resour Announc. 2020b;9:e00478\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBolger AM, Lohse M, Usadel B, Trimmomatic. A flexible trimmer for Illumina Sequence Data. Bioinformatics. 2014;30:2114\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, et al. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 2012;19:455\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H. 2013. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv preprint arXiv:1303.3997.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWalker BJ, Abeel T, Shea T, Priest M, Abouelliel A, et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS ONE. 2014;9:e112963.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMikheenko A, Prjibelski A, Saveliev V, Antipov D, Gurevich A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics. 2018;34:142\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalmer JM, Stajich J. Funannotate v1. 8.1: Eukaryotic genome annotation. Zenodo 2020; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo\u003c/span\u003e\u003cspan address=\"10.5281/zenodo\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, 4054262.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBao W, Kojima KK, Kohany O. Repbase Update, a database of repetitive elements in eukaryotic genomes. Mob Dna. 2015;6:1\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWicklow DT. Sporogenic germination of sclerotia of \u003cem\u003eAspergillus flavus\u003c/em\u003e and \u003cem\u003eAspergillus parasiticus\u003c/em\u003e. Trans Br Mycol Soc. 1984;82:621\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEhrlich KC, Chang PK, Yu J, Cotty PJ. Aflatoxin biosynthesis cluster gene \u003cem\u003ecypA\u003c/em\u003e is required for G- aflatoxin formation. Appl Environ Microbiol. 2004;70:6518\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang PK, Cary JW, Bhatnagar D, Cleveland TE, Bennett JW, et al. Cloning of the \u003cem\u003eAspergillus parasiticus apa-2\u003c/em\u003e gene associated with the regulation of aflatoxin biosynthesis. Appl Environ Microbiol. 1993;59:3273\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoore GG, Elliott JL, Singh R, Horn BW, Dorner JW, Stone EA, et al. Sexuality generates diversity in the aflatoxin gene cluster: evidence on a global scale. PLoS Pathog. 2013;9:e1003574.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamirez-Prado JH, Moore GG, Horn BW, Carbone I. Characterization and population analysis of the mating-type genes in \u003cem\u003eAspergillus flavus\u003c/em\u003e and \u003cem\u003eAspergillus parasiticus\u003c/em\u003e. Fungal Genet Biol. 2008;45:1292\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. Genome Project Data Processing Subgroup, the sequence alignment/map format and SAMtools. Bioinformatics. 2009;25:2078\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, et al. From FastQ data to high-confidence variant calls: the genome analysis toolkit best practices pipeline. Curr Protoc Bioinf. 2013;43:11\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStamatakis A, RAxML-VI-HPC. Maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22:2688\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLetunic I, Bork P. Interactive Tree Of Life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 2021;49:293\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJombart T, Devillard S, Balloux F. Discriminant analysis of principal components: a new method for the analysis of genetically structured populations. BMC Genet. 2010;11:94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWickham H. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York. 2016; ISBN 978-3-319-24277-4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFalush D, Stephens M, Pritchard JK. Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies. Genetics. 2003;164:1567\u0026ndash;87.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTajima F. Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics. 1989;123:585\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDanecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA. The variant call format and VCFtools. Bioinformatics. 2011;27:2156\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEmms DM, Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019;20:238.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGasteiger E, Gattiker A, Hoogland C, Ivanyi I, Appel RD, Bairoch A, ExPASy. The proteomics server for in-depth protein knowledge and analysis. Nucleic Acids Res. 2003;31:3784\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBasenko EY, Pulman JA, Shanmugasundram A, Harb OS, Crouch K, Starns D. FungiDB: an integrated bioinformatic resource for fungi and oomycetes. J Fungi. 2018;20:39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJones P, Binns D, Chang HY, Fraser M, Li W, McAnulla C, McWilliam H, et al. InterProScan 5: genome-scale protein function classification. Bioinformatics. 2014;30:1236\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBradbury PJ, Zhang Z, Kroon DE, Casstevens TM, Ramdoss Y, Buckler ES. TASSEL: software for association mapping of complex traits in diverse samples. Bioinformatics. 2007;23:2633\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHenderson CR. Best linear unbiased estimation and prediction under a selection model. Biometrics. 1975;31:423\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGangurde SS, Wang H, Yaduru S, Pandey MK, Fountain JC, Chu Y, et al. Nested-association mapping (NAM)‐based genetic dissection uncovers candidate genes for seed and pod weights in peanut (\u003cem\u003eArachis hypogaea\u003c/em\u003e). Plant Biotechnol J. 2020;18:1457\u0026ndash;71.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYin L. R package CMplot v 3.1.3 2016; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/YinLiLin/CMplot\u003c/span\u003e\u003cspan address=\"https://github.com/YinLiLin/CMplot\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGebru ST, Mammel MK, Gangiredla J, Tartera C, Cary JW, Moore GG, et al. Draft genome sequences of 20 \u003cem\u003eAspergillus flavus\u003c/em\u003e isolates from corn kernels and cornfield soils in Louisiana. Microbiol Resour Announc. 2020;9:e00826\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTePaske MR, Gloer JB, Wicklow DT, Dowd PF. Aflavarin and β-Aflatrem: new anti-insectan metabolites from the sclerotia of \u003cem\u003eAspergillus flavus\u003c/em\u003e. J Nat Prod. 1992;55:1080\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVinokurova NG, Ivanushkina NE, Khmel\u0026rsquo;Nitskaya II, Arinbasarov MU. Synthesis of α-cyclopiazonic acid by fungi of the genus \u003cem\u003eAspergillus\u003c/em\u003e. Appl Biochem Microbiol. 2007;43:435\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosfarizan M, Mohd SM, Nurashikin S, Madihah MS, Arbakariya BA. Kojic acid: Applications and development of fermentation process for production. Biotechnol Mol Biol Rev. 2010;5:24\u0026ndash;37.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChandrika TN, Shrestha SK, Ngo HX, Tsodikov OV, Howard KC, Garneau-Tsodikova S. Alkylated piperazines and piperazine-azole hybrids as antifungal agents. J Med Chem. 2018;61:158\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRichard GF, Cham. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-3-030-38281-0_12\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-38281-0_12\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBosi E, Fani R, Fondi M. Defining orthologs and pangenome size metrics. Bacterial Pangenomics: Methods Mol Biol. 2015;1231:191\u0026ndash;202.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBadet T, Oggenfuss U, Abraham L, McDonald BA, Croll D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen \u003cem\u003eZymoseptoria tritici\u003c/em\u003e. BMC Biol. 2020;18:1\u0026ndash;18.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgarwal G, Gitaitis RD, Dutta B. Pan-genome of novel \u003cem\u003ePantoea stewartii\u003c/em\u003e subsp. indologenes reveals genes involved in onion pathogenicity and evidence of lateral gene transfer. Microorganisms. 2021;9:1761.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKelkar YD, Ochman H. Causes and consequences of genome expansion in fungi. Genome Biol Evol. 2012;4:13\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeorgianna DR, Payne GA. Genetic regulation of aflatoxin biosynthesis: from gene to genome. Fungal Genet Biol. 2009;46:113\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrice MS, Yu J, Nierman WC, Kim HS, Pritchard B, Jacobus CA, et al. The aflatoxin pathway regulator AflR induces gene transcription inside and outside of the aflatoxin biosynthetic cluster. FEMS Microbiol Lett. 2006;255:275\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang PK. Lack of interaction between AFLR and AFLJ contributes to nonaflatoxigenicity of \u003cem\u003eAspergillus sojae\u003c/em\u003e. J Biotechnol. 2004a;107:245\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang PK. The \u003cem\u003eAspergillus parasiticus\u003c/em\u003e protein AFLJ interacts with the aflatoxin pathway-specific regulator AFLR. Mol Genet Genomics. 2003;268:711\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWen Y, Hatabayashi H, Arai H, Kitamoto HK, Yabe K. Function of the \u003cem\u003ecypX\u003c/em\u003e and \u003cem\u003emoxY\u003c/em\u003e Genes in Aflatoxin Biosynthesis in \u003cem\u003eAspergillus parasiticus\u003c/em\u003e. Appl Environ Microbiol. 2005;71:3192\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang PK, Yu J, Yu JH. \u003cem\u003eaflT\u003c/em\u003e, a MFS transporter-encoding gene located in the aflatoxin gene cluster, does not have a significant role in aflatoxin secretion. Fungal Genet Biol 2004b; 41:911\u0026ndash;920.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoze LV, Hong SY, Linz JE. Aflatoxin Biosynthesis: Current Frontiers. Annu Rev Food Sci Technol. 2013;4:293\u0026ndash;311.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKelkar HS, Keller NP, Adams TH. \u003cem\u003eAspergillus nidulans\u003c/em\u003e stcP encodes an O-methyltransferase that is required for sterigmatocystin biosynthesis. Appl Environ Microbiol. 1996;62:4296\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu J. Current understanding on aflatoxin biosynthesis and future perspective in reducing aflatoxin contamination. Toxins. 2012;4:1024\u0026ndash;57.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSharma KK, Pothana A, Prasad K, Shah D, Kaur J, Bhatnagar D, et al. Peanuts that keep aflatoxin at bay: a threshold that matters. Plant Biotechnol J. 2018;16:1024\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRaruang Y, Omolehin O, Hu D, Wei Q, Han ZQ, Rajasekaran K. Host induced gene silencing targeting \u003cem\u003eAspergillus flavus aflM\u003c/em\u003e reduced aflatoxin contamination in transgenic maize under field conditions. Front Microbiol. 2020;11:754.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCary JW, Gilbert MK, Lebar MD, Majumdar R, Calvo AM. \u003cem\u003eAspergillus flavus\u003c/em\u003e secondary metabolites: more than just aflatoxins. Food Saf. 2018;6:7\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003cdiv class=\"colspec\" align=\"char\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSummary of newly sequenced genome assemblies from corn and soil samples used for development of an \u003cspan class=\"Italic\"\u003eAspergillus flavus\u003c/span\u003e pangenome (AflaPan).\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth rowspan=\"2\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFeature\u003c/div\u003e\n \u003c/th\u003e\n \u003cth colspan=\"3\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSoil isolates\u003c/div\u003e\n \u003c/th\u003e\n \u003cth colspan=\"3\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCorn isolates\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMinimum\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMaximum\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAverage\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMinimum\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMaximum\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAverage\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eScaffolds\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e149\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7,949\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e696.8\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e111.0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1230\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e269.7\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eContigs\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e170\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e8,104\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e726.3\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e146.0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,436.0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e295.7\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eGenome size (Mb)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e36.6\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e43.2\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e37.6\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e36.5\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e38.5\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e37.0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAverage scaffold size (Kb)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e5.4\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e248.6\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e98.0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e27.8\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e274.1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e161.5\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAverage contig size (Kb)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e5.3\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e217.9\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e87.8\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e26.8\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e253.1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e144.4\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eContig N50 (Kb)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e11.1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,439.6\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e307.4\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e346.3\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,054.5\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e603.7\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eScaffold N50 (Kb)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e11.1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2,316.2\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e406.2\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e378.6\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,367.5\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e775.6\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003cdiv class=\"colspec\" align=\"char\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003ctable id=\"Tab2\" style=\"width: 1003px;\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ePan-GWAS identified orthogroups associated with aflatoxin (ppb) production in \u003cspan class=\"Italic\"\u003eAspergillus flavus\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr style=\"height: 15px;\"\u003e\n \u003cth style=\"height: 15px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOrthogroup \u003csup\u003ea\u003c/sup\u003e\u003c/div\u003e\n \u003c/th\u003e\n \u003cth style=\"height: 15px; width: 99.7662px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e-log10(p) \u003csup\u003eb\u003c/sup\u003e\u003c/div\u003e\n \u003c/th\u003e\n \u003cth style=\"height: 15px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eGenome \u003csup\u003ec\u003c/sup\u003e\u003c/div\u003e\n \u003c/th\u003e\n \u003cth style=\"height: 15px; width: 76px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eLength \u003csup\u003ed\u003c/sup\u003e\u003c/div\u003e\n \u003c/th\u003e\n \u003cth style=\"height: 15px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAnnotation \u003csup\u003ee\u003c/sup\u003e\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012151\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e444\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAflatoxin biosynthesis protein R (\u003cspan class=\"Italic\"\u003eaflR\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012150\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e438\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAflatoxin biosynthesis protein S (\u003cspan class=\"Italic\"\u003eaflS\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0011833\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.06\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e458\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFusaristatin A biosynthesis cluster protein\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0000052\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9.97\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSoft-core\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e764\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAspyridones biosynthesis protein F\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012617\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.33\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e493\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eNeosartiricin biosynthesis protein R\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012284\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.54\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e317\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCyclochlorotine biosynthesis protein O\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012070\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.16\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e534\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCytochrome P450 monooxygenase; Acurin A biosynthesis cluster protein (\u003cspan class=\"Italic\"\u003eacrF\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012018\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.27\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e508\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCytochrome P450 monooxygenase aflV; Aflatoxin biosynthesis protein (\u003cspan class=\"Italic\"\u003eaflV\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012566\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.12\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e470\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCytochrome P450 monooxygenase; AK-toxin biosynthesis protein 7 (\u003cspan class=\"Italic\"\u003eAKT7\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0000168\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.55\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e521\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCytochrome P450 monooxygenase aneD; Aculenes biosynthesis cluster protein D (\u003cspan class=\"Italic\"\u003eaneD\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0010996\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e16.78\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSoft-core\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e215\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCytochrome P450 monooxygenase apdE; Aspyridones biosynthesis protein E (\u003cspan class=\"Italic\"\u003eapdE\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0013427\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e22.40\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e341\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDehydrogenase azaJ; Azaphilone biosynthesis cluster protein azaJ (\u003cspan class=\"Italic\"\u003eazaJ\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0011609\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.22\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e345\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDehydrogenase RED2; T-toxin biosynthesis protein (RED2)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012391\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.34\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e116\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDemethylsterigmatocystin 6-O-methyltransferase; Aflatoxin biosynthesis protein O (\u003cspan class=\"Italic\"\u003eaflO\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012144\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.47\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e548\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEfflux pump aflT; Aflatoxin biosynthesis protein T (\u003cspan class=\"Italic\"\u003eaflT\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0013752\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.10\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e261\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEfflux pump atB; Terreic acid biosynthesis cluster protein B (\u003cspan class=\"Italic\"\u003eatB\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012168\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.89\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e541\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEfflux pump mlcE; Compactin biosynthesis protein E (\u003cspan class=\"Italic\"\u003emlcE\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0011998\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.55\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e651\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEfflux pump radE; Radicicol biosynthesis cluster protein (\u003cspan class=\"Italic\"\u003eradE\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012554\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.92\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,102\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEfflux pump terG; Terrein biosynthesis cluster protein (\u003cspan class=\"Italic\"\u003eterG\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012658\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.81\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e508\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEfflux pump terG; Terrein biosynthesis cluster protein (\u003cspan class=\"Italic\"\u003eterG\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13.588px;\"\u003e\n \u003ctd style=\"height: 13.588px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012155\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13.588px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9.12\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13.588px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13.588px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e208\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13.588px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEnterobactin synthase component B; Enterobactin biosynthesis bifunctional protein (\u003cspan class=\"Italic\"\u003eEntB\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012119\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,737\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFatty acid synthase alpha subunit aflA; Aflatoxin biosynthesis protein A (\u003cspan class=\"Italic\"\u003eaflA\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012283\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.46\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e636\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFumagillin dodecapentaenoate synthase; Fumagillin biosynthesis polyketide synthase\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0014250\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e10.44\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e220\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMethyltransferase pytC; Pyranterreones biosynthesis cluster protein C (\u003cspan class=\"Italic\"\u003epytC\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012525\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.62\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e330\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMFS gliotoxin efflux transporter gliA; Gliotoxin biosynthesis protein A (\u003cspan class=\"Italic\"\u003egliA\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012623\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.44\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e497\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMFS transporter cpaT; Cyclopiazonic acid biosynthesis cluster protein T (\u003cspan class=\"Italic\"\u003ecpaT\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0013082\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.78\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e317\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eNotoamide biosynthesis activator; Notoamide biosynthesis cluster protein L (\u003cspan class=\"Italic\"\u003eNotL\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0011985\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.13\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e643\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDehydrogenase patE; Patulin biosynthesis cluster protein E (\u003cspan class=\"Italic\"\u003epatE\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0000472\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.52\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSoft-core\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e720\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ePeptide transporter imqD; Imizoquin biosynthesis cluster protein D (\u003cspan class=\"Italic\"\u003eimqD\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012742\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.29\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e707\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ePrenyltransferase phnF; Phenalenone biosynthesis cluster protein F (\u003cspan class=\"Italic\"\u003ephnF\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012565\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e7.81\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e251\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eProbable tetra hydroxynaphthalene reductase; Conidial pigment biosynthesis cluster\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012428\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.26\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e286\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAspercryptin biosynthesis cluster protein L (\u003cspan class=\"Italic\"\u003eatnL\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012853\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e17.19\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e726\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCyclochlorotine biosynthesis protein T (\u003cspan class=\"Italic\"\u003ecctT\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012547\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.95\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1,060\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMelleolides biosynthesis cluster protein\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012333\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.85\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e599\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFusicoccin A biosynthetic gene clusters protein\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 13px;\"\u003e\n \u003ctd style=\"height: 13px; width: 119px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOG0012630\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 99.7662px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.81\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 88.2338px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eShell Novel\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 76px;\" align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e451\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd style=\"height: 13px; width: 589px;\" align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSatratoxin biosynthesis SC3 cluster protein 17 (\u003cspan class=\"Italic\"\u003eSAT17\u003c/span\u003e)\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ctfoot\u003e\n \u003ctr style=\"height: 15px;\"\u003e\n \u003ctd style=\"height: 15px; width: 972px;\" colspan=\"5\"\u003e\u003csup\u003ea\u003c/sup\u003e Orthogroups of AflaPan significantly associated with toxin production\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 15px;\"\u003e\n \u003ctd style=\"height: 15px; width: 972px;\" colspan=\"5\"\u003e\u003csup\u003eb\u003c/sup\u003e -log10 transformation of \u003cspan class=\"Italic\"\u003ep\u003c/span\u003e values identified in Pan-GWAS analysis\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 28px;\"\u003e\n \u003ctd style=\"height: 28px; width: 972px;\" colspan=\"5\"\u003e\u003csup\u003ec\u003c/sup\u003e Genome (Core, soft-core, shell or cloud) of the orthogroups. Where, the new genes are denoted by suffix novel. For instance, shell novel means, the new orthogroups in the shell genome. Whereas, \u0026lsquo;Shell\u0026rsquo; means the orthogroups already existing in previous reference genomes.\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 15px;\"\u003e\n \u003ctd style=\"height: 15px; width: 972px;\" colspan=\"5\"\u003e\u003csup\u003ed\u003c/sup\u003e Protein sequence length of orthogroup\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr style=\"height: 15px;\"\u003e\n \u003ctd style=\"height: 15px; width: 972px;\" colspan=\"5\"\u003e\u003csup\u003ee\u003c/sup\u003e Annotation of orthogroups\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tfoot\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-plant-biology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pbio","sideBox":"Learn more about [BMC Plant Biology](http://bmcplantbiol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/pbio/default.aspx","title":"BMC Plant Biology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Aspergillus flavus, aflatoxin, core genome, accessory genome, genomic diversity, pan-GWAS","lastPublishedDoi":"10.21203/rs.3.rs-3958535/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3958535/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003e \u003cem\u003eAspergillus flavus\u003c/em\u003e is an important agricultural and food safety threat due to its production of carcinogenic aflatoxins. It has high level of genetic diversity that is adapted to various environments. Recently, we reported two reference genomes of \u003cem\u003eA. flavus\u003c/em\u003e isolates, AF13 (\u003cem\u003eMAT1-2\u003c/em\u003e and highly aflatoxigenic isolate) and NRRL3357 (\u003cem\u003eMAT1-1\u003c/em\u003e and moderate aflatoxin producer). Where, an insertion of 310 kb in AF13 included an aflatoxin producing gene bZIP transcription factor, named \u003cem\u003eatfC\u003c/em\u003e. Observations of significant genomic variants between these isolates of contrasting phenotypes prompted an investigation into variation among other agricultural isolates of \u003cem\u003eA. flavus\u003c/em\u003e with the goal of discovering novel genes potentially associated with aflatoxin production regulation. Present study was designed with three main objectives: (1) collection of large number of \u003cem\u003eA. flavus\u003c/em\u003e isolates from diverse sources including maize plants and field soils; (2) whole genome sequencing of collected isolates and development of a pangenome; and (3) pangenome-wide association study (Pan-GWAS) to identify novel secondary metabolite cluster genes.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003ePangenome analysis of 346 \u003cem\u003eA. flavus\u003c/em\u003e isolates identified a total of 17,855 unique orthologous gene clusters, with mere 41% (7,315) core genes and 59% (10,540) accessory genes indicating accumulation of high genomic diversity during domestication. 5,994 orthologous gene clusters in accessory genome not annotated in either the \u003cem\u003eA. flavus\u003c/em\u003e AF13 or NRRL3357 reference genomes. Pan-genome wide association analysis of the genomic variations identified 391 significant associated pan-genes associated with aflatoxin production. Interestingly, most of the significantly associated pan-genes (94%; 369 associations) belonged to accessory genome indicating that genome expansion has resulted in the incorporation of new genes associated with aflatoxin and other secondary metabolites.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eIn summary, this study provides complete pangenome framework for the species of \u003cem\u003eAspergillus flavus\u003c/em\u003e along with associated genes for pathogen survival and aflatoxin production. The large accessory genome indicated large genome diversity in the species \u003cem\u003eA. flavus\u003c/em\u003e, however AflaPan is a closed pangenome represents optimum diversity of species \u003cem\u003eA. flavus\u003c/em\u003e. Most importantly, the newly identified aflatoxin producing gene clusters will be a new source for seeking aflatoxin mitigation strategies and needs new attention in research.\u003c/p\u003e","manuscriptTitle":"Aspergillus flavus pangenome (AflaPan) uncovers novel aflatoxin and secondary metabolite associated gene clusters","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-22 08:55:49","doi":"10.21203/rs.3.rs-3958535/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-03-12T15:58:10+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-03-12T09:08:56+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-02-26T05:53:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"f60ad9fb-1c20-405f-8701-eb5f6de05170","date":"2024-02-23T15:17:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"f9f0c548-e5a5-49ce-a3ab-5ef1aa5a1155","date":"2024-02-23T09:21:51+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-02-23T04:07:03+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-02-20T22:14:53+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-02-20T22:14:16+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Plant Biology","date":"2024-02-15T11:42:39+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-plant-biology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pbio","sideBox":"Learn more about [BMC Plant Biology](http://bmcplantbiol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/pbio/default.aspx","title":"BMC Plant Biology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"253e7420-01a9-468d-a542-4239234e1ca7","owner":[],"postedDate":"February 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-05-02T00:46:08+00:00","versionOfRecord":{"articleIdentity":"rs-3958535","link":"https://doi.org/10.1186/s12870-024-04950-8","journal":{"identity":"bmc-plant-biology","isVorOnly":false,"title":"BMC Plant Biology"},"publishedOn":"2024-05-01 00:46:08","publishedOnDateReadable":"May 1st, 2024"},"versionCreatedAt":"2024-02-22 08:55:49","video":"","vorDoi":"10.1186/s12870-024-04950-8","vorDoiUrl":"https://doi.org/10.1186/s12870-024-04950-8","workflowStages":[]},"version":"v1","identity":"rs-3958535","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3958535","identity":"rs-3958535","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.