Assessment of plasmids for relating the 2020 Salmonella enterica serovar Newport onion outbreak to farms implicated by the outbreak investigation

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-14

This study found that shared plasmids between clinical Salmonella Newport isolates and farm isolates provide evidence connecting the 2020 outbreak strain to implicated farms.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

This preprint investigated the 2020 Salmonella enterica serovar Newport red onion outbreak, using whole-genome sequencing from clinical isolates and environmental Salmonella isolates collected near two implicated California farms (Holtville and Bakersfield). Because standard SNP phylogenetics had found no close genomic relatives among farm-region isolates, the authors tested an alternative plasmid-driven hypothesis, quantifying plasmid content and comparing putative plasmids between clinical and farm isolates, including pangenome-based estimates of accessory genes. They found the clinical isolates formed a highly related clade with a large conserved core genome and many accessory genes (with at least 64% on plasmids), and observed high similarity between multiple plasmids from farm isolates and clinical isolates, with phylogenetic signals suggesting recent common ancestry and possible transfer via intermediary species. A major caveat was that plasmid “promiscuity” limited conclusions about exact geography, source, and time since transfer, and the authors emphasized needs for more detailed metadata and more extensive environmental sampling. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background: The Salmonella enterica serovar Newport red onion outbreak of 2020 was the largest foodborne outbreak of Salmonella in over a decade. The epidemiological investigation suggested two farms as the likely source of contamination. However, single nucleotide polymorphism (SNP) analysis of the whole genome sequencing data did not find any Salmonella isolates from the farm regions that were closely related to the clinical isolates—preventing the use of phylogenetics in source identification. Here, we explored an alternative method for analyzing the whole genome sequencing data driven by the hypothesis that if the outbreak strain had come from the farm regions, then the clinical isolates would disproportionately contain plasmids found in isolates from the farm regions due to recent horizontal transfer. Results SNP analysis confirmed that the clinical isolates formed a highly related clade with evidence for ancestry in California going back a decade. The clinical isolates not only had a large and highly conserved core genome (4,399 genes), but also 2,577 sparsely distributed accessory genes—at least 64% of which were carried on plasmids. Amongst the clinical isolates and Salmonella isolates from the farm regions were 2,187 and 503 putative plasmids, respectively. High similarity was observed between 17 plasmids from 8 farm isolates and 14 plasmids from 13 clinical isolates. Phylogenetic analysis suggested the highly similar plasmids shared a recent common ancestor and might have been transferred via intermediary species, but the seeming promiscuity of the plasmids prevented any conclusions about geographic location, isolation source, and time since transfer. Our sampling analysis suggested that observing a similar number and combination of highly similar plasmids in random samples of environmental Salmonella enterica within NCBI Pathogen Detection database was unlikely, supporting a connection between the outbreak strain and the farms implicated by the epidemiological investigation. Conclusion Horizontally transferred plasmids provided evidence for a connection between clinical isolates and the farms implicated as the source of the outbreak. Our case study suggests that such analyses might add a new dimension to source tracking investigations, but highlights the need for detailed and accurate metadata, more extensive environmental sampling, and a better understanding of plasmid molecular evolution.
Full text 197,022 characters · extracted from preprint-html · click to expand
Assessment of plasmids for relating the 2020 Salmonella enterica serovar Newport onion outbreak to farms implicated by the outbreak investigation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Assessment of plasmids for relating the 2020 Salmonella enterica serovar Newport onion outbreak to farms implicated by the outbreak investigation Seth Commichaux, Hugh Rand, Kiran Javkar, Erin K. Molloy, James B. Pettengill, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2166997/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 04 Apr, 2023 Read the published version in BMC Genomics → Version 1 posted 9 You are reading this latest preprint version Abstract Background The Salmonella enterica serovar Newport red onion outbreak of 2020 was the largest foodborne outbreak of Salmonella in over a decade. The epidemiological investigation suggested two farms as the likely source of contamination. However, single nucleotide polymorphism (SNP) analysis of the whole genome sequencing data did not find any Salmonella isolates from the farm regions that were closely related to the clinical isolates—preventing the use of phylogenetics in source identification. Here, we explored an alternative method for analyzing the whole genome sequencing data driven by the hypothesis that if the outbreak strain had come from the farm regions, then the clinical isolates would disproportionately contain plasmids found in isolates from the farm regions due to recent horizontal transfer. Results SNP analysis confirmed that the clinical isolates formed a highly related clade with evidence for ancestry in California going back a decade. The clinical isolates not only had a large and highly conserved core genome (4,399 genes), but also 2,577 sparsely distributed accessory genes—at least 64% of which were carried on plasmids. Amongst the clinical isolates and Salmonella isolates from the farm regions were 2,187 and 503 putative plasmids, respectively. High similarity was observed between 17 plasmids from 8 farm isolates and 14 plasmids from 13 clinical isolates. Phylogenetic analysis suggested the highly similar plasmids shared a recent common ancestor and might have been transferred via intermediary species, but the seeming promiscuity of the plasmids prevented any conclusions about geographic location, isolation source, and time since transfer. Our sampling analysis suggested that observing a similar number and combination of highly similar plasmids in random samples of environmental Salmonella enterica within NCBI Pathogen Detection database was unlikely, supporting a connection between the outbreak strain and the farms implicated by the epidemiological investigation. Conclusion Horizontally transferred plasmids provided evidence for a connection between clinical isolates and the farms implicated as the source of the outbreak. Our case study suggests that such analyses might add a new dimension to source tracking investigations, but highlights the need for detailed and accurate metadata, more extensive environmental sampling, and a better understanding of plasmid molecular evolution. Source tracking molecular epidemiology Salmonella enterica Salmonella enterica Newport pangenome mobilome plasmid Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction In 2020, the United States Food and Drug Administration (FDA) and the Centers for Disease Control and Prevention (CDC) responded to the largest Salmonella outbreak in over a decade. The outbreak was caused by a strain of Salmonella enterica subspecies enterica serovar Newport (from lineage III), referred to here as Salmonella Newport. Salmonella Newport is one of the top five serovars contributing to the approximately 80.3 million foodborne cases of Salmonellosis each year [ 1 ]. Salmonella Newport is composed of three polyphyletic lineages and is frequently associated with cattle [ 2 ]; however, it can colonize a wide range of animal (wild and domesticated) and plant species, providing it multiple reservoirs and multiple transmission routes to humans [ 3 ]. Salmonella Newport can also persist in the environment (e.g., in manure) for months [ 4 ]. The first cases from the outbreak were reported in June of 2020 and the CDC declared the end of the outbreak in October 2020, after causing nearly 2,000 illnesses in the United States and Canada. The investigation traced the outbreak from the food history of the sick patients to red onions grown on two farms in Holtville and Bakersfield California [ 5 ]. This Salmonella outbreak was soon followed in 2021 by another bulb onion outbreak of Salmonella enterica serovar Oranienburg, lending urgency to understanding the chain of events that led to the outbreaks. The leading hypothesis of the FDA from the on-site investigation of the onion farms was that contaminated irrigation water was used to grow the onions at the Holtville California farm [ 7 ]. Plausible sources of contamination were identified, such as sheep grazing on adjacent land and signs of animal intrusion (e.g., scat and large flocks of birds). After harvesting, the outbreak strain could have been transmitted to Bakersfield because many of the Holtville onions had been shipped to Bakersfield for packaging and distribution. Although the investigation occurred after the onions were packaged and distributed, visual observations of the packing house confirmed numerous opportunities for contamination, including signs of animal and pest intrusion, as well as food contact surfaces which had not been properly inspected, maintained, or cleaned [ 7 ]. In molecular epidemiology, if the whole genome sequencing data (WGS) of clinical and environmental isolates collected during an outbreak share a recent common ancestor, they are likely linked in the transmission chain [ 8 ]. Single nucleotide polymorphism (SNP) analysis of the WGS data with the Center for Food Safety and Nutrition (CFSAN) SNP pipeline [ 9 ] determined that the clinical isolates formed a single, highly-related clade (further referred to as the clinical clade), consistent with a single point source of contamination. However, SNP analysis also determined that none of the environmental Salmonella enterica , from 30 serovars, collected near and on the farms (further referred to as the farm isolates), were closely related to the outbreak strain—preventing the conclusive identification of the outbreak source with the WGS data. Here, we explored an alternative method for analyzing the WGS data driven by the hypothesis that isolates that had recently coexisted in the same microbiome (i.e., the outbreak strain and environmental isolates) might share plasmids related by recent horizontal transfer. Horizontal transfer is the movement of genetic material between the genomes of organisms, turning the tree of life into an evolutionary network [ 10 , 11 ]. A common bioinformatic approach for detecting horizontal transfer events is to identify incongruences between the phylogeny of the lineages being analyzed and the phylogenies of the sequences suspected to be horizontally transferred [ 12 ]. Horizontal transfer is a fundamental source of genetic adaptation (such as antimicrobial resistance genes and metabolic pathways) to new environments and conditions for bacteria [ 10 , 11 ]. Plasmids are small (generally between 1 and 200 kbp) intracellular genetic elements that can semi-independently replicate and are thought to be the most impactful source of rapid horizontal transfer in microbial communities [ 13 , 14 ]. Multiple plasmids, representing various incompatibility types (Inc), have been observed in Salmonella Newport (i.e., IncA/C, IncR, IncI1, IncN, IncH, IncF, ColE1, IncP); however, plasmids are likely under-identified given that most were found while looking for genes conferring antimicrobial resistance [ 15 – 21 ]. As part of our study, we analyzed the pangenome—the totality of gene families found in a clade [ 22 ]—of the clinical isolates. The pangenome can be subdivided into the core genes, which are common to all the genomes and the accessory genes which are only found in a subset. Further, pangenomes can be described as closed if the number of observed gene families approaches a limit as more genomes are sampled. In contrast, in an open pangenome, the number of genes continues to grow as more genomes are sampled. A recent study of the three polyphyletic lineages of Salmonella Newport found that each lineage had a closed pangenome with a large number of core genes (ranging from 3,489 to 3,820), indicating the lineages underwent limited horizontal transfer [ 23 ]. We hypothesized that if the reservoir population of the clinical clade existed in the Holtville and Bakersfield farm regions, there was the possibility that it interacted with the local microbial communities via horizontally transferred plasmids. Through the analysis of the 2020 Salmonella Newport onion outbreak, we first explored whether a highly related lineage of Salmonella Newport had a substantial accessory genome and plasmid diversity. Secondly, we sought to detect horizontally transferred plasmids that provided evidence the clinical and farm isolates had coexisted in the same microbiota. Results Whole genome SNP analysis of the clinical isolates Whole genome, reference-based Single nucleotide polymorphism (SNP) analysis confirmed the previous findings of the FDA and CDC outbreak investigation—the clinical clade formed a single, low-diversity clade. The pairwise number of SNP differences ranged from 0 to 16 with a median of 0 (Fig. 1 A). For reference, the median assembly length for the clinical isolates was 4.8 Mbp. The phylogeny constructed from the SNP results revealed a single large polytomy which precluded any observations about the clustering of the clinical isolates by collection date or geographic location. To assess if the clinical clade was associated with a geographical location, the previous SNP analysis was extended to include all closely related (within 1,000 cgMLST alleles of the clinical isolates) environmental isolates from the National Center for Biotechnology Information (NCBI) Pathogen Detection database (Fig. 2 ). The most closely related isolates were mainly from California, New Mexico, and Mexico. The clinical isolates formed a subclade within the SNP phylogeny with 10 environmental isolates from California (seven isolated from almonds, one from parsley, two from environmental swabs) and one from Washington state (isolated from pistachios likely grown in California [ 24 ]). These isolates had been collected between 2010 and 2017 and were between 8 and 39 SNPs distant from the clinical isolates. Sister to this subclade were environmental isolates from California (comminuted beef), New Mexico (environmental swabs from an unknown source), and Mexico (mostly water samples from rivers, canals, and a reservoir, with one from a chicken caecum) collected between 2014 and 2020. The pangenome of the clinical clade had a substantial accessory genome The core genome of the clinical clade was highly conserved, consistent with the low diversity observed in the SNP analysis. There were 4,399 core genes (genes occurring in 99% of the clinical isolates) which was near the median number of genes per genome (4,512 genes). The large and conserved core genome was accompanied by a substantial accessory genome containing 2,577 genes. These accessory genes were sparsely distributed, mostly occurring in 6% or less of the isolates (Fig. 1 B). Plasmids of the clinical and farm isolates Pangenome analysis of the clinical clade revealed a substantial but sparsely distributed accessory genome, indicative of horizontal transfer. It was found that ≥ 64% of the accessory genes were carried on plasmids. For the clinical isolates, 1,814 putative plasmid contigs were identified, belonging to 20 known plasmid types (Table 1 ). We also identified 373 putative plasmid contigs of unknown type. The typeable plasmids accounted for 1,590 genes of the clinical clade pangenome (16 core and 1,574 accessory genes) and ranged in size from approximately 1.6 kbp to 85 kbp long. The most observed plasmid type was the IncFII(S), occurring in all clinical isolates. Other plasmid types occurred infrequently. For example, the second most abundant plasmid type, IncI1-I(Gamma), was only identified in 27 (1.6%) clinical isolates. No clustering of the plasmids was observed in the whole genome SNP phylogeny due to it being a large, unresolved polytomy. Additionally, no plasmid was found to be significantly associated with hospitalization. Table 1 Summary of the plasmids found in the clinical and farm isolates. Plasmid types observed in both the clinical and farm isolates are bolded and in red. Plasmid type Number of clinical isolates with plasmid Number of Holtville isolates with plasmid Number of Bakersfield isolates with plasmid Median contig length (bp) Median number of genes IncFII(S) 1728 3 7 72766 76 Unclassified 373 235 42 6170 7 Col440I 3 60 8 4322 5 IncI1-I(Gamma) 27 13 3 85161 93 IncQ1 0 20 6 8259 10 IncFIB(pB171) 0 25 0 28614 33 Col(pHAD28) 1 12 4 4751 7 IncFII 1 15 0 70562 85 ColpVC 12 4 0 2223 1 IncX3(pEC14) 0 14 0 38955 51 IncFII(pCRY) 0 3 10 49311 62 Col(BS512) 12 0 0 2216 2 ColRNAI 0 1 10 9712 11 IncI2(Delta) 6 0 0 58474 75 IncI(Gamma) 4 0 0 84482 95 Col(MG828) 4 0 0 1672 1 IncFII(Yp) 0 4 0 171753 190 IncR 0 4 0 13431 15 pXuzhou21 3 0 0 40119 54 IncL 3 0 0 59108 81 Col156 2 0 0 4937 5 IncX4 2 0 0 30161 44 Col8282 1 0 0 4218 3 IncB/O/K/Z 1 0 0 21481 25 IncN 1 0 0 38347 49 IncX1 1 0 0 40944 51 pSL483 1 0 0 38560 53 IncM1 1 0 0 58540 76 IncFII(SARC14) 0 1 0 31141 24 Total 2187 413 90 The 512 farm isolates, from 30 different Salmonella serotypes, contained 227 putative extrachromosomal plasmid contigs belonging to 14 known plasmid types and 277 putative plasmid contigs of unknown type (Table 1 ). No plasmid type was common to all isolates and 346 isolates had no identified plasmid contigs. The most commonly observed plasmid types were the Col440I (68 isolates) and the IncQ1 (26 isolates). Six plasmid types were observed in both the clinical and farm isolates: IncFII(S), IncFII, IncI1-I(Gamma), ColpVC, Col440I, and Col(pHAD28). There were also many uncharacterized plasmids observed in the clinical and farm isolates. Analysis of plasmids sharing high similarity in the the clinical and farm isolates For our analysis, we sought to identify plasmids that might have undergone recent horizontal transfer between the clinical and farm isolates. To do this, we restricted our analysis to clinical and farm plasmids sharing high similarity (≥ 95% identity and 90% alignment coverage) i.e., those more likely to be related by recent horizontal transfer. Amongst the six plasmid types and the uncharacterized plasmids observed in both the clinical and farm isolates, there were 14 plasmids from 13 clinical isolates with high similarity to 17 plasmids from 8 farm isolates (Fig. 3 ). The highly similar clinical and farm plasmids belonged to the Col440I, ColpVC, and IncI1-I(Gamma) plasmid types as well as one uncharacterized plasmid type (Table 2 ). The eight farm isolates with these plasmids had been collected from 2 sampling sites separated by 409 km (254 miles): 1) water samples from the New River in Seeley, California (about 20 miles West of Holtville); 2) soil samples near an irrigation filling station next to the Bakersfield onion farm. Five isolates (four Salmonella Corvallis and one Salmonella Liverpool) were collected from the first site and each of them carried two or three high similarity plasmids. The other three isolates, two from Salmonella Idikan and one from Salmonella Typhimurium were collected from the second site and only carried one high similarity plasmid per isolate. Amongst the clinical isolates, all carried a single plasmid with high similarity to a farm plasmid except isolate SRR12424118 which carried two. Table 2 Summary of metadata for farm or best BLAST hit NCBI plasmids with high similarity (≥ 95% identity and ≥ 90% alignment coverage) to a clinical plasmid. Only NCBI plasmids from environmental sources were included. The asterisk for the IncFII(S) plasmid indicates an exception where the pMLST genes were identical between the clinical and farm plasmids, but the percentage of shared gene cargo was less than 40%. Here, phylogenetic neighborhood refers to all NCBI Pathogen Detection isolates within 1000 core gene alleles of the clinical clade. Plasmid type Number of isolates with plasmid (clinical, Holtville, Bakersfield) Number of highly similar plasmids (clinical, Holtville, Bakersfield) Geographic extent of NCBI best hits Environmental isolation sources of NCBI best hits Temporal range of NCBI best hits Bacterial host range (genera) of NCBI best hits IncFII(S)* 1728 3 7 1728 3* 7* United States, Mexico River water, soil, almonds, pistachios, chicken 2010 to 2020 Salmonella IncI1-I(Gamma) 27 13 3 10 1 1 United States, South Korea, Canada, United Kingdom, Denmark, Switzerland, Cow, sheep, dog, pig, chicken, catfish, horse, iguana, lettuce, soil 2002 to 2022 Escherichia , Salmonella , Shigella ColpVC 12 3 0 1 4 0 United States, Canada Chicken, turkey, pig 2007 to 2020 Salmonella Col440I 3 60 8 1 4 2 United States, Germany, Mexico, Venezuela, Ecuador Chicken, cow, pig, papaya, river water 2013 to 2020 Escherichia , Salmonella Unclassified 373 235 42 2 5 0 United States, Australia, United Kingdom, India, Japan Dog, cow, pig, chicken, wastewater 2014 to 2021 Escherichia , Salmonella , Klebsiella Next, we sought to include closely related plasmids from the NCBI for phylogenetic analysis. The 14 clinical plasmids sharing high similarity with the 17 farm plasmids were used to recruit the best basic local alignment search tool (BLAST) hits from the NCBI Pathogen Detection and Nucleotide databases. The best BLAST hits were filtered for plasmids from environmental isolates that shared as high of sequence similarity as that between the clinical and corresponding farm plasmids. We restricted our analysis to environmental isolates to increase the chance of discovering signal from geographic and isolation sources. With the clinical, farm, and NCBI plasmid sequences, multiple sequence alignments were built, and phylogenetic analysis was performed. In some cases, the phylogenies lacked resolution, i.e., formed large polytomies, and so graphs (referred to here as mutation graphs) were built to visualize the mutational similarities and differences between the plasmid sequences. It is important to note that mutation graphs are only meant to visualize mutational differences and are not phylogenetic hypotheses [ 25 ]. For the phylogenetic analysis of the plasmids, we hypothesized that clinical and farm plasmids related by recent horizontal transfer would belong to the same subclade within the phylogeny. However, it was uncertain how close they would cluster in the phylogeny because plasmids can be rapidly transmitted between species and the transmission chain between clinical and farm isolates might have involved intermediate species [ 26 – 28 ]. A brief description of each plasmid as well as the results from the multiple sequence alignments, phylogenies, and mutation graphs are described in the paragraphs below. IncI1-I(gamma) plasmid IncI1 plasmids have been widely observed in Enterobacteriaceae and isolated from many animals. This plasmid type often confers antimicrobial resistance and/or colicin production [ 29 ]. Those analyzed here had a mean length of 85 kbp and 90 predicted genes. Most of the genes could not be functionally annotated (72 genes), however, genes with predicted functions were involved with the Type II secretion system, plasmid segregation and partitioning, bacterial outer membrane adhesion, pilus formation, DNA polymerase IV, and a Colicin-Ia operon. The IncI1-I(gamma) plasmid type occurred in 27 clinical isolates and 16 farm isolates, but only ten clinical and two farm plasmids (1 Holtville Salmonella Liverpool isolate and 1 Bakersfield Salmonella Typhimurium isolate) shared high similarity. The whole plasmids could not be aligned due to differences in gene content and order, so the concatenated core genes were used for the multiple sequence alignment. Analysis of the multiple sequence alignment revealed that the Bakersfield and Holtville plasmids were more distant from each other (276 SNPs) than to any clinical plasmid. Further, the clinical plasmids were more similar to the Holtville (median = 162 SNPs) than the Bakersfield (median = 230 SNPs) plasmids. Additionally, the median SNP distance between the clinical isolates (146 SNPs) was similar to the distance between the clinical and Holtville farm plasmids (155 SNPs). According to metadata held internally by the FDA, the clinical isolates carrying these plasmids had been collected from patients over the duration of a month in multiple states in the USA as well as Canada. The phylogeny of the clinical and farm plasmids and the best BLAST hits from the NCBI showed that eight of the clinical plasmids belonged to a subclade with the two farm isolates (Fig. 4 ). This subclade also contained plasmids from four Escherichia coli , one Escherichia fergusonii , one Shigella flexneri , one Salmonella Newport, and one Salmonella Derby isolates. These isolates were collected from 2002 to 2019, from the United States, Canada, the United Kingdom, South Korea, and Denmark, from cows, sheep, dogs, pigs, and soil samples. Uncharacterized plasmid There were 373 and 277 putative plasmid contigs that could not be classified as a known plasmid type in the clinical and farm isolates, respectively. Amongst these, two clinical and five farm plasmid sequences shared high sequence similarity. These sequences corresponded to a cryptic plasmid previously observed in species across the Enterobacteriaceae family [ 30 – 32 ]. The sequences analyzed here were approximately 4,200 bp long with six predicted genes including a repB replication gene, a mobQ relaxase, a conjugal transfer gene, and three hypothetical genes. The two clinical plasmids differed by 6 SNPs and had been collected from two patients in Oregon and Missouri separated by 13 days. All five of the farm plasmids were collected on the same day at the Holtville collection site, with four in Salmonella Corvallis isolates and one in a Salmonella Liverpool isolate. The four Holtville Salmonella Corvallis plasmids differed from each other by 2 SNPs and the Salmonella Liverpool plasmid by 62 SNPs. Sixteen highly similar plasmids were found in the NCBI Nucleotide and Pathogen databases for which there was no consistent name or type. The phylogeny showed that the two clinical plasmids and four Holtville Salmonella Corvallis plasmids belonged to sister subclades within a larger subclade. These clinical and farm plasmids differed by 29 to 41 SNPs (Fig. 5 ). The subclade of the clinical plasmids also contained plasmids from three E. coli and one Salmonella Heidelberg isolates collected from several animals (dog, cow, pig, chicken) in the United States between 2016 to 2021. The subclade of the four Holtville plasmids contained one Salmonella Typhimurium isolate collected from a pig in Australia in 2014. Col440I plasmid The Col440I plasmid has been observed in species across the Enterobacteriaceae family [ 30 , 31 ] and those analyzed here had three predicted genes: qnrB (quinolone resistance gene [ 33 ]), a phage shock protein transcription activator ( psp gene), and a pentapeptide repeat protein. The plasmid type was only found in three clinical isolates but was the most frequently observed plasmid type in the farm isolates, occurring in 68. However, only one clinical and six farm plasmids shared high similarity. For the highly similar farm plasmids, four were found in Salmonella Corvallis isolates at the collection site near the Holtville farm and two were found in Salmonella Idikan isolates at the collection site near the Bakersfield farm. In the NCBI Nucleotide and Pathogen databases there were 19 highly similar plasmids. These had been collected from 2013 to 2020 from diverse isolation sources (chickens, cows, pigs, papaya, river water, wastewater) and geographic locations (USA, Germany, Canada, Mexico, Venezuela, Ecuador), as well as multiple Salmonella enterica serovars and E. coli . The phylogeny formed a single, unresolved polytomy so a mutation graph was made to visualize the differences between the sequences (Supplementary Fig. 1). The mutation graph showed that there were only 3 SNP differences between the clinical plasmid and the other plasmids, which were all identical. ColpVC plasmid The ColpVC plasmid is a cryptic plasmid [ 34 ] that has been observed in species across the Enterobacteriaceae family [ 30 , 31 ]. Those analyzed here had one replication gene (pfam01446) and one hypothetical gene. The plasmid type was observed in 12 clinical isolates and 4 Salmonella Corvallis farm isolates from the Holtville collection site. Only one of the clinical ColpVC plasmids was highly similar to the 4 farm plasmids. The clinical plasmid differed by 1 or 3 SNPs from the farm plasmids. The four farm plasmids had been collected on the same day, from the same site and serovar, but differed by 0 to 4 SNPs. Within the NCBI Nucleotide and Pathogen databases there were 100 highly similar plasmids found. These had been collected from the USA, UK, and Canada from 1999 to 2021 and were mostly from chickens and turkeys. These were mainly from Salmonella enterica serovars Reading, Kentucky, and Enteritidis, as well as E. coli . The phylogeny formed a single, unresolved polytomy so a mutation graph was made to visualize the differences between the sequences (Supplementary Fig. 2). The mutation graph revealed 57 plasmids that were identical to the clinical plasmid. These had been collected from the USA between 2007 and 2020, from various Salmonella enterica serovars (mostly Reading, Kentucky, and Enteritidis), and were mainly from chickens or turkeys. None of the plasmids from the NCBI were identical to the farm plasmids, but several shared one or two of the same SNP differences with the clinical plasmid. Analysis of the IncFII(S) plasmid Although no high similarity IncFII(S) plasmids were found in the farm isolates, we performed a more granular version of the previous analysis for this plasmid type because it was carried by all clinical isolates. The IncFII(S) plasmid was predicted to be host restricted to the Salmonella genus [ 30 , 31 ]. The clinical versions of this plasmid were approximately 73 kbp long with 77 genes. The genes coded for a Type-F conjugative transfer system, plasmid replication and persistence (e.g., ccdA / ccdB toxin-antitoxin system), as well as a saf fimbrial operon which is strongly correlated with increased virulence in humans [ 35 ]. Only two annotated genes appeared to confer metabolic functions to the bacterial host i.e., ammonia monooxygenase and succinate dehydrogenase flavoprotein. Although the IncFII(S) had no full-length, high similarity matches between the clinical and farm plasmids, further analysis revealed 10 farm isolates with identical genes for the pMLST profile (FIC_5, FIIS_1, FIIY_10). Three were Salmonella arizonae isolates collected near the Holtville farm which shared 29 genes (out of a median of 77 genes per plasmid) with the clinical plasmid. The other seven isolates were Salmonella Typhimurium collected from Bakersfield which only shared four genes with the clinical plasmid. Within the NCBI Pathogen Detection and Nucleotide databases, there were thirty full length, highly similar instances found in environmental isolates (25 Salmonella Newport and 5 Salmonella Javiana). The five Salmonella Javiana isolates had been collected from leafy greens and poultry—the only connection to an animal source. The thirty environmental isolates were all from California or Mexico, except one from Arizona and two from Washington. Twenty of the environmental isolates had been previously identified by the whole genome SNP analysis as highly related to the clinical clade (8 to 39 SNPs different). The phylogeny of the IncFII(S) plasmid mostly formed a single unresolved polytomy except for three isolates which branched into another polytomy. As such, a mutation graph was constructed to better observe the differences between the isolates (Supplementary Fig. 3). The most similar plasmids (having one or zero SNP difference) were from the same California isolates that had formed a subclade with the clinical isolates in the SNP phylogeny. Assessing the significance of observing the highly similar plasmids We sought to assess through a sampling experiment if the number of highly similar plasmids in the clinical and farm isolates (Table 2 ) was due to originating from the same regional microbiota. The null hypothesis for the experiment was that just as many high similarity ColpVC, Col440I, and IncI1-I(Gamma) plasmids could be found in randomly sampled environmental Salmonella enterica isolates—which had been collected across the United States and internationally—as were found in the farm isolates. As a side note, for this analysis the uncharacterized plasmid type was excluded because we could not confidently identify all instances in the clinical and farm isolates. The sampling experiment revealed the frequency of observing high similarity matches to clinical plasmids in random samples of environmental Salmonella isolates (Fig. 6 ). The frequency of observing at least one highly similar ColpVC, Col440I, or IncI1-I(Gamma) plasmid was very high (100%) but varied by plasmid type: 99.7% (ColpVC), 85.1% (Col440I), and 20.7% (IncI1-I(Gamma)). The frequency of observing the outbreak counts or higher for individual plasmid types was very high for the ColpVC plasmids (99.1%), but low for the IncI1-I(Gamma) (2.8%) and Col440I (0%) plasmid types. The phenomena of 3 isolates, each carrying highly similar plasmids from two or more types, as was seen in the outbreak, was never observed. The frequency of observing a similar number of Col440I and IncI1-I(Gamma) plasmids as in the outbreak (6 and 2, respectively) was assessed—excluding the ColpVC plasmids because they were nearly always observed and, thus, not informative (Fig. 6 E). There was a moderate frequency of observing at least one Col440I and one IncI1-I(Gamma) plasmid (18%) in the randomly sample environmental isolates and a very low frequency (0.6%) of observing two of each plasmid type. The frequency of observing at least the same sum of highly similar Col440I and IncI1-I(Gamma) plasmids as in the outbreak was very low (0.01%). Discussion The epidemiological investigation of the 2020 Salmonella Newport onion outbreak strongly implicated two onion farms in Holtville and Bakersfield California as the likely source of the outbreak [ 5 , 7 ]. SNP analysis supported that the clinical isolates likely originated from central California because the most closely related environmental isolates in the NCBI Pathogen Detection database were mainly collected from California and almonds; nearly 100% of the almonds grown in the United states come from California and ~ 75% are produced in five counties (Stanislaus, Fresno, Kern, Merced, and Madera) in central California where Bakersfield is located [ 36 , 37 ]. We had hypothesized that if the outbreak strain had existed in the Holtville and Bakersfield farm regions, it might have interacted with the local microbial communities via horizontally transferred plasmids. Alignment-based and phylogenetic analyses identified highly similar plasmids in the clinical and farm isolates that were likely related by recent horizontal transfer. However, the apparent high promiscuity of the plasmids (e.g., isolated from multiple species within Enterobacteriaceae , isolated from multiple animal and environmental sources, an international distribution, all within the last 20 years) prevented the extraction of information about geographic location, isolation source, the time of transfer, and if there had been direct transfer between the clinical and farm Salmonella serovars or if there were intermediary species. Although promiscuity limited the conclusions drawn for individual plasmids, we argue that the number and combination of highly similar plasmids carried by an outbreak strain and microbiota from a suspected environmental source is a unique and valuable epidemiological marker. Our sampling analysis indicated it was unlikely to observe the number and combination of highly similar plasmids in random samples of environmental Salmonella enterica . This evidence supports that the outbreak strain had interacted, via horizontally transferred plasmids, with the microbiota from the implicated farm regions. An important consideration for sampling analyses, highlighted by our analysis, is that the relevant plasmids were rare in the clinical and farm isolates and thousands of isolates were necessary to identify them. Nonetheless, identifying 20 plasmid types in a very homogeneous population of Salmonella Newport lineage III was surprising given that our literature review indicated that only 11 plasmid types were previously observed in Salmonella Newport [ 15 – 21 ]. This indicates a much greater diversity of plasmids in Salmonella Newport than currently reported and warrants further investigation. Importantly, because many of the observed plasmid types were rare, future studies should examine large sets of isolates. The large diversity of plasmids might be explained by the dynamics of the outbreak, e.g., a population that rapidly grew and changed environments. For example, horizontal transfer can increase as a result of host transition e.g., transitioning from the environment or reservoir host to humans [ 38 ]. Future work should further explore how outbreaks affect plasmid diversity within pathogen populations. Additionally, although we focused on plasmids because they constituted most of the accessory genome, future studies should explore other types of mobile genetic elements such as phages and transposons. Studies should also continue to assess the burden of evidence needed to relate a pathogen to an environmental microbiome using horizontal transfer events. A limitation of our study was that the farm isolates were the only representatives of the farm microbiomes. Although plasmids can be directly transferred between Salmonella enterica serovars, conjugation is dependent upon the proximity and density of donor and recipient cells. Therefore, the local microbiota often serves both as a reservoir for plasmids and as an important intermediary [ 26 – 28 ]. As such, ideally, other species of Enterobacteriaceae would have been collected as well as metagenomic data to increase the chance of detecting horizontally transferred plasmids. It also would have been informative to sample the gut microbiomes of local animals like cattle, which are thought to be a reservoir of Salmonella Newport [ 2 ], because the gut microbiome presents ideal conditions for the horizontal transfer and persistence of plasmids [ 26 – 28 ]. Ideally, there would have also been Enterobacteriaceae isolates and metagenomes from the gut microbiomes of patients to differentiate plasmids that were acquired there from other sources. These considerations point out the complexities of sampling adequately to understand plasmids observed in outbreak isolates. The analysis of the plasmid multiple sequence alignments, phylogenies, and mutation graphs highlighted important gaps in our knowledge about specific plasmid types—especially at very granular levels of microevolution. For example, identical instances of the clinical ColpVC plasmid (Supplementary Fig. 2) could be found over the last 15 years. And yet the farm ColpVC plasmids differed by 0 to 4 SNPs despite being collected from the same sampling site on the same day. Similarly, the two clinical instances of the uncharacterized plasmid (Fig. 5 ) differed by 6 SNPs despite being collected within 13 days of each other (from different patients in different states). And yet, within the plasmid phylogeny the clinical isolates were separated by environmental isolates which had been collected up to six years before. Potentially, the clinical instances represent different lineages of the plasmid and were independently acquired, but the four Salmonella Corvallis farm isolates showed a similar pattern—each differed from the others by two SNPs—despite being collected from the same sampling site on the same day. An alternative explanation is plasmid heterozygosity in the clinical and farm isolates i.e., copies of a plasmid with sequence variations within a single cell or population [ 39 ]. It is also possible that the plasmids might have a high mutation rate but the rate at which they are fixed in the population is low [ 40 ]. Adding further complexity, the Salmonella Liverpool instance of the uncharacterized plasmid differed from the Salmonella Corvallis instances by 62 SNPs despite being collected from the same sampling site and the same day, indicating potential barriers to conjugation despite proximity. Our work also stresses the need for improving the comprehensiveness of plasmid classification. Commonly used plasmid typing methods (e.g., MOB and incompatibility typing) can fail to classify 50% of the plasmids in highly curated datasets (e.g., RefSeq plasmids) much less environmental samples [ 6 ]. Further, current typing methods are most comprehensive for Enterobacteriaceae , but we still observed hundreds of putative plasmid contigs that could not be classified. Related to classification, it can be difficult to assess the evolutionary relatedness of plasmid sequences [ 6 , 41 ]. For example, the three pMLST genes of the IncFII(S) plasmid were identical in the clinical isolates as well as several farm and environmental isolates going back a decade. When using the replicon-based pMLST system [ 42 ], these plasmids would be considered highly related. However, the amount of shared gene cargo for these plasmids ranged from 5–100%—in all cases, the shared gene cargo had high sequence similarity. Here, using the proportion of shared genes as a distance metric [ 41 ] would find the plasmids with less shared gene cargo as more divergent. Methods are needed to reconcile the molecular evolution of core genes and cargo genes when analyzing plasmids. Together, these issues underscore that more work is needed to characterize the molecular evolution of specific plasmid types. Additionally, more environmental sampling is needed to better understand the distribution, population dynamics, and transmission dynamics of plasmids in the environment. Essential to this effort are comprehensive sequence databases and quality metadata. Additional sequence data and metadata might have revealed more details about the plasmid phylogenies e.g., the transmission chain between bacterial species, the geographic source, the isolation source, timeframe of transmission. On a cautionary note, it is important that genome databases not remove “redundant” sequences based upon core genome similarity because the clinical clade had a large, highly conserved core genome, but it was the sparsely distributed accessory genome that made this analysis possible. Conclusion Alignment-based and phylogenetic analyses of highly similar plasmids found in clinical isolates and environmental Salmonella collected from the farm regions implicated in the 2020 Salmonella Newport onion outbreak suggested they were related by recent horizontal transfer. Alignment-based and phylogenetic analyses identified highly similar plasmids in the clinical and farm isolates that were likely related by recent horizontal transfer. However, the promiscuity of the plasmids prevented the extraction of information about geographic location, isolation source, the time of transfer, and if there had been direct transfer between the clinical and farm Salmonella serovars or if there were intermediary species. Despite this promiscuity, our sampling analysis showed that the number and combination of highly similar plasmids carried by an outbreak strain and microbiota from a suspected environmental source might be a unique and valuable epidemiological marker. The sampling analysis indicated it was unlikely to observe the number and combination of highly similar plasmids in random samples of environmental Salmonella enterica . Together our results support the findings of the FDA and CDC investigation i.e., that the outbreak strain had likely originated from two onion farm regions in California. Horizontally transferred plasmids can potentially provide information about microbial interactions, connecting outbreak strains to environments and microbiomes. However, although such analyses might add a new dimension to source tracking investigations, they are dependent upon detailed and accurate metadata, more extensive environmental sampling, and a better understanding of plasmid molecular evolution. Methods Data and sequencing The clinical clade consisted of 1,728 clinical Salmonella enterica (serovar Newport) isolates collected during the 2020 onion outbreak (June to October) in the United States (1,173 isolates) and Canada (555 isolates) [ 43 ]. The farm isolates consisted of 512 environmental Salmonella enterica isolates, from 30 Salmonella enterica serovars, collected from 49 locations, between 2020 and 2021, on and near the onion farms implicated by the epidemiological traceback as well as the nearby irrigation district and public lands. Illumina MiSeq short read sequencing data was available for all clinical and environmental isolates (Supplementary File 1). In addition to the short-read sequencing data, 23 isolates (15 clinical, 8 environmental) were chosen for long-read sequencing so that closed genomes could be acquired. The closed genomes were used to provide more detailed information about the genome biology of the isolates and to validate the presence of plasmids. The 15 clinical isolates were sequenced with Oxford Nanopore technology and the 8 environmental isolates were sequenced with the Pacific Bioscience (PacBio) technology (Supplementary File 1). The clinical isolates were chosen to maximize gene coverage of the pangenome (covered 74.4% of the pangenome) estimated using the 1,728 clinical short read assemblies. The bacteria were grown overnight in tryptic soy broth (TSB) at 37˚C and genomic DNA was extracted using the Maxwell RSC cultured cell DNA kit (Promega, Madison, WI) following the manufacturer’s protocols. The DNA was used to construct libraries for long-read sequencing on the GridIon (Oxford Nanopore Technologies, Oxford, UK) using the rapid sequencing kit RBK004 and run on a MIN106D flow cell (R9.4.1) for 48 hours according to the manufacturer’s instructions. For the 8 environmental isolates, multiplexed microbial SMRTbell libraries were prepared using the SMRTbell Template Prep Kit 2.0 according to PacBio protocol “Preparing Multiplexed Microbial Libraries Using SMRTbell Express Template Prep Kit 2.0” (PacBio, Menlo Park, CA, November 2021). The multiplexed SMRTbell library was then sequenced on a PacBio Sequel IIe sequencer (PacBio, Menlo Park, CA) using Binding Kit 2.2 and Sequel II sequencing Kit 2.0 on one SMRT cell 8M (PacBio, Menlo Park, CA), with 30 hours collection time. Phylogenetic analysis of the clinical clade The SNP matrix was generated by reference-based SNP analysis implemented in the CFSAN SNP Pipeline [ 9 ] with default parameters. The reference genome used for SNP analysis by this study was one of the clinical isolates from the outbreak. It was selected from the clinical isolates for having one of the highest N50’s and lowest number of contigs. The maximum likelihood phylogeny was inferred from the SNP matrix, using GARLI v2.01 [ 44 ] under the General Time-Reversible (GTR) model with Gamma distributed rate heterogeneity, estimate invariant sites, 1000 bootstraps, and 2 categories of variable rates. The phylogeny was rooted using NCBI SRA isolate SRR13685683—a Salmonella Newport clinical isolate collected in 2018—as the outgroup. To explore the phylogenetic neighborhood of the clinical clade (Fig. 2 ), all environmental isolates in the NCBI Pathogen Detection database that were within 1,000 core gene alleles of the clinical clade (using the 1,152 gene cgMLST scheme developed for Salmonella enterica at the FDA [ 45 ]) were recruited for phylogenetic analysis as described in the previous paragraph. The phylogenies were visualized with FigTree (v1.4.4) [ 46 ]. Genome assembly, gene prediction and annotation, and pangenome estimation The short reads were assembled with SPAdes (v3.13.0) [ 47 ] using default settings. For quality control, contigs shorter than 500bp or with less than 10X depth of coverage were removed. Genes in the contigs were predicted and annotated with Prokka (v1.14.5) [ 48 ]. Further annotations were obtained running eggNOG-Mapper (v2.1.6) [ 49 ], using default parameters, on the predicted genes. The predicted genes were clustered with Roary (v3.12.0) [ 50 ] to identify the pangenome of the clinical clade using 90% identity for clustering and requiring that 99% of isolates possess a gene to be considered core. The PacBio sequencing data was demultiplexed by running the Demultiplex Barcodes application and de novo assembly was done using the Microbial Assembly in SMRTLink v.10 (PacBio, Menlo Park, CA). The nanopore sequencing data, and their corresponding MiSeq short reads, were de novo assembled with Unicycler v.0.4.8 [ 51 ]. Nanopore reads shorter than 5 kbp were not used for constructing the hybrid assembly. The long-read sequencing datasets were used to generate 23 complete genomes. These were circularized, oriented to start at the dnaA gene, uploaded to NCBI (Supplementary File 1), and annotated using the NCBI Prokaryotic Genome Annotation Pipeline (PGAP) v5.3. Identification and annotation of plasmids Platon (v1.6) [ 52 ], using default settings, was used to identify and annotate the plasmids in the assemblies. Platon is a tool for the identification of extrachromosomal plasmid contigs in short read draft assemblies. Contigs are characterized by testing for circularization; the detection of incompatibility groups; the detection of rRNA genes; the detection of antimicrobial resistance genes; a homology search against reference plasmid sequences; the detection of oriT sequences; the detection of plasmid replication genes; the detection of mobilization genes; and the detection of conjugation genes [ 52 ]. The online COPLA server [ 30 ] was used to determine the host range of the plasmids. The most observed plasmid type was the IncFII(S), annotated as present by Platon in 98.5% of the isolates. The IncFII(S) plasmids, identified by Platon, were BLAST aligned against the pMLST [ 42 ] database to identify the sequence profiles. To determine if the IncFII(S) plasmid was present in all clinical isolates, the raw reads of all isolates were mapped to the IncFII(S) plasmid of isolate SRR12199170 using Bowtie2 [ 53 ] with default settings. A custom Python script was used to calculate the breadth of coverage of mapped reads for each isolate. Additionally, to identify all contigs belonging to IncFII(S) in all isolate assemblies, the contigs were BLAST aligned (95% identity) to the SRR12199170 plasmid. For the 24 isolates with no identified IncFII(S) plasmid, 23 had reads mapping, with > 92% breadth of coverage, to the complete IncFII(S) plasmid in isolate SRR12199170, and one isolate had 7% breadth of coverage. Validation of plasmids identified in the short-read assemblies The closed genomes that had been long read sequenced with nanopore and PacBio technologies were used to validate the presence of plasmids in the short-read assemblies. Platon was used to confirm that all the DNA sequences shorter than the chromosome in the closed genomes were plasmids. The plasmids identified in the short-read assemblies were BLAST aligned to those in the nanopore and Pacbio assemblies. Plasmids that aligned with ≥ 99% identity and ≥ 99% query coverage were counted as present. This approach confirmed the presence of all plasmids (21 in the environmental isolates and 36 in the clinical isolates) except for a IncFII(S) plasmid that only partially assembled in the short-read assembly of a clinical isolate (only ~ 10% was found in the assembly and the short reads) and a ColpVC plasmid which was not found in the short reads or assembly of an environmental isolate. Comparison of clinical and farm plasmids and finding the best BLAST hits in NCBI The clinical and farm isolate plasmids were pairwise BLAST aligned (using 95% identity). The farm plasmids were filtered for those that aligned with at least 90% coverage of a clinical plasmid using a custom Python script. Metadata internal to the FDA was used to determine the GPS coordinates of the sampling locations and the Salmonella serovars. Farm and clinical plasmids that BLAST aligned, were then queried against the NCBI Pathogen Detection and Nucleotide databases using BLAST (≥ 95% identity). The farm plasmids were filtered for those that aligned with at least 90% coverage of a clinical plasmid using a custom Python script. The results were additionally filtered with a custom Python script for hits that were at least as similar as the farm and clinical plasmids were to each other. The metadata for the best BLAST hits was either parsed from the GenBank file online, if from NCBI Nucleotide database, or from the metadata provided on the NCBI Pathogen Detection website. The metadata was used to select for environmental isolates and to filter out clinical isolates so that information about geographic location and the isolation source could be analyzed. Plasmid phylogenies, mutation graphs, and relatedness The matching clinical and farm plasmids, as well as the filtered BLAST hits of the clinical plasmids from the NCBI Pathogen Detection and Nucleotide databases, as described earlier, were used to build maximum likelihood phylogenies—this was also done for the IncFII(S) plasmid and its best BLAST hits, although there was no matching farm plasmid. For the Col440I, ColpVC, and the uncharacterized plasmids (Fig. 5 ), a custom Python script was used to trim the sequences if they were over-circularized and to reorient the sequences to a common start locus. Then whole plasmid alignment was performed with Muscle [ 54 ] within the MEGA software (version 11) [ 55 ]. For the IncI1-I(Gamma) and IncFII(S) plasmids (Fig. 4 and Supplementary Fig. 3, respectively), recombination, gene gain and loss, and potentially assembly errors made it difficult to perform whole plasmid alignment, so Roary was used to find the core genes (> 95% identity, 100% of isolates must have gene to be core), which were then concatenated and aligned by Roary using MAFFT (v7.305b) [ 56 ]. There were 41 and 33 core genes for the IncFII(S) and IncI1-I(Gamma) plasmids, respectively. For the IncFII(S) plasmids, only one clinical instance was used (isolate SRR12199170) because all the clinical instances clustered at > 99% identity and > 95% coverage with MMSeqs2 [ 57 ]. Maximum likelihood phylogenies were built from the multiple sequence alignments using Mega with the Tamura-Nei nucleotide substitution model; having invariant sites; using the nearest-neighbor-interchange maximum likelihood heuristic method; using the default neighbor joining method to make the initial tree; the bootstrap method as the test of phylogeny with 1000 bootstraps; using the isolate with the oldest collection date as the root; and collapsing all branches with less than 0.75 bootstrap support. The phylogenies were visualized in FigTree. In some cases, the phylogenies lacked resolution, i.e., formed large polytomies, and so graphs (referred to here as mutation graphs) were built to visualize the mutational similarities and differences between the plasmid sequences. Mutation graphs were constructed for the Col440I (Supplementary Fig. 1), ColpVC (Supplementary Fig. 2), and IncFII(S) (Supplementary Fig. 3) plasmids by identifying all variations in the multiple sequence alignments between the sequences and the respective clinical plasmids using a custom Python script. It is important to note that mutation graphs are only meant to visualize mutational differences and are not phylogenetic hypotheses [ 25 ]. Snippy [ 58 ] was used to determine the number of SNP differences between the uncharacterized plasmids. A custom Python script was used to measure the total number of pairwise nucleotide differences for the sequences in the concatenated core gene alignment of the IncI1-I(Gamma) plasmid. In this case, differences were measured per column of the multiple sequence alignment for every two sequence combinations. Here a nucleotide difference could be a mismatch between different nucleotides or a nucleotide and a gap—caused by an insertion or deletion. Statistical analysis for assessing frequency of observing highly similar plasmids in environmental Salmonella enterica We sought to assess through a sampling experiment if the number of highly similar plasmids in the clinical and farm isolates (Table 2 ) was due to originating from the same regional microbiome. The null hypothesis for the experiment was that just as many high similarity ColpVC, Col440I, and IncI1-I(Gamma) plasmids could be found in randomly sampled environmental Salmonella enterica isolates—which had been collected across the United States and internationally—as were found in the farm isolates. For this analysis the uncharacterized plasmid type was excluded because we could not confidently identify all instances in the clinical and farm isolates. To test the null hypothesis, 516 isolates (the number of farm isolates) were randomly sampled from the 95,284 environmental Salmonella enterica in the NCBI Pathogen Detection database (excluding those from the outbreak). All the clinical ColpVC, Col440I, and IncI1-I(Gamma) plasmids—not just those that were highly similar to the farm isolate plasmids—were then aligned to the assemblies of the 516 isolates. Assemblies with plasmids sharing as high a sequence similarity as the highly similar clinical and farm plasmids were counted (Supplementary Table 1). This was repeated 10,000 times and the frequency of observing the plasmids was recorded. The metadata for Salmonella enterica was downloaded from the NCBI Pathogen Detection ftp site on March 28, 2022. A custom Python script was used to randomly select 516 environmental isolates. All the clinical ColpVC, Col440I, and IncI1-I(Gamma) plasmids were BLAST aligned to the 516 assemblies of the randomly selected environmental isolates. A custom Python script was used to parse the BLAST results. To calculate the percent identity of all local alignments to a clinical plasmid, the clinical plasmid was initially represented as an array of zeros. Every position within the clinical plasmid was updated with the maximum percent identity from all local alignments. If the average percent identity of the array representing the alignment to the clinical plasmid was as high as the minimum percent identity for the relevant plasmid type (Supplementary Table 1), it was counted as a highly similar. Random sampling of the environmental isolates and finding highly similar plasmids was repeated 10,000 times to build distributions. The distributions were used to assess the frequency of observing the highly similar plasmids. A custom Python script was used to visualize the results. Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials The read sets used for this analysis are all publically available from NCBI. The SRA accessions are listed in Supplementary File 1. The metadata for the isolates was downloaded and extracted from the NCBI Pathogen Detection ftp site (https://ftp.ncbi.nlm.nih.gov/pathogen/Results/Salmonella/PDG000000002.2417/Metadata/PDG000000002.2417.metadata.tsv). The reference genome used for SNP analysis by this study was one of the clinical isolates from the outbreak (NCBI SRA accession SRR12199170). Competing interests The findings and conclusions presented in this article are those of the authors and do not necessarily represent the view of the US Food and Drug Administration. All the authors, except MP, are either employees of or funded by the Food and Drug Administration. Additionally, HR, JP, and YL directly worked on the investigation of the Salmonella Newport onion outbreak in 2020. Funding KJ was supported by the Joint Institute for Food Safety and Applied Nutrition at the University of Maryland through the cooperative agreement #5U01-FD001418, provided by the Food and Drug Administration, Center for Food Safety and Applied Nutrition. Authors' contributions SC, HR, YL, and KJ designed the study. SC, YL, KJ, and AP analyzed the data. SC, HR, KJ, EKM, JBP, AP, MP, SF, and YL interpreted the data. SC, HR, and YL wrote the manuscript. MH and VJ performed all sequencing and wet lab work. All authors read, provided feedback, and approved the final manuscript. Acknowledgements Not applicable. References Majowicz SE, Musto J, Scallan E, Angulo FJ, Kirk M, O'Brien SJ, Jones TF, Fazil A, Hoekstra RM, for the International Collaboration on Enteric Disease “Burden of Illness” S: The Global Burden of Nontyphoidal Salmonella Gastroenteritis . Clinical Infectious Diseases 2010, 50 (6):882-889. Cao G, Meng J, Strain E, Stones R, Pettengill J, Zhao S, McDermott P, Brown E, Allard M: Phylogenetics and Differentiation of Salmonella Newport Lineages by Whole Genome Sequencing . PLOS ONE 2013, 8 (2):e55687. Pan H, Paudyal N, Li X, Fang W, Yue M: Multiple Food-Animal-Borne Route in Transmission of Antibiotic-Resistant Salmonella Newport to Humans . Frontiers in Microbiology 2018, 9 :23. You Y, Rankin Shelley C, Aceto Helen W, Benson Charles E, Toth John D, Dou Z: Survival of Salmonella enterica Serovar Newport in Manure and Manure-Amended Soils . Applied and Environmental Microbiology 2006, 72 (9):5777-5783. Outbreak of Salmonella Newport Infections Linked to Onions [https://www.cdc.gov/salmonella/newport-07-20/index.html] Acman M, van Dorp L, Santini JM, Balloux F: Large-scale network analysis captures biological features of bacterial plasmids . Nat Commun 2020, 11 (1):2452-2452. Factors Potentially Contributing to the Contamination of Red Onions Implicated in the Summer 2020 Outbreak of Salmonella Newport [https://www.fda.gov/food/outbreaks-foodborne-illness/factors-potentially-contributing-contamination-red-onions-implicated-summer-2020-outbreak-salmonella] Blanc DS, Magalhães B, Koenig I, Senn L, Grandbastien B: Comparison of Whole Genome (wg-) and Core Genome (cg-) MLST (BioNumericsTM) Versus SNP Variant Calling for Epidemiological Investigation of Pseudomonas aeruginosa . Frontiers in Microbiology 2020, 11 :1729. Davis S, Pettengill JB, Luo Y, Payne J, Shpuntoff A, Rand H, Strain E: CFSAN SNP Pipeline: an automated method for constructing SNP matrices from next-generation sequence data . PeerJ Computer Science 2015, 1 :e20. Ochman H, Lawrence JG, Groisman EA: Lateral gene transfer and the nature of bacterial innovation . Nature 2000, 405 (6784):299-304. Arnold BJ, Huang IT, Hanage WP: Horizontal gene transfer and adaptive evolution in bacteria . Nature Reviews Microbiology 2021. Sevillya G, Adato O, Snir S: Detecting horizontal gene transfer: a probabilistic approach . BMC Genomics 2020, 21 (1):106. Pinilla-Redondo R, Cyriaque V, Jacquiod S, Sørensen SJ, Riber L: Monitoring plasmid-mediated horizontal gene transfer in microbiomes: recent advances and future perspectives . Plasmid 2018, 99 :56-67. Hall JPJ, Brockhurst MA, Dytham C, Harrison E: The evolution of plasmid stability: Are infectious transmission and compensatory evolution competing evolutionary trajectories? Plasmid 2017, 91 :90-95. McMillan EA, Jackson CR, Frye JG: Transferable Plasmids of Salmonella enterica Associated With Antibiotic Resistance Genes . Frontiers in Microbiology 2020, 11 :2497. Cao G, Allard M, Hoffmann M, Muruvanda T, Luo Y, Payne J, Meng K, Zhao S, McDermott P, Brown E et al : Sequence Analysis of IncA/C and IncI1 Plasmids Isolated from Multidrug-Resistant Salmonella Newport Using Single-Molecule Real-Time Sequencing . Foodborne Pathogens and Disease 2018, 15 (6):361-371. Poole TL, Callaway TR, Norman KN, Scott HM, Loneragan GH, Ison SA, Beier RC, Harhay DM, Norby B, Nisbet DJ: Transferability of antimicrobial resistance from multidrug-resistant Escherichia coli isolated from cattle in the USA to E. coli and Salmonella Newport recipients . Journal of Global Antimicrobial Resistance 2017, 11 :123-132. Elbediwi M, Pan H, Biswas S, Li Y, Yue M: Emerging colistin resistance in Salmonella enterica serovar Newport isolates from human infections . Emerging Microbes & Infections 2020, 9 (1):535-538. Zheng J, Luo Y, Reed E, Bell R, Brown EW, Hoffmann M: Whole-Genome Comparative Analysis of Salmonella enterica Serovar Newport Strains Reveals Lineage-Specific Divergence . Genome Biology and Evolution 2017, 9 (4):1047-1050. Chen C-Y, Strobaugh TP, Jr., Nguyen L-HT, Abley M, Lindsey RL, Jackson CR: Isolation and characterization of two novel groups of kanamycin-resistance ColE1-like plasmids in Salmonella enterica serotypes from food animals . PLOS ONE 2018, 13 (3):e0193435. Campbell D, Tagg K, Bicknese A, McCullough A, Chen J, Karp BE, Folster JP: Identification and Characterization of Salmonella enterica Serotype Newport Isolates with Decreased Susceptibility to Ciprofloxacin in the United States . Antimicrobial Agents and Chemotherapy , 62 (7):e00653-00618. Tettelin H, Masignani V, Cieslewicz MJ, Donati C, Medini D, Ward NL, Angiuoli SV, Crabtree J, Jones AL, Durkin AS et al : Genome analysis of multiple pathogenic isolates of Streptococcus agalactiae: implications for the microbial "pan-genome" . Proc Natl Acad Sci U S A 2005, 102 (39):13950-13955. de Moraes MH, Soto EB, Salas González I, Desai P, Chu W, Porwollik S, McClelland M, Teplitski M: Genome-Wide Comparative Functional Analyses Reveal Adaptations of Salmonella sv. Newport to a Plant Colonization Lifestyle . Frontiers in Microbiology 2018, 9 :877. National Agricultural Statistics Service [https://quickstats.nass.usda.gov/] Sánchez-Pacheco Santiago J, Kong S, Pulido-Santacruz P, Murphy Robert W, Kubatko L: Median-joining network analysis of SARS-CoV-2 genomes is neither phylogenetic nor evolutionary . Proceedings of the National Academy of Sciences 2020, 117 (23):12518-12519. Tanner JR, Kingsley RA: Evolution of Salmonella within Hosts . Trends in Microbiology 2018, 26 (12):986-998. Scott KP: The role of conjugative transposons in spreading antibiotic resistance between bacteria that inhabit the gastrointestinal tract . Cellular and Molecular Life Sciences CMLS 2002, 59 (12):2071-2082. Aviv G, Rahav G, Gal-Mor O, Davies Julian E: Horizontal Transfer of the Salmonella enterica Serovar Infantis Resistance and Virulence Plasmid pESI to the Gut Microbiota of Warm-Blooded Hosts . mBio , 7 (5):e01395-01316. Foley Steven L, Kaldhone Pravin R, Ricke Steven C, Han J: Incompatibility Group I1 (IncI1) Plasmids: Their Genetics, Biology, and Public Health Relevance . Microbiology and Molecular Biology Reviews , 85 (2):e00031-00020. Redondo-Salvo S, Bartomeus-Peñalver R, Vielva L, Tagg KA, Webb HE, Fernández-López R, de la Cruz F: COPLA, a taxonomic classifier of plasmids . BMC Bioinformatics 2021, 22 (1):390. Redondo-Salvo S, Fernández-López R, Ruiz R, Vielva L, de Toro M, Rocha EPC, Garcillán-Barcia MP, de la Cruz F: Pathways for horizontal gene transfer in bacteria revealed by a global map of their plasmids . Nat Commun 2020, 11 (1):3602. Zaleski P, Wolinowska R, Strzezek K, Lakomy A, Plucienniczak A: The complete sequence and segregational stability analysis of a new cryptic plasmid pIGWZ12 from a clinical strain of Escherichia coli . Plasmid 2006, 56 (3):228-232. Hooper DC, Jacoby GA: Mechanisms of drug resistance: quinolone resistance . Ann N Y Acad Sci 2015, 1354 (1):12-31. Oladeinde A, Cook K, Orlek A, Zock G, Herrington K, Cox N, Plumblee Lawrence J, Hall C: Hotspot mutations and ColE1 plasmids contribute to the fitness of Salmonella Heidelberg in poultry litter . PLOS ONE 2018, 13 (8):e0202286. Folkesson A, Advani A, Sukupolvi S, Pfeifer JD, Normark S, Löfdahl S: Multiple insertions of fimbrial operons correlate with the evolution of Salmonella serovars responsible for human disease . Molecular Microbiology 1999, 33 (3):612-622. Almond production in California [https://apps1.cdfa.ca.gov/FertilizerResearch/docs/Almond_Production_CA.pdf] California Almond Facts [https://www.almonds.com/sites/default/files/content/attachments/almond_industry_-_kern_county.pdf] Sheppard SK, Guttman DS, Fitzgerald JR: Population genomics of bacterial host adaptation . Nature Reviews Genetics 2018, 19 (9):549-565. Bedhomme S, Perez Pantoja D, Bravo IG: Plasmid and clonal interference during post horizontal gene transfer evolution . Mol Ecol 2017, 26 (7):1832-1847. Hughes JM, Lohman BK, Deckert GE, Nichols EP, Settles M, Abdo Z, Top EM: The role of clonal interference in the evolutionary dynamics of plasmid-host adaptation . mBio 2012, 3 (4):e00077-e00012. Suzuki M, Doi Y, Arakawa Y: ORF-based binarized structure network analysis of plasmids (OSNAp), a novel approach to core gene-independent plasmid phylogeny . Plasmid 2020, 108 :102477. Jolley KA, Bray JE, Maiden MCJ: Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications . Wellcome Open Res 2018, 3 :124. [https://www.cdc.gov/salmonella/newport-07-20/index.html] Bazinet AL, Zwickl DJ, Cummings MP: A gateway for phylogenetic analysis powered by grid computing featuring GARLI 2.0 . Syst Biol 2014, 63 (5):812-818. Pettengill JB, Pightling AW, Baugher JD, Rand H, Strain E: Real-Time Pathogen Detection in the Era of Whole-Genome Sequencing and Big Data: Comparison of k-mer and Site-Based Methods for Inferring the Genetic Distances among Tens of Thousands of Salmonella Samples . PLOS ONE 2016, 11 (11):e0166162. Figtree [https://github.com/rambaut/figtree] Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD et al : SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing . J Comput Biol 2012, 19 (5):455-477. Seemann T: Prokka: rapid prokaryotic genome annotation . Bioinformatics 2014, 30 (14):2068-2069. Huerta-Cepas J, Forslund K, Coelho LP, Szklarczyk D, Jensen LJ, von Mering C, Bork P: Fast Genome-Wide Functional Annotation through Orthology Assignment by eggNOG-Mapper . Molecular Biology and Evolution 2017, 34 (8):2115-2122. Page AJ, Cummins CA, Hunt M, Wong VK, Reuter S, Holden MTG, Fookes M, Falush D, Keane JA, Parkhill J: Roary: rapid large-scale prokaryote pan genome analysis . Bioinformatics 2015, 31 (22):3691-3693. Wick RR, Judd LM, Gorrie CL, Holt KE: Unicycler: Resolving bacterial genome assemblies from short and long sequencing reads . PLoS Comput Biol 2017, 13 (6):e1005595-e1005595. Schwengers O, Barth P, Falgenhauer L, Hain T, Chakraborty T, Goesmann A: Platon: identification and characterization of bacterial plasmid contigs in short-read draft assemblies exploiting protein sequence-based replicon distribution scores . Microb Genom 2020, 6 (10):mgen000398. Langmead B, Salzberg SL: Fast gapped-read alignment with Bowtie 2 . Nature methods 2012, 9 (4):357-359. Edgar RC: MUSCLE: multiple sequence alignment with high accuracy and high throughput . Nucleic acids research 2004, 32 (5):1792-1797. Tamura K, Stecher G, Kumar S: MEGA11: Molecular Evolutionary Genetics Analysis Version 11 . Molecular Biology and Evolution 2021, 38 (7):3022-3027. Katoh K, Misawa K, Kuma Ki, Miyata T: MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform . Nucleic Acids Research 2002, 30 (14):3059-3066. Steinegger M, Söding J: MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets . Nature Biotechnology 2017, 35 (11):1026-1028. Snippy [https://github.com/tseemann/snippy] Additional Declarations Competing interest reported. The findings and conclusions presented in this article are those of the authors and do not necessarily represent the view of the US Food and Drug Administration. All the authors, except MP, are either employees of or funded by the Food and Drug Administration. Additionally, HR, JP, and YL directly worked on the investigation of the Salmonella Newport onion outbreak in 2020. Supplementary Files SupplementaryMaterial.docx Cite Share Download PDF Status: Published Journal Publication published 04 Apr, 2023 Read the published version in BMC Genomics → Version 1 posted Editorial decision: Major revision 14 Dec, 2022 Reviews received at journal 28 Nov, 2022 Reviewers agreed at journal 07 Nov, 2022 Reviewers agreed at journal 01 Nov, 2022 Reviewers invited by journal 01 Nov, 2022 Editor assigned by journal 31 Oct, 2022 Editor invited by journal 31 Oct, 2022 Submission checks completed at journal 31 Oct, 2022 First submitted to journal 14 Oct, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2166997","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":148233125,"identity":"3751530f-e0db-42de-932a-afbcb1b98d33","order_by":0,"name":"Seth Commichaux","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABIUlEQVRIie2PMUvEMBiGvxJolnTPcXL5BUKKcC7d/REulUK7WNDF6TjiDekScdWf4XK4WQl0KrreZuGg0w23OhyY1uEGU28VzDO9Q5683wvgcPxBmAAfYGYS6gLABIsYkCeGFV4i87LaKyekPKSAUbxFn3vlQh1U8H3VeCI6Py5Q23xKnSmcNOj6eQ7sTHHrFqUx90SaL7V/Gt7JNFek5eix1hDWtVWBVeJTT2ijEJ8GMspfaMxRIEsIH9LYqnys98poJ2lGaLY1ynxYWSGjLL6VsWmJCb3sWhAwmpTW+XUypXHVb5mOj97SUJHNlQ6kJpxoq8KK15ZuZ1G+fNftaHOjGcHZ09ocNmHFrbAe1vHj5O53wq0dv8KGOxwOh+N/8QXnhF/H8r5KrAAAAABJRU5ErkJggg==","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":true,"prefix":"","firstName":"Seth","middleName":"","lastName":"Commichaux","suffix":""},{"id":148233128,"identity":"22b0a297-8dcd-4566-97b1-ae287ed5055b","order_by":1,"name":"Hugh Rand","email":"","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":false,"prefix":"","firstName":"Hugh","middleName":"","lastName":"Rand","suffix":""},{"id":148233130,"identity":"f8c15d5e-9f6c-4730-a1fa-3dfb8a044515","order_by":2,"name":"Kiran Javkar","email":"","orcid":"","institution":"Department of Computer Science, University of Maryland, MD","correspondingAuthor":false,"prefix":"","firstName":"Kiran","middleName":"","lastName":"Javkar","suffix":""},{"id":148233132,"identity":"e454b589-d0b6-4206-b47d-cd3a4b8ee906","order_by":3,"name":"Erin K. Molloy","email":"","orcid":"","institution":"Department of Computer Science, University of Maryland, MD","correspondingAuthor":false,"prefix":"","firstName":"Erin","middleName":"K.","lastName":"Molloy","suffix":""},{"id":148233134,"identity":"dabdbbd7-4ec3-48b7-8b10-fe5b500ef5d6","order_by":4,"name":"James B. Pettengill","email":"","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":false,"prefix":"","firstName":"James","middleName":"B.","lastName":"Pettengill","suffix":""},{"id":148233136,"identity":"e9e4d852-3db8-4c60-8c99-ba5a903d2562","order_by":5,"name":"Arthur Pightling","email":"","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":false,"prefix":"","firstName":"Arthur","middleName":"","lastName":"Pightling","suffix":""},{"id":148233138,"identity":"7ede5931-2d80-487e-8fa4-307335697f98","order_by":6,"name":"Maria Hoffmann","email":"","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":false,"prefix":"","firstName":"Maria","middleName":"","lastName":"Hoffmann","suffix":""},{"id":148233140,"identity":"0110fa94-c561-42c6-8fcf-4b37d0765f1e","order_by":7,"name":"Mihai Pop","email":"","orcid":"","institution":"Department of Computer Science, University of Maryland, MD","correspondingAuthor":false,"prefix":"","firstName":"Mihai","middleName":"","lastName":"Pop","suffix":""},{"id":148233142,"identity":"97ce771b-b6c0-4797-8064-9eaa8ffb077e","order_by":8,"name":"Victor Jayeola","email":"","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":false,"prefix":"","firstName":"Victor","middleName":"","lastName":"Jayeola","suffix":""},{"id":148233143,"identity":"ac10b55e-869b-47c3-a5e6-dea038954e00","order_by":9,"name":"Steven Foley","email":"","orcid":"","institution":"National Center for Toxicological Research, Food and Drug Administration, AR","correspondingAuthor":false,"prefix":"","firstName":"Steven","middleName":"","lastName":"Foley","suffix":""},{"id":148233144,"identity":"0b1c8c1e-4fb5-4522-886a-e5a3041b9e31","order_by":10,"name":"Yan Luo","email":"","orcid":"","institution":"Center for Food Safety and Nutrition, Food and Drug Administration, MD","correspondingAuthor":false,"prefix":"","firstName":"Yan","middleName":"","lastName":"Luo","suffix":""}],"badges":[],"createdAt":"2022-10-14 15:59:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2166997/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2166997/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12864-023-09245-0","type":"published","date":"2023-04-04T20:24:28+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":28574528,"identity":"2407d14c-7c54-49b9-b0ee-3b65a2ee305f","added_by":"auto","created_at":"2022-11-02 17:35:31","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":39945,"visible":true,"origin":"","legend":"\u003cp\u003eGenetic characteristics of the 1,728 clinical isolates. A) The median number of pairwise SNP differences. B) A histogram showing the number of isolates each of the 6,976 pangenome genes occurs in.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/a2f94f7ed206653dc45fab4f.png"},{"id":28574822,"identity":"7f5a2b4c-82d5-4190-9558-d56f79c0a2a9","added_by":"auto","created_at":"2022-11-02 17:43:31","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":343343,"visible":true,"origin":"","legend":"\u003cp\u003eGeography of the phylogenetic neighborhood of the clinical clade. The SNP-based maximum likelihood phylogeny contains all environmental isolates identified in the NCBI Pathogen Detection database within 1000 cgMLST alleles of the clinical clade. Only select bootstrap values that highlight the separation between the clinical and environmental isolates are shown for clarity. The clinical clade (blue), which is only represented by ten isolates here, is nested within a larger clade (magenta) of ten isolates collected from California and one isolate from Washington, collected between 2010 and 2017. The sister clade (orange) contains isolates from California, New Mexico, and Mexico collected between 2014 and 2020. The phylogeny was inferred using GARLI (with 1,000 bootstraps) with the SNP matrix generated by the CFSAN SNP Pipeline.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/1af15782206fc86cbc282b54.png"},{"id":28574532,"identity":"54775b00-0f9c-44e8-a6fe-eba68457f868","added_by":"auto","created_at":"2022-11-02 17:35:32","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":326211,"visible":true,"origin":"","legend":"\u003cp\u003eHigh similarity plasmids shared by the farm isolates and the clinical clade. A) The map (made in Google My Maps) shows the 49 sampling sites where farm isolates were recovered and the two sites, near Holtville (orange) and Bakersfield (magenta), where high similarity plasmids to those in the clinical clade were found. B) The chart shows the relationship of similar plasmids in the clinical clade and farm isolates. The magenta and orange boxes list the \u003cem\u003eSalmonella enterica\u003c/em\u003eserovars that were collected from the Bakersfield and Holtville regions, respectively. Edges connect the serovar or clinical clade to the plasmid they share. Edge labels indicate the number of isolates with the shared plasmid. The grey boxes indicate the plasmid type that is shared.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/e67144c36f2e0acdf1a4da4a.png"},{"id":28574824,"identity":"970773db-9678-4c5c-83f3-be67a79b2e9d","added_by":"auto","created_at":"2022-11-02 17:43:32","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":169898,"visible":true,"origin":"","legend":"\u003cp\u003eA) The unrooted maximum likelihood phylogeny of the IncI1-I(Gamma) plasmid (branch lengths are not shown). The phylogeny was built from the multiple sequence alignment of the concatenated core genes (33 genes and 24,173 bp). There were 23,180 invariant sites in the multiple sequence alignment. The bootstrap values are provided as internal node labels. Of the 27 clinical isolates and 16 farm isolates that had this plasmid, only ten clinical isolates (blue) and two farm isolates (orange) shared high similarity.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/c4e956628aab058ff7ee1a75.png"},{"id":28574529,"identity":"c8cf7b57-64a0-4450-893d-8e7457a339a3","added_by":"auto","created_at":"2022-11-02 17:35:32","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":59053,"visible":true,"origin":"","legend":"\u003cp\u003eThe unrooted maximum likelihood phylogeny of the uncharacterized plasmid (branch lengths are not shown). The phylogeny was built using the multiple sequence alignment of the whole plasmids (4,201 bp) which had 3,956 invariant sites. The Holtville isolates are in orange and the clinical isolates are in blue. Four of five Holtville isolates formed a clade, each differing by 2 SNPs from the rest. These Holtville isolates were 2 SNPs different from a \u003cem\u003eSalmonella \u003c/em\u003eTyphimurium isolate collected from a pig in Australia in 2014. The clinical plasmids (which differed from each other by 6 SNPs) were in a sister clade to the four Holtville plasmids, differing by 29 to 41 SNPs. Bootstrap values are provided as internal node labels.\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/015391aeaa1164b5ab2f742f.png"},{"id":28574823,"identity":"6162161c-6ecd-4919-a4e2-fa1ecd31c0bd","added_by":"auto","created_at":"2022-11-02 17:43:32","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":100336,"visible":true,"origin":"","legend":"\u003cp\u003eFrequency of observing highly similar plasmids in 10,000 random samples of 516 environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e isolates. The histograms show the number of A) highly similar ColpVC plasmids; B) highly similar Col440I plasmids; C) highly similar IncI1-I(Gamma) plasmids; D) isolates with multiple highly similar plasmids; E) the sum of highly similar Col440I and IncI1-I(Gamma) plasmids. The red vertical lines show the numbers observed in the outbreak. Frequencies show the proportion of 10,000 samples with at least the same number of occurrences as the outbreak.\u003c/p\u003e","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/92d88184941e68a847b809fa.png"},{"id":44724514,"identity":"dc42fde6-6586-4a9f-a9da-66ace85d14ca","added_by":"auto","created_at":"2023-10-16 20:31:30","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2637237,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/906f91ed-c5dc-4665-a27b-f82b7a30e60b.pdf"},{"id":28574531,"identity":"c6ff98e7-b411-4d78-8253-ecc42b46d477","added_by":"auto","created_at":"2022-11-02 17:35:32","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":496926,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-2166997/v1/4ad23248c4558b3acc06e233.docx"}],"financialInterests":"Competing interest reported. The findings and conclusions presented in this article are those of the authors and do not necessarily represent the view of the US Food and Drug Administration. All the authors, except MP, are either employees of or funded by the Food and Drug Administration. Additionally, HR, JP, and YL directly worked on the investigation of the Salmonella Newport onion outbreak in 2020.","formattedTitle":"Assessment of plasmids for relating the 2020 Salmonella enterica serovar Newport onion outbreak to farms implicated by the outbreak investigation","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn 2020, the United States Food and Drug Administration (FDA) and the Centers for Disease Control and Prevention (CDC) responded to the largest \u003cem\u003eSalmonella\u003c/em\u003e outbreak in over a decade. The outbreak was caused by a strain of \u003cem\u003eSalmonella enterica subspecies enterica\u003c/em\u003e serovar Newport (from lineage III), referred to here as \u003cem\u003eSalmonella\u003c/em\u003e Newport. \u003cem\u003eSalmonella\u003c/em\u003e Newport is one of the top five serovars contributing to the approximately 80.3\u0026nbsp;million foodborne cases of Salmonellosis each year [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. \u003cem\u003eSalmonella\u003c/em\u003e Newport is composed of three polyphyletic lineages and is frequently associated with cattle [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]; however, it can colonize a wide range of animal (wild and domesticated) and plant species, providing it multiple reservoirs and multiple transmission routes to humans [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. \u003cem\u003eSalmonella\u003c/em\u003e Newport can also persist in the environment (e.g., in manure) for months [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe first cases from the outbreak were reported in June of 2020 and the CDC declared the end of the outbreak in October 2020, after causing nearly 2,000 illnesses in the United States and Canada. The investigation traced the outbreak from the food history of the sick patients to red onions grown on two farms in Holtville and Bakersfield California [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. This \u003cem\u003eSalmonella\u003c/em\u003e outbreak was soon followed in 2021 by another bulb onion outbreak of \u003cem\u003eSalmonella enterica\u003c/em\u003e serovar Oranienburg, lending urgency to understanding the chain of events that led to the outbreaks.\u003c/p\u003e \u003cp\u003eThe leading hypothesis of the FDA from the on-site investigation of the onion farms was that contaminated irrigation water was used to grow the onions at the Holtville California farm [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Plausible sources of contamination were identified, such as sheep grazing on adjacent land and signs of animal intrusion (e.g., scat and large flocks of birds). After harvesting, the outbreak strain could have been transmitted to Bakersfield because many of the Holtville onions had been shipped to Bakersfield for packaging and distribution. Although the investigation occurred after the onions were packaged and distributed, visual observations of the packing house confirmed numerous opportunities for contamination, including signs of animal and pest intrusion, as well as food contact surfaces which had not been properly inspected, maintained, or cleaned [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn molecular epidemiology, if the whole genome sequencing data (WGS) of clinical and environmental isolates collected during an outbreak share a recent common ancestor, they are likely linked in the transmission chain [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Single nucleotide polymorphism (SNP) analysis of the WGS data with the Center for Food Safety and Nutrition (CFSAN) SNP pipeline [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] determined that the clinical isolates formed a single, highly-related clade (further referred to as the clinical clade), consistent with a single point source of contamination. However, SNP analysis also determined that none of the environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e, from 30 serovars, collected near and on the farms (further referred to as the farm isolates), were closely related to the outbreak strain\u0026mdash;preventing the conclusive identification of the outbreak source with the WGS data.\u003c/p\u003e \u003cp\u003eHere, we explored an alternative method for analyzing the WGS data driven by the hypothesis that isolates that had recently coexisted in the same microbiome (i.e., the outbreak strain and environmental isolates) might share plasmids related by recent horizontal transfer. Horizontal transfer is the movement of genetic material between the genomes of organisms, turning the tree of life into an evolutionary network [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. A common bioinformatic approach for detecting horizontal transfer events is to identify incongruences between the phylogeny of the lineages being analyzed and the phylogenies of the sequences suspected to be horizontally transferred [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Horizontal transfer is a fundamental source of genetic adaptation (such as antimicrobial resistance genes and metabolic pathways) to new environments and conditions for bacteria [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Plasmids are small (generally between 1 and 200 kbp) intracellular genetic elements that can semi-independently replicate and are thought to be the most impactful source of rapid horizontal transfer in microbial communities [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Multiple plasmids, representing various incompatibility types (Inc), have been observed in \u003cem\u003eSalmonella\u003c/em\u003e Newport (i.e., IncA/C, IncR, IncI1, IncN, IncH, IncF, ColE1, IncP); however, plasmids are likely under-identified given that most were found while looking for genes conferring antimicrobial resistance [\u003cspan additionalcitationids=\"CR16 CR17 CR18 CR19 CR20\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAs part of our study, we analyzed the pangenome\u0026mdash;the totality of gene families found in a clade [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]\u0026mdash;of the clinical isolates. The pangenome can be subdivided into the core genes, which are common to all the genomes and the accessory genes which are only found in a subset. Further, pangenomes can be described as closed if the number of observed gene families approaches a limit as more genomes are sampled. In contrast, in an open pangenome, the number of genes continues to grow as more genomes are sampled. A recent study of the three polyphyletic lineages of \u003cem\u003eSalmonella\u003c/em\u003e Newport found that each lineage had a closed pangenome with a large number of core genes (ranging from 3,489 to 3,820), indicating the lineages underwent limited horizontal transfer [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe hypothesized that if the reservoir population of the clinical clade existed in the Holtville and Bakersfield farm regions, there was the possibility that it interacted with the local microbial communities via horizontally transferred plasmids. Through the analysis of the 2020 \u003cem\u003eSalmonella\u003c/em\u003e Newport onion outbreak, we first explored whether a highly related lineage of \u003cem\u003eSalmonella\u003c/em\u003e Newport had a substantial accessory genome and plasmid diversity. Secondly, we sought to detect horizontally transferred plasmids that provided evidence the clinical and farm isolates had coexisted in the same microbiota.\u003c/p\u003e"},{"header":"Results","content":"\u003ch2\u003eWhole genome SNP analysis of the clinical isolates\u003c/h2\u003e\n\u003cp\u003eWhole genome, reference-based Single nucleotide polymorphism (SNP) analysis confirmed the previous findings of the FDA and CDC outbreak investigation\u0026mdash;the clinical clade formed a single, low-diversity clade. The pairwise number of SNP differences ranged from 0 to 16 with a median of 0 (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eA). For reference, the median assembly length for the clinical isolates was 4.8 Mbp. The phylogeny constructed from the SNP results revealed a single large polytomy which precluded any observations about the clustering of the clinical isolates by collection date or geographic location.\u003c/p\u003e\n\u003cp\u003eTo assess if the clinical clade was associated with a geographical location, the previous SNP analysis was extended to include all closely related (within 1,000 cgMLST alleles of the clinical isolates) environmental isolates from the National Center for Biotechnology Information (NCBI) Pathogen Detection database (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). The most closely related isolates were mainly from California, New Mexico, and Mexico. The clinical isolates formed a subclade within the SNP phylogeny with 10 environmental isolates from California (seven isolated from almonds, one from parsley, two from environmental swabs) and one from Washington state (isolated from pistachios likely grown in California [\u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e]). These isolates had been collected between 2010 and 2017 and were between 8 and 39 SNPs distant from the clinical isolates. Sister to this subclade were environmental isolates from California (comminuted beef), New Mexico (environmental swabs from an unknown source), and Mexico (mostly water samples from rivers, canals, and a reservoir, with one from a chicken caecum) collected between 2014 and 2020.\u003c/p\u003e\n\u003ch2\u003eThe pangenome of the clinical clade had a substantial accessory genome\u003c/h2\u003e\n\u003cp\u003eThe core genome of the clinical clade was highly conserved, consistent with the low diversity observed in the SNP analysis. There were 4,399 core genes (genes occurring in 99% of the clinical isolates) which was near the median number of genes per genome (4,512 genes). The large and conserved core genome was accompanied by a substantial accessory genome containing 2,577 genes. These accessory genes were sparsely distributed, mostly occurring in 6% or less of the isolates (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eB).\u003c/p\u003e\n\u003ch2\u003ePlasmids of the clinical and farm isolates\u003c/h2\u003e\n\u003cp\u003ePangenome analysis of the clinical clade revealed a substantial but sparsely distributed accessory genome, indicative of horizontal transfer. It was found that \u0026ge;\u0026thinsp;64% of the accessory genes were carried on plasmids. For the clinical isolates, 1,814 putative plasmid contigs were identified, belonging to 20 known plasmid types (Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). We also identified 373 putative plasmid contigs of unknown type. The typeable plasmids accounted for 1,590 genes of the clinical clade pangenome (16 core and 1,574 accessory genes) and ranged in size from approximately 1.6 kbp to 85 kbp long. The most observed plasmid type was the IncFII(S), occurring in all clinical isolates. Other plasmid types occurred infrequently. For example, the second most abundant plasmid type, IncI1-I(Gamma), was only identified in 27 (1.6%) clinical isolates. No clustering of the plasmids was observed in the whole genome SNP phylogeny due to it being a large, unresolved polytomy. Additionally, no plasmid was found to be significantly associated with hospitalization.\u003c/p\u003e\u0026nbsp;\u003ctable border=\"1\" id=\"Tab1\"\u003e\n \u003ccaption language=\"En\"\u003eTable 1\u0026nbsp;\u003cp\u003eSummary of the plasmids found in the clinical and farm isolates. Plasmid types observed in both the clinical and farm isolates are bolded and in red.\u003c/p\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"6\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003e\n \u003cp\u003ePlasmid type\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eNumber of clinical\u003c/p\u003e\n \u003cp\u003eisolates with plasmid\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eNumber of Holtville isolates with plasmid\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eNumber of Bakersfield isolates with plasmid\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eMedian contig length (bp)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eMedian number of genes\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eIncFII(S)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e1728\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e7\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e72766\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e76\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eUnclassified\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e373\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e235\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6170\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCol440I\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e60\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e8\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4322\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eIncI1-I(Gamma)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e27\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e13\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e85161\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e93\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncQ1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e8259\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncFIB(pB171)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e28614\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e33\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eCol(pHAD28)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e12\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4751\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eIncFII\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e15\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e70562\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e85\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eColpVC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e12\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2223\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncX3(pEC14)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e38955\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e51\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncFII(pCRY)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49311\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e62\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCol(BS512)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2216\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eColRNAI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e9712\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e11\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncI2(Delta)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e58474\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e75\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncI(Gamma)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e84482\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e95\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCol(MG828)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1672\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncFII(Yp)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e171753\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e190\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e13431\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003epXuzhou21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e40119\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e54\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncL\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e59108\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e81\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCol156\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4937\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncX4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e30161\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e44\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCol8282\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4218\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncB/O/K/Z\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e21481\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e38347\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncX1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e40944\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e51\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003epSL483\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e38560\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncM1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e58540\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e76\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncFII(SARC14)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e31141\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e24\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eTotal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2187\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e413\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eThe 512 farm isolates, from 30 different \u003cem\u003eSalmonella\u003c/em\u003e serotypes, contained 227 putative extrachromosomal plasmid contigs belonging to 14 known plasmid types and 277 putative plasmid contigs of unknown type (Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). No plasmid type was common to all isolates and 346 isolates had no identified plasmid contigs. The most commonly observed plasmid types were the Col440I (68 isolates) and the IncQ1 (26 isolates).\u003c/p\u003e\n\u003cp\u003eSix plasmid types were observed in both the clinical and farm isolates: IncFII(S), IncFII, IncI1-I(Gamma), ColpVC, Col440I, and Col(pHAD28). There were also many uncharacterized plasmids observed in the clinical and farm isolates.\u003c/p\u003e\n\u003ch2\u003eAnalysis of plasmids sharing high similarity in the the clinical and farm isolates\u003c/h2\u003e\n\u003cp\u003eFor our analysis, we sought to identify plasmids that might have undergone recent horizontal transfer between the clinical and farm isolates. To do this, we restricted our analysis to clinical and farm plasmids sharing high similarity (\u0026ge;\u0026thinsp;95% identity and 90% alignment coverage) i.e., those more likely to be related by recent horizontal transfer. Amongst the six plasmid types and the uncharacterized plasmids observed in both the clinical and farm isolates, there were 14 plasmids from 13 clinical isolates with high similarity to 17 plasmids from 8 farm isolates (Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). The highly similar clinical and farm plasmids belonged to the Col440I, ColpVC, and IncI1-I(Gamma) plasmid types as well as one uncharacterized plasmid type (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). The eight farm isolates with these plasmids had been collected from 2 sampling sites separated by 409 km (254 miles): 1) water samples from the New River in Seeley, California (about 20 miles West of Holtville); 2) soil samples near an irrigation filling station next to the Bakersfield onion farm. Five isolates (four \u003cem\u003eSalmonella\u003c/em\u003e Corvallis and one \u003cem\u003eSalmonella\u003c/em\u003e Liverpool) were collected from the first site and each of them carried two or three high similarity plasmids. The other three isolates, two from \u003cem\u003eSalmonella\u003c/em\u003e Idikan and one from \u003cem\u003eSalmonella\u003c/em\u003e Typhimurium were collected from the second site and only carried one high similarity plasmid per isolate. Amongst the clinical isolates, all carried a single plasmid with high similarity to a farm plasmid except isolate SRR12424118 which carried two.\u003c/p\u003e\u0026nbsp;\u003ctable border=\"1\" id=\"Tab2\"\u003e\n \u003ccaption language=\"En\"\u003eTable 2\u0026nbsp;\u003cp\u003eSummary of metadata for farm or best BLAST hit NCBI plasmids with high similarity (\u0026ge;\u0026thinsp;95% identity and \u0026ge;\u0026thinsp;90% alignment coverage) to a clinical plasmid. Only NCBI plasmids from environmental sources were included. The asterisk for the IncFII(S) plasmid indicates an exception where the pMLST genes were identical between the clinical and farm plasmids, but the percentage of shared gene cargo was less than 40%. Here, phylogenetic neighborhood refers to all NCBI Pathogen Detection isolates within 1000 core gene alleles of the clinical clade.\u003c/p\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"11\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003e\n \u003cp\u003ePlasmid type\u003c/p\u003e\n \u003c/th\u003e\n \u003cth colspan=\"3\"\u003e\n \u003cp\u003eNumber of\u003c/p\u003e\n \u003cp\u003eisolates with plasmid\u003c/p\u003e\n \u003cp\u003e(clinical, Holtville, Bakersfield)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth colspan=\"3\"\u003e\n \u003cp\u003eNumber of highly similar plasmids\u003c/p\u003e\n \u003cp\u003e(clinical, Holtville, Bakersfield)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eGeographic extent of NCBI best hits\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eEnvironmental isolation sources of NCBI best hits\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eTemporal range of NCBI best hits\u003c/p\u003e\n \u003c/th\u003e\n \u003cth\u003e\n \u003cp\u003eBacterial host range (genera) of NCBI best hits\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncFII(S)*\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e3*\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e7*\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eUnited States, Mexico\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eRiver water, soil, almonds, pistachios, chicken\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2010 to 2020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cem\u003eSalmonella\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIncI1-I(Gamma)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eUnited States, South Korea, Canada, United Kingdom, Denmark, Switzerland,\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCow, sheep, dog, pig, chicken, catfish, horse, iguana, lettuce, soil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2002 to 2022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cem\u003eEscherichia\u003c/em\u003e, \u003cem\u003eSalmonella\u003c/em\u003e, \u003cem\u003eShigella\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eColpVC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eUnited States, Canada\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eChicken, turkey, pig\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2007 to 2020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cem\u003eSalmonella\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCol440I\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e60\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eUnited States, Germany, Mexico, Venezuela, Ecuador\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eChicken, cow, pig, papaya, river water\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2013 to 2020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cem\u003eEscherichia\u003c/em\u003e, \u003cem\u003eSalmonella\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eUnclassified\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e373\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e235\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eUnited States, Australia, United Kingdom, India, Japan\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eDog, cow, pig, chicken, wastewater\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e2014 to 2021\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cem\u003eEscherichia\u003c/em\u003e, \u003cem\u003eSalmonella\u003c/em\u003e, \u003cem\u003eKlebsiella\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eNext, we sought to include closely related plasmids from the NCBI for phylogenetic analysis. The 14 clinical plasmids sharing high similarity with the 17 farm plasmids were used to recruit the best basic local alignment search tool (BLAST) hits from the NCBI Pathogen Detection and Nucleotide databases. The best BLAST hits were filtered for plasmids from environmental isolates that shared as high of sequence similarity as that between the clinical and corresponding farm plasmids. We restricted our analysis to environmental isolates to increase the chance of discovering signal from geographic and isolation sources. With the clinical, farm, and NCBI plasmid sequences, multiple sequence alignments were built, and phylogenetic analysis was performed. In some cases, the phylogenies lacked resolution, i.e., formed large polytomies, and so graphs (referred to here as mutation graphs) were built to visualize the mutational similarities and differences between the plasmid sequences. It is important to note that mutation graphs are only meant to visualize mutational differences and are not phylogenetic hypotheses [\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eFor the phylogenetic analysis of the plasmids, we hypothesized that clinical and farm plasmids related by recent horizontal transfer would belong to the same subclade within the phylogeny. However, it was uncertain how close they would cluster in the phylogeny because plasmids can be rapidly transmitted between species and the transmission chain between clinical and farm isolates might have involved intermediate species [\u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e]. A brief description of each plasmid as well as the results from the multiple sequence alignments, phylogenies, and mutation graphs are described in the paragraphs below.\u003c/p\u003e\n\u003ch2\u003eIncI1-I(gamma) plasmid\u003c/h2\u003e\n\u003cp\u003eIncI1 plasmids have been widely observed in \u003cem\u003eEnterobacteriaceae\u003c/em\u003e and isolated from many animals. This plasmid type often confers antimicrobial resistance and/or colicin production [\u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e]. Those analyzed here had a mean length of 85 kbp and 90 predicted genes. Most of the genes could not be functionally annotated (72 genes), however, genes with predicted functions were involved with the Type II secretion system, plasmid segregation and partitioning, bacterial outer membrane adhesion, pilus formation, DNA polymerase IV, and a Colicin-Ia operon.\u003c/p\u003e\n\u003cp\u003eThe IncI1-I(gamma) plasmid type occurred in 27 clinical isolates and 16 farm isolates, but only ten clinical and two farm plasmids (1 Holtville \u003cem\u003eSalmonella\u003c/em\u003e Liverpool isolate and 1 Bakersfield \u003cem\u003eSalmonella\u003c/em\u003e Typhimurium isolate) shared high similarity. The whole plasmids could not be aligned due to differences in gene content and order, so the concatenated core genes were used for the multiple sequence alignment. Analysis of the multiple sequence alignment revealed that the Bakersfield and Holtville plasmids were more distant from each other (276 SNPs) than to any clinical plasmid. Further, the clinical plasmids were more similar to the Holtville (median\u0026thinsp;=\u0026thinsp;162 SNPs) than the Bakersfield (median\u0026thinsp;=\u0026thinsp;230 SNPs) plasmids. Additionally, the median SNP distance between the clinical isolates (146 SNPs) was similar to the distance between the clinical and Holtville farm plasmids (155 SNPs). According to metadata held internally by the FDA, the clinical isolates carrying these plasmids had been collected from patients over the duration of a month in multiple states in the USA as well as Canada.\u003c/p\u003e\n\u003cp\u003eThe phylogeny of the clinical and farm plasmids and the best BLAST hits from the NCBI showed that eight of the clinical plasmids belonged to a subclade with the two farm isolates (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e). This subclade also contained plasmids from four \u003cem\u003eEscherichia coli\u003c/em\u003e, one \u003cem\u003eEscherichia fergusonii\u003c/em\u003e, one \u003cem\u003eShigella flexneri\u003c/em\u003e, one \u003cem\u003eSalmonella\u003c/em\u003e Newport, and one \u003cem\u003eSalmonella\u003c/em\u003e Derby isolates. These isolates were collected from 2002 to 2019, from the United States, Canada, the United Kingdom, South Korea, and Denmark, from cows, sheep, dogs, pigs, and soil samples.\u003c/p\u003e\n\u003ch2\u003eUncharacterized plasmid\u003c/h2\u003e\n\u003cp\u003eThere were 373 and 277 putative plasmid contigs that could not be classified as a known plasmid type in the clinical and farm isolates, respectively. Amongst these, two clinical and five farm plasmid sequences shared high sequence similarity. These sequences corresponded to a cryptic plasmid previously observed in species across the \u003cem\u003eEnterobacteriaceae\u003c/em\u003e family [\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e]. The sequences analyzed here were approximately 4,200 bp long with six predicted genes including a \u003cem\u003erepB\u003c/em\u003e replication gene, a \u003cem\u003emobQ\u003c/em\u003e relaxase, a conjugal transfer gene, and three hypothetical genes.\u003c/p\u003e\n\u003cp\u003eThe two clinical plasmids differed by 6 SNPs and had been collected from two patients in Oregon and Missouri separated by 13 days. All five of the farm plasmids were collected on the same day at the Holtville collection site, with four in \u003cem\u003eSalmonella\u003c/em\u003e Corvallis isolates and one in a \u003cem\u003eSalmonella\u003c/em\u003e Liverpool isolate. The four Holtville \u003cem\u003eSalmonella\u003c/em\u003e Corvallis plasmids differed from each other by 2 SNPs and the \u003cem\u003eSalmonella\u003c/em\u003e Liverpool plasmid by 62 SNPs. Sixteen highly similar plasmids were found in the NCBI Nucleotide and Pathogen databases for which there was no consistent name or type.\u003c/p\u003e\n\u003cp\u003eThe phylogeny showed that the two clinical plasmids and four Holtville \u003cem\u003eSalmonella\u003c/em\u003e Corvallis plasmids belonged to sister subclades within a larger subclade. These clinical and farm plasmids differed by 29 to 41 SNPs (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e). The subclade of the clinical plasmids also contained plasmids from three \u003cem\u003eE. coli\u003c/em\u003e and one \u003cem\u003eSalmonella\u003c/em\u003e Heidelberg isolates collected from several animals (dog, cow, pig, chicken) in the United States between 2016 to 2021. The subclade of the four Holtville plasmids contained one \u003cem\u003eSalmonella\u003c/em\u003e Typhimurium isolate collected from a pig in Australia in 2014.\u003c/p\u003e\n\u003ch2\u003eCol440I plasmid\u003c/h2\u003e\n\u003cp\u003eThe Col440I plasmid has been observed in species across the \u003cem\u003eEnterobacteriaceae\u003c/em\u003e family [\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e] and those analyzed here had three predicted genes: \u003cem\u003eqnrB\u003c/em\u003e (quinolone resistance gene [\u003cspan class=\"CitationRef\"\u003e33\u003c/span\u003e]), a phage shock protein transcription activator (\u003cem\u003epsp\u003c/em\u003e gene), and a pentapeptide repeat protein. The plasmid type was only found in three clinical isolates but was the most frequently observed plasmid type in the farm isolates, occurring in 68. However, only one clinical and six farm plasmids shared high similarity. For the highly similar farm plasmids, four were found in \u003cem\u003eSalmonella\u003c/em\u003e Corvallis isolates at the collection site near the Holtville farm and two were found in \u003cem\u003eSalmonella\u003c/em\u003e Idikan isolates at the collection site near the Bakersfield farm.\u003c/p\u003e\n\u003cp\u003eIn the NCBI Nucleotide and Pathogen databases there were 19 highly similar plasmids. These had been collected from 2013 to 2020 from diverse isolation sources (chickens, cows, pigs, papaya, river water, wastewater) and geographic locations (USA, Germany, Canada, Mexico, Venezuela, Ecuador), as well as multiple \u003cem\u003eSalmonella enterica\u003c/em\u003e serovars and \u003cem\u003eE. coli\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003eThe phylogeny formed a single, unresolved polytomy so a mutation graph was made to visualize the differences between the sequences (Supplementary Fig. 1). The mutation graph showed that there were only 3 SNP differences between the clinical plasmid and the other plasmids, which were all identical.\u003c/p\u003e\n\u003ch2\u003eColpVC plasmid\u003c/h2\u003e\n\u003cp\u003eThe ColpVC plasmid is a cryptic plasmid [\u003cspan class=\"CitationRef\"\u003e34\u003c/span\u003e] that has been observed in species across the \u003cem\u003eEnterobacteriaceae\u003c/em\u003e family [\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e]. Those analyzed here had one replication gene (pfam01446) and one hypothetical gene. The plasmid type was observed in 12 clinical isolates and 4 \u003cem\u003eSalmonella\u003c/em\u003e Corvallis farm isolates from the Holtville collection site. Only one of the clinical ColpVC plasmids was highly similar to the 4 farm plasmids. The clinical plasmid differed by 1 or 3 SNPs from the farm plasmids. The four farm plasmids had been collected on the same day, from the same site and serovar, but differed by 0 to 4 SNPs.\u003c/p\u003e\n\u003cp\u003eWithin the NCBI Nucleotide and Pathogen databases there were 100 highly similar plasmids found. These had been collected from the USA, UK, and Canada from 1999 to 2021 and were mostly from chickens and turkeys. These were mainly from \u003cem\u003eSalmonella enterica\u003c/em\u003e serovars Reading, Kentucky, and Enteritidis, as well as \u003cem\u003eE. coli\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003eThe phylogeny formed a single, unresolved polytomy so a mutation graph was made to visualize the differences between the sequences (Supplementary Fig. 2). The mutation graph revealed 57 plasmids that were identical to the clinical plasmid. These had been collected from the USA between 2007 and 2020, from various \u003cem\u003eSalmonella enterica\u003c/em\u003e serovars (mostly Reading, Kentucky, and Enteritidis), and were mainly from chickens or turkeys. None of the plasmids from the NCBI were identical to the farm plasmids, but several shared one or two of the same SNP differences with the clinical plasmid.\u003c/p\u003e\n\u003ch2\u003eAnalysis of the IncFII(S) plasmid\u003c/h2\u003e\n\u003cp\u003eAlthough no high similarity IncFII(S) plasmids were found in the farm isolates, we performed a more granular version of the previous analysis for this plasmid type because it was carried by all clinical isolates. The IncFII(S) plasmid was predicted to be host restricted to the \u003cem\u003eSalmonella\u003c/em\u003e genus [\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e]. The clinical versions of this plasmid were approximately 73 kbp long with 77 genes. The genes coded for a Type-F conjugative transfer system, plasmid replication and persistence (e.g., \u003cem\u003eccdA\u003c/em\u003e/\u003cem\u003eccdB\u003c/em\u003e toxin-antitoxin system), as well as a \u003cem\u003esaf\u003c/em\u003e fimbrial operon which is strongly correlated with increased virulence in humans [\u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e]. Only two annotated genes appeared to confer metabolic functions to the bacterial host i.e., ammonia monooxygenase and succinate dehydrogenase flavoprotein.\u003c/p\u003e\n\u003cp\u003eAlthough the IncFII(S) had no full-length, high similarity matches between the clinical and farm plasmids, further analysis revealed 10 farm isolates with identical genes for the pMLST profile (FIC_5, FIIS_1, FIIY_10). Three were \u003cem\u003eSalmonella\u003c/em\u003e arizonae isolates collected near the Holtville farm which shared 29 genes (out of a median of 77 genes per plasmid) with the clinical plasmid. The other seven isolates were \u003cem\u003eSalmonella\u003c/em\u003e Typhimurium collected from Bakersfield which only shared four genes with the clinical plasmid.\u003c/p\u003e\n\u003cp\u003eWithin the NCBI Pathogen Detection and Nucleotide databases, there were thirty full length, highly similar instances found in environmental isolates (25 \u003cem\u003eSalmonella\u003c/em\u003e Newport and 5 \u003cem\u003eSalmonella\u003c/em\u003e Javiana). The five \u003cem\u003eSalmonella\u003c/em\u003e Javiana isolates had been collected from leafy greens and poultry\u0026mdash;the only connection to an animal source. The thirty environmental isolates were all from California or Mexico, except one from Arizona and two from Washington. Twenty of the environmental isolates had been previously identified by the whole genome SNP analysis as highly related to the clinical clade (8 to 39 SNPs different).\u003c/p\u003e\n\u003cp\u003eThe phylogeny of the IncFII(S) plasmid mostly formed a single unresolved polytomy except for three isolates which branched into another polytomy. As such, a mutation graph was constructed to better observe the differences between the isolates (Supplementary Fig. 3). The most similar plasmids (having one or zero SNP difference) were from the same California isolates that had formed a subclade with the clinical isolates in the SNP phylogeny.\u003c/p\u003e\n\u003ch2\u003eAssessing the significance of observing the highly similar plasmids\u003c/h2\u003e\n\u003cp\u003eWe sought to assess through a sampling experiment if the number of highly similar plasmids in the clinical and farm isolates (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e) was due to originating from the same regional microbiota. The null hypothesis for the experiment was that just as many high similarity ColpVC, Col440I, and IncI1-I(Gamma) plasmids could be found in randomly sampled environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e isolates\u0026mdash;which had been collected across the United States and internationally\u0026mdash;as were found in the farm isolates. As a side note, for this analysis the uncharacterized plasmid type was excluded because we could not confidently identify all instances in the clinical and farm isolates.\u003c/p\u003e\n\u003cp\u003eThe sampling experiment revealed the frequency of observing high similarity matches to clinical plasmids in random samples of environmental \u003cem\u003eSalmonella\u003c/em\u003e isolates (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e). The frequency of observing at least one highly similar ColpVC, Col440I, or IncI1-I(Gamma) plasmid was very high (100%) but varied by plasmid type: 99.7% (ColpVC), 85.1% (Col440I), and 20.7% (IncI1-I(Gamma)). The frequency of observing the outbreak counts or higher for individual plasmid types was very high for the ColpVC plasmids (99.1%), but low for the IncI1-I(Gamma) (2.8%) and Col440I (0%) plasmid types. The phenomena of 3 isolates, each carrying highly similar plasmids from two or more types, as was seen in the outbreak, was never observed.\u003c/p\u003e\n\u003cp\u003eThe frequency of observing a similar number of Col440I and IncI1-I(Gamma) plasmids as in the outbreak (6 and 2, respectively) was assessed\u0026mdash;excluding the ColpVC plasmids because they were nearly always observed and, thus, not informative (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003eE). There was a moderate frequency of observing at least one Col440I and one IncI1-I(Gamma) plasmid (18%) in the randomly sample environmental isolates and a very low frequency (0.6%) of observing two of each plasmid type. The frequency of observing at least the same sum of highly similar Col440I and IncI1-I(Gamma) plasmids as in the outbreak was very low (0.01%).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe epidemiological investigation of the 2020 \u003cem\u003eSalmonella\u003c/em\u003e Newport onion outbreak strongly implicated two onion farms in Holtville and Bakersfield California as the likely source of the outbreak [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. SNP analysis supported that the clinical isolates likely originated from central California because the most closely related environmental isolates in the NCBI Pathogen Detection database were mainly collected from California and almonds; nearly 100% of the almonds grown in the United states come from California and ~\u0026thinsp;75% are produced in five counties (Stanislaus, Fresno, Kern, Merced, and Madera) in central California where Bakersfield is located [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe had hypothesized that if the outbreak strain had existed in the Holtville and Bakersfield farm regions, it might have interacted with the local microbial communities via horizontally transferred plasmids. Alignment-based and phylogenetic analyses identified highly similar plasmids in the clinical and farm isolates that were likely related by recent horizontal transfer. However, the apparent high promiscuity of the plasmids (e.g., isolated from multiple species within \u003cem\u003eEnterobacteriaceae\u003c/em\u003e, isolated from multiple animal and environmental sources, an international distribution, all within the last 20 years) prevented the extraction of information about geographic location, isolation source, the time of transfer, and if there had been direct transfer between the clinical and farm \u003cem\u003eSalmonella\u003c/em\u003e serovars or if there were intermediary species.\u003c/p\u003e \u003cp\u003eAlthough promiscuity limited the conclusions drawn for individual plasmids, we argue that the number and combination of highly similar plasmids carried by an outbreak strain and microbiota from a suspected environmental source is a unique and valuable epidemiological marker. Our sampling analysis indicated it was unlikely to observe the number and combination of highly similar plasmids in random samples of environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e. This evidence supports that the outbreak strain had interacted, via horizontally transferred plasmids, with the microbiota from the implicated farm regions.\u003c/p\u003e \u003cp\u003eAn important consideration for sampling analyses, highlighted by our analysis, is that the relevant plasmids were rare in the clinical and farm isolates and thousands of isolates were necessary to identify them. Nonetheless, identifying 20 plasmid types in a very homogeneous population of \u003cem\u003eSalmonella\u003c/em\u003e Newport lineage III was surprising given that our literature review indicated that only 11 plasmid types were previously observed in \u003cem\u003eSalmonella\u003c/em\u003e Newport [\u003cspan additionalcitationids=\"CR16 CR17 CR18 CR19 CR20\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. This indicates a much greater diversity of plasmids in \u003cem\u003eSalmonella\u003c/em\u003e Newport than currently reported and warrants further investigation. Importantly, because many of the observed plasmid types were rare, future studies should examine large sets of isolates. The large diversity of plasmids might be explained by the dynamics of the outbreak, e.g., a population that rapidly grew and changed environments. For example, horizontal transfer can increase as a result of host transition e.g., transitioning from the environment or reservoir host to humans [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Future work should further explore how outbreaks affect plasmid diversity within pathogen populations. Additionally, although we focused on plasmids because they constituted most of the accessory genome, future studies should explore other types of mobile genetic elements such as phages and transposons. Studies should also continue to assess the burden of evidence needed to relate a pathogen to an environmental microbiome using horizontal transfer events.\u003c/p\u003e \u003cp\u003eA limitation of our study was that the farm isolates were the only representatives of the farm microbiomes. Although plasmids can be directly transferred between \u003cem\u003eSalmonella enterica\u003c/em\u003e serovars, conjugation is dependent upon the proximity and density of donor and recipient cells. Therefore, the local microbiota often serves both as a reservoir for plasmids and as an important intermediary [\u003cspan additionalcitationids=\"CR27\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. As such, ideally, other species of \u003cem\u003eEnterobacteriaceae\u003c/em\u003e would have been collected as well as metagenomic data to increase the chance of detecting horizontally transferred plasmids. It also would have been informative to sample the gut microbiomes of local animals like cattle, which are thought to be a reservoir of \u003cem\u003eSalmonella\u003c/em\u003e Newport [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], because the gut microbiome presents ideal conditions for the horizontal transfer and persistence of plasmids [\u003cspan additionalcitationids=\"CR27\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Ideally, there would have also been \u003cem\u003eEnterobacteriaceae\u003c/em\u003e isolates and metagenomes from the gut microbiomes of patients to differentiate plasmids that were acquired there from other sources. These considerations point out the complexities of sampling adequately to understand plasmids observed in outbreak isolates.\u003c/p\u003e \u003cp\u003eThe analysis of the plasmid multiple sequence alignments, phylogenies, and mutation graphs highlighted important gaps in our knowledge about specific plasmid types\u0026mdash;especially at very granular levels of microevolution. For example, identical instances of the clinical ColpVC plasmid (Supplementary Fig.\u0026nbsp;2) could be found over the last 15 years. And yet the farm ColpVC plasmids differed by 0 to 4 SNPs despite being collected from the same sampling site on the same day. Similarly, the two clinical instances of the uncharacterized plasmid (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e) differed by 6 SNPs despite being collected within 13 days of each other (from different patients in different states). And yet, within the plasmid phylogeny the clinical isolates were separated by environmental isolates which had been collected up to six years before. Potentially, the clinical instances represent different lineages of the plasmid and were independently acquired, but the four \u003cem\u003eSalmonella\u003c/em\u003e Corvallis farm isolates showed a similar pattern\u0026mdash;each differed from the others by two SNPs\u0026mdash;despite being collected from the same sampling site on the same day. An alternative explanation is plasmid heterozygosity in the clinical and farm isolates i.e., copies of a plasmid with sequence variations within a single cell or population [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. It is also possible that the plasmids might have a high mutation rate but the rate at which they are fixed in the population is low [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. Adding further complexity, the \u003cem\u003eSalmonella\u003c/em\u003e Liverpool instance of the uncharacterized plasmid differed from the \u003cem\u003eSalmonella\u003c/em\u003e Corvallis instances by 62 SNPs despite being collected from the same sampling site and the same day, indicating potential barriers to conjugation despite proximity.\u003c/p\u003e \u003cp\u003eOur work also stresses the need for improving the comprehensiveness of plasmid classification. Commonly used plasmid typing methods (e.g., MOB and incompatibility typing) can fail to classify 50% of the plasmids in highly curated datasets (e.g., RefSeq plasmids) much less environmental samples [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Further, current typing methods are most comprehensive for \u003cem\u003eEnterobacteriaceae\u003c/em\u003e, but we still observed hundreds of putative plasmid contigs that could not be classified. Related to classification, it can be difficult to assess the evolutionary relatedness of plasmid sequences [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. For example, the three pMLST genes of the IncFII(S) plasmid were identical in the clinical isolates as well as several farm and environmental isolates going back a decade. When using the replicon-based pMLST system [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e], these plasmids would be considered highly related. However, the amount of shared gene cargo for these plasmids ranged from 5\u0026ndash;100%\u0026mdash;in all cases, the shared gene cargo had high sequence similarity. Here, using the proportion of shared genes as a distance metric [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e] would find the plasmids with less shared gene cargo as more divergent. Methods are needed to reconcile the molecular evolution of core genes and cargo genes when analyzing plasmids.\u003c/p\u003e \u003cp\u003eTogether, these issues underscore that more work is needed to characterize the molecular evolution of specific plasmid types. Additionally, more environmental sampling is needed to better understand the distribution, population dynamics, and transmission dynamics of plasmids in the environment. Essential to this effort are comprehensive sequence databases and quality metadata. Additional sequence data and metadata might have revealed more details about the plasmid phylogenies e.g., the transmission chain between bacterial species, the geographic source, the isolation source, timeframe of transmission. On a cautionary note, it is important that genome databases not remove \u0026ldquo;redundant\u0026rdquo; sequences based upon core genome similarity because the clinical clade had a large, highly conserved core genome, but it was the sparsely distributed accessory genome that made this analysis possible.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eAlignment-based and phylogenetic analyses of highly similar plasmids found in clinical isolates and environmental \u003cem\u003eSalmonella\u003c/em\u003e collected from the farm regions implicated in the 2020 \u003cem\u003eSalmonella\u003c/em\u003e Newport onion outbreak suggested they were related by recent horizontal transfer. Alignment-based and phylogenetic analyses identified highly similar plasmids in the clinical and farm isolates that were likely related by recent horizontal transfer. However, the promiscuity of the plasmids prevented the extraction of information about geographic location, isolation source, the time of transfer, and if there had been direct transfer between the clinical and farm \u003cem\u003eSalmonella\u003c/em\u003e serovars or if there were intermediary species. Despite this promiscuity, our sampling analysis showed that the number and combination of highly similar plasmids carried by an outbreak strain and microbiota from a suspected environmental source might be a unique and valuable epidemiological marker. The sampling analysis indicated it was unlikely to observe the number and combination of highly similar plasmids in random samples of environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e. Together our results support the findings of the FDA and CDC investigation i.e., that the outbreak strain had likely originated from two onion farm regions in California. Horizontally transferred plasmids can potentially provide information about microbial interactions, connecting outbreak strains to environments and microbiomes. However, although such analyses might add a new dimension to source tracking investigations, they are dependent upon detailed and accurate metadata, more extensive environmental sampling, and a better understanding of plasmid molecular evolution.\u003c/p\u003e "},{"header":"Methods","content":"\u003ch2\u003eData and sequencing\u003c/h2\u003e\n\u003cp\u003eThe clinical clade consisted of 1,728 clinical \u003cem\u003eSalmonella enterica\u003c/em\u003e (serovar Newport) isolates collected during the 2020 onion outbreak (June to October) in the United States (1,173 isolates) and Canada (555 isolates) [\u003cspan class=\"CitationRef\"\u003e43\u003c/span\u003e]. The farm isolates consisted of 512 environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e isolates, from 30 \u003cem\u003eSalmonella enterica\u003c/em\u003e serovars, collected from 49 locations, between 2020 and 2021, on and near the onion farms implicated by the epidemiological traceback as well as the nearby irrigation district and public lands. Illumina MiSeq short read sequencing data was available for all clinical and environmental isolates (Supplementary File 1).\u003c/p\u003e\n\u003cp\u003eIn addition to the short-read sequencing data, 23 isolates (15 clinical, 8 environmental) were chosen for long-read sequencing so that closed genomes could be acquired. The closed genomes were used to provide more detailed information about the genome biology of the isolates and to validate the presence of plasmids. The 15 clinical isolates were sequenced with Oxford Nanopore technology and the 8 environmental isolates were sequenced with the Pacific Bioscience (PacBio) technology (Supplementary File 1). The clinical isolates were chosen to maximize gene coverage of the pangenome (covered 74.4% of the pangenome) estimated using the 1,728 clinical short read assemblies.\u003c/p\u003e\n\u003cp\u003eThe bacteria were grown overnight in tryptic soy broth (TSB) at 37˚C and genomic DNA was extracted using the Maxwell RSC cultured cell DNA kit (Promega, Madison, WI) following the manufacturer\u0026rsquo;s protocols. The DNA was used to construct libraries for long-read sequencing on the GridIon (Oxford Nanopore Technologies, Oxford, UK) using the rapid sequencing kit RBK004 and run on a MIN106D flow cell (R9.4.1) for 48 hours according to the manufacturer\u0026rsquo;s instructions.\u003c/p\u003e\n\u003cp\u003eFor the 8 environmental isolates, multiplexed microbial SMRTbell libraries were prepared using the SMRTbell Template Prep Kit 2.0 according to PacBio protocol \u0026ldquo;Preparing Multiplexed Microbial Libraries Using SMRTbell Express Template Prep Kit 2.0\u0026rdquo; (PacBio, Menlo Park, CA, November 2021). The multiplexed SMRTbell library was then sequenced on a PacBio Sequel IIe sequencer (PacBio, Menlo Park, CA) using Binding Kit 2.2 and Sequel II sequencing Kit 2.0 on one SMRT cell 8M (PacBio, Menlo Park, CA), with 30 hours collection time.\u003c/p\u003e\n\u003ch2\u003ePhylogenetic analysis of the clinical clade\u003c/h2\u003e\n\u003cp\u003eThe SNP matrix was generated by reference-based SNP analysis implemented in the CFSAN SNP Pipeline [\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e] with default parameters. The reference genome used for SNP analysis by this study was one of the clinical isolates from the outbreak. It was selected from the clinical isolates for having one of the highest N50\u0026rsquo;s and lowest number of contigs. The maximum likelihood phylogeny was inferred from the SNP matrix, using GARLI v2.01 [\u003cspan class=\"CitationRef\"\u003e44\u003c/span\u003e] under the General Time-Reversible (GTR) model with Gamma distributed rate heterogeneity, estimate invariant sites, 1000 bootstraps, and 2 categories of variable rates. The phylogeny was rooted using NCBI SRA isolate SRR13685683\u0026mdash;a \u003cem\u003eSalmonella\u003c/em\u003e Newport clinical isolate collected in 2018\u0026mdash;as the outgroup.\u003c/p\u003e\n\u003cp\u003eTo explore the phylogenetic neighborhood of the clinical clade (Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e), all environmental isolates in the NCBI Pathogen Detection database that were within 1,000 core gene alleles of the clinical clade (using the 1,152 gene cgMLST scheme developed for \u003cem\u003eSalmonella enterica\u003c/em\u003e at the FDA [\u003cspan class=\"CitationRef\"\u003e45\u003c/span\u003e]) were recruited for phylogenetic analysis as described in the previous paragraph. The phylogenies were visualized with FigTree (v1.4.4) [\u003cspan class=\"CitationRef\"\u003e46\u003c/span\u003e].\u003c/p\u003e\n\u003ch2\u003eGenome assembly, gene prediction and annotation, and pangenome estimation\u003c/h2\u003e\n\u003cp\u003eThe short reads were assembled with SPAdes (v3.13.0) [\u003cspan class=\"CitationRef\"\u003e47\u003c/span\u003e] using default settings. For quality control, contigs shorter than 500bp or with less than 10X depth of coverage were removed. Genes in the contigs were predicted and annotated with Prokka (v1.14.5) [\u003cspan class=\"CitationRef\"\u003e48\u003c/span\u003e]. Further annotations were obtained running eggNOG-Mapper (v2.1.6) [\u003cspan class=\"CitationRef\"\u003e49\u003c/span\u003e], using default parameters, on the predicted genes. The predicted genes were clustered with Roary (v3.12.0) [\u003cspan class=\"CitationRef\"\u003e50\u003c/span\u003e] to identify the pangenome of the clinical clade using 90% identity for clustering and requiring that 99% of isolates possess a gene to be considered core.\u003c/p\u003e\n\u003cp\u003eThe PacBio sequencing data was demultiplexed by running the Demultiplex Barcodes application and \u003cem\u003ede novo\u003c/em\u003e assembly was done using the Microbial Assembly in SMRTLink v.10 (PacBio, Menlo Park, CA). The nanopore sequencing data, and their corresponding MiSeq short reads, were \u003cem\u003ede novo\u003c/em\u003e assembled with Unicycler v.0.4.8 [\u003cspan class=\"CitationRef\"\u003e51\u003c/span\u003e]. Nanopore reads shorter than 5 kbp were not used for constructing the hybrid assembly.\u003c/p\u003e\n\u003cp\u003eThe long-read sequencing datasets were used to generate 23 complete genomes. These were circularized, oriented to start at the \u003cem\u003ednaA\u003c/em\u003e gene, uploaded to NCBI (Supplementary File 1), and annotated using the NCBI Prokaryotic Genome Annotation Pipeline (PGAP) v5.3.\u003c/p\u003e\n\u003ch2\u003eIdentification and annotation of plasmids\u003c/h2\u003e\n\u003cp\u003ePlaton (v1.6) [\u003cspan class=\"CitationRef\"\u003e52\u003c/span\u003e], using default settings, was used to identify and annotate the plasmids in the assemblies. Platon is a tool for the identification of extrachromosomal plasmid contigs in short read draft assemblies. Contigs are characterized by testing for circularization; the detection of incompatibility groups; the detection of rRNA genes; the detection of antimicrobial resistance genes; a homology search against reference plasmid sequences; the detection of oriT sequences; the detection of plasmid replication genes; the detection of mobilization genes; and the detection of conjugation genes [\u003cspan class=\"CitationRef\"\u003e52\u003c/span\u003e]. The online COPLA server [\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e] was used to determine the host range of the plasmids.\u003c/p\u003e\n\u003cp\u003eThe most observed plasmid type was the IncFII(S), annotated as present by Platon in 98.5% of the isolates. The IncFII(S) plasmids, identified by Platon, were BLAST aligned against the pMLST [\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e] database to identify the sequence profiles. To determine if the IncFII(S) plasmid was present in all clinical isolates, the raw reads of all isolates were mapped to the IncFII(S) plasmid of isolate SRR12199170 using Bowtie2 [\u003cspan class=\"CitationRef\"\u003e53\u003c/span\u003e] with default settings. A custom Python script was used to calculate the breadth of coverage of mapped reads for each isolate. Additionally, to identify all contigs belonging to IncFII(S) in all isolate assemblies, the contigs were BLAST aligned (95% identity) to the SRR12199170 plasmid. For the 24 isolates with no identified IncFII(S) plasmid, 23 had reads mapping, with \u0026gt;\u0026thinsp;92% breadth of coverage, to the complete IncFII(S) plasmid in isolate SRR12199170, and one isolate had 7% breadth of coverage.\u003c/p\u003e\n\u003ch2\u003eValidation of plasmids identified in the short-read assemblies\u003c/h2\u003e\n\u003cp\u003eThe closed genomes that had been long read sequenced with nanopore and PacBio technologies were used to validate the presence of plasmids in the short-read assemblies. Platon was used to confirm that all the DNA sequences shorter than the chromosome in the closed genomes were plasmids. The plasmids identified in the short-read assemblies were BLAST aligned to those in the nanopore and Pacbio assemblies. Plasmids that aligned with \u0026ge;\u0026thinsp;99% identity and \u0026ge;\u0026thinsp;99% query coverage were counted as present. This approach confirmed the presence of all plasmids (21 in the environmental isolates and 36 in the clinical isolates) except for a IncFII(S) plasmid that only partially assembled in the short-read assembly of a clinical isolate (only\u0026thinsp;~\u0026thinsp;10% was found in the assembly and the short reads) and a ColpVC plasmid which was not found in the short reads or assembly of an environmental isolate.\u003c/p\u003e\n\u003ch2\u003eComparison of clinical and farm plasmids and finding the best BLAST hits in NCBI\u003c/h2\u003e\n\u003cp\u003eThe clinical and farm isolate plasmids were pairwise BLAST aligned (using 95% identity). The farm plasmids were filtered for those that aligned with at least 90% coverage of a clinical plasmid using a custom Python script. Metadata internal to the FDA was used to determine the GPS coordinates of the sampling locations and the \u003cem\u003eSalmonella\u003c/em\u003e serovars.\u003c/p\u003e\n\u003cp\u003eFarm and clinical plasmids that BLAST aligned, were then queried against the NCBI Pathogen Detection and Nucleotide databases using BLAST (\u0026ge;\u0026thinsp;95% identity). The farm plasmids were filtered for those that aligned with at least 90% coverage of a clinical plasmid using a custom Python script. The results were additionally filtered with a custom Python script for hits that were at least as similar as the farm and clinical plasmids were to each other.\u003c/p\u003e\n\u003cp\u003eThe metadata for the best BLAST hits was either parsed from the GenBank file online, if from NCBI Nucleotide database, or from the metadata provided on the NCBI Pathogen Detection website. The metadata was used to select for environmental isolates and to filter out clinical isolates so that information about geographic location and the isolation source could be analyzed.\u003c/p\u003e\n\u003ch2\u003ePlasmid phylogenies, mutation graphs, and relatedness\u003c/h2\u003e\n\u003cp\u003eThe matching clinical and farm plasmids, as well as the filtered BLAST hits of the clinical plasmids from the NCBI Pathogen Detection and Nucleotide databases, as described earlier, were used to build maximum likelihood phylogenies\u0026mdash;this was also done for the IncFII(S) plasmid and its best BLAST hits, although there was no matching farm plasmid.\u003c/p\u003e\n\u003cp\u003eFor the Col440I, ColpVC, and the uncharacterized plasmids (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e), a custom Python script was used to trim the sequences if they were over-circularized and to reorient the sequences to a common start locus. Then whole plasmid alignment was performed with Muscle [\u003cspan class=\"CitationRef\"\u003e54\u003c/span\u003e] within the MEGA software (version 11) [\u003cspan class=\"CitationRef\"\u003e55\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eFor the IncI1-I(Gamma) and IncFII(S) plasmids (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e and Supplementary Fig. 3, respectively), recombination, gene gain and loss, and potentially assembly errors made it difficult to perform whole plasmid alignment, so Roary was used to find the core genes (\u0026gt;\u0026thinsp;95% identity, 100% of isolates must have gene to be core), which were then concatenated and aligned by Roary using MAFFT (v7.305b) [\u003cspan class=\"CitationRef\"\u003e56\u003c/span\u003e]. There were 41 and 33 core genes for the IncFII(S) and IncI1-I(Gamma) plasmids, respectively. For the IncFII(S) plasmids, only one clinical instance was used (isolate SRR12199170) because all the clinical instances clustered at \u0026gt;\u0026thinsp;99% identity and \u0026gt;\u0026thinsp;95% coverage with MMSeqs2 [\u003cspan class=\"CitationRef\"\u003e57\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eMaximum likelihood phylogenies were built from the multiple sequence alignments using Mega with the Tamura-Nei nucleotide substitution model; having invariant sites; using the nearest-neighbor-interchange maximum likelihood heuristic method; using the default neighbor joining method to make the initial tree; the bootstrap method as the test of phylogeny with 1000 bootstraps; using the isolate with the oldest collection date as the root; and collapsing all branches with less than 0.75 bootstrap support. The phylogenies were visualized in FigTree.\u003c/p\u003e\n\u003cp\u003eIn some cases, the phylogenies lacked resolution, i.e., formed large polytomies, and so graphs (referred to here as mutation graphs) were built to visualize the mutational similarities and differences between the plasmid sequences. Mutation graphs were constructed for the Col440I (Supplementary Fig.\u0026nbsp;1), ColpVC (Supplementary Fig.\u0026nbsp;2), and IncFII(S) (Supplementary Fig.\u0026nbsp;3) plasmids by identifying all variations in the multiple sequence alignments between the sequences and the respective clinical plasmids using a custom Python script. It is important to note that mutation graphs are only meant to visualize mutational differences and are not phylogenetic hypotheses [\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eSnippy [\u003cspan class=\"CitationRef\"\u003e58\u003c/span\u003e] was used to determine the number of SNP differences between the uncharacterized plasmids. A custom Python script was used to measure the total number of pairwise nucleotide differences for the sequences in the concatenated core gene alignment of the IncI1-I(Gamma) plasmid. In this case, differences were measured per column of the multiple sequence alignment for every two sequence combinations. Here a nucleotide difference could be a mismatch between different nucleotides or a nucleotide and a gap\u0026mdash;caused by an insertion or deletion.\u003c/p\u003e\n\u003ch2\u003eStatistical analysis for assessing frequency of observing highly similar plasmids in environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eWe sought to assess through a sampling experiment if the number of highly similar plasmids in the clinical and farm isolates (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e) was due to originating from the same regional microbiome. The null hypothesis for the experiment was that just as many high similarity ColpVC, Col440I, and IncI1-I(Gamma) plasmids could be found in randomly sampled environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e isolates\u0026mdash;which had been collected across the United States and internationally\u0026mdash;as were found in the farm isolates. For this analysis the uncharacterized plasmid type was excluded because we could not confidently identify all instances in the clinical and farm isolates.\u003c/p\u003e\n\u003cp\u003eTo test the null hypothesis, 516 isolates (the number of farm isolates) were randomly sampled from the 95,284 environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e in the NCBI Pathogen Detection database (excluding those from the outbreak). All the clinical ColpVC, Col440I, and IncI1-I(Gamma) plasmids\u0026mdash;not just those that were highly similar to the farm isolate plasmids\u0026mdash;were then aligned to the assemblies of the 516 isolates. Assemblies with plasmids sharing as high a sequence similarity as the highly similar clinical and farm plasmids were counted (Supplementary Table 1). This was repeated 10,000 times and the frequency of observing the plasmids was recorded.\u003c/p\u003e\n\u003cp\u003eThe metadata for \u003cem\u003eSalmonella enterica\u003c/em\u003e was downloaded from the NCBI Pathogen Detection ftp site on March 28, 2022. A custom Python script was used to randomly select 516 environmental isolates. All the clinical ColpVC, Col440I, and IncI1-I(Gamma) plasmids were BLAST aligned to the 516 assemblies of the randomly selected environmental isolates. A custom Python script was used to parse the BLAST results. To calculate the percent identity of all local alignments to a clinical plasmid, the clinical plasmid was initially represented as an array of zeros. Every position within the clinical plasmid was updated with the maximum percent identity from all local alignments. If the average percent identity of the array representing the alignment to the clinical plasmid was as high as the minimum percent identity for the relevant plasmid type (Supplementary Table 1), it was counted as a highly similar. Random sampling of the environmental isolates and finding highly similar plasmids was repeated 10,000 times to build distributions. The distributions were used to assess the frequency of observing the highly similar plasmids. A custom Python script was used to visualize the results.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eEthics approval and consent to participate\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003eConsent for publication\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003eAvailability of data and materials\u003c/p\u003e\n\u003cp\u003eThe read sets used for this analysis are all publically available from NCBI. The SRA accessions are listed in Supplementary File 1. The metadata for the isolates was downloaded and extracted from the NCBI Pathogen Detection ftp site (https://ftp.ncbi.nlm.nih.gov/pathogen/Results/Salmonella/PDG000000002.2417/Metadata/PDG000000002.2417.metadata.tsv). The reference genome used for SNP analysis by this study was one of the clinical isolates from the outbreak (NCBI SRA accession SRR12199170).\u003c/p\u003e\n\u003cp\u003eCompeting interests\u003c/p\u003e\n\u003cp\u003eThe findings and conclusions presented in this article are those of the authors and do not necessarily represent the view of the US Food and Drug Administration. All the authors, except MP, are either employees of or funded by the Food and Drug Administration. Additionally, HR, JP, and YL directly worked on the investigation of the \u003cem\u003eSalmonella\u0026nbsp;\u003c/em\u003eNewport onion outbreak in 2020.\u003c/p\u003e\n\u003cp\u003eFunding\u003c/p\u003e\n\u003cp\u003eKJ was supported by the Joint Institute for Food Safety and Applied Nutrition at the University of Maryland through the cooperative agreement #5U01-FD001418, provided by the Food and Drug Administration, Center for Food Safety and Applied Nutrition.\u003c/p\u003e\n\u003cp\u003eAuthors\u0026apos; contributions\u003c/p\u003e\n\u003cp\u003eSC, HR, YL, and KJ designed the study. SC, YL, KJ, and AP analyzed the data. SC, HR, KJ, EKM, JBP, AP, MP, SF, and YL interpreted the data. SC, HR, and YL wrote the manuscript. MH and VJ performed all sequencing and wet lab work. All authors read, provided feedback, and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003eAcknowledgements\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMajowicz SE, Musto J, Scallan E, Angulo FJ, Kirk M, O\u0026apos;Brien SJ, Jones TF, Fazil A, Hoekstra RM, for the International Collaboration on Enteric Disease \u0026ldquo;Burden of Illness\u0026rdquo; S: \u003cstrong\u003eThe Global Burden of Nontyphoidal Salmonella Gastroenteritis\u003c/strong\u003e. \u003cem\u003eClinical Infectious Diseases \u003c/em\u003e2010, \u003cstrong\u003e50\u003c/strong\u003e(6):882-889.\u003c/li\u003e\n\u003cli\u003eCao G, Meng J, Strain E, Stones R, Pettengill J, Zhao S, McDermott P, Brown E, Allard M: \u003cstrong\u003ePhylogenetics and Differentiation of Salmonella Newport Lineages by Whole Genome Sequencing\u003c/strong\u003e. \u003cem\u003ePLOS ONE \u003c/em\u003e2013, \u003cstrong\u003e8\u003c/strong\u003e(2):e55687.\u003c/li\u003e\n\u003cli\u003ePan H, Paudyal N, Li X, Fang W, Yue M: \u003cstrong\u003eMultiple Food-Animal-Borne Route in Transmission of Antibiotic-Resistant Salmonella Newport to Humans\u003c/strong\u003e. \u003cem\u003eFrontiers in Microbiology \u003c/em\u003e2018, \u003cstrong\u003e9\u003c/strong\u003e:23.\u003c/li\u003e\n\u003cli\u003eYou Y, Rankin Shelley C, Aceto Helen W, Benson Charles E, Toth John D, Dou Z: \u003cstrong\u003eSurvival of Salmonella enterica Serovar Newport in Manure and Manure-Amended Soils\u003c/strong\u003e. \u003cem\u003eApplied and Environmental Microbiology \u003c/em\u003e2006, \u003cstrong\u003e72\u003c/strong\u003e(9):5777-5783.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOutbreak of Salmonella Newport Infections Linked to Onions \u003c/strong\u003e[https://www.cdc.gov/salmonella/newport-07-20/index.html]\u003c/li\u003e\n\u003cli\u003eAcman M, van Dorp L, Santini JM, Balloux F: \u003cstrong\u003eLarge-scale network analysis captures biological features of bacterial plasmids\u003c/strong\u003e. \u003cem\u003eNat Commun \u003c/em\u003e2020, \u003cstrong\u003e11\u003c/strong\u003e(1):2452-2452.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eFactors Potentially Contributing to the Contamination of Red Onions Implicated in the Summer 2020 Outbreak of Salmonella Newport \u003c/strong\u003e[https://www.fda.gov/food/outbreaks-foodborne-illness/factors-potentially-contributing-contamination-red-onions-implicated-summer-2020-outbreak-salmonella]\u003c/li\u003e\n\u003cli\u003eBlanc DS, Magalh\u0026atilde;es B, Koenig I, Senn L, Grandbastien B: \u003cstrong\u003eComparison of Whole Genome (wg-) and Core Genome (cg-) MLST (BioNumericsTM) Versus SNP Variant Calling for Epidemiological Investigation of Pseudomonas aeruginosa\u003c/strong\u003e. \u003cem\u003eFrontiers in Microbiology \u003c/em\u003e2020, \u003cstrong\u003e11\u003c/strong\u003e:1729.\u003c/li\u003e\n\u003cli\u003eDavis S, Pettengill JB, Luo Y, Payne J, Shpuntoff A, Rand H, Strain E: \u003cstrong\u003eCFSAN SNP Pipeline: an automated method for constructing SNP matrices from next-generation sequence data\u003c/strong\u003e. \u003cem\u003ePeerJ Computer Science \u003c/em\u003e2015, \u003cstrong\u003e1\u003c/strong\u003e:e20.\u003c/li\u003e\n\u003cli\u003eOchman H, Lawrence JG, Groisman EA: \u003cstrong\u003eLateral gene transfer and the nature of bacterial innovation\u003c/strong\u003e. \u003cem\u003eNature \u003c/em\u003e2000, \u003cstrong\u003e405\u003c/strong\u003e(6784):299-304.\u003c/li\u003e\n\u003cli\u003eArnold BJ, Huang IT, Hanage WP: \u003cstrong\u003eHorizontal gene transfer and adaptive evolution in bacteria\u003c/strong\u003e. \u003cem\u003eNature Reviews Microbiology \u003c/em\u003e2021.\u003c/li\u003e\n\u003cli\u003eSevillya G, Adato O, Snir S: \u003cstrong\u003eDetecting horizontal gene transfer: a probabilistic approach\u003c/strong\u003e. \u003cem\u003eBMC Genomics \u003c/em\u003e2020, \u003cstrong\u003e21\u003c/strong\u003e(1):106.\u003c/li\u003e\n\u003cli\u003ePinilla-Redondo R, Cyriaque V, Jacquiod S, S\u0026oslash;rensen SJ, Riber L: \u003cstrong\u003eMonitoring plasmid-mediated horizontal gene transfer in microbiomes: recent advances and future perspectives\u003c/strong\u003e. \u003cem\u003ePlasmid \u003c/em\u003e2018, \u003cstrong\u003e99\u003c/strong\u003e:56-67.\u003c/li\u003e\n\u003cli\u003eHall JPJ, Brockhurst MA, Dytham C, Harrison E: \u003cstrong\u003eThe evolution of plasmid stability: Are infectious transmission and compensatory evolution competing evolutionary trajectories?\u003c/strong\u003e \u003cem\u003ePlasmid \u003c/em\u003e2017, \u003cstrong\u003e91\u003c/strong\u003e:90-95.\u003c/li\u003e\n\u003cli\u003eMcMillan EA, Jackson CR, Frye JG: \u003cstrong\u003eTransferable Plasmids of Salmonella enterica Associated With Antibiotic Resistance Genes\u003c/strong\u003e. \u003cem\u003eFrontiers in Microbiology \u003c/em\u003e2020, \u003cstrong\u003e11\u003c/strong\u003e:2497.\u003c/li\u003e\n\u003cli\u003eCao G, Allard M, Hoffmann M, Muruvanda T, Luo Y, Payne J, Meng K, Zhao S, McDermott P, Brown E\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eSequence Analysis of IncA/C and IncI1 Plasmids Isolated from Multidrug-Resistant Salmonella Newport Using Single-Molecule Real-Time Sequencing\u003c/strong\u003e. \u003cem\u003eFoodborne Pathogens and Disease \u003c/em\u003e2018, \u003cstrong\u003e15\u003c/strong\u003e(6):361-371.\u003c/li\u003e\n\u003cli\u003ePoole TL, Callaway TR, Norman KN, Scott HM, Loneragan GH, Ison SA, Beier RC, Harhay DM, Norby B, Nisbet DJ: \u003cstrong\u003eTransferability of antimicrobial resistance from multidrug-resistant Escherichia coli isolated from cattle in the USA to E. coli and Salmonella Newport recipients\u003c/strong\u003e. \u003cem\u003eJournal of Global Antimicrobial Resistance \u003c/em\u003e2017, \u003cstrong\u003e11\u003c/strong\u003e:123-132.\u003c/li\u003e\n\u003cli\u003eElbediwi M, Pan H, Biswas S, Li Y, Yue M: \u003cstrong\u003eEmerging colistin resistance in Salmonella enterica serovar Newport isolates from human infections\u003c/strong\u003e. \u003cem\u003eEmerging Microbes \u0026amp; Infections \u003c/em\u003e2020, \u003cstrong\u003e9\u003c/strong\u003e(1):535-538.\u003c/li\u003e\n\u003cli\u003eZheng J, Luo Y, Reed E, Bell R, Brown EW, Hoffmann M: \u003cstrong\u003eWhole-Genome Comparative Analysis of Salmonella enterica Serovar Newport Strains Reveals Lineage-Specific Divergence\u003c/strong\u003e. \u003cem\u003eGenome Biology and Evolution \u003c/em\u003e2017, \u003cstrong\u003e9\u003c/strong\u003e(4):1047-1050.\u003c/li\u003e\n\u003cli\u003eChen C-Y, Strobaugh TP, Jr., Nguyen L-HT, Abley M, Lindsey RL, Jackson CR: \u003cstrong\u003eIsolation and characterization of two novel groups of kanamycin-resistance ColE1-like plasmids in Salmonella enterica serotypes from food animals\u003c/strong\u003e. \u003cem\u003ePLOS ONE \u003c/em\u003e2018, \u003cstrong\u003e13\u003c/strong\u003e(3):e0193435.\u003c/li\u003e\n\u003cli\u003eCampbell D, Tagg K, Bicknese A, McCullough A, Chen J, Karp BE, Folster JP: \u003cstrong\u003eIdentification and Characterization of Salmonella enterica Serotype Newport Isolates with Decreased Susceptibility to Ciprofloxacin in the United States\u003c/strong\u003e. \u003cem\u003eAntimicrobial Agents and Chemotherapy\u003c/em\u003e, \u003cstrong\u003e62\u003c/strong\u003e(7):e00653-00618.\u003c/li\u003e\n\u003cli\u003eTettelin H, Masignani V, Cieslewicz MJ, Donati C, Medini D, Ward NL, Angiuoli SV, Crabtree J, Jones AL, Durkin AS\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eGenome analysis of multiple pathogenic isolates of Streptococcus agalactiae: implications for the microbial \u0026quot;pan-genome\u0026quot;\u003c/strong\u003e. \u003cem\u003eProc Natl Acad Sci U S A \u003c/em\u003e2005, \u003cstrong\u003e102\u003c/strong\u003e(39):13950-13955.\u003c/li\u003e\n\u003cli\u003ede Moraes MH, Soto EB, Salas Gonz\u0026aacute;lez I, Desai P, Chu W, Porwollik S, McClelland M, Teplitski M: \u003cstrong\u003eGenome-Wide Comparative Functional Analyses Reveal Adaptations of Salmonella sv. Newport to a Plant Colonization Lifestyle\u003c/strong\u003e. \u003cem\u003eFrontiers in Microbiology \u003c/em\u003e2018, \u003cstrong\u003e9\u003c/strong\u003e:877.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eNational Agricultural Statistics Service \u003c/strong\u003e[https://quickstats.nass.usda.gov/]\u003c/li\u003e\n\u003cli\u003eS\u0026aacute;nchez-Pacheco Santiago J, Kong S, Pulido-Santacruz P, Murphy Robert W, Kubatko L: \u003cstrong\u003eMedian-joining network analysis of SARS-CoV-2 genomes is neither phylogenetic nor evolutionary\u003c/strong\u003e. \u003cem\u003eProceedings of the National Academy of Sciences \u003c/em\u003e2020, \u003cstrong\u003e117\u003c/strong\u003e(23):12518-12519.\u003c/li\u003e\n\u003cli\u003eTanner JR, Kingsley RA: \u003cstrong\u003eEvolution of Salmonella within Hosts\u003c/strong\u003e. \u003cem\u003eTrends in Microbiology \u003c/em\u003e2018, \u003cstrong\u003e26\u003c/strong\u003e(12):986-998.\u003c/li\u003e\n\u003cli\u003eScott KP: \u003cstrong\u003eThe role of conjugative transposons in spreading antibiotic resistance between bacteria that inhabit the gastrointestinal tract\u003c/strong\u003e. \u003cem\u003eCellular and Molecular Life Sciences CMLS \u003c/em\u003e2002, \u003cstrong\u003e59\u003c/strong\u003e(12):2071-2082.\u003c/li\u003e\n\u003cli\u003eAviv G, Rahav G, Gal-Mor O, Davies Julian E: \u003cstrong\u003eHorizontal Transfer of the Salmonella enterica Serovar Infantis Resistance and Virulence Plasmid pESI to the Gut Microbiota of Warm-Blooded Hosts\u003c/strong\u003e. \u003cem\u003emBio\u003c/em\u003e, \u003cstrong\u003e7\u003c/strong\u003e(5):e01395-01316.\u003c/li\u003e\n\u003cli\u003eFoley Steven L, Kaldhone Pravin R, Ricke Steven C, Han J: \u003cstrong\u003eIncompatibility Group I1 (IncI1) Plasmids: Their Genetics, Biology, and Public Health Relevance\u003c/strong\u003e. \u003cem\u003eMicrobiology and Molecular Biology Reviews\u003c/em\u003e, \u003cstrong\u003e85\u003c/strong\u003e(2):e00031-00020.\u003c/li\u003e\n\u003cli\u003eRedondo-Salvo S, Bartomeus-Pe\u0026ntilde;alver R, Vielva L, Tagg KA, Webb HE, Fern\u0026aacute;ndez-L\u0026oacute;pez R, de la Cruz F: \u003cstrong\u003eCOPLA, a taxonomic classifier of plasmids\u003c/strong\u003e. \u003cem\u003eBMC Bioinformatics \u003c/em\u003e2021, \u003cstrong\u003e22\u003c/strong\u003e(1):390.\u003c/li\u003e\n\u003cli\u003eRedondo-Salvo S, Fern\u0026aacute;ndez-L\u0026oacute;pez R, Ruiz R, Vielva L, de Toro M, Rocha EPC, Garcill\u0026aacute;n-Barcia MP, de la Cruz F: \u003cstrong\u003ePathways for horizontal gene transfer in bacteria revealed by a global map of their plasmids\u003c/strong\u003e. \u003cem\u003eNat Commun \u003c/em\u003e2020, \u003cstrong\u003e11\u003c/strong\u003e(1):3602.\u003c/li\u003e\n\u003cli\u003eZaleski P, Wolinowska R, Strzezek K, Lakomy A, Plucienniczak A: \u003cstrong\u003eThe complete sequence and segregational stability analysis of a new cryptic plasmid pIGWZ12 from a clinical strain of Escherichia coli\u003c/strong\u003e. \u003cem\u003ePlasmid \u003c/em\u003e2006, \u003cstrong\u003e56\u003c/strong\u003e(3):228-232.\u003c/li\u003e\n\u003cli\u003eHooper DC, Jacoby GA: \u003cstrong\u003eMechanisms of drug resistance: quinolone resistance\u003c/strong\u003e. \u003cem\u003eAnn N Y Acad Sci \u003c/em\u003e2015, \u003cstrong\u003e1354\u003c/strong\u003e(1):12-31.\u003c/li\u003e\n\u003cli\u003eOladeinde A, Cook K, Orlek A, Zock G, Herrington K, Cox N, Plumblee Lawrence J, Hall C: \u003cstrong\u003eHotspot mutations and ColE1 plasmids contribute to the fitness of Salmonella Heidelberg in poultry litter\u003c/strong\u003e. \u003cem\u003ePLOS ONE \u003c/em\u003e2018, \u003cstrong\u003e13\u003c/strong\u003e(8):e0202286.\u003c/li\u003e\n\u003cli\u003eFolkesson A, Advani A, Sukupolvi S, Pfeifer JD, Normark S, L\u0026ouml;fdahl S: \u003cstrong\u003eMultiple insertions of fimbrial operons correlate with the evolution of Salmonella serovars responsible for human disease\u003c/strong\u003e. \u003cem\u003eMolecular Microbiology \u003c/em\u003e1999, \u003cstrong\u003e33\u003c/strong\u003e(3):612-622.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAlmond production in California \u003c/strong\u003e[https://apps1.cdfa.ca.gov/FertilizerResearch/docs/Almond_Production_CA.pdf]\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCalifornia Almond Facts \u003c/strong\u003e[https://www.almonds.com/sites/default/files/content/attachments/almond_industry_-_kern_county.pdf]\u003c/li\u003e\n\u003cli\u003eSheppard SK, Guttman DS, Fitzgerald JR: \u003cstrong\u003ePopulation genomics of bacterial host adaptation\u003c/strong\u003e. \u003cem\u003eNature Reviews Genetics \u003c/em\u003e2018, \u003cstrong\u003e19\u003c/strong\u003e(9):549-565.\u003c/li\u003e\n\u003cli\u003eBedhomme S, Perez Pantoja D, Bravo IG: \u003cstrong\u003ePlasmid and clonal interference during post horizontal gene transfer evolution\u003c/strong\u003e. \u003cem\u003eMol Ecol \u003c/em\u003e2017, \u003cstrong\u003e26\u003c/strong\u003e(7):1832-1847.\u003c/li\u003e\n\u003cli\u003eHughes JM, Lohman BK, Deckert GE, Nichols EP, Settles M, Abdo Z, Top EM: \u003cstrong\u003eThe role of clonal interference in the evolutionary dynamics of plasmid-host adaptation\u003c/strong\u003e. \u003cem\u003emBio \u003c/em\u003e2012, \u003cstrong\u003e3\u003c/strong\u003e(4):e00077-e00012.\u003c/li\u003e\n\u003cli\u003eSuzuki M, Doi Y, Arakawa Y: \u003cstrong\u003eORF-based binarized structure network analysis of plasmids (OSNAp), a novel approach to core gene-independent plasmid phylogeny\u003c/strong\u003e. \u003cem\u003ePlasmid \u003c/em\u003e2020, \u003cstrong\u003e108\u003c/strong\u003e:102477.\u003c/li\u003e\n\u003cli\u003eJolley KA, Bray JE, Maiden MCJ: \u003cstrong\u003eOpen-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications\u003c/strong\u003e. \u003cem\u003eWellcome Open Res \u003c/em\u003e2018, \u003cstrong\u003e3\u003c/strong\u003e:124.\u003c/li\u003e\n\u003cli\u003e[https://www.cdc.gov/salmonella/newport-07-20/index.html]\u003c/li\u003e\n\u003cli\u003eBazinet AL, Zwickl DJ, Cummings MP: \u003cstrong\u003eA gateway for phylogenetic analysis powered by grid computing featuring GARLI 2.0\u003c/strong\u003e. \u003cem\u003eSyst Biol \u003c/em\u003e2014, \u003cstrong\u003e63\u003c/strong\u003e(5):812-818.\u003c/li\u003e\n\u003cli\u003ePettengill JB, Pightling AW, Baugher JD, Rand H, Strain E: \u003cstrong\u003eReal-Time Pathogen Detection in the Era of Whole-Genome Sequencing and Big Data: Comparison of k-mer and Site-Based Methods for Inferring the Genetic Distances among Tens of Thousands of Salmonella Samples\u003c/strong\u003e. \u003cem\u003ePLOS ONE \u003c/em\u003e2016, \u003cstrong\u003e11\u003c/strong\u003e(11):e0166162.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eFigtree \u003c/strong\u003e[https://github.com/rambaut/figtree]\u003c/li\u003e\n\u003cli\u003eBankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eSPAdes: a new genome assembly algorithm and its applications to single-cell sequencing\u003c/strong\u003e. \u003cem\u003eJ Comput Biol \u003c/em\u003e2012, \u003cstrong\u003e19\u003c/strong\u003e(5):455-477.\u003c/li\u003e\n\u003cli\u003eSeemann T: \u003cstrong\u003eProkka: rapid prokaryotic genome annotation\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2014, \u003cstrong\u003e30\u003c/strong\u003e(14):2068-2069.\u003c/li\u003e\n\u003cli\u003eHuerta-Cepas J, Forslund K, Coelho LP, Szklarczyk D, Jensen LJ, von Mering C, Bork P: \u003cstrong\u003eFast Genome-Wide Functional Annotation through Orthology Assignment by eggNOG-Mapper\u003c/strong\u003e. \u003cem\u003eMolecular Biology and Evolution \u003c/em\u003e2017, \u003cstrong\u003e34\u003c/strong\u003e(8):2115-2122.\u003c/li\u003e\n\u003cli\u003ePage AJ, Cummins CA, Hunt M, Wong VK, Reuter S, Holden MTG, Fookes M, Falush D, Keane JA, Parkhill J: \u003cstrong\u003eRoary: rapid large-scale prokaryote pan genome analysis\u003c/strong\u003e. \u003cem\u003eBioinformatics \u003c/em\u003e2015, \u003cstrong\u003e31\u003c/strong\u003e(22):3691-3693.\u003c/li\u003e\n\u003cli\u003eWick RR, Judd LM, Gorrie CL, Holt KE: \u003cstrong\u003eUnicycler: Resolving bacterial genome assemblies from short and long sequencing reads\u003c/strong\u003e. \u003cem\u003ePLoS Comput Biol \u003c/em\u003e2017, \u003cstrong\u003e13\u003c/strong\u003e(6):e1005595-e1005595.\u003c/li\u003e\n\u003cli\u003eSchwengers O, Barth P, Falgenhauer L, Hain T, Chakraborty T, Goesmann A: \u003cstrong\u003ePlaton: identification and characterization of bacterial plasmid contigs in short-read draft assemblies exploiting protein sequence-based replicon distribution scores\u003c/strong\u003e. \u003cem\u003eMicrob Genom \u003c/em\u003e2020, \u003cstrong\u003e6\u003c/strong\u003e(10):mgen000398.\u003c/li\u003e\n\u003cli\u003eLangmead B, Salzberg SL: \u003cstrong\u003eFast gapped-read alignment with Bowtie 2\u003c/strong\u003e. \u003cem\u003eNature methods \u003c/em\u003e2012, \u003cstrong\u003e9\u003c/strong\u003e(4):357-359.\u003c/li\u003e\n\u003cli\u003eEdgar RC: \u003cstrong\u003eMUSCLE: multiple sequence alignment with high accuracy and high throughput\u003c/strong\u003e. \u003cem\u003eNucleic acids research \u003c/em\u003e2004, \u003cstrong\u003e32\u003c/strong\u003e(5):1792-1797.\u003c/li\u003e\n\u003cli\u003eTamura K, Stecher G, Kumar S: \u003cstrong\u003eMEGA11: Molecular Evolutionary Genetics Analysis Version 11\u003c/strong\u003e. \u003cem\u003eMolecular Biology and Evolution \u003c/em\u003e2021, \u003cstrong\u003e38\u003c/strong\u003e(7):3022-3027.\u003c/li\u003e\n\u003cli\u003eKatoh K, Misawa K, Kuma Ki, Miyata T: \u003cstrong\u003eMAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform\u003c/strong\u003e. \u003cem\u003eNucleic Acids Research \u003c/em\u003e2002, \u003cstrong\u003e30\u003c/strong\u003e(14):3059-3066.\u003c/li\u003e\n\u003cli\u003eSteinegger M, S\u0026ouml;ding J: \u003cstrong\u003eMMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets\u003c/strong\u003e. \u003cem\u003eNature Biotechnology \u003c/em\u003e2017, \u003cstrong\u003e35\u003c/strong\u003e(11):1026-1028.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eSnippy \u003c/strong\u003e[https://github.com/tseemann/snippy]\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Source tracking, molecular epidemiology, Salmonella enterica, Salmonella enterica Newport, pangenome, mobilome, plasmid","lastPublishedDoi":"10.21203/rs.3.rs-2166997/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2166997/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe \u003cem\u003eSalmonella enterica\u003c/em\u003e serovar Newport red onion outbreak of 2020 was the largest foodborne outbreak of \u003cem\u003eSalmonella\u003c/em\u003e in over a decade. The epidemiological investigation suggested two farms as the likely source of contamination. However, single nucleotide polymorphism (SNP) analysis of the whole genome sequencing data did not find any \u003cem\u003eSalmonella\u003c/em\u003e isolates from the farm regions that were closely related to the clinical isolates\u0026mdash;preventing the use of phylogenetics in source identification. Here, we explored an alternative method for analyzing the whole genome sequencing data driven by the hypothesis that if the outbreak strain had come from the farm regions, then the clinical isolates would disproportionately contain plasmids found in isolates from the farm regions due to recent horizontal transfer.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eSNP analysis confirmed that the clinical isolates formed a highly related clade with evidence for ancestry in California going back a decade. The clinical isolates not only had a large and highly conserved core genome (4,399 genes), but also 2,577 sparsely distributed accessory genes\u0026mdash;at least 64% of which were carried on plasmids. Amongst the clinical isolates and \u003cem\u003eSalmonella\u003c/em\u003e isolates from the farm regions were 2,187 and 503 putative plasmids, respectively. High similarity was observed between 17 plasmids from 8 farm isolates and 14 plasmids from 13 clinical isolates. Phylogenetic analysis suggested the highly similar plasmids shared a recent common ancestor and might have been transferred via intermediary species, but the seeming promiscuity of the plasmids prevented any conclusions about geographic location, isolation source, and time since transfer. Our sampling analysis suggested that observing a similar number and combination of highly similar plasmids in random samples of environmental \u003cem\u003eSalmonella enterica\u003c/em\u003e within NCBI Pathogen Detection database was unlikely, supporting a connection between the outbreak strain and the farms implicated by the epidemiological investigation.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eHorizontally transferred plasmids provided evidence for a connection between clinical isolates and the farms implicated as the source of the outbreak. Our case study suggests that such analyses might add a new dimension to source tracking investigations, but highlights the need for detailed and accurate metadata, more extensive environmental sampling, and a better understanding of plasmid molecular evolution.\u003c/p\u003e","manuscriptTitle":"Assessment of plasmids for relating the 2020 Salmonella enterica serovar Newport onion outbreak to farms implicated by the outbreak investigation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-11-02 17:35:26","doi":"10.21203/rs.3.rs-2166997/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2022-12-14T18:54:20+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-11-28T17:18:20+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"e15f1c11-c48a-4167-ab82-775c8d653b4b","date":"2022-11-08T03:56:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"2fb58051-d2b2-44ed-b6cb-4ba59e97831d","date":"2022-11-01T13:37:18+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-11-01T13:33:00+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-10-31T20:47:17+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2022-10-31T10:37:52+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-10-31T10:35:22+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Genomics","date":"2022-10-14T15:53:08+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a98eed33-6936-4042-acdd-4c92c784dd5a","owner":[],"postedDate":"November 2nd, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T20:28:32+00:00","versionOfRecord":{"articleIdentity":"rs-2166997","link":"https://doi.org/10.1186/s12864-023-09245-0","journal":{"identity":"bmc-genomics","isVorOnly":false,"title":"BMC Genomics"},"publishedOn":"2023-04-04 20:24:28","publishedOnDateReadable":"April 4th, 2023"},"versionCreatedAt":"2022-11-02 17:35:26","video":"","vorDoi":"10.1186/s12864-023-09245-0","vorDoiUrl":"https://doi.org/10.1186/s12864-023-09245-0","workflowStages":[]},"version":"v1","identity":"rs-2166997","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2166997","identity":"rs-2166997","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00