Genomic constraints shape the evolution of alternative routes to drug resistance in prokaryotes

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

ABSTRACT Antimicrobial resistance (AMR) is often modelled as the accumulation of resistance genes leading to multidrug resistance (MDR). We show that gene co-occurrence patterns in two opportunistic pathogens are consistent with fitness trade-offs that constrain which combinations of resistance mechanisms coexist. We applied a combined pangenomic and machine-learning analysis to 9,584 Escherichia coli genomes (99.2% phylogroup B2) and 7,057 Pseudomonas aeruginosa genomes. In E. coli, we identified eight cases of mutually exclusive gene pairs that independently predicted the same MDR phenotype, suggesting alternative routes to resistance whose components are typically not co-inherited. In a separate dataset of 352 strains with paired minimum inhibitory concentration (MIC) data, these dissociated combinations co-occurred more often in resistant than susceptible strains, consistent with the constraints being conditional on antibiotic selection. 33 gene pairs showed opposing association patterns between the two species, with combinations significantly associated in one species and significantly dissociated in the other (e.g. associated in E. coli and dissociated in P. aeruginosa, or vice versa). This indicates that genomic context modifies the contribution of individual genes to resistance phenotypes, and offers one explanation for the observation that 106 ARGs are present in >95% of strains yet do not predict resistance phenotype on their own. The findings are consistent with resistance evolution being shaped by fitness trade-offs and suggest that the dissociation patterns we identify could be targets for follow-up experimental work on resistance-associated fitness costs. IMPACT STATEMENT Multidrug resistance threatens global health; however, resistance acquisition remains poorly understood. We show that genomes cannot simply accumulate all available resistance mechanisms. Instead, species-specific grammatical rules render certain gene combinations typically incompatible. Analysing 9,584 Escherichia coli genomes, we identify eight mutually exclusive routes to resistance, demonstrating that resistance is more constrained than previously thought. Strikingly, 33 gene pairs that cooperate in E. coli actively disassociate in Pseudomonas aeruginosa , revealing that genomic context determines resistance outcomes. This explains why 106 resistance genes present in all strains fail to confer universal resistance, challenging current diagnostics. Our findings transform our understanding of resistance evolution from an unlimited accumulation model to a grammatically constrained pathways model, presenting new therapeutic strategies that exploit evolutionary dead ends.
Full text 74,862 characters · extracted from preprint-html · click to expand
Genomic constraints shape the evolution of alternative routes to drug resistance in prokaryotes | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Genomic constraints shape the evolution of alternative routes to drug resistance in prokaryotes View ORCID Profile Lucy Dillon , James O. McInerney , View ORCID Profile Christopher J. Creevey doi: https://doi.org/10.1101/2025.08.20.671315 Lucy Dillon 1 School of Biological Sciences, Queen’s University Belfast , Belfast, United Kingdom Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Lucy Dillon For correspondence: l.dillon{at}qub.ac.uk James O. McInerney 2 Department of Evolution, Ecology and Behaviour, Institute of Infection, Veterinary & Ecological Sciences, University of Liverpool , Liverpool, United Kingdom Find this author on Google Scholar Find this author on PubMed Search for this author on this site Christopher J. Creevey 1 School of Biological Sciences, Queen’s University Belfast , Belfast, United Kingdom Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Christopher J. Creevey Abstract Full Text Info/History Metrics Data/Code Preview PDF ABSTRACT Background Variation within the prokaryotic pangenome is not random, and natural selection that favours particular combinations of genes appears to dominate over random drift. What is less clear is whether similar phenotypes tend to arise independently within pangenomes, or whether there is a single route through which they evolve. We have previously shown that the wider genomic context is vital to understanding specific phenotypes, such as antimicrobial resistance. Considering the amount of genomic variation we see in many pathogens, the question of routes to resistance and how they arise is particularly urgent. Methods We hypothesised that mutually exclusive routes to multidrug antimicrobial resistance evolve through distinct yet predictable pathways dependent upon the species-specific pangenomic background. To test this, we developed an integrated pangenomic and machine learning framework to analyse the emergence of multidrug resistance (MDR) in two clinically significant pathogens, Pseudomonas aeruginosa and Escherichia coli . Machine learning models uncovered genes statistically linked to particular MDR phenotypes. We then examined whether these genes were significantly associated or dissociated with one another in gene co-occurrence networks and evaluated the importance of genomic context on pairs of genes. Results We demonstrate three key findings on the evolution of multidrug resistance in our dataset. First, antimicrobial resistance genes (ARGs) were distributed across both the core and accessory genomes, challenging the prevailing view that ARGs are typically confined to the accessory genome. Second, we demonstrate that E. coli possesses mutually exclusive pathways to MDR, with several resistance genes showing distinct dissociation patterns, suggesting alternative evolutionary routes to identical MDR phenotypes. Third, the network of ARG coassociations differed significantly between the P. aeruginosa and E. coli pangenomes with 33 gene pairs showing opposite association patterns between species, revealing species-specific genomic constraints on MDR emergence. These findings emphasise that bacterial evolution is more constrained than previously thought and that designing intervention therapies requires consideration of both historical evolutionary trajectories and the specific genomic context in which the resistance phenotypes arose. INTRODUCTION The relationship between phenotype and genotype is rarely as straightforward as the presence of a single gene determining the expression of a particular phenotype ( 1 ) and this complexity is particularly evident in antimicrobial resistance. Many phenotypes arise from the interplay of multiple genes, which construct complex mechanisms ( 2 ). Furthermore, gene expressions can be altered by conditions such as stress ( 3 ), for example, exposure to antibiotics, which can introduce various changes to gene expression regulation (e.g. regulation of porins) ( 4 ). Understanding these complex relationships is essential for predicting and combatting the evolution of multidrug resistance, yet we lack the fundamental knowledge of the constraints that shape resistance evolution. Complicating this further, prokaryotic genomes exhibit considerable variation within a species, which can be represented as a pangenome ( 5 ). Prokaryotic pangenome variation is not random but shows highly structured associations in gene content, likely driven by natural selection ( 6 ). For instance, the presence of any gene is often related to the presence of its functional partners ( 7 ), and the wider genomic content can determine the essentiality of a gene to cell viability ( 8 ). How these pangenomic fitness effects relate to an organism’s traits and phenotypes is still unclear, but it follows that the benefit of any gene to a particular phenotype may also be conditioned on the wider genetic environment. This means that the predictability of a trait or phenotype of an organism may depend not only on the presence or absence of specific genes but also on the wider genomic context in which they are found, for example, the link between AMR phenotype and ARG presence might not be straightforward. The complexity of these dependencies is key to understanding how AMR mechanisms are constructed, and this understanding, in turn, offers a possible avenue to combating resistance. The current trajectory of AMR predicts that there will be 10 million deaths annually as a result of AMR infections by the year 2050 ( 9 ). Therefore, urgency is required to understand how AMR arises and how this develops into multidrug resistance (MDR). Our previous work demonstrated that accurate prediction of AMR phenotypes involves more than simply identifying known antibiotic resistance genes (ARGs) and depends on understanding the complex interactions between known ARGs and the network of accessory non-canonical genes that potentiate or ameliorate resistance ( 2 ). Genes can be incompatible with each other for various reasons, for example, gene toxicity ( 10 ), issues with genome stability ( 11 , 12 ), functional redundancy ( 13 ), and changes in gene expression level ( 14 ). As an example, multiple plasmids often cause stability issues, such as I-complex plasmids ( 15 ). This raises the question of whether genetic incompatibilities and genomic context can also prohibit or constrain how traits and phenotypes can arise. In this study, we examined the available pangenome of two key opportunistic pathogens, P. aeruginosa and E. coli ( 16 , 17 ), to deepen our understanding of their relationship with AMR and MDR, and in particular, the evolutionary pathways through which they arrive at these phenotypes. While MDR is often thought to arise thanks to either MDR-specific genes, plasmids or the accumulation of multiple ARGs ( 18 , 19 ), we have previously shown that the broader genomic background is crucial for understanding which routes to resistance are available to which species ( 2 ). We tested two specific hypotheses using a combination of pangenomics and machine learning: first, that bacteria can evolve MDR through mutually exclusive evolutionary pathways; and second, that the contribution of individual resistance genes to AMR is contingent on their genomic context. Using the E. coli and P. aeruginosa pangenomes as model systems, we identified 106 unique ARGs across both species that were highly conserved (present in >95% of genomes). Importantly, many of these ARGs (40.1%) could target multiple drug classes, providing numerous potential routes to MDR. Supporting our first hypothesis, we discovered mutually exclusive evolutionary trajectories to resistance within E. coli , demonstrating that certain combinations of resistance mechanisms cannot coexist. In line with our second hypothesis, we found that specific ARGs produced opposing effects in E. coli versus P. aeruginosa , indicating that, in these cases, genomic background fundamentally alters the route to resistance. These findings demonstrate that the evolution of antimicrobial resistance is not completely deterministic, but rather can be contingent upon the broader genomic context in which resistance genes operate. This suggests that future efforts to combat antimicrobial drug resistance should broaden to encompass the entire genomic context. RESULTS MDR genes were found to be present in the core genome of both E. coli & P. aeruginosa Using the Roary ( 20 ) software pipeline, we built pangenome presence-absence matrices for 7,057 P. aeruginosa genomes and 9,584 E. coli genomes (see Methods for details). The pangenome analysis revealed that both species have an “open” pangenome ( Fig.1A , Fig.2A )) ( 21 ), consistent with previous studies ( 22 – 24 ). Download figure Open in new tab Figure 1. Summary of methods used in this study. Part A shows an overview of the pangenome analysis and annotations that were mapped back to the pangenome. Part B shows a brief description of how the machine learning models were built to predict AMR phenotypes of two antibiotics based on the MICs of the antibiotics which were then classified by EUCAST breakpoints. Part C shows a brief explanation of how significantly linked genes to a particular AMR phenotype were then linked to the pangenome Coinfinder outputs (further details are in Fig.5 ). Download figure Open in new tab Figure 2. Pangenome results and mapping antimicrobial resistance genes (ARGs) to the pangenome. A. Pangenome results of E. coli (i) and P. aeruginosa (ii). Both plots display the core genome (genes present in >99% of genomes), softcore genome (≥95% - <99%), shell genome (≥15% - <95%), and cloud (≥0%-<15%). B. The ARGs present in ≥95% of genomes of E. coli (i) and P. aeruginosa (ii). The ARGs are shown to correspond to particular drug classes (as documented in the CARD database). The colour bar represents the percentage of the genomes across the pangenome that the gene was found in, specifically highlighting only the softcore and core genomes. C. The ARG distribution of E. coli (i) and P. aeruginosa (ii) across the pangenome. The ARG names are not displayed due to the large number of ARGs detected, see Table S1 for full details. Similar to part B, the ARGs correspond to the drug classes, and the colour bar represents the percentage of the genomes across the pangenome that which the gene was found. First, we interrogated the pangenome matrices to understand whether ARGs are typically found in the accessory genome. While the majority of ARGs (414/466 in E. coli and 360/416 in P. aeruginosa ) were found in the accessory genome ( < 95% of genomes) ( Fig.2C ), we identified 106 unique defined ARGs that were present in the core or softcore, which we define as being at least 95% of genomes (56 unique genes for P. aeruginosa and 51 unique genes for E. coli ) ( Fig.2B )). Additionally, 43/106 of the unique ARGs found within the softcore/core genome of both species had MDR genotypes (targeted ≥3 drug classes). The ARGs found within the soft core or core genome consisted mainly of efflux pumps (94/106 unique ARGs) (Table S1). These efflux-related genes may be traditional housekeeping genes; however, the fact that we were able to identify ARGs within the core genome highlights that ARG identification tools should be used cautiously and not used to predict phenotypic resistance. ARGs can have alternative functions within the cell, in addition to being involved with drug resistance, for example efflux pumps transport heavy metals out of the cell ( 25 ). We have previously shown that ARG identification tools have limits in predicting AMR phenotypes, and more sophisticated techniques, such as machine learning, can provide higher accuracy in predicting AMR phenotypes ( 2 ). Interestingly, 11% of the ARGs found in the softcore and core genome had an alternative mechanism (to efflux) such as antibiotic inactivation or antibiotic target alteration (Table S1). Mutually exclusive routes to resistance identified in E. coli We constructed 55 decision tree models to predict MDR in E. coli using an indepdendent dataset of 352 genomes (deliberately withheld from the E. coli pangenome analysis) with corresponding minimum inhibitory concentration (MIC) values for 16 antibiotics (see Materials and Methods) ( Fig.1B ). The decision tree models provide insight into what genes are key to conferring resistance to antibiotic(s) and independently validate gene associations from the pangenome analysis (see Fig.S1 for an example decision tree). The decision trees had an average accuracy of 76.2% across all 55 models (Table S2) (Fig.S2). We define accuracy as the sum of true positives and true negatives divided by the total number of instances (Table S2). The average precision, recall and F1 scores were significantly lower in the decision tree models (52.7%, 45.4%, 54.7%, respectively) when compared to the accuracy values (Fig.S2). Note that precision, recall and F1 could not be calculated for several of the decision tree models due to division by zero (resulting in an NA value, see confusion matrix Table S2). To understand the routes to MDR in E. coli , we used our MDR decision tree models to identify gene families that were inferred to have significantly contributed to a specific AMR phenotype. We found 362 unique eggNOG gene families across all 55 decision tree models. Using these decision tree models, we systematically evaluated the contribution of each gene family to particular AMR phenotypes using a combination of Chi-squared tests, Fisher’s exact tests, and post hoc tests. We identified 326 unique gene families (from 519 total instances across the 55 decision trees) that showed significant associations with specific MDR phenotypes (i.e. susceptible to antibiotic A, susceptible to antibiotic B (SS); susceptible to antibiotic A, resistant to antibiotic B (SR); resistant to antibiotic A, susceptible to antibiotic B (RS); resistant to antibiotic A, resistant to antibiotic B (RR)) (Table S3). To identify mutually exclusive pathways to resistance, we combined decision tree analysis with Coinfinder ( 7 ) dissociation networks. We specifically tested whether gene pairs that strongly predicted the same resistance phenotype on our decision tree analysis were less likely to co-occur in bacterial genomes than chance expectation. The Coinfinder analysis independently confirmed the predictions from our association analysis for 317 of the gene pairs across the association or dissociation networks. Of these, we uncovered 35 pairs of genes phenotypically important to the same MDR phenotype exhibiting significant genomic disassociation ( Fig.3B , Table S3). This suggests that there are mutually exclusive evolutionary routes to equivalent resistance phenotypes. Download figure Open in new tab Figure 3. Heatmaps of genes that were statistically linked to a phenotype in the decision trees and then were found to be statistically associated or disassociated with another gene in the Coinfinder networks. Part A shows genes that were associated, and part B shows the disassociations. The phenotype combinations found in the E. coli Coinfinder networks are shown. All combinations of MDR phenotypes were observed in both the association and disassociation networks. For instance, we observed 20 gene families statistically linked to the SS phenotype that were significantly disassociated (in Coinfinder) with gene families statistically linked to the RR phenotype (see Table S3). For example, in the ceftriaxone-ciprofloxacin resistance network the gene family COG3316, which encodes a putative transposon, is significantly linked to dual resistance (RR phenotype) in our decision tree model. On the Coinfinder dissociation network, this gene family showed significant dissociation from the 2AV54 gene family, which itself strongly predicted dual susceptibility (SS phenotype). This pattern of mutual exclusivity between genes predicting opposing phenotypes provides a biological validation of our approach. We also found eight examples of putative mutual exclusion. For example, the two gene families COG3316 and COG3677 are both associated with the RR phenotype (MDR) in the ceftriaxone and ciprofloxacin decision tree model. Nonetheless, the gene families are found to be significantly disassociated in the Coinfinder disassociated network, which indicates that there are mutually exclusive routes to MDR phenotypes. The gene families COG3316 and COG3677 are putative transposases (identified using the eggNOG database and searching sequences in the BLAST NT database ( 26 )), this could suggest they are involved in horizontal gene transfer of resistance within E. coli . In the results in Fig.3 , we can see the different gene families statistically associated with particular phenotypes, from the decision tree models, being compared with the Coinfinder network in which the gene family pairs are found. These data illustrate how they interact. For example, if two genes are both linked to RR and are found to be in the Coinfinder disassociated network, this suggests mutually exclusive routes to resistance. However, if they are found in the associated network in Coinfinder, they could be part of the same route to resistance. Consequently, based on this understanding, we were able to identify 27 potential mutually exclusive routes and 121 gene pairs that may be contributing to the same AMR mechanism ( Fig.3 ). Nevertheless, we cannot explain all instances, such as genes associated with SS-linked genes, since it is difficult to explain associations between genes and the absence of a phenotype that could be beneficial to the organism. The gene families that had a statistical link to a phenotype within a decision tree model were analysed for a significant link to a specific country. This was to determine if this could be a possible cause of disassociation by evolution in isolation. We analysed the gene families that were labelled as ‘RR’ and disassociated with other gene families labelled as ‘RR’, and we were able to identify five gene family pairs that were significantly linked with a specific country. This resulted in 12 gene family pairs from the same geographic location being disassociated and 13 gene family pairs that had different geographic locations, meaning that geographic location is not likely to be involved in these genes being associated/disassociated (Table S4). These findings align more closely with genomic incompatibility, rather than evolution by isolation. We noticed that there are Coinfinder associations between genes linked to susceptible AMR phenotypes and genes statistically linked with resistant phenotypes ( Fig.3A ). We struggled to explain this biologically. The number of isolates used to build our decision trees was fairly low, meaning that the numbers in the exit nodes that were statistically evaluated to assign AMR phenotypes were also quite low. The decision tree models are also built using eggNOG gene families, which are broad gene groups, which could mean that there may be differences within the same eggNOG gene family. Another factor relating to the eggNOG gene families is that they do not take point mutations into consideration, which are documented to be involved in AMR mechanisms ( 27 ). These caveats could explain the unusual Coinfinder association between SS and RR linked genes. AMR phenotypes are dependent on genomic context In order to address whether the Coinfinder networks were stable across the two different pangenomes, we matched the statistically significantly associated gene family pairs found in E. coli to the Coinfinder networks of P. aeruginosa ( Fig.1C ). We found 69 matches (gene pair combinations) in the P. aeruginosa Coinfinder networks, however, contrary to our expectations, not all the gene family matches were in the same networks for the two species (i.e. associate and disassociate) ( Table 1 , Table S5). This suggests that similar mutually exclusive routes to resistance may exist in both P. aeruginosa and E. coli, and that these mutual exclusions may be general trends and not specific to species. However, there were also 33 genes (gene combinations) which had different networks in the two species ( Fig.4 , Table S5). Figure 4 clearly shows that genes which are associated in E. coli may be disassociated in P. aeruginosa . We searched these key gene families for ARGs and were able to identify 15 unique ARGs (Table S6) conferring resistance to 10 different drug classes. We detected these ARGs within three gene families, as seen in Figure 4 highlighted in purple. We can see that COG2602 (Beta-lactamase class D) is associated with COG0515 (Carbamoylphosphate synthase small subunit) and, COG0789 (DNA-binding transcriptional regulator, MerR family) is associated with COG1190 (Lysyl-tRNA synthetase, class II) in both P. aeruginosa and E. coli ( Fig.4A ). These examples imply that some mechanisms of resistance are conserved across different species (genomic environments). In contrast, COG2271 (a sugar phosphate permease) is associated with COG4388 (Mu-like prophage I protein) in P. aeruginosa , but they significantly avoid one another in E. coli . This indicates that some mechanisms are unique to particular species, are specific to genomic contexts, or are very plastic in the process leading to resistance. These associations allow us to examine how specific genes contribute to resistance phenotypes in different antibiotic combinations. Download figure Open in new tab Figure 4. Gene interactions in Coinfinder networks compared in E. coli and P. aeruginosa . A. All against all gene associations for P. aeruginosa and E. coli . These genes are statistically linked to MDR phenotypes in the decision tree models (e.g. SS, SR, RS, RR). The plot is split into two triangles split by the grey squares, the lower triangle shows P. aeruginosa genes interactions, and the upper triangle shows the E. coli gene interactions. The red boxes indicate genes which are associated in Coinfinder and the blue are disassociated genes in Coinfinder (in the respective Coinfinder species networks). The boxes with a point in the centre represent gene family pairs which share the same association in both species. We can see three gene families highlighted in purple, when annotated with RGI we identified these gene families to have ARGs present. Note - there are no instances of disassociated genes in both species. Parts B-D show genes involved in specific antibiotic models. View this table: View inline View popup Download powerpoint Table 1. A summary of gene matches from P. aeruginosa ’s Coinfinder network to E. coli ’s key genes. See Table S5 for full details. The model with the largest number of key genes involved in both E. coli and P. aeruginosa’s Coinfinder networks was the Amikacin-Cefepime model, with a total of seven genes. Notably, 10 of the gene pair combinations manifested contrasting patterns, being disassociated in E. coli but associated in P. aeruginosa ( Fig. 4B ). COG0859 an ADP-heptose:LPS heptosyltransferase is associated in P. aeruginosa for every gene combination available, yet in E. coli , three gene families are disassociated (COG3177 (a Fic family signalling protein), COG1051 (ADP-ribose pyrophosphatase), and COG0677 (UDP-N-acetyl-D-mannosaminuronate dehydrogenase)). COG0859 is associated with the phenotype SR (involved in resistance to cefepime, not amikacin). COG3177, COG0677, and COG1051 are all linked to an SS phenotype in E. coli . This is most likely why they are disassociated with a gene family that is involved in resistance to cefepime. In the cefepime-levofloxacin model, three genes are found in both the E. coli and P. aeruginosa Coinfinder networks. We identified that while COG0243 and COG4388 are associated in both species; COG4388 and COG2271 are associated in P. aeruginosa and disassociated in E. coli ( Fig. 4C ). Furthermore, we analysed the gene family sequences through the Resistance Gene Identifier (RGI) ( 28 ) and identified COG2271 to be linked to AMR, since we identified 15 unique ARG(s) which correspond to 10 drug classes (Table S7). In the eggNOG database, COG2271 is labelled as being involved in transmembrane transporter activity. Furthermore, the ARGs predicted within COG2271 are involved in efflux transporters ( SoxR ) ( 29 ) and importers within the cell ( GlpT ) ( 30 ) as defined in the CARD database online (see Table S6 for full ARG names). Unfortunately, the function is not available for COG4388, but the phenotype is labelled as SR and is the same for COG2271, suggesting it does have a role in resistance to levofloxacin but not cefepime. COG1263 is labelled in the eggNOG database to be involved in the phosphotransferase system (PTS), this is linked to being involved in AMR phenotype, specifically when the promoter was deleted it was linked to resistance in beta-lactams and aminoglycosides. Amikacin is an aminoglycoside, and cefepime is a beta-lactam (cephalosporin). COG1051 is labelled in the eggNOG database as being involved in GDP-mannose mannosyl hydrolase activity, this has been linked to L-fucose production ( 31 ) which has a role in antimicrobial activity. L-fucose is necessary for cell adhesion and biofilm formation ( 32 ), and both of these mechanisms are common in opportunistic infections ( 33 , 34 ). Nevertheless, the phenotype of COG1051 is SS. This is most likely disassociated with COG1063, which has an SR phenotype. However, in P. aeruginosa, these genes are associated, which may suggest that these genes have a role in a mechanism together, unlike in E. coli . In the ampicillin-levofloxacin model, there are three genes which are found in both E. coli’s and P. aeruginosa’s Coinfinder networks. We identified that while COG1203 and COG0673 are associated in both species, COG0449 and COG0673 are associated in P. aeruginosa and disassociated in E. coli ( Fig. 4D ). COG0449 and COG0673 are both associated with the RR phenotype in the ampicillin-levofloxacin decision tree model, however, these genes are disassociated in E. coli , despite this, these genes are associated in P. aeruginosa . COG0449 is involved in glutamine-fructose-6-phosphate transaminase (isomerizing) activity, which is part of the hexosamine pathway ( 34 ). This is a potential therapeutic target for antimicrobials ( 36 ), which suggests it does have a key role in AMR. COG0673 is thought to be involved in inositol 2-dehydrogenase activity. Inositol has been shown to influence microbial physiology, such as motility and extracellular matrix phenotypes ( 37 ). These features are known to increase the pathogenicity of bacteria ( 38 ) which could aid with antimicrobial treatment. DISCUSSION The integration of pangenomics with machine learning has revealed unexpected complexity in how bacteria evolve antimicrobial resistance. Our findings challenge the simplistic models of resistance acquisition and demonstrate that resistance evolution is constrained by species-specific genomic incompatibilities. Addressing this requires moving beyond simple gene cataloguing to understanding the wider interactions across the pangenome. High levels of diversity of MDR mechanisms exist within and between E. coli and P. aeruginosa In this study, we illustrate the complexity of resistance mechanisms both within E. coli and between E. coli and P. aeruginosa . This complexity challenges the “one gene, one resistance” paradigm that underlies many diagnostic and surveillance approaches ( 39 – 43 ). The presence of 106 ARGs in the core pangenome, with 56.8% targeting multiple drug classes, demonstrates that resistance cannot be reduced to simple presence/absense matrices. ( Fig. 2 ). If we assume a simplistic correlation between ARG presence and resistance phenotype, this would suggest that all 9,584 E. coli and 7,057 P. aeruginosa isolates have MDR phenotypes. Understanding this requires investigating genome diversity and the interactions between genes. Our pangenome analysis of E. coli and P. aeruginosa highlights the extraordinary plasticity of these species. The minimal core genomes, representing less than 1% of both species was notably small, even compared to previous studies on these organisms ( 22 , 44 – 47 ). This underscores that each isolate represents a unique resistance evolution experiment. Furthermore, the 317 significant gene associations and dissociations show that this plasticity is not random noise, but structured variation that follows rules imposed by functional constraints. While no two species may share the same pangenome, the evolutionary constraints they face are the same, resulting in independent genetic solutions to the same phenotypic problem. The wider genomic background determines MDR expression We have previously shown that there are often multiple routes to resistance to a single antibiotic, and often can be species-specific ( 2 ). Here we have shown that this extends to MDR phenotypes. Our integrated pangenome-machine learning approach identified that not all resistance mechanisms are compatible within E. coli ( Fig.3 ). Additionally, while certain genes exhibit a specific Coinfinder association in E. coli , the same genes may show the opposite association in P. aeruginosa , indicating that they do not share the same mechanism across species ( Fig.4 ). This highlights the critical role of genomic context in understanding AMR phenotypes. The importance of genomic context extends beyond individual gene function to entire resistance strategies. As recently emphasised for biosynthetic gene clusters ( 48 ), the genomic neighbourhood fundamentally shapes functional outcomes. Our finding that 33 gene pairs show opposite associations between E coli and P. aeruginosa demonstrates this principle, that the same genetic elements can promote resistance in one genomic context while being deleterious in another. This context-dependency means that resistance markers identified in one species cannot be directly applied to others and that species or even strain specific understanding of resistance drivers is required. The patterns of mutual exclusivity we observe suggest several possible mechanistic explanations. The incompatibility between COG3316 and COG3677 which are both putative transposases that confer MDR phenotypes, may arise from competition for cellular resources required for transposition, or from destabilising effects of multiple active transposable elements ( 10 – 12 ). Similarly, the 94 efflux-related genes we identified in the core genome likely face constraints on simultaneous expression due to membrane space limitations or energetic costs of maintaining multiple ATP-dependent transport systems. The striking difference in gene associations between E. coli and P. aeruginosa, particularly for genes like COG2602 (β-lactamase) and COG0515 (carbamoylphosphate synthase), suggests that regulatory architecture unique to each species determines which resistance combinations are viable. These mechanistic constraints, whether metabolic, regulatory, or structural, shape the evolutionary landscape of resistance and explain why bacteria cannot simply accumulate all available resistance mechanisms. Despite this complexity arising from genomic constraints rather than geographic isolation we found no significant effect of location, consistent with findings from previous studies ( 49 – 52 ). This may be due to the limitation that genomes can only be linked to a country rather than a specific ecological niche, which could play a more significant role in driving disassociation. Other studies have suggested that gene clusters of secondary metabolites are more closely associated with ecological functions, particularly in response to environmental stressors within specific habitats ( 13 , 53 – 58 ). Not all routes to MDR resistance are compatible Toxicity ( 10 ), genome stability ( 11 , 12 ), functional redundancy ( 13 ), and changes in gene expression level ( 14 ) can result from incompatibilities between a variety of genes ( 15 ). However, little is known about how the incompatibility of ARGs may influence MDR phenotypes. We identified 8 instances in E. coli where the genes forming a route to an MDR phenotype were mutually exclusive with genes that formed a different route to the same MDR phenotype. This suggests that different pathways to resistance may be mutually exclusive. These gene disassociations indicate that specific constraints and trade-offs shape the presence of certain genes, influencing how particular phenotypes emerge in specific organisms. The mutually exclusive routes may also be a result of functional redundancy if an isolate already has a resistance mechanism, it might be a genomic burden to have another resistance mechanism. An example of this from our data is the observed dissociation between COG3316 and COG3677 transposases, despite conferring MDR, may actively conflict with each other. These trade-offs represent evolutionary bottlenecks (for instance if bacteria cannot maintain all possible resistance mechanisms simultaneously) that could be exploited therapeutically by forcing bacteria towards less favourable evolutionary paths. The wider genomic context must be considered in MDR treatment Many hospitals incorporate whole genome sequencing in the diagnostic stages of diseases ( 59 , 60 ). For example, real-time polymerase chain reaction (PCR) can be used in the clinical setting to identify ARGs in patient samples ( 61 ). When using pathogen genomes to predict antibiotic resistance, the accuracy of these predictions depends on accounting for taxonomy and other genes that may influence the phenotype. Our findings have immediate implications for clinical practice. First, diagnostic platforms must move beyond binary resistance gene identification to incorporate genomic incompatibilities. Knowing that some gene combinations are not possible could improve prediction accuracy. Second, our identification of mutually exclusive pathways suggests that carefully designed sequential treatments could trap bacteria in evolutionary dead ends. Third, as resistance mechanisms can vary significantly within a species and may not be transferable across species, resistance interpretation must be tailored to the pathogen. This emphasises the complexity of resistance mechanisms and the challenges they pose for accurate resistance prediction. To improve AMR phenotype predictions, factors such as taxonomy and key accessory genes influencing resistance should be considered. The lack of applicability of “universal” resistance markers across species may explain cases of treatment failure despite “susceptible” genotypes and false resistance predictions despite ARG presence. Methodological innovation enables biological discovery The biological insights presented here were made possible by the convergence of two complementary approaches. Traditional pangenome analysis alone would have identified gene presence/absence patterns but not their phenotypic relevance. Machine learning alone would have predicted resistance but not the evolutionary constraints underlying them. By integrating Coinfinder’s unsupervised network analysis with supervised decision tree learning across 55 antibiotic combinations, we could both identify statistical associations and validate their phenotypic outcomes, revealing patterns missed by either approach alone. The 317 gene pairs identified by both methods demonstrate that novel biological insights can be observed when we move away from single analytical frameworks to approaches that capture both evolutionary and functional outcomes. Conclusion While we cannot establish causation without experimental validation and model improvements are possible, our findings provide a robust framework for understanding resistance evolution that can guide future experimental work. Indeed, our findings have identified 27 specific gene combinations that should be monitored in surveillance programmes, provide testable hypotheses for 121 gene pairs that may share resistance mechanisms and identify 317 unique gene pairs with significant associations or dissociation that should be considered in resistance prediction models for E. coli . Our work establishes a foundation for future critical research areas through experimental validation by knock-out studies of the identified mutually exclusive gene pairs, expansion of our framework to additional priority pathogens and development of the next generation of evolution-aware surveillance tools that account for genomic incompatibilities. Our computational pipelines ( https://github.com/LucyDillon/MDR_pangenome ) and underlying data and results are freely available to facilitate these efforts. Finally, our findings can inform clinical decision-making by reinforcing the importance of accounting for genomic background and mutually exclusive resistance mechanisms when designing MDR treatment strategies. Overall, the identification of species-specific mutually exclusive routes to multidrug resistance demonstrates that bacterial evolution is more constrained than previously thought. Understanding not just what genes confer resistance, but which combinations are evolutionarily forbidden, may prove key to predicting, preventing and ultimately controlling the evolution of multidrug resistance in bacterial pathogens. MATERIALS AND METHODS Within this study, we investigated AMR phenotype within a species ( E. coli ) and between species ( P. aeruginosa ) using the genomic context provided by pangenomics and machine learning models. Data availability To ensure full reproducibility of our pangenomic findings, we provide: (1) all genome assemblies analysed (accession lists at https://github.com/LucyDillon/MDR_pangenome/ ), (2) Roary pangenome matrices for both species and (3) Coinfinder association/dissociation gene pairs are available on OSF, (4) trained decision tree models for all 55 antibiotic combinations (.arff format), (5) RGI annotations mapped to pangenome positions (Table S1) ( https://osf.io/95k6z/?view_only=b46e8760a93b4e70a849fe4d2210165f ) , and (6) complete analysis pipeline including sourmash clustering parameters ( https://github.com/LucyDillon/MDR_pangenome ). View supplementary information at this link: https://osf.io/95k6z/?view_only=b46e8760a93b4e70a849fe4d2210165f CODE AVAILABILITY Please see https://github.com/LucyDillon/MDR_pangenome for all files and scripts used in this study. This includes scripts to make figures in R (see the “Figures” directory in the GitHub link). Apply decision tree software: https://github.com/ChrisCreevey/apply_decision_tree Data for Analysis To carry out the pangenome association/disassociation analysis, genomes were downloaded from BV-BRC ( 62 ), 37,451 E. coli and 7,109 P. aeruginosa , respectively. These genomes were selected from BV-BRC using the “complete” or “WGS” genomes and “good quality”, and all genomes were filtered to have less or equal to 500 contigs (to reduce the number of genomes and increase the quality). Genomes used in our previous models ( 2 ) were removed from the list (see E_coli_MDR_tree_genome_ids.txt). Using Sourmash (version 4.6.1) ( 63 ), we clustered the genomes to remove redundancy and misclassified species, using a Kmer size of 31. Since there was a large amount of E. coli genomes, a sub-cluster was chosen; these genomes were more genetically similar to each other and a more manageable number (9,584 genomes) (Fig.S3). The advantage of this approach is that a smaller dataset of closely related taxonomy for the pangenome will be more appropriate, resulting in a larger core genome to analyse. This resulted in 9,584 E. coli and 7,057 P. aeruginosa genomes. We chose both species due to their clinical relevance. To create MDR machine learning models of E. coli, we used 350 E. coli genomes with corresponding minimum inhibitory concentration (MIC) values to multiple antibiotics from our previous dataset, sourced from BV-BRC ( 2 ), to build the MDR models. The MIC values were matched to European Committee on Antimicrobial Susceptibility Testing (EUCAST) breakpoints ( 64 ) to determine if isolates were resistant or susceptible. Isolates with intermediate phenotypes were removed from the dataset. The genome IDs for the E. coli genomes can be found in E_coli_MDR_tree_genome_ids.txt. Pangenome analysis The genomes were annotated with Prokka (v1.14.5) with default parameters ( 65 ). Roary (version 3.13.0) ( 20 ) was used to create the pangenome and create the core genome alignment (Roary.sh). To create a Newick file (necessary for Coinfinder input), FastTree (v2.1.11) ( 66 ) was used with the core gene alignment and an R script remove_zero_branch_length.R was used to edit the tree to remove any zero branch lengths in the tree. The Roary output file: gene_presence_absence.csv file was edited to remove any spaces from the index of the file before using Coinfinder (see remove_spaces_help.md for details). Coinfinder (v1.2.1) was then used to find significantly associated and disassociated co-occurring genes across the pangenome using default settings ( 7 ). We analysed both associate and disassociate analyses separately on our data (Coinfinder_associate.sh, Coinfinder_disassociate.sh). Gephi was used to visualise the network.gexf files using the Fruchterman Reingold layout ( 67 ). Functional and ARGs annotation of the pangenome eggNOG mapper (v2.1.6) was used with default parameters to functionally annotate all the genomes ( 68 ) (see Pseudomonas_eggnog.sh and E_coli_eggnog.sh). To match the eggNOG mapper results to the pangenome, the IDs generated using Prokka can be used to link the Roary output (gene presence absence.csv) to the eggNOG file (pipeline commands: eggNOG in pangenome.md). As can be seen in the scripts, once the ID is matched, the number of isolates the gene is present in (in column 4 of gene presence absence.csv) can be returned. This can be matched to the eggNOG COG (gene family). By using the number of isolates, we can determine if the gene is present in the core, soft core, shell or cloud portion of the pangenome. Resistance Gene Identifier (RGI) (v5.1.1) ( 69 ) was used with mainly default parameters to identify ARGs across the pangenome, which uses the CARD database (v3.1.1) ( 70 ) (see E_coli_RGI.sh and Pseudomonas_RGI.sh). We used a ‘strict’ identification cut-off when running RGI to ensure we reduced the chance of identifying false positives. We also analysed the most common sequence of the protein ORF matched to the RGI gene (column 19 in RGI output file) for ARGs present in >95% of genomes using BLASTp to ensure the sequences were not matching to similar genes that may have a more generic housekeeping function. The ARGs were matched to the pangenome by following the commands RGI_AMR_analysis.md file. In summary, all unique ARGs were identified across all the isolates, then each unique ARG was searched for the number of times it appeared in the isolates (only counting the first copy in the event that multiple copies of the same ARG were present in the same isolate), and then a percentage across all isolates was calculated. MDR decision tree models To investigate key genes involved in multi-drug resistance (MDR), additional data from BV-BRC were downloaded to build machine-learning models, consisting of 352 E. coli genomes with corresponding MIC values to at least two different antibiotics. The genome IDs can be found in E coli genomes MDR models.csv. J48 decision tree models were created using E. coli genomes in WEKA ( 71 ) for 55 pairwise comparisons between 16 antibiotics (Table S2). Even though there were a possible 240 pairwise comparisons, data only existed for 55. These models represent how MDR arises (we used pairs of antibiotics; therefore, the models can only predict MDR for two antibiotics). The genomes were categorised into “Susceptible” and “Resistant” for individual antibiotics using EUCAST breakpoints ( 64 ). The model can predict various phenotype combinations for different antibiotics. For example, a pairwise comparison between two antibiotics could have four different combinations: SS, SR, RS, RR (where S: Susceptible and R: Resistant; and the first letter would correspond to the phenotype of the antibiotic A and the second letter of each pair would correspond to the phenotype to antibiotic B). The data input for WEKA is an Attribute-Relation File Format (ARFF) file. The attributes (i.e., all eggNOG gene families and AMR phenotype) are listed in the ARFF file. The data follows the attributes in the file and is in the format of the copy number of the gene family or absence (0), and the MDR phenotype (either SS, SR, RS, or RR). Each training file had genomes across all four phenotypes, so the model could predict all four phenotypes. A custom Python script Make_MDR_arff_files.py was used to make the training data files to build the MDR decision tree models. The models were trained using the default parameters and to ensure robustness, we performed multiple validation steps: (1) 10-fold cross validation for all machine-learning models, (2) permutation testing (n-1000) to confirm statistical significance and (3) independent validation using the held-out dataset of 352 genomes described earlier and which were not included in the intital pangenome analysis. We calculated the precision, recall, and F1 scores for susceptible and resistant predictions separately and averaged the values. This ensured the entire confusion matrix was considered when we evaluated the models’ statistics. Combining pangenomic and decision tree results We investigated the genes involved in particular MDR phenotypes within the decision trees that are significantly associated or disassociated with each other. To achieve this, we searched for each gene family involved in the 55 MDR decision trees in the Coinfinder output files: “*pairs.tsv” for evidence of association and disassociation. This provided a list of gene families involved in significant associations or disassociation between at least one gene family in the MDR decision trees. To focus this list on significant occurrences between gene families only in the decision tree models, we combined both reduced *pairs.tsv files for associated and disassociated gene families and opened in Cytoscape ( 72 ). Gene families in the respective models were labelled as “Tree” and “NA” if not in a model. Then, only gene families in the decision tree models were selected. To identify which AMR phenotype (‘SS’, ‘SR’, ‘RS’, and ‘RR’) the gene family was most associated with, a Chi-squared test was applied to the matrix of positive and negative influences the gene had on different phenotypes (please see schematic Fig.5 for reference). Our decision tree models are always binary, representing branch points distinguishing between the phenotypic consequences of having more or fewer of a gene, we counted how many genomes in our models were found to be resistant or susceptible on either side of the branch point associated with each gene in the decision tree. This gave an overview of the “positive” phenotypic associations with a gene (i.e. the phenotypic consequences of having more of this gene) and or the “negative” phenotypic associations with a gene (i.e. the phenotypic consequences of having fewer of this gene). For example, if the presence of a gene is always associated with resistance, then we should observe in the branch associated with having more of this gene, a larger number of genomes with the resistance phenotype than the susceptible phenotype. Meanwhile, in the branch associated with having fewer (or zero) copies of the gene, we should see no particular association with resistant or susceptible genomes, and its absence cannot be selected against. This would leave a tell-tale difference in the ratios of S:R genomes in the negative branch compared to the positive branch, which could be identified with a simple Chi-squared, Fisher’s exact test, and a post hoc test (to confirm the phenotype with which the gene was associated) were performed on each gene matrix for each gene family in each model. This also holds for genes whose presence is associated with susceptibility, allowing the identification of genes important to each phenotype. In the case of where there are 4 phenotypes (SS, SR, RS and RR), the same principle holds, and a post hoc test can identify which of the phenotypes is most associated with the more (positive) or fewer (negative) copies of the gene. See Figure 5 for a detailed example. The matrix files were calculated using apply decision tree (see code availability). We then merged the positive and negative interaction matrices into one file using a custom Python script merge_presence_absence_matrix.py (Table S8). The overall matrix file was then read into R to perform the stats (this can be found in Gene_phenotype_association.R). Download figure Open in new tab Figure 5. Schematic representation of how gene families within the decision trees are mapped to the Coinfinder networks. The top left shows a representative decision tree, the gene families are analysed in a matrix by assessing how many genomes it is associated with (top right). The matrix for each gene family within each model is then analysed by using a Chi-squared and Fisher’s exact test. Gene families which were found to be significantly associated with a particular AMR phenotype within the decision tree models were then searched in the Coinfinder associated/disassociated networks. The gene families were searched as a pair from one MDR decision tree model in the CoinFinder associated/disassociated networks. See Table S3 for gene families matched to the Coinfinder. Then, we selected pairs of gene families which were found to be significantly associated (P value = < 0.05) with a particular MDR phenotype and compared them to the Coinfinder significantly associated gene network and disassociated gene networks. Lists of gene families present in each decision tree model were created. Then, using a custom Python script (match_DT_to_coinfinder.py), we extracted gene family matches in the Coinfinder network (in which the source and target nodes in the Coinfinder network corresponded to gene families in one decision tree). The results of gene families were manually searched and compared to the decision trees and statistical results. We analysed whether gene families were significantly linked to specific geographic locations to determine if geographic isolation could be a potential reason why gene families linked to a specific phenotype in the decision trees may be disassociated with another gene of the same phenotype. We were able to capture the geographic information from the isolates from the metadata available in BV-BRC. Then we performed a Chi-squared and Fisher’s exact test to determine if there was a significant link between geographic isolation and disassociation (see geo_stats.py for details Fig.S4-S5). We then searched for the significantly associated/disassociated gene pairs that we found in E. coli in the P. aeruginosa Coinfinder networks. AUTHOR CONTRIBUTIONS L Dillon: conceptualization, software, formal analysis, investigation, visualization, methodology, and writing—original draft, review, and editing. JO McInerney: conceptualization, methodology, and writing, review, and editing. CJ Creevey: conceptualization, data curation, software, supervision, investigation, visualization, methodology, and writing—original draft, review, and editing. ACKNOWLEDGEMENTS We acknowledge PhD funding from the Department for the Economy, Northern Ireland, to L Dillon. CJ Creevey wishes to acknowledge funding from UKRI MR/Y015223/1 and EU via Horizon 2020 (818368 MASTER and 101000213 Holoruminant). Funder Information Declared We acknowledge PhD funding from the Department for the Economy, Northern Ireland, to L Dillon. EU via Horizon , 818368 , 101000213 UKRI , MR/Y015223/1 Footnotes https://osf.io/95k6z/?view_only=b46e8760a93b4e70a849fe4d2210165f https://github.com/LucyDillon/MDR_pangenome REFERENCES 1. ↵ Orgogozo V , Morizot B , Martin A. The differential view of genotype–phenotype relationships . Front Genet [Internet] . 2015 May 19 [cited 2025 Feb 18]; 6 . Available from: https://www.frontiersin.org/journals/genetics/articles/10.3389/fgene.2015.00179/full 2. ↵ Dillon L , Dimonaco NJ , Creevey CJ . Accessory genes define species-specific routes to antibiotic resistance . Life Sci Alliance [Internet] . 2024 Apr 1 [cited 2024 Jan 16]; 7 ( 4 ). Available from: https://www.life-science-alliance.org/content/7/4/e202302420 3. ↵ Opijnen T van , Camilli A. A fine scale phenotype–genotype virulence map of a bacterial pathogen . Genome Res . 2012 Jan 12; 22 ( 12 ): 2541 – 51 . OpenUrl Abstract / FREE Full Text 4. ↵ Dam S , Pagès JM , Masi M . Stress responses, outer membrane permeability control and antimicrobial resistance in Enterobacteriaceae . Microbiology . 2018 ; 164 ( 3 ): 260 – 7 . OpenUrl CrossRef PubMed 5. ↵ Robinson LA , Collins ACZ , Murphy RA , Davies JC , Allsopp LP. Diversity and prevalence of type VI secretion system effectors in clinical Pseudomonas aeruginosa isolates . Front Microbiol [Internet] . 2023 [cited 2023 Feb 7]; 13 . Available from: https://www.frontiersin.org/articles/10.3389/fmicb.2022.1042505 6. ↵ Beavan A , Domingo-Sananes MR , McInerney J. Contingency, Repeatability and Predictability in the Evolution of a Prokaryotic Pangenome. [Internet] . bioRxiv ; 2023 [cited 2023 Mar 22]. p. 2023.03.20.533463 . Available from: https://www.biorxiv.org/content/10.1101/2023.03.20.533463v1 7. ↵ Whelan FJ , Rusilowicz M , McInerney JO . Coinfinder: detecting significant associations and dissociations in pangenomes . Microb Genomics . 2020 Feb 25; 6 ( 3 ): e000338 . OpenUrl 8. ↵ Rosconi F , Rudmann E , Li J , Surujon D , Anthony J , Frank M , et al. A bacterial pan-genome makes gene essentiality strain-dependent and evolvable . Nat Microbiol . 2022 Oct; 7 ( 10 ): 1580 – 92 . OpenUrl PubMed 9. ↵ Naghavi M , Vollset SE , Ikuta KS , Swetschinski LR , Gray AP , Wool EE , et al. Global burden of bacterial antimicrobial resistance 1990–2021: a systematic analysis with forecasts to 2050 . The Lancet . 2024 Sept 28; 404 ( 10459 ): 1199 – 226 . OpenUrl 10. ↵ Sorek R , Zhu Y , Creevey CJ , Francino MP , Bork P , Rubin EM . Genome-Wide Experimental Determination of Barriers to Horizontal Gene Transfer . Science . 2007 Nov 30; 318 ( 5855 ): 1449 – 52 . OpenUrl Abstract / FREE Full Text 11. ↵ Ammar AM , Abd El-Aziz NK , Aggour MG , Ahmad AAM , Abdelkhalek A , Muselin F , et al. A Newly Incompatibility F Replicon Allele (FIB81) in Extensively Drug-Resistant Escherichia coli Isolated from Diseased Broilers . Int J Mol Sci . 2024 Jan; 25 ( 15 ): 8347 . OpenUrl PubMed 12. ↵ Uhlin BE , Nordström K . Plasmid incompatibility and control of replication: copy mutants of the R-factor R1 in Escherichia coli K-12 . J Bacteriol . 1975 Nov; 124 ( 2 ): 641 – 9 . OpenUrl Abstract / FREE Full Text 13. ↵ Bruns H , Crüsemann M , Letzel AC , Alanjary M , McInerney JO , Jensen PR , et al. Function-related replacement of bacterial siderophore pathways . ISME J . 2018 Feb 1; 12 ( 2 ): 320 – 9 . OpenUrl CrossRef PubMed 14. ↵ Sloan DB , Warren JM , Williams AM , Kuster SA , Forsythe ES . Incompatibility and Interchangeability in Molecular Evolution . Genome Biol Evol . 2022 Dec 30; 15 ( 1 ): evac184 . OpenUrl 15. ↵ Rozwandowicz M , Hordijk J , Bossers A , Zomer AL , Wagenaar JA , Mevius DJ , et al. Incompatibility and phylogenetic relationship of I-complex plasmids . Plasmid . 2020 May 1; 109 : 102502 . OpenUrl PubMed 16. ↵ Qin S , Xiao W , Zhou C , Pu Q , Deng X , Lan L , et al. Pseudomonas aeruginosa: pathogenesis, virulence factors, antibiotic resistance, interaction with host, technology advances and emerging therapeutics . Signal Transduct Target Ther . 2022 June 25; 7 ( 1 ): 1 – 27 . OpenUrl CrossRef PubMed 17. ↵ Denamur E , Clermont O , Bonacorsi S , Gordon D . The population genetics of pathogenic Escherichia coli . Nat Rev Microbiol . 2021 Jan; 19 ( 1 ): 37 – 54 . OpenUrl CrossRef PubMed 18. ↵ Mehrotra T , Konar D , Pragasam AK , Kumar S , Jana P , Babele P , et al. Antimicrobial resistance heterogeneity among multidrug-resistant Gram-negative pathogens: Phenotypic, genotypic, and proteomic analysis . Proc Natl Acad Sci . 2023 Aug 15; 120 ( 33 ): e2305465120 . OpenUrl PubMed 19. ↵ Shankar C , Jacob JJ , Vasudevan K , Biswas R , Manesh A , Sethuvel DPM , et al. Emergence of Multidrug Resistant Hypervirulent ST23 Klebsiella pneumoniae: Multidrug Resistant Plasmid Acquisition Drives Evolution . Front Cell Infect Microbiol [Internet] . 2020 [cited 2024 Feb 14]; 10 . Available from: https://www.frontiersin.org/articles/10.3389/fcimb.2020.575289 20. ↵ Page AJ , Cummins CA , Hunt M , Wong VK , Reuter S , Holden MTG , et al. Roary: rapid large-scale prokaryote pan genome analysis . Bioinformatics . 2015 Nov 15; 31 ( 22 ): 3691 – 3 . OpenUrl CrossRef PubMed 21. ↵ McInerney JO , McNally A , O’Connell MJ . Why prokaryotes have pangenomes . Nat Microbiol . 2017 Apr; 2 ( 4 ): 17040 . OpenUrl PubMed 22. ↵ Hall RJ , Whelan FJ , Cummins EA , Connor C , McNally A , McInerney JO . Gene-gene relationships in an Escherichia coli accessory genome are linked to function and mobility . Microb Genomics . 2021 ; 7 ( 9 ): 000650 . OpenUrl 23. Rasko DA , Rosovitz MJ , Myers GSA , Mongodin EF , Fricke WF , Gajer P , et al. The Pangenome Structure of Escherichia coli: Comparative Genomic Analysis of E. coli Commensal and Pathogenic Isolates . J Bacteriol . 2008 Oct 15; 190 ( 20 ): 6881 – 93 . OpenUrl Abstract / FREE Full Text 24. ↵ Valot B , Guyeux C , Rolland JY , Mazouzi K , Bertrand X , Hocquet D . What It Takes to Be a Pseudomonas aeruginosa? The Core Genome of the Opportunistic Pathogen Updated . PLoS ONE . 2015 May 11; 10 ( 5 ): e0126468 . OpenUrl CrossRef PubMed 25. ↵ Lekshmi M , Ortiz-Alegria A , Kumar S , Varela MF . Major facilitator superfamily efflux pumps in human pathogens: Role in multidrug resistance and beyond . Curr Res Microb Sci . 2024 Jan 1; 7 : 100248 . OpenUrl PubMed 26. ↵ Altschul SF , Gish W , Miller W , Myers EW , Lipman DJ . Basic local alignment search tool . J Mol Biol . 1990 Oct 5; 215 ( 3 ): 403 – 10 . OpenUrl CrossRef PubMed Web of Science 27. ↵ Zarske M , Werckenthin C , Golz JC , Stingl K . The point mutation A1387G in the 16S rRNA gene confers aminoglycoside resistance in Campylobacter jejuni and Campylobacter coli . Antimicrob Agents Chemother . 2024 Oct 15; 68 ( 11 ): e00833 – 24 . OpenUrl PubMed 28. ↵ Alcock BP , Huynh W , Chalil R , Smith KW , Raphenya AR , Wlodarski MA , et al. CARD 2023: expanded curation, support for machine learning, and resistome prediction at the Comprehensive Antibiotic Resistance Database . Nucleic Acids Res . 2023 Jan 6; 51 ( D1 ): D690 – 9 . OpenUrl CrossRef PubMed 29. ↵ The Comprehensive Antibiotic Resistance Database soxR [Internet] . [cited 2024 Nov 12]. Available from: https://card.mcmaster.ca/ontology/39965 30. ↵ The Comprehensive Antibiotic Resistance Database-GlpT [Internet] . [cited 2024 Nov 12]. Available from: https://card.mcmaster.ca/ontology/40588 31. ↵ Fu C , Xu X , Xie Y , Liu Y , Liu M , Chen A , et al. Rational design of GDP-d-mannose mannosyl hydrolase for microbial l-fucose production . Microb Cell Factories . 2023 Mar 24; 22 ( 1 ): 56 . OpenUrl 32. ↵ Spaulding CN , Klein RD , Schreiber HL , Janetka JW , Hultgren SJ . Precision antimicrobial therapeutics: the path of least resistance? Npj Biofilms Microbiomes . 2018 Feb 27; 4 ( 1 ): 1 – 7 . OpenUrl PubMed 33. ↵ Flores-Mireles AL , Walker JN , Caparon M , Hultgren SJ . Urinary tract infections: epidemiology, mechanisms of infection and treatment options . Nat Rev Microbiol . 2015 May; 13 ( 5 ): 269 – 84 . OpenUrl CrossRef PubMed 34. ↵ Ciofu O , Moser C , Jensen PØ , Høiby N . Tolerance and resistance of microbial biofilms . Nat Rev Microbiol . 2022 Oct; 20 ( 10 ): 621 – 35 . OpenUrl CrossRef PubMed 35. GFPT1 - Glutamine--fructose-6-phosphate aminotransferase [isomerizing] 1 - Homo sapiens (Human) | UniProtKB | UniProt [Internet] . [cited 2024 Oct 29]. Available from: https://www.uniprot.org/uniprotkb/Q06210/entry 36. ↵ Soni V , Rosenn EH , Venkataraman R . Insights into the central role of N-acetyl-glucosamine-1-phosphate uridyltransferase (GlmU) in peptidoglycan metabolism and its potential as a therapeutic target . Biochem J . 2023 July 27; 480 ( 15 ): 1147 – 64 . OpenUrl PubMed 37. ↵ O’Banion BS , Jones P , Demetros AA , Kelley BR , Knoor LH , Wagner AS , et al. Plant myo -inositol transport influences bacterial colonization phenotypes . Curr Biol . 2023 Aug 7; 33 ( 15 ): 3111 – 3124 .e5. OpenUrl CrossRef PubMed 38. ↵ Rajput A , Chauhan SM , Mohite OS , Hyun JC , Ardalani O , Jahn LJ , et al. Pangenome analysis reveals the genetic basis for taxonomic classification of the Lactobacillaceae family . Food Microbiol . 2023 Oct; 115 : 104334 . OpenUrl CrossRef PubMed 39. ↵ Zaheer R , Noyes N , Ortega Polo R , Cook SR , Marinier E , Van Domselaar G , et al. Impact of sequencing depth on the characterization of the microbiome and resistome . Sci Rep . 2018 Apr 12; 8 ( 1 ): 5890 . OpenUrl CrossRef PubMed 40. Jankowski P , Gan J , Le T , McKennitt M , Garcia A , Yanaç K , et al. Metagenomic community composition and resistome analysis in a full-scale cold climate wastewater treatment plant . Environ Microbiome . 2022 Jan 15; 17 : 3 . OpenUrl PubMed 41. Tan S , Dvorak CMT , Estrada AA , Gebhart C , Marthaler DG , Murtaugh MP . MinION sequencing of Streptococcus suis allows for functional characterization of bacteria by multilocus sequence typing and antimicrobial resistance profiling . J Microbiol Methods . 2020 Feb 1; 169 : 105817 . OpenUrl CrossRef PubMed 42. Bortolaia V , Kaas RS , Ruppe E , Roberts MC , Schwarz S , Cattoir V , et al. ResFinder 4.0 for predictions of phenotypes from genotypes . J Antimicrob Chemother . 2020 Dec 1; 75 ( 12 ): 3491 – 500 . OpenUrl CrossRef PubMed 43. ↵ Verschuuren T , Bosch T , Mascaro V , Willems R , Kluytmans J . External validation of WGS-based antimicrobial susceptibility prediction tools, KOVER-AMR and ResFinder 4.1, for Escherichia coli clinical isolates . Clin Microbiol Infect . 2022 Nov 1; 28 ( 11 ): 1465 – 70 . OpenUrl PubMed 44. ↵ Freschi L , Vincent AT , Jeukens J , Emond-Rheault JG , Kukavica-Ibrulj I , Dupont MJ , et al. The Pseudomonas aeruginosa Pan-Genome Provides New Insights on Its Population Structure, Horizontal Gene Transfer, and Pathogenicity . Genome Biol Evol . 2019 Jan 1; 11 ( 1 ): 109 – 20 . OpenUrl CrossRef PubMed 45. Wang C , Ye Q , Jiang A , Zhang J , Shang Y , Li F , et al. Pseudomonas aeruginosa Detection Using Conventional PCR and Quantitative Real-Time PCR Based on Species-Specific Novel Gene Targets Identified by Pangenome Analysis . Front Microbiol [Internet] . 2022 [cited 2023 Feb 7]; 13 . Available from: https://www.frontiersin.org/articles/10.3389/fmicb.2022.820431 46. Mosquera-Rendón J , Rada-Bravo AM , Cárdenas-Brito S , Corredor M , Restrepo-Pineda E , Benítez-Páez A . Pangenome-wide and molecular evolution analyses of the Pseudomonas aeruginosa species . BMC Genomics . 2016 Jan 12; 17 ( 1 ): 45 . OpenUrl CrossRef PubMed 47. ↵ Subedi D , Vijay AK , Kohli GS , Rice SA , Willcox M . Comparative genomics of clinical strains of Pseudomonas aeruginosa strains isolated from different geographic sites . Sci Rep . 2018 Oct 23; 8 ( 1 ): 15668 . OpenUrl PubMed 48. ↵ Salamzade R , Kalan LR . Context matters: assessing the impacts of genomic background and ecology on microbial biosynthetic gene cluster evolution . mSystems . 2025 Feb 24; 10 ( 3 ): e01538 – 24 . OpenUrl PubMed 49. ↵ Green JL , Bohannan BJM , Whitaker RJ . Microbial biogeography: from taxonomy to traits . Science . 2008 May 23; 320 ( 5879 ): 1039 – 43 . OpenUrl Abstract / FREE Full Text 50. Martiny JBH , Bohannan BJM , Brown JH , Colwell RK , Fuhrman JA , Green JL , et al. Microbial biogeography: putting microorganisms on the map . Nat Rev Microbiol . 2006 Feb; 4 ( 2 ): 102 – 12 . OpenUrl CrossRef PubMed Web of Science 51. Salamzade R , Manson AL , Walker BJ , Brennan-Krohn T , Worby CJ , Ma P , et al. Inter-species geographic signatures for tracing horizontal gene transfer and long-term persistence of carbapenem resistance . Genome Med . 2022 Apr 5; 14 : 37 . OpenUrl CrossRef PubMed 52. ↵ Chase AB , Bogdanov A , Demko AM , Jensen PR . Biogeographic patterns of biosynthetic potential and specialized metabolites in marine sediments . ISME J . 2023 July; 17 ( 7 ): 976 – 83 . OpenUrl PubMed 53. ↵ van Bergeijk DA , Terlouw BR , Medema MH , van Wezel GP . Ecology and genomics of Actinobacteria: new concepts for natural product discovery . Nat Rev Microbiol . 2020 Oct; 18 ( 10 ): 546 – 58 . OpenUrl CrossRef PubMed 54. Cordero OX , Polz MF . Explaining microbial genomic diversity in light of evolutionary ecology . Nat Rev Microbiol . 2014 Apr; 12 ( 4 ): 263 – 73 . OpenUrl CrossRef PubMed 55. Crits-Christoph A , Bhattacharya N , Olm MR , Song YS , Banfield JF . Transporter genes in biosynthetic gene clusters predict metabolite characteristics and siderophore activity . Genome Res . 2021 Feb; 31 ( 2 ): 239 – 50 . OpenUrl Abstract / FREE Full Text 56. Zipperer A , Konnerth MC , Laux C , Berscheid A , Janek D , Weidenmaier C , et al. Human commensals producing a novel antibiotic impair pathogen colonization . Nature . 2016 July 28; 535 ( 7613 ): 511 – 6 . OpenUrl CrossRef PubMed 57. Bosak T , Losick RM , Pearson A . A polycyclic terpenoid that alleviates oxidative stress . Proc Natl Acad Sci . 2008 May 6; 105 ( 18 ): 6725 – 9 . OpenUrl Abstract / FREE Full Text 58. ↵ Chevrette MG , Thomas CS , Hurley A , Rosario-Meléndez N , Sankaran K , Tu Y , et al. Microbiome composition modulates secondary metabolism in a multispecies bacterial community . Proc Natl Acad Sci U S A . 2022 Oct 18; 119 ( 42 ): e2212930119 . OpenUrl CrossRef PubMed 59. ↵ Hilt EE , Ferrieri P . Next Generation and Other Sequencing Technologies in Diagnostic Microbiology and Infectious Diseases . Genes . 2022 Sept; 13 ( 9 ): 1566 . OpenUrl CrossRef 60. ↵ Kumar P , Sundermann AJ , Martin EM , Snyder GM , Marsh JW , Harrison LH , et al. Method for Economic Evaluation of Bacterial Whole Genome Sequencing Surveillance Compared to Standard of Care in Detecting Hospital Outbreaks . Clin Infect Dis . 2021 July 1; 73 ( 1 ): e9 – 18 . OpenUrl CrossRef PubMed 61. ↵ Cason C , D’Accolti M , Soffritti I , Mazzacane S , Comar M , Caselli E. Next-generation sequencing and PCR technologies in monitoring the hospital microbiome and its drug resistance . Front Microbiol [Internet] . 2022 July 28 [cited 2025 Mar 19]; 13 . Available from: https://www.frontiersin.org/journals/microbiology/articles/10.3389/fmicb.2022.969863/full 62. ↵ Olson RD , Assaf R , Brettin T , Conrad N , Cucinell C , Davis JJ , et al. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR . Nucleic Acids Res . 2022 Nov 9; gkac1003 . 63. ↵ Brown CT , Irber L. sourmash: a library for MinHash sketching of DNA . J Open Source Softw . 2016 Sept 14; 1 ( 5 ): 27 . OpenUrl CrossRef 64. ↵ eucast: Clinical breakpoints and dosing of antibiotics [Internet] . [cited 2022 Oct 16]. Available from: https://www.eucast.org/clinical_breakpoints 65. ↵ Seemann T . Prokka: rapid prokaryotic genome annotation . Bioinformatics . 2014 July 15; 30 ( 14 ): 2068 – 9 . OpenUrl CrossRef PubMed Web of Science 66. ↵ Price MN , Dehal PS , Arkin AP . FastTree: Computing Large Minimum Evolution Trees with Profiles instead of a Distance Matrix . Mol Biol Evol . 2009 July 1; 26 ( 7 ): 1641 – 50 . OpenUrl CrossRef PubMed Web of Science 67. ↵ Bastian M , Heymann S , Jacomy M . Gephi: An Open Source Software for Exploring and Manipulating Networks . Proc Int AAAI Conf Web Soc Media . 2009 Mar 19; 3 ( 1 ): 361 – 2 . OpenUrl CrossRef 68. ↵ Cantalapiedra CP , Hernández-Plaza A , Letunic I , Bork P , Huerta-Cepas J . eggNOG-mapper v2: Functional Annotation, Orthology Assignments, and Domain Prediction at the Metagenomic Scale . Mol Biol Evol . 2021 Dec 9; 38 ( 12 ): 5825 – 9 . OpenUrl CrossRef PubMed 69. ↵ Jia B , Raphenya AR , Alcock B , Waglechner N , Guo P , Tsang KK , et al. CARD 2017: expansion and model-centric curation of the comprehensive antibiotic resistance database . Nucleic Acids Res . 2017 Jan 4; 45 ( D1 ): D566 – 73 . OpenUrl CrossRef PubMed 70. ↵ McArthur AG , Waglechner N , Nizam F , Yan A , Azad MA , Baylay AJ , et al. The comprehensive antibiotic resistance database . Antimicrob Agents Chemother . 2013 July; 57 ( 7 ): 3348 – 57 . OpenUrl Abstract / FREE Full Text 71. ↵ Hall M , Frank E , Holmes G , Pfahringer B , Reutemann P , Witten IH . The WEKA data mining software: an update . ACM SIGKDD Explor Newsl . 2009 Nov 16; 11 ( 1 ): 10 – 8 . OpenUrl CrossRef 72. ↵ Shannon P , Markiel A , Ozier O , Baliga NS , Wang JT , Ramage D , et al. Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks . Genome Res . 2003 Nov; 13 ( 11 ): 2498 – 504 . OpenUrl Abstract / FREE Full Text View the discussion thread. Back to top Previous Next Posted August 21, 2025. Download PDF Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Genomic constraints shape the evolution of alternative routes to drug resistance in prokaryotes Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Genomic constraints shape the evolution of alternative routes to drug resistance in prokaryotes Lucy Dillon , James O. McInerney , Christopher J. Creevey bioRxiv 2025.08.20.671315; doi: https://doi.org/10.1101/2025.08.20.671315 Share This Article: Copy Citation Tools Genomic constraints shape the evolution of alternative routes to drug resistance in prokaryotes Lucy Dillon , James O. McInerney , Christopher J. Creevey bioRxiv 2025.08.20.671315; doi: https://doi.org/10.1101/2025.08.20.671315 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Bioinformatics Subject Areas All Articles Animal Behavior and Cognition (7619) Biochemistry (17642) Bioengineering (13865) Bioinformatics (41863) Biophysics (21410) Cancer Biology (18548) Cell Biology (25437) Clinical Trials (138) Developmental Biology (13359) Ecology (19863) Epidemiology (2067) Evolutionary Biology (24288) Genetics (15587) Genomics (22467) Immunology (17704) Microbiology (40301) Molecular Biology (17142) Neuroscience (88447) Paleontology (666) Pathology (2825) Pharmacology and Toxicology (4815) Physiology (7634) Plant Biology (15109) Scientific Communication and Education (2042) Synthetic Biology (4285) Systems Biology (9812) Zoology (2268)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0