A graph-based approach for the visualisation and analysis of bacterial pangenomes | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A graph-based approach for the visualisation and analysis of bacterial pangenomes Joshua D Harling-Lee, Jamie Gorzynski, Gonzalo Yebra, Tim Angus, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1406006/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. Starting with the output of standard pangenome analysis tools, we provide scripts to generate input files for a new open-source network analysis platform called Graphia. We describe the generation and analysis of interactive, 3D network graphs to explore the structure of a bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenomes datasets, thereby enhancing interpretation of these complex data. Bacteria Pangenome Accessory genes Network graphs Data visualisation Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Full Text Tables Table 1 is available in the Supplementary Files section. Supplementary Files Table1legionellametadata.csv suppfiguresandinfo.docx supplementarydata.zip Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major revision 05 May, 2022 Reviews received at journal 11 Apr, 2022 Reviewers invited by journal 11 Apr, 2022 Editor invited by journal 23 Mar, 2022 Editor assigned by journal 01 Mar, 2022 Submission checks completed at journal 28 Feb, 2022 First submitted to journal 28 Feb, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1406006","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":97647566,"identity":"7a7a0b7b-fc60-450a-94c1-974a84e240e6","order_by":0,"name":"Joshua D Harling-Lee","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABN0lEQVRIie2RPWrDMBSAnxHIi+JZpcG+go2hU+hZZALxkoAhUDIEajDIW2cXeghnakeBIV1M6QmKTS/gKbRLWslp2oLizh30Dfp77+PpSQAGwz/ESsFHch5jtWsAXHUomNqJPoMNKQR/hcM++S9FclCO4Sj9Th5QUIpKlKxfiHOePwNbTeL7PG9E+wCuI1DbWXymXwwnqNguCR7XCbB6trira19ENYRnAofU4nNdIT4imBFM5/FbxKtFQecgIg5RKeACLL46rex7hUG0/4gpjRulXJfC3g0qI66UWMjeBaOU+UqRI1FV9ItlOKlGN30V+TrbaVAQ1QunwW1FlpQ9ae0HebZ5JTvmekXcQLe+9Kidt+07n3jOY77puquppmTH9wfi/w5Q+QFw8iO9n6Xd6GGDwWAwSD4BWJdo7QKkJV4AAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-1568-8250","institution":"The University of Edinburgh The Roslin Institute","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Joshua","middleName":"D","lastName":"Harling-Lee","suffix":""},{"id":97647567,"identity":"af45088f-846e-49af-b47b-872be01496bb","order_by":1,"name":"Jamie Gorzynski","email":"","orcid":"","institution":"The University of Edinburgh The Roslin Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jamie","middleName":"","lastName":"Gorzynski","suffix":""},{"id":97647568,"identity":"0b01a854-7ec9-42e2-bb58-3438a5da74d1","order_by":2,"name":"Gonzalo Yebra","email":"","orcid":"https://orcid.org/0000-0002-3472-3667","institution":"The University of Edinburgh The Roslin Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Gonzalo","middleName":"","lastName":"Yebra","suffix":""},{"id":97647569,"identity":"2dc99f43-42fb-4e4f-857a-6e9d6352d5cf","order_by":3,"name":"Tim Angus","email":"","orcid":"https://orcid.org/0000-0003-1736-9456","institution":"Janssen Immunology, PA, US","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tim","middleName":"","lastName":"Angus","suffix":""},{"id":97647570,"identity":"ecf89d59-0a58-4546-aca7-4f9a668643b6","order_by":4,"name":"J Ross Fitzgerald","email":"","orcid":"https://orcid.org/0000-0002-9233-8468","institution":"The University of Edinburgh The Roslin Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"J","middleName":"Ross","lastName":"Fitzgerald","suffix":""},{"id":97647571,"identity":"9706d6ff-d379-42ad-beae-06069ecadb52","order_by":5,"name":"Tom C Freeman","email":"","orcid":"https://orcid.org/0000-0001-5235-8483","institution":"Janssen Immunology, PA, US","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tom","middleName":"C","lastName":"Freeman","suffix":""}],"badges":[],"createdAt":"2022-03-01 01:52:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1406006/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1406006/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":20262946,"identity":"7741dc29-c2ac-4e61-ad4c-0704d29b2b75","added_by":"auto","created_at":"2022-04-12 19:25:12","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":70855,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of typical analysis workflow. Starting with whole genome sequences, the first step is genome annotation, then identification of elements of the pangenome, for which common tools are shown. The main input to the GraPPLE script library is the gene presence/absence matrix (in binary format). From this, pairwise Jaccard similarities are calculated, then converted to a list with annotations (compliant with the “.layout” file schema for Graphia load). Suggested filters and transformations to apply within Graphia are listed in order. GraPPLE contains scripts to change the file formats of synteny graphs produced by common tools (where necessary) to allow for loading into Graphia, again with suggested transformations listed.\u003c/p\u003e","description":"","filename":"Figure1WorkProcess.png","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/49605b797d3394268455d5cd.png"},{"id":20262699,"identity":"faf9f2fe-0cca-40dc-ac8d-5c3182b4c91a","added_by":"auto","created_at":"2022-04-12 19:20:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":6628888,"visible":true,"origin":"","legend":"\u003cp\u003eGenome-to-genome networks provide insight into population structure and associations. (A) Phylogenetic tree of all S. aureus isolates produced using ParSNP v1.2. Largest clonal complexes (CCs) are coloured and labelled, with minor CCs in white. (B) Relationship network between genomes based on the accessory genome visualised in 3D. Nodes represent individual isolates, edges represent shared accessory gene content (JSC \u0026gt; 0.8 filter applied), and a k-NN (k = 8) edge filter has been applied. The largest CCs are coloured and labelled, corresponding to the phylogeny in (A). (C) The same network as in (B) but visualised using a 2D layout. (Ci) shows only CC5 isolates, demonstrating sub-structures within the network with nodes coloured by sequence type. (D) The same network as in (C) and (Di) the same network as in (Ci), coloured by host.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/cd67eb3c475a5f1b8ff60300.png"},{"id":20262945,"identity":"432a4367-7924-4ece-bcb1-5669c5e2862e","added_by":"auto","created_at":"2022-04-12 19:25:11","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":3463668,"visible":true,"origin":"","legend":"\u003cp\u003eGene association networks provide relationship between gene presence and host genome characteristics. (A) A filtered gene-association graph visualised in 2D space. Each node represents a gene and weighted edges correspond to the Jaccard similarity coefficient between each pair of genes (JSC \u0026gt;0.550). High and low frequency genes have been removed (n \u0026lt;7; n \u0026gt;770), as have smaller components (n \u0026lt;10), for visual clarity. Nodes are coloured by cluster (MCLi = 1.50). (B) Four gene distribution profile plots, representing commonly observed distributions: Cluster 1, near-core genes; Cluster 3, set of genes showing a near lineage-specific distribution (CC45); Cluster 25, host-associated (aves) set of genes; Cluster 28, widely distributed but not lineage-specific set of genes. Each bar in a plot represents an individual genome, with bar height equal to the proportion of genes in the cluster present in that genome. Coloured horizontal bars below the x-axis represent the host (upper) and clonal complex (lower).\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/574416be60df5fd780cdc9b3.png"},{"id":20262705,"identity":"6ee3a7dd-0d90-44c2-8f24-a99a040739a0","added_by":"auto","created_at":"2022-04-12 19:20:12","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1179681,"visible":true,"origin":"","legend":"\u003cp\u003eSyntenic connections within the S. aureus pangenome shows location of gene sets. A) Full synteny network of S. aureus at 90% gene identity. Nodes represent genes and are sized according to the number of isolates in which they occur. Edges are weighted as the number of times two genes occur next to one another. Coloured by Louvain Cluster (LCi = 0.400). B) “Phage” region (right-hand side of network in (A)), selected by retaining only nodes in clusters 1, 5, 6, 7 and 9. Coloured as in (A). C) 2D layout of the same region as in B, with common annotation highlighted alongside core genes. This network allows detailed inspection of phage integration sites with respect to core genes, accessory genes shared between common phage types, and putative novel factors carried on phage variants. Gene cluster 28 (see Figure 3) is also found in this region. D) Region, in context, that carries gene cluster 25 (orange; see Figure 3) as a single syntenic set, inserted next to a near-core gene. Examples of gene cluster 3 (yellow) sets are also seen. E) 2D layout of region from D (clusters 2 and 18). Smaller, focused networks aid navigation and investigation of specific gene locations and neighbours.\u003c/p\u003e","description":"","filename":"Figure4resub.png","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/03e6f730c4330bf222f835e4.png"},{"id":20262703,"identity":"6954b3bb-31f0-4725-9fea-754bee5e06f3","added_by":"auto","created_at":"2022-04-12 19:20:12","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":633657,"visible":true,"origin":"","legend":"\u003cp\u003eInvestigating the L. pneumophila pangenome through network graphs. (A) Genome-genome similarity network, coloured by ST. Nodes represent genomes, and edges are weighted by pairwise JSC. Two key ST36 clades are highlighted. Edges are filtered by JSC \u0026gt;0.5 and using a k-NN = 10. (B) Same network as in (A) but coloured by sampling location. Same two key ST36 clades are highlighted. (C) Gene-gene similarity network; nodes represent genes, edges weighted by pairwise JSC. Edges are filtered by JSC \u0026gt;0.55, and k-NN = 20. Clustering with MCLi = 1.50. Components of n \u0026lt;10 removed for visual clarity. (D) Accessory gene distribution plots of the largest gene clusters associated with the division in ST36 between Swiss (reds and orange) and US (blues) genomes. Plots as described in Figure 3. (E) 3D rendering of the pangenome synteny graph for L. pneumophila; coloured nodes correspond to the clusters in (C, D). (F) Filtered and 2D view of the key region, with convergence of both Swiss and US-associated regions to the same core region.\u003c/p\u003e","description":"","filename":"Figure5resub.png","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/762051fbd1a02b027ae4d1c8.png"},{"id":20262947,"identity":"6d25866e-68c3-4066-a418-1f947bb54011","added_by":"auto","created_at":"2022-04-12 19:25:18","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":388168,"visible":true,"origin":"","legend":"","description":"","filename":"finalmanuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1_covered.pdf"},{"id":20262700,"identity":"4a17755c-580e-4907-bc41-4d3c0ad404fb","added_by":"auto","created_at":"2022-04-12 19:20:11","extension":"csv","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":16984,"visible":true,"origin":"","legend":"","description":"","filename":"Table1legionellametadata.csv","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/71286fae6e224293086348a4.csv"},{"id":20262702,"identity":"248d69ad-563d-4998-8cff-0576a375081b","added_by":"auto","created_at":"2022-04-12 19:20:11","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":639665,"visible":true,"origin":"","legend":"","description":"","filename":"suppfiguresandinfo.docx","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/a067e20d32a855d062d85f0e.docx"},{"id":20262706,"identity":"64009bd8-e3de-4119-b74f-b6d07e062c99","added_by":"auto","created_at":"2022-04-12 19:20:12","extension":"zip","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":14752570,"visible":true,"origin":"","legend":"","description":"","filename":"supplementarydata.zip","url":"https://assets-eu.researchsquare.com/files/rs-1406006/v1/3f2f241aca259ce310b435c9.zip"}],"financialInterests":"","formattedTitle":"A graph-based approach for the visualisation and analysis of bacterial pangenomes","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-1406006/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."},{"header":"Tables","content":"\u003cp\u003eTable 1 is available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-bioinformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"binf","sideBox":"Learn more about [BMC Bioinformatics](http://bmcbioinformatics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/binf","title":"BMC Bioinformatics","twitterHandle":"@BMC_Bioinformatics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Bacteria, Pangenome, Accessory genes, Network graphs, Data visualisation","lastPublishedDoi":"10.21203/rs.3.rs-1406006/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1406006/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. Starting with the output of standard pangenome analysis tools, we provide scripts to generate input files for a new open-source network analysis platform called Graphia. We describe the generation and analysis of interactive, 3D network graphs to explore the structure of a bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenomes datasets, thereby enhancing interpretation of these complex data.","manuscriptTitle":"A graph-based approach for the visualisation and analysis of bacterial pangenomes","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-04-12 19:20:09","doi":"10.21203/rs.3.rs-1406006/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2022-05-05T06:46:24+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-04-11T12:28:02+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-04-11T06:59:11+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"BMC Bioinformatics","date":"2022-03-23T14:19:17+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-03-02T03:04:15+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-02-28T23:00:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Bioinformatics","date":"2022-02-28T20:51:32+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-bioinformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"binf","sideBox":"Learn more about [BMC Bioinformatics](http://bmcbioinformatics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/binf","title":"BMC Bioinformatics","twitterHandle":"@BMC_Bioinformatics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9ce857a9-8362-46c3-bc4a-52b552a04463","owner":[],"postedDate":"April 12th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2022-08-22T14:50:23+00:00","versionOfRecord":[],"versionCreatedAt":"2022-04-12 19:20:09","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1406006","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1406006","identity":"rs-1406006","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.