Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model

preprint OA: closed
View at publisher

Abstract

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked six representative methods spanning gene-cluster, compacted coloured de Bruijn graph (ccDBG), multiple sequence alignment, and hybrid approaches. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas unitig graphs expanded, with distinct degree–prevalence fingerprints across tools. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one ccDBG implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy. Authors Summary Bacterial populations are often described using “pangenome graphs,” which aim to capture all genetic variation across many genomes in a single structure. However, different tools build these graphs in fundamentally different ways, and little is known about how those differences affect the results. In this study, we systematically compared several widely used approaches using a clinically important strain of Escherichia coli that is rich in repeated and mobile DNA. We found that the size, connectivity, and overall structure of the resulting graphs varied dramatically depending on the method used. Importantly, we also show that incomplete genome assemblies (common in large sequencing studies) strongly alter graph structure, and that different tools respond to incomplete data in different ways. In some cases, this affects the detection of medically relevant genes, including Shiga toxin genes linked to severe disease. Our results demonstrate that pangenome graphs are not interchangeable representations of bacterial diversity. Instead, their structure depends on both the method and the quality of the input data. We argue that researchers should choose graph-building tools carefully and report structural properties explicitly to ensure reproducible and interpretable results.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00