A Conditional Random Field approach for de novo reconstruction of bacterial haplotypes from a de Bruijn graph representation

preprint OA: closed
Full text JSON View at publisher

Abstract

Background Detecting distinct bacterial strains in a mixed sample is an important, yet less well-developed aspect of metagenomic research. Several methods exist that successfully retrieve a de novo reconstruction of viral strains. However, the reconstruction of bacterial haplotypes poses its own distinct challenges, and methods that successfully reconstruct full genome-length bacterial strains de novo are scarce. Here, we develop HaploDetox, a method for de novo bacterial haplotype reconstruction from short reads. We use a de Bruijn graph representation of the reads in which nodes correspond with k-mers from the read set and arcs represent overlap between two nodes’ sequences. Our aim is to accurately assign labels to each node and arc in the graph to reveal the presence or absence of their corresponding sequence in individual strains. Results Using a negative binomial mixture model, we model the relationship between the read coverage of nodes and arcs in the graph and their presence in a strain. We achieve improved labelling accuracy by including contextual information from neighbouring nodes and arcs with a Conditional Random Field. These labels are used to extract strain-specific de Bruijn graphs from the original graph. Additionally, we allow users to assess the number of strains present in the dataset based on model selection criteria. We evaluate our node/arc labelling accuracy on simulated datasets and in silico mixes of real datasets containing different numbers of strains, as well as on in vitro mixed real datasets. Existing de novo haplotype reconstruction methods present their reconstruction as strain-specific sets of SNPs. We demonstrate that HaploDetox assigns strain-specific SNPs with a higher recall and similar precision than existing methods, by aligning the unitigs from strain-specific graphs to a reference genome. Conclusions We achieve improved strain-specific SNP phasing accuracy as compared to existing methods for de novo bacterial haplotype reconstruction. Additionally, HaploDetox is not limited to the determination of strain-specific SNPs, and other types of variant calls can be obtained through reference alignment. Finally, strain-specific de Bruijn graphs are an important first step towards full genome-length bacterial haplotype-aware assembly.
Full text 2,589 characters · extracted from oa-doi-fallback · 3 sections · click to expand

Abstract

Background Detecting distinct bacterial strains in a mixed sample is an important, yet less well-developed aspect of metagenomic research. Several methods exist that successfully retrieve a de novo reconstruction of viral strains. However, the reconstruction of bacterial haplotypes poses its own distinct challenges, and methods that successfully reconstruct full genome-length bacterial strains de novo are scarce. Here, we develop HaploDetox, a method for de novo bacterial haplotype reconstruction from short reads. We use a de Bruijn graph representation of the reads in which nodes correspond with k-mers from the read set and arcs represent overlap between two nodes’ sequences. Our aim is to accurately assign labels to each node and arc in the graph to reveal the presence or absence of their corresponding sequence in individual strains.

Results

Using a negative binomial mixture model, we model the relationship between the read coverage of nodes and arcs in the graph and their presence in a strain. We achieve improved labelling accuracy by including contextual information from neighbouring nodes and arcs with a Conditional Random Field. These labels are used to extract strain-specific de Bruijn graphs from the original graph. Additionally, we allow users to assess the number of strains present in the dataset based on model selection criteria. We evaluate our node/arc labelling accuracy on simulated datasets and in silico mixes of real datasets containing different numbers of strains, as well as on in vitro mixed real datasets. Existing de novo haplotype reconstruction methods present their reconstruction as strain-specific sets of SNPs. We demonstrate that HaploDetox assigns strain-specific SNPs with a higher recall and similar precision than existing methods, by aligning the unitigs from strain-specific graphs to a reference genome.

Conclusions

We achieve improved strain-specific SNP phasing accuracy as compared to existing methods for de novo bacterial haplotype reconstruction. Additionally, HaploDetox is not limited to the determination of strain-specific SNPs, and other types of variant calls can be obtained through reference alignment. Finally, strain-specific de Bruijn graphs are an important first step towards full genome-length bacterial haplotype-aware assembly. Competing Interest Statement The authors have declared no competing interest. List of abbreviations - BIC - Bayesian Information Criterion - CRF - Conditional Random Field - EM - Expectation-Maximisation - MAE - Mean Absolute Error - SNP - Single-Nucleotide Polymorphism

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00