A Systematic Comparison of Human Mitochondrial Genome Assembly Tools

preprint OA: closed
View at publisher

Abstract

Background: Mitochondria are the cell organelles that produce the majority of the chemical energy required to power the biochemical reactions of the cell. Despite being a part of a eukaryotic host cell, the mitochondria contain a separate genome whose origin is linked with the endocytosis of a prokaryotic cell by the eukaryotic host cell and encodes separate genomic information throughout their genomes. Mitochondrial genomes accommodate essential genes and are regularly utilized in biotechnology and phylogenetics. Various assemblers capable of generating full mitochondrial genomes are being continuously developed. These tools often use whole-genome sequencing data as an input containing reads from the mitochondrial genome. Till now no published work has explored the systematic comparison of all the available tools for assembling mitochondrial genome using short-read sequencing data. This evaluation is required in order to identify the best tool that can be well optimized for small-scale projects or even national-level research. Results Here we present a benchmark study of ten mitochondrial assembly tools capable of producing mitochondrial genomes for whole genome paired-end sequencing data. Simulated and real whole genome sequencing data was used as an input for these assemblers. Each of these publicly accessible tools are containerized as docker images to ensure the reproducibility. Our findings demonstrate that the examined assemblers have various computing requirements and degrees of success with the input datasets. Conclusions Based on the overall performance metrics and consistency in assembly quality for all sequencing data, MToolBox performed the best. However, among all the assemblers for simulated datasets, NOVOPlasty consumed the smallest amount of runtime and processing resources. Therefore, NOVOPlasty may be more practical to use when there is a big sample size and a lack of computational resources. Besides, as long read sequencing gains popularity, mitochondrial genome assemblers that can use long read sequencing data must be developed.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00