ViReaDB: A user-friendly database for compactly storing viral sequence data and rapidly computing consensus genome sequences
preprint
OA: gold
CC-BY-ND-4.0
AI-generated summary
ViReaDB is a user-friendly database system that compactly stores viral sequence data and rapidly computes consensus genomes, demonstrating efficiency with SARS-CoV-2 reads.
One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works
Abstract
Motivation In viral molecular epidemiology, reconstruction of consensus genomes from sequence data is critical for tracking mutations and variants of concern. However, storage of the raw sequence data can become prohibitively large, and computing consensus genome from sequence data can be slow and requires bioinformatics expertise. Results ViReaDB is a user-friendly database system for compactly storing viral sequence data and rapidly computing consensus genome sequences. From a dataset of 1 million trimmed mapped SARS-CoV-2 reads, it is able to compute the base counts and the consensus genome in 16 minutes, store the reads alongside the base counts and consensus in 50 MB, and optionally store just the base counts and consensus (without the reads) in 300 KB. Availability ViReaDB is freely available on PyPI ( https://pypi.org/project/vireadb ) and on GitHub ( https://github.com/niemasd/ViReaDB ) as an open-source Python software project. Contact [email protected]
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-ND-4.0