Efficient cardinality estimation for k-mers in large DNA sequencing data sets

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This paper introduces an efficient C++ implementation with a Python interface of the HyperLogLog cardinality estimation sketch for counting DNA k-mers, distributed within the khmer software package.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

We present an open implementation of the HyperLogLog cardinality estimation sketch for counting fixed-length substrings of DNA strings (“k-mers”). The HyperLogLog sketch implementation is in C++ with a Python interface, and is distributed as part of the khmer software package. khmer is freely available from https://github.com/dib-lab/khmer under a BSD License. The features presented here are included in version 1.4 and later.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0