Pre-processing of paleogenomes: Mitigating reference bias and postmortem damage in ancient genome data

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

ABSTRACT Ancient DNA analysis is subject to various technical challenges, including bias towards the reference allele (“reference bias”), postmortem damage (PMD) that confounds real variants, and limited coverage. Here, we conduct a systematic comparison of alternative approaches against reference bias and against PMD. To reduce reference bias, we either (a) mask variable sites before alignment or (b) align the data to a graph genome representing all variable sites. Compared to alignment to the linear reference genome, both masking and graph alignment effectively remove allelic bias when using simulated or real ancient human genome data, but only if sequencing data is available in FASTQ or unfiltered BAM format. Reference bias remains indelible in quality-filtered BAM files and in 1240K-capture data. We next study three approaches to overcome postmortem damage: (a) trimming, (b) rescaling base qualities, and (c) a new algorithm we present here, bamRefine , which masks only PMD-vulnerable polymorphic sites. We find that bamRefine is optimal in increasing the number of genotyped loci up to 20% compared to trimming and in improving accuracy compared to rescaling. We propose graph alignment coupled with bamRefine to minimise data loss and bias. We also urge the paleogenomics community to publish FASTQ files.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0