diempy: fast and reference-free genome polarisation

preprint OA: closed
Full text JSON View at publisher

Abstract

Most ancestry-assignment methods rely on putatively pure reference panels, which are often unrealistic and bias inference. The genome polarisation algorithm diem , introduced previously, avoids reference panels by jointly inferring the polarity of common allelic states and quantifying variant diagnosticity via an expectation–maximisation procedure. Here we present diempy , an efficient python implementation of diem coupled with tools that turn polarised calls into analysis-ready outputs. diempy offers lossless VCF-to- diem BED conversion; ploidy-aware handling of individuals and chromosomes; flexible masking of sites, regions and individuals; and interactive visualisation of polarised genomes, hybrid indices, clines and ternary plots. Post-processing functions include DI thresholding, kernel smoothing, and automatic detection and run-length encoding of contiguous ancestry tracts. BED-based I/O facilitates integration with population-genomic workflows (e.g. filtering by annotation or ploidy). These features make reference-free genome polarisation with diempy practical and reproducible for studies of population structure, admixture and species barriers.
Full text 1,443 characters · extracted from oa-doi-fallback · click to expand
Abstract Most ancestry-assignment methods rely on putatively pure reference panels, which are often unrealistic and bias inference. The genome polarisation algorithm diem, introduced previously, avoids reference panels by jointly inferring the polarity of common allelic states and quantifying variant diagnosticity via an expectation–maximisation procedure. Here we present diempy, an efficient python implementation of diem coupled with tools that turn polarised calls into analysis-ready outputs. diempy offers lossless VCF-to-diem BED conversion; ploidy-aware handling of individuals and chromosomes; flexible masking of sites, regions and individuals; and interactive visualisation of polarised genomes, hybrid indices, clines and ternary plots. Post-processing functions include DI thresholding, kernel smoothing, and automatic detection and run-length encoding of contiguous ancestry tracts. BED-based I/O facilitates integration with population-genomic workflows (e.g. filtering by annotation or ploidy). These features make reference-free genome polarisation with diempy practical and reproducible for studies of population structure, admixture and species barriers. Competing Interest Statement The authors have declared no competing interest. Footnotes Figure 2: The labels and ordering of individuals was wrong in two subpanels. This has been fixed. There is no change to the data or to the results, simply fixing a plotting error.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00