Estimating Error Models for Whole Genome Sequencing Using Mixtures of Dirichlet-Multinomial Distributions

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This study models next-generation sequencing read depth using a mixture of Dirichlet-multinomial distributions to improve genotyping accuracy, identifying a minor component enriched for copy number variants.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Motivation Accurate identification of genotypes is critical in identifying de novo mutations, linking mutations with disease, and determining mutation rates. Because de novo mutations are rare, even low levels of genotyping error can cause a large fraction of false positive de novo mutations. Biological and technical processes that adversely affect genotyping include copy-number-variation, paralogous sequences, library preparation, sequencing error, and reference-mapping biases, among others. Results We modeled the read depth for all data as a mixture of Dirichlet-multinomial distributions, resulting in significant improvements over previously used models. In most cases the best model was comprised of two distributions. The major-component distribution is similar to a binomial distribution with low error and low reference bias. The minor-component distribution is overdispersed with higher error and reference bias. We also found that sites fitting the minor component are enriched for copy number variants and low complexity region. We expect that this approach to modeling the distribution of NGS data, will lead to improved genotyping. For example, this approach provides an expected distribution of reads that can be incorporated into a model to estimate de novo mutations using reads across a pedigree. Availability Methods and data files are available at https://github.com/CartwrightLab/WuEtAl2016/ . Contact [email protected]

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-NC-ND-4.0