A Pan-Genome Data Structure Induced by Pooled Sequencing Facilitates Variant Mining in Heterogeneous Germplasm

preprint OA: closed
View at publisher

Abstract

Abstract Valuable genetic variation lies unused in gene banks due to the difficulty of exploiting heterogeneous germplasm accessions. Advances in molecular breeding, including transgenics and genome editing, present the opportunity to exploit hidden sequence variation directly. Here we describe the pangenome data structure induced by wholegenome sequencing of pooled individuals from wild populations of Patellifolia spp., a source of disease resistance genes for the related crop species sugar beet (Beta vulgaris). We represent the pangenome of a heterogeneous population sample as a set of phased reads mapped to a reference genome assembly derived from the sequence pool or other sources. We show that this basic data structure can be queried by homology to identify short haplotypic variants present in the wild relative, at genes of agronomic interest in the crop. Further we demonstrate the possibility of cataloging short haplotypic variation in all Patellifolia genomic regions that have corresponding single copy orthologous regions in sugar beet. The data structure, termed a "phased read archive", can be produced, altered, and queried using standard tools to facilitate discovery of agronomicallyimportant sequence variation in heterogeneous germplasm.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00