Near-perfect identification of half-sibling versus niece/nephew–avuncular pairs without pedigree information or genotyped relatives

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study developed a computational framework using haplotype-level sharing to accurately distinguish half-sibling from avuncular pairs in large genomic biobanks without pedigree data.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-17 · read from full text

The paper studies how to identify second-degree relative pairs—specifically half-siblings versus avuncular (aunt/uncle–niece/nephew) pairs—using only genotype data in large biobanks that lack pedigree information and may have cryptic relatedness. The authors develop a genotype-only framework that uses haplotype-level sharing features from across-chromosome phasing and models them with a multivariate Gaussian mixture model to separate these relationship types, also distinguishing nieces/nephews from aunts/uncles within avuncular pairs. Validation is performed against ground-truth labels inferred via a multi-step process from a family graph built using high-confidence first-degree relationships, achieving 100% classification of avuncular pairs and identification of 60/61 half-sibling pairs (98.4% sensitivity, 100% specificity). The study’s main caveat is that the ground-truth relies on inferred labels from first-degree relationship inference rather than experimentally established pedigrees. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Motivation Large-scale genomic biobanks contain thousands of second-degree relatives without pedigree data. Accurately distinguishing half-sibling and avuncular pairs–both of which share approximately 25% of the genome–remains a significant challenge. Current SNP-based methods rely on aggregate Identical-By-Descent (IBD) segment counts and age differences, but substantial overlap in these distributions leads to high misclassification rates. There is a need for a scalable, genotype-only method that can resolve these second-degree ambiguities without requiring observed pedigrees or phenotypic information. Results We present a novel computational framework that achieves complete separation of half-siblings and avuncular pairs using haplotype-level sharing features based on data that has been computationally phased across-chromosomes. These features also allow differentiation of nieces/nephews from aunts/uncles within avuncular pairs. By modeling these features with a multivariate Gaussian mixture model, we achieve exceptional classification performance in biobank-scale data. Ground-truth labels for validation were established through a multi-step inference process within a family graph constructed from high-confidence first, second and third-degree relationships. All 352 ground-truth avuncular pairs and all but one of the 80 half-sibling pairs were correctly classified. Treating avuncular status as the positive class, this implies a specificity of 98.75% and sensitivity of 1. The mixture-implied Bayes error of those without ground-truth labels was higher, estimated at 9.3% under the our modeling assumptions. Our results suggest that the framework could perform comparably on other datasets of similar size. This method provides a robust, scalable solution for pedigree reconstruction and the control of cryptic relatedness in large-scale genomic studies. Contact [email protected] , [email protected] , and [email protected]
Full text 1,974 characters · extracted from oa-doi-fallback · 2 sections · click to expand

Abstract

Motivation Large-scale genomic biobanks contain thousands of second-degree relatives without pedigree data. Accurately distinguishing half-siblings from avuncular pairs–both sharing approximately 25% of the genome–remains a significant challenge. Current SNP-based methods rely on aggregate Identical-By-Descent (IBD) segment counts and age differences, but substantial overlap in these distributions leads to high misclassification rates. There is a need for a scalable, genotype-only method that can resolve these second-degree ambiguities without requiring observed pedigrees or phenotypic information.

Results

We present a novel computational framework that achieves near-complete separation of half-siblings and avuncular pairs using haplotype-level sharing features derived from across-chromosome phasing. These features also allow differentiation of nieces/nephews from aunts/uncles within avuncular pairs. By modeling these features with a multivariate Gaussian mixture model (GMM), we demonstrate exceptional classification performance in biobank-scale data. Ground-truth labels for validation were established through a multi-step inference process within a family graph constructed from high-confidence first-degree relationships. All 395 ground-truth avuncular pairs were correctly classified, and 60 of 61 half-sibling pairs were correctly identified. Treating (e.g.) half-sibling status as the positive class, this corresponds to 98.4% sensitivity and 100% specificity. Our results suggest that the framework could perform comparably on other datasets of similar size. This method provides a robust, scalable solution for pedigree reconstruction and the control of cryptic relatedness in large-scale genomic studies. Contact emmanuel.sapin{at}colorado.edu, kristen.kelly{at}colorado.edu, and matthew.c.keller{at}colorado.edu Competing Interest Statement The authors have declared no competing interest. Footnotes the gaussian mixture model was updated

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00