Performance of Information Theory Derived Semantic Similarity Algorithms for Differential Diagnosis and Clustering

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

ABSTRACT Semantic similarity analysis with Human Phenotype Ontology (HPO) enables fuzzy, specificity weighted comparisons of clinical manifestations of individuals and diseases and can be used to support differential diagnostics or to stratify cohorts. Many methods have been proposed to calculate semantic similarity for various applications, including the Phenomizer, which calculates the average best match over all terms in the query and disease, and set-based methods ranging from the Jaccard Intersection to methods that leverage the conditional information content to calculate similarity. However, these methods have not been described under a single mathematical model or robustly compared using a comprehensive data set. Here, we describe several semantic similarity algorithms using derivations based on information theory, propose three of our own variations to these models, and compare the performance of each approach for differential diagnostic ranking and phenotypic clustering. We find that Phenomizer performs better when diseases are ranked by similarity alone, without generating p-values. Additionally, non-normalized algorithms that use conditional information perform similarly to Phenomizer for differential diagnosis. In contrast, normalized algorithms perform best when clustering cohorts. Availability Data is available through the Phenopacket-Store ( https://github.com/monarch-initiative/phenopacket-store ). Algorithms are implemented in the Python package SetSim ( https://github.com/P2GX/setsim ).
Full text 1,604 characters · extracted from oa-doi-fallback · click to expand
ABSTRACT Semantic similarity analysis with Human Phenotype Ontology (HPO) enables fuzzy, specificity weighted comparisons of clinical manifestations of individuals and diseases and can be used to support differential diagnostics or to stratify cohorts. Many methods have been proposed to calculate semantic similarity for various applications, including the Phenomizer, which calculates the average best match over all terms in the query and disease, and set-based methods ranging from the Jaccard Intersection to methods that leverage the conditional information content to calculate similarity. However, these methods have not been described under a single mathematical model or robustly compared using a comprehensive data set. Here, we describe several semantic similarity algorithms using derivations based on information theory, propose three of our own variations to these models, and compare the performance of each approach for differential diagnostic ranking and phenotypic clustering. We find that Phenomizer performs better when diseases are ranked by similarity alone, without generating p-values. Additionally, non-normalized algorithms that use conditional information perform similarly to Phenomizer for differential diagnosis. In contrast, normalized algorithms perform best when clustering cohorts. Availability Data is available through the Phenopacket-Store (https://github.com/monarch-initiative/phenopacket-store). Algorithms are implemented in the Python package SetSim (https://github.com/P2GX/setsim). Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00