The effects of biological knowledge graph topology on embedding-based link prediction
preprint
OA: closed
CC-BY-ND-4.0
Abstract
Due to the limited information available about rare diseases and their causal variants, knowledge graphs are often used to augment our understanding and make inferences about new gene-disease connections. Knowledge graph embedding methods have been successfully applied to various biomedical link prediction tasks but have yet to be adopted for rare disease variant prioritization. Here, we explore the effect of knowledge graph topology on knowledge graph embedding link prediction performance and challenge the assumption that massively aggregating knowledge graphs is beneficial in deciphering rare disease cases and improving prediction outcomes. We find that using a filtered version of the Monarch knowledge graph with only 11% of the original size results in notably improved model predictive performance. Additionally, these findings suggest that successful KG optimization depends on selecting high-quality information rather than simply maximizing the amount of data included.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-26T02:00:01.498150+00:00
License: CC-BY-ND-4.0