Deep latent variable modelling reveals clinically significant subgroups among transfusion recipients

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

A bstract Background Transfusion recipients are a heterogeneous group of patients, yet the identification of these groups has traditionally relied on human-driven univariate analyses and domain knowledge instead of analyzing multivariate characteristics of individuals. Electronic health records (EHR) combined with unsupervised machine learning enables robust, data-driven way for phenotyping patient populations, providing finer-grained view on subgroup characteristics. Materials and Methods We introduce an extension to the Variational Autoencoder (VAE) framework and apply the model to EHR data of 19,629 adult transfusion recipients. The latent representation of VAEs approximates a low-dimensional manifold of input data, in which patients with similar characteristics are embedded close to one another. The model integrates clustering via a Gaussian Mixture Model (GMM) prior to identify clinically relevant patient subgroups from diagnosis codes, laboratory values and demographics, while simultaneously classifying the type of transfused products. The final clusters are derived with a modified consensus clustering approach. Results We identified six patient groups with distinct diagnosis, laboratory, demographic, and transfusion profiles. These novel clusters provide a refined characterization of transfusion-related phenotypes, revealing more detailed distinctions among patient subgroups. Our model achieved moderate classification accuracy, with AUROC of 0.879, 0.806 and 0.861, and PR-AUC of 0.448, 0.357 and 0.492 for red blood cells (RBC), plasma and platelets, respectively. Clustering accuracy remains consistent across training and test sets. Conclusions Data-driven phenotyping of transfusion recipients revealed previously unexplored patient phenotypes differing in multiple characteristics. The model helps to understand the heterogeneous nature of patients requiring transfusion and provides insights on how different blood product profiles shift cluster assignments. These findings underscore the utility of latent variable modelling for population characterization and suggest potential applications in optimizing transfusion strategies as well as blood supply chain management. Validation in external cohorts remains unestablished.
Full text 4,232 characters · extracted from oa-doi-fallback · 4 sections · click to expand

Abstract

Background Transfusion recipients are a heterogeneous group of patients, yet the identification of these groups has traditionally relied on human-driven univariate analyses and domain knowledge instead of analyzing multivariate characteristics of individuals. Electronic health records (EHR) combined with unsupervised machine learning enables robust, data-driven way for phenotyping patient populations, providing finer-grained view on subgroup characteristics.

Materials and methods

We introduce an extension to the Variational Autoencoder (VAE) framework and apply the model to EHR data of 19,629 adult transfusion recipients. The latent representation of VAEs approximates a low-dimensional manifold of input data, in which patients with similar characteristics are embedded close to one another. The model integrates clustering via a Gaussian Mixture Model (GMM) prior to identify clinically relevant patient subgroups from diagnosis codes, laboratory values and demographics, while simultaneously classifying the type of transfused products. The final clusters are derived with a modified consensus clustering approach.

Results

We identified six patient groups with distinct diagnosis, laboratory, demographic, and transfusion profiles. These novel clusters provide a refined characterization of transfusion-related phenotypes, revealing more detailed distinctions among patient subgroups. Our model achieved moderate classification accuracy, with AUROC of 0.879, 0.806 and 0.861, and PR-AUC of 0.448, 0.357 and 0.492 for red blood cells (RBC), plasma and platelets, respectively. Clustering accuracy remains consistent across training and test sets.

Conclusions

Data-driven phenotyping of transfusion recipients revealed previously unexplored patient phenotypes differing in multiple characteristics. The model helps to understand the heterogeneous nature of patients requiring transfusion and provides insights on how different blood product profiles shift cluster assignments. These findings underscore the utility of latent variable modelling for population characterization and suggest potential applications in optimizing transfusion strategies as well as blood supply chain management. Validation in external cohorts remains unestablished. Competing Interest Statement The authors have declared no competing interest. Funding Statement The study was supported by The Research Fund of the Finnish Red Cross Blood Service. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: In accordance with Finnish legislation (Act on the Secondary Use of Health and Social Data 552/2019), ethical committee approval or informed consent is not required for retrospective registry studies. This study complied with the General Data Protection Regulation (GDPR) and received approval from HUS Helsinki University Hospital (HUS/579/2022, 17 October 2022). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data Availability EU and national legislation restrict the distribution of personal health data. No data is made available outside authorized use.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00