Abstract
Background Transfusion recipients are a heterogeneous group of patients, yet the identification of these groups has traditionally relied on human-driven univariate analyses and domain knowledge instead of analyzing multivariate characteristics of individuals. Electronic health records (EHR) combined with unsupervised machine learning enables robust, data-driven way for phenotyping patient populations, providing finer-grained view on subgroup characteristics.
Materials and methods
We introduce an extension to the Variational Autoencoder (VAE) framework and apply the model to EHR data of 19,629 adult transfusion recipients. The latent representation of VAEs approximates a low-dimensional manifold of input data, in which patients with similar characteristics are embedded close to one another. The model integrates clustering via a Gaussian Mixture Model (GMM) prior to identify clinically relevant patient subgroups from diagnosis codes, laboratory values and demographics, while simultaneously classifying the type of transfused products. The final clusters are derived with a modified consensus clustering approach.
Results
We identified six patient groups with distinct diagnosis, laboratory, demographic, and transfusion profiles. These novel clusters provide a refined characterization of transfusion-related phenotypes, revealing more detailed distinctions among patient subgroups. Our model achieved moderate classification accuracy, with AUROC of 0.879, 0.806 and 0.861, and PR-AUC of 0.448, 0.357 and 0.492 for red blood cells (RBC), plasma and platelets, respectively. Clustering accuracy remains consistent across training and test sets.
Conclusions
Data-driven phenotyping of transfusion recipients revealed previously unexplored patient phenotypes differing in multiple characteristics. The model helps to understand the heterogeneous nature of patients requiring transfusion and provides insights on how different blood product profiles shift cluster assignments. These findings underscore the utility of latent variable modelling for population characterization and suggest potential applications in optimizing transfusion strategies as well as blood supply chain management. Validation in external cohorts remains unestablished.
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
The study was supported by The Research Fund of the Finnish Red Cross Blood Service.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
In accordance with Finnish legislation (Act on the Secondary Use of Health and Social Data 552/2019), ethical committee approval or informed consent is not required for retrospective registry studies. This study complied with the General Data Protection Regulation (GDPR) and received approval from HUS Helsinki University Hospital (HUS/579/2022, 17 October 2022).
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.
Yes
Data Availability
EU and national legislation restrict the distribution of personal health data. No data is made available outside authorized use.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.