Evaluating distributed-learning algorithms on real-world healthcare data

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

Abstract With the application of intelligent data analysis in healthcare, the need to protect patients’ data increases as well. The centralised approach, where data from several institutions are aggregated into a single server renders good knowledge discovery opportunities but is a huge data breach liability. Distributed learning can help reduce this risk as data is exploited inside each institution and only derived models are shared among peers, but there is no clear assessment of the performance gap between approaches. In this work, we used real clinical data from nine Portuguese hospitalar obstetrics departments, to compare the predictive performance of a) local, b) distributed and c) globally centralised models, focusing on predicting delivery and patient-related outcomes. Our evidence suggests that distributed learning is a viable tool to ap- ply machine learning without a significant decrease in performance while promoting a more robust privacy-concerning methodology.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0