Predicting Intentional Self-Harm Following Psychiatric Discharge in Catalonia, Spain: Machine Learning Models from Linked Registry Data

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

ABSTRACT Introduction Patients recently discharged from psychiatric hospitalization are at increased risk of intentional self-harm, including suicide. Using linked population-based registry data from Catalonia, Spain, we developed machine learning-based prediction models for post-discharge intentional self-harm across different follow-up horizons, sex, and age groups, and evaluated their generalizability and robustness with multiple validation strategies. Methods Retrospective cohort study including 41,827 individuals accounting for 71,865 psychiatric hospitalizations with discharge at age ≥10 years, between January 1, 2015, and December 31, 2018, in Catalonia, Spain, with follow-up until December 31, 2019. Primary outcome was intentional self-harm (fatal or non-fatal) within 7, 30, 90, 180, and 365 days post-discharge. Models incorporated 247 predictors from electronic health records, including sociodemographic characteristics, mental and physical disorder categories, categories of dispensed psychotropic medication, and history of self-harm and psychiatric hospitalization. Model performance was evaluated using the area under the receiver operating characteristic curve (AUCROC) and the area under the precision-recall curve (AUCPR). Predictor importance was assessed using Shapley Additive Explanations (SHAP). Results Within 365 days, 4,901 hospitalizations (6.8%) were followed by intentional self-harm. The 365-day model trained on the full cohort achieved a AUCROC of 0.819, in the test sample with adjusted AUCPR indicating a median 5.4-fold improvement over baseline prevalence. This model generalized well across event horizons and sex–age strata, outperforming subgroup-specific models when data sparsity limited performance. Separate models trained by event horizons, and stratified by sex, and sex–age groups achieved a median AUCROC of 0.775 (IQR 0.764–0.808), with adjusted AUCPR indicating a median 5.4-fold improvement over baseline prevalence (IQR 4.5–6.2). Key predictors included the recency of the last registered diagnosis of depressive episodes, recurrent depression, adjustment disorders, and schizophrenia, as well as recent SSRI dispensation and the number of childhood-onset disorder and musculoskeletal disease diagnoses in the previous five years. Predictor importance varied considerably across sex–age strata, with smaller differences across horizons. Subject-level and temporal split validation strategies reduced performance (AUCROC 0.711–0.746), though estimates remained clinically informative (2.8–3.1-fold improvement over baseline prevalence). Conclusions Machine learning models using routinely collected health records predicted intentional self-harm after psychiatric hospitalization with good discrimination and clinically meaningful precision–recall performance. A single 365-day model generalized well across horizons and demographic groups, suggesting that one broadly trained model may provide a pragmatic and scalable approach for clinical implementation.
Full text 7,742 characters · extracted from oa-doi-fallback · 4 sections · click to expand

Abstract

Introduction Patients recently discharged from psychiatric hospitalization are at increased risk of intentional self-harm, including suicide. Using linked population-based registry data from Catalonia, Spain, we developed machine learning-based prediction models for post-discharge intentional self-harm across different follow-up horizons, sex, and age groups, and evaluated their generalizability and robustness with multiple validation strategies.

Methods

Retrospective cohort study including 41,827 individuals accounting for 71,865 psychiatric hospitalizations with discharge at age ≥10 years, between January 1, 2015, and December 31, 2018, in Catalonia, Spain, with follow-up until December 31, 2019. Primary outcome was intentional self-harm (fatal or non-fatal) within 7, 30, 90, 180, and 365 days post-discharge. Models incorporated 247 predictors from electronic health records, including sociodemographic characteristics, mental and physical disorder categories, categories of dispensed psychotropic medication, and history of self-harm and psychiatric hospitalization. Model performance was evaluated using the area under the receiver operating characteristic curve (AUCROC) and the area under the precision-recall curve (AUCPR). Predictor importance was assessed using Shapley Additive Explanations (SHAP).

Results

Within 365 days, 4,901 hospitalizations (6.8%) were followed by intentional self-harm. The 365-day model trained on the full cohort achieved a AUCROC of 0.819, in the test sample with adjusted AUCPR indicating a median 5.4-fold improvement over baseline prevalence. This model generalized well across event horizons and sex–age strata, outperforming subgroup-specific models when data sparsity limited performance. Separate models trained by event horizons, and stratified by sex, and sex–age groups achieved a median AUCROC of 0.775 (IQR 0.764–0.808), with adjusted AUCPR indicating a median 5.4-fold improvement over baseline prevalence (IQR 4.5–6.2). Key predictors included the recency of the last registered diagnosis of depressive episodes, recurrent depression, adjustment disorders, and schizophrenia, as well as recent SSRI dispensation and the number of childhood-onset disorder and musculoskeletal disease diagnoses in the previous five years. Predictor importance varied considerably across sex–age strata, with smaller differences across horizons. Subject-level and temporal split validation strategies reduced performance (AUCROC 0.711–0.746), though estimates remained clinically informative (2.8–3.1-fold improvement over baseline prevalence).

Conclusions

Machine learning models using routinely collected health records predicted intentional self-harm after psychiatric hospitalization with good discrimination and clinically meaningful precision–recall performance. A single 365-day model generalized well across horizons and demographic groups, suggesting that one broadly trained model may provide a pragmatic and scalable approach for clinical implementation. Competing Interest Statement In the past 3 years, Ronald C. Kessler was a consultant for Cambridge Health Alliance, Child Mind Institute, Massachusetts General Hospital, RallyPoint LLC., Sage Therapeutics, University of Michigan, and University of North Carolina. He has stock options in Cerebral Inc., Mirah, PYM (Prepare Your Mind), and Verisense Health. He owns an interest in Menssano LLC. Diego Palao has received grants and also served as consultant or advisor for Rovi, Janssen, and Lundbeck with no financial or other relationship relevant to the subject of this article. The other authors have no conflict of interest to declare. Funding Statement Philippe Mortier is supported by Miguel Servet grant CP21/00078 co–financed by the Instituto de Salud Carlos III (ISCIII) and co–funded by the European Union; grant PI22/00107 funded by ISCIII and co–funded by the European Union; and grant 202220–30–31 from the Fundació la Marató de TV3. Ana Portillo–Van Diest is supported by grant FI23/00004 funded by ISCIII and co–funded by the European Union. This work was further supported by ISCIII PI17/00521 (CODIRISC/CSRC–Epi) con fondos FEDER (Jordi Alonso); Spanish Ministry of Science and Innovation/ISCIII/FEDER PI21/01148 (Diego Palao); the Secretaria d'Universitats i Recerca del Departament d'Economia i Coneixement of the Generalitat de Catalunya AGAUR 2021 SGR 00624 (Jordi Alonso) and AGAUR 2021 SGR 01431 (Diego Palao); the CERCA program of the I3PT (Diego Palao); Centro de Investigación Biomédica en Red de Salud Mental, Instituto de Salud Carlos III (CIBERSAM, ISCIII), and grant CB06/02/0046 from the Centro de Investigación Biomédica en Red de Epidemiología y Salud Pública (CIBERESP), ISCIII (Jordi Alonso). Funding agencies for this study had no role in study design; in the collection, analysis, and interpretation of data; in the writing of the report; and in the decision to submit the paper for publication. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The research protocol received approval from the Parc de Salut Mar clinical research ethics committee (CEIC protocol 2017/7431/I), which granted a waiver for individual informed consent due to the fully anonymized nature of all registry data and comprehensive reidentification risk assessment procedures. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data availability The primary data, including healthcare, mortality, and administrative records, were provided by a third party, the Agency for Quality and Assessment of Catalonia (Agència de Qualitat i Avaluació Sanitàries de Catalunya; AQuAS), under the PADRIS (Programa d’Analítica de Dades per a la Recerca i la Innovació en Salut) framework. Access to these data is restricted, and must comply with PADRIS’s legal and ethical requirements. Interested parties can obtain access to the data, code, and documentation upon request, in accordance with the agreement’s provisions. The minimum dataset needed to replicate the analyses underlying this study, including the anonymized individual-level registry data, data dictionaries, and statistical code, are available upon reasonable request from the corresponding author (Gemma Vilagut: gvilagut{at}researchmar.net), provided (a) the purpose is to replicate our analysis and results without additional investigator support, (b) access is granted following approval of a brief proposal and the signing of a Data Access Agreement, and (c) the request aligns with the terms of our agreement with PADRIS/AQuAS and is approved by the PADRIS legal representative.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0