Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Background Pretraining electronic health record (EHR) data using language models by treating patient trajectories as natural language sentences has enhanced performance across various medical tasks. However, EHR pretraining models have never been utilized in adverse drug event (ADE) prediction. Methods A retrospective study was conducted on observational medical outcomes partnership (OMOP)-common data model (CDM) based EHR data from two separate tertiary hospitals. The data included patient information in various domains such as diagnosis, prescription, measurement, and procedure. For pretraining, codes were randomly masked, and the model was trained to infer the masked tokens utilizing preceding and following history. In this process, we introduced domain embedding (DE) to provide information about the domain of the masked token, preventing the model from finding codes from irrelevant domains. For qualitative analysis, we identified important features using the attention matrix from each finetuned model. Results 510,879 and 419,505 adult inpatients from two separate tertiary hospitals were included in internal and external datasets. EHR pretraining model with DE outperformed all the other baselines in all cohorts. For feature importance analysis, we demonstrated that the results were consistent with priorly reported background clinical knowledge. In addition to cohort-level interpretation, patient-level interpretation was also available. Conclusions CDM-based EHR pretraining model with DE is a proper model for various ADE prediction tasks. The results of the qualitative analysis with feature importance were consistent with background clinical knowledge. Plain language summary Patient history is like natural language; each medical code corresponds to a word, and the sequence of medical codes corresponds to a sentence. Language models learn from sentences by inferring masked words using remaining unmasked words, and language model-based EHR pretraining models comprehend medical context similarly. As several studies that have utilized EHR pretraining models have achieved great success in various tasks, we applied EHR pretraining models for adverse drug event (ADE) prediction. For better inference of the EHR pretraining model, we introduced domain embedding (DE) to provide a hint of the domain for each masked code. The model pretrained with DE performed the best in various ADE tasks, and regarding background clinical knowledge was well-reflected in our feature importance-based qualitative analysis.
Full text 4,581 characters · extracted from oa-doi-fallback · 4 sections · click to expand

Abstract

Background Pretraining electronic health record (EHR) data using language models by treating patient trajectories as natural language sentences has enhanced performance across various medical tasks. However, EHR pretraining models have never been utilized in adverse drug event (ADE) prediction.

Methods

A retrospective study was conducted on observational medical outcomes partnership (OMOP)-common data model (CDM) based EHR data from two separate tertiary hospitals. The data included patient information in various domains such as diagnosis, prescription, measurement, and procedure. For pretraining, codes were randomly masked, and the model was trained to infer the masked tokens utilizing preceding and following history. In this process, we introduced domain embedding (DE) to provide information about the domain of the masked token, preventing the model from finding codes from irrelevant domains. For qualitative analysis, we identified important features using the attention matrix from each finetuned model.

Results

510,879 and 419,505 adult inpatients from two separate tertiary hospitals were included in internal and external datasets. EHR pretraining model with DE outperformed all the other baselines in all cohorts. For feature importance analysis, we demonstrated that the results were consistent with priorly reported background clinical knowledge. In addition to cohort-level interpretation, patient-level interpretation was also available.

Conclusions

CDM-based EHR pretraining model with DE is a proper model for various ADE prediction tasks. The results of the qualitative analysis with feature importance were consistent with background clinical knowledge. Plain language summary Patient history is like natural language; each medical code corresponds to a word, and the sequence of medical codes corresponds to a sentence. Language models learn from sentences by inferring masked words using remaining unmasked words, and language model-based EHR pretraining models comprehend medical context similarly. As several studies that have utilized EHR pretraining models have achieved great success in various tasks, we applied EHR pretraining models for adverse drug event (ADE) prediction. For better inference of the EHR pretraining model, we introduced domain embedding (DE) to provide a hint of the domain for each masked code. The model pretrained with DE performed the best in various ADE tasks, and regarding background clinical knowledge was well-reflected in our feature importance-based qualitative analysis. Competing Interest Statement The authors have declared no competing interest. Funding Statement This study did not receive any funding. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Institutional Review Board (IRB) of Seoul National University Hospital (IRB approval No. 2406-060-1543) approved the study with a waiver of informed consent, considering that our study used retrospective and observational EHR data. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Footnotes We revised IRB approval number. No additional changes in this manuscript. Data Availability All data produced in the present study are available upon reasonable request to the authors. Data availability The datasets of SNUH and AUMC are not publicly available due to patient privacy and are available from the corresponding author on reasonable request and IRB approvals.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00