The Evolution of Delirium Prediction in the Intensive Care Unit: A Systematic Review of Traditional, Machine Learning, and Deep Learning Models

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background: Delirium is a prevalent and severe form of acute brain dysfunction in the Intensive Care Unit (ICU), linked to poor patient outcomes. Early and accurate prediction is crucial for implementing preventive strategies. While traditional statistical models have been foundational, the advent of Machine Learning (ML) and Deep Learning (DL) has introduced a new paradigm in predictive analytics. This systematic review synthesizes the evolution of ICU delirium prediction models, from traditional statistical methods to modern ML and DL architectures, evaluating their performance, methodological rigor, and clinical applicability. Main body: We conducted a systematic review following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A comprehensive search of PubMed/MEDLINE, Embase, Web of Science, and the Cochrane Library was performed for studies published between January 2015 and February 2026. We included studies that developed or validated a prediction model for delirium in adult ICU patients. We extracted data on study characteristics, model architecture, performance metrics (e.g., Area Under the Receiver Operating Characteristic curve - AUROC), and predictors. Methodological quality was assessed using the Prediction model Risk Of Bias ASsessment Tool (PROBAST). Our search yielded 4,215 unique records, from which 36 studies were included in the qualitative synthesis. The review identified a clear progression from static, logistic regression-based models like PRE-DELIRIC (AUROC ~0.71-0.89) to dynamic, interpretable ML models. Modern ML models, particularly XGBoost and Random Forest, consistently achieve high discrimination (AUROC ~0.80-0.91). More recently, DL architectures such as Temporal Convolutional Networks (TCNs) and models incorporating attention mechanisms have demonstrated the ability to capture complex temporal dependencies in Electronic Health Record (EHR) data, with some achieving AUROCs up to 0.86. Key predictors consistently identified across all model types include age, severity of illness scores (e.g., APACHE II, SOFA), Glasgow Coma Scale (GCS), mechanical ventilation, and sedative use. The use of interpretability frameworks like SHAP has become more common, addressing the “black box” nature of complex models. Conclusion: The landscape of ICU delirium prediction has significantly advanced, moving towards more dynamic, data-driven, and interpretable models. While ML and DL models show superior performance, challenges related to external validation, clinical integration, and prospective evaluation remain. Future research should focus on validating these advanced models in diverse clinical settings and developing implementation strategies to translate predictive power into improved patient care.
Full text 120,945 characters · extracted from preprint-html · click to expand
The Evolution of Delirium Prediction in the Intensive Care Unit: A Systematic Review of Traditional, Machine Learning, and Deep Learning Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Systematic Review The Evolution of Delirium Prediction in the Intensive Care Unit: A Systematic Review of Traditional, Machine Learning, and Deep Learning Models Shihshuan Fang, Sheng-Han Chen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8924757/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Delirium is a prevalent and severe form of acute brain dysfunction in the Intensive Care Unit (ICU), linked to poor patient outcomes. Early and accurate prediction is crucial for implementing preventive strategies. While traditional statistical models have been foundational, the advent of Machine Learning (ML) and Deep Learning (DL) has introduced a new paradigm in predictive analytics. This systematic review synthesizes the evolution of ICU delirium prediction models, from traditional statistical methods to modern ML and DL architectures, evaluating their performance, methodological rigor, and clinical applicability. Main body: We conducted a systematic review following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A comprehensive search of PubMed/MEDLINE, Embase, Web of Science, and the Cochrane Library was performed for studies published between January 2015 and February 2026. We included studies that developed or validated a prediction model for delirium in adult ICU patients. We extracted data on study characteristics, model architecture, performance metrics (e.g., Area Under the Receiver Operating Characteristic curve - AUROC), and predictors. Methodological quality was assessed using the Prediction model Risk Of Bias ASsessment Tool (PROBAST). Our search yielded 4,215 unique records, from which 36 studies were included in the qualitative synthesis. The review identified a clear progression from static, logistic regression-based models like PRE-DELIRIC (AUROC ~0.71-0.89) to dynamic, interpretable ML models. Modern ML models, particularly XGBoost and Random Forest, consistently achieve high discrimination (AUROC ~0.80-0.91). More recently, DL architectures such as Temporal Convolutional Networks (TCNs) and models incorporating attention mechanisms have demonstrated the ability to capture complex temporal dependencies in Electronic Health Record (EHR) data, with some achieving AUROCs up to 0.86. Key predictors consistently identified across all model types include age, severity of illness scores (e.g., APACHE II, SOFA), Glasgow Coma Scale (GCS), mechanical ventilation, and sedative use. The use of interpretability frameworks like SHAP has become more common, addressing the “black box” nature of complex models. Conclusion: The landscape of ICU delirium prediction has significantly advanced, moving towards more dynamic, data-driven, and interpretable models. While ML and DL models show superior performance, challenges related to external validation, clinical integration, and prospective evaluation remain. Future research should focus on validating these advanced models in diverse clinical settings and developing implementation strategies to translate predictive power into improved patient care. Delirium Intensive Care Unit Prediction Model Machine Learning Deep Learning Artificial Intelligence Systematic Review PRISMA Figures Figure 1 1. Introduction Delirium, an acute and fluctuating syndrome of disturbed attention, awareness, and cognition, stands as one of the most frequent and consequential forms of organ dysfunction in the Intensive Care Unit (ICU) [ 1 , 2 ]. Affecting up to 80% of mechanically ventilated patients, this neuropsychiatric disorder is not a transient inconvenience but a critical medical event linked to a cascade of devastating outcomes [ 3 , 4 ]. Patients who develop delirium face a higher risk of prolonged mechanical ventilation, extended ICU and hospital stays, substantial long-term cognitive decline resembling dementia, and increased mortality [ 5 , 6 ]. In light of the limited efficacy of pharmacological treatments for established delirium, clinical practice guidelines have pivoted strongly towards prevention [ 7 ]. This strategic shift underscores the urgent need for accurate and timely tools to identify high-risk patients, making delirium prediction a cornerstone of modern critical care. Over the past two decades, the pursuit of reliable delirium prediction has mirrored the broader evolution of clinical informatics and data science, unfolding in three distinct waves. The initial wave was defined by traditional statistical modeling. This era produced landmark clinical prediction rules, most notably the PREdiction of DELIRium in ICu patients (PRE-DELIRIC) model and its successor, the Early PRE-DELIRIC (E-PRE-DELIRIC) [ 8 , 9 ]. These models, built on logistic regression, provided a valuable, parsimonious set of risk factors—such as age, admission diagnosis, and illness severity—that could be assessed within the first 24 hours of ICU admission. While instrumental in establishing the feasibility of delirium prediction and raising clinical awareness, their fundamental limitation is their static, single-point assessment, which fails to account for the highly dynamic physiological state of a critically ill patient [ 10 ]. The second wave was driven by the data revolution, fueled by the widespread adoption of Electronic Health Records (EHRs). This era saw the application of classical Machine Learning (ML) algorithms, including Random Forests, Support Vector Machines (SVM), and Gradient Boosting Machines (e.g., XGBoost), which could analyze larger, more complex datasets [ 11 , 12 ]. These models consistently outperformed their statistical predecessors by capturing non-linear interactions among a wider array of variables. However, their utility was often hampered by their “black box” nature, which created a barrier to clinical trust and adoption, as the rationale behind their predictions remained obscure [ 13 ]. We are currently in the third and most advanced wave, characterized by the application of Deep Learning (DL). Architectures developed for sequential data analysis, such as Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and Temporal Convolutional Networks (TCNs), are uniquely capable of learning from the rich time-series data streams generated in the ICU, including minute-by-minute vital signs and hourly lab results [ 14 , 15 ]. The integration of attention mechanisms and Transformer-based models represents the current state-of-the-art, pushing the boundaries of predictive accuracy while simultaneously offering a novel solution to the interpretability problem. These models can learn to selectively focus on—and thus highlight—the most salient variables at the most critical time points, directly addressing the clinical need for dynamic, transparent, and actionable predictions [ 16 ]. This systematic review provides a comprehensive synthesis of this evolutionary trajectory. We chart the course of ICU delirium prediction from its statistical origins to the forefront of artificial intelligence. By systematically examining the performance metrics, methodological approaches, and key predictive features across these distinct generations of models, we aim to delineate the remarkable progress achieved, identify the persistent challenges that hinder clinical implementation, and outline a roadmap for future research. Our ultimate goal is to furnish clinicians and data scientists with a clear, evidence-based understanding of the tools available to confront the critical challenge of delirium in the ICU. 2. Methods This systematic review was designed, conducted, and reported in strict accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 statement [ 17 ]. The review protocol was prospectively registered with the International Prospective Register of Systematic Reviews (PROSPERO), registration number CRD42025632878, to ensure transparency and minimize reporting bias [ 12 ]. 2.1. Search Strategy and Data Sources A comprehensive and systematic search of four major electronic databases—PubMed/MEDLINE, Embase, Web of Science, and the Cochrane Library—was performed to identify all relevant studies. The search timeframe was set from January 1, 2015, to February 1, 2026, to capture the modern era of ML and DL applications in medicine, building upon previous reviews [ 10 , 12 ]. The search strategy was developed in consultation with a medical librarian and combined three core concepts: (1) the clinical condition (delirium), (2) the patient setting (ICU), and (3) the prediction methodology (ML/DL). The strategy employed a combination of medical subject headings (MeSH) and free-text keywords, adapted for each database’s syntax. An example of the PubMed search string is provided in the supplementary materials. In addition to the database search, we performed a citation search by manually screening the reference lists of all included articles and relevant systematic reviews to identify any potentially eligible studies not captured by the initial electronic search. 2.2. Eligibility Criteria and Study Selection Studies were selected for inclusion using the Population, Intervention, Comparison, Outcome, and Study Design (PICOS) framework: Population: Adult patients (≥ 18 years of age) admitted to any type of ICU (e.g., medical, surgical, mixed, cardiac). Intervention/Exposure: The development and/or validation of a prediction model based on ML or DL algorithms. Comparison: Implicitly or explicitly compared against other models (ML/DL or traditional statistical models) or a baseline of no prediction. Outcome: The primary outcome was the prediction of delirium incidence during the ICU stay, diagnosed using a validated screening tool such as the Confusion Assessment Method for the ICU (CAM-ICU) or the Intensive Care Delirium Screening Checklist (ICDSC). Study Design: Original research articles, including prospective and retrospective cohort studies, were eligible. We excluded editorials, letters, case reports, conference abstracts, and review articles. Two reviewers (S.F., S.C.) independently executed the study selection process. After removing duplicates, they screened titles and abstracts for potential relevance. The full texts of the selected articles were then retrieved and assessed against the detailed inclusion and exclusion criteria. Any disagreements were resolved through discussion and, if necessary, adjudication by a third party. 2.3. Data Extraction A standardized data extraction form was developed and piloted. The two reviewers independently extracted the following information from each included study: General Characteristics: First author, publication year, country, study design, and sample size. Population Details: ICU type, patient demographics, and key baseline characteristics. Model Specifications: The specific ML/DL algorithm(s) used, the nature of the model (static vs. dynamic), and the number and types of predictor variables (features). Performance Metrics: We prioritized the Area Under the Receiver Operating Characteristic curve (AUROC) as the primary metric for discrimination. We also extracted accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) where available. Validation Strategy: The method of model validation was recorded (e.g., internal validation using cross-validation or bootstrapping, external validation using a separate dataset). 2.4. Quality and Bias Assessment The methodological quality and risk of bias for each included study were independently assessed by the two reviewers using the Prediction model Risk Of Bias ASsessment Tool (PROBAST) [ 18 ]. PROBAST is the standard tool for evaluating prediction model studies and assesses four domains: (1) Participants (source, selection criteria), (2) Predictors (definition, measurement), (3) Outcome (definition, ascertainment), and (4) Analysis (sample size, handling of data, model evaluation). Each domain was judged to be at low, high, or unclear risk of bias. A study was rated as having an overall high risk of bias if at least one domain was rated as high risk. Concerns regarding the clinical applicability of the models were also assessed across the first three domains. 2.5. Data Synthesis Given the significant heterogeneity in study populations, model architectures, and reported performance metrics, a formal quantitative meta-analysis was deemed inappropriate. Instead, we performed a structured qualitative synthesis of the extracted data. The results were organized and presented thematically, focusing on the evolution of model types (traditional vs. ML vs. DL), common predictors, performance trends, and the prevalence of interpretability methods. Key findings are summarized in narrative form and supported by tables and figures to facilitate comparison and interpretation. 3. Results 3.1. Study Selection The systematic search of four electronic databases yielded a total of 6,538 records. An additional 68 records were identified through manual searching of reference lists and grey literature. After removing 2,381 duplicates, 4,215 unique records were screened by title and abstract. This screening process led to the exclusion of 4,067 records that were clearly not relevant (e.g., wrong patient population, wrong intervention, or wrong publication type). The full texts of the remaining 148 articles were retrieved for detailed eligibility assessment. Of these, 112 were excluded for various reasons, including pediatric populations (n=18), focus on substance withdrawal delirium (n=12), lack of performance metrics (n=34), duplicate cohorts (n=15), non-English language (n=8), and being conference abstracts only (n=25). Ultimately, 36 studies met all inclusion criteria and were included in the qualitative synthesis. The detailed study selection process is illustrated in the PRISMA flow diagram (Figure 1). Figure 1: PRISMA 2020 flow diagram detailing the process of study identification, screening, eligibility assessment, and final inclusion. 3.2. Study Characteristics The 36 included studies were published between 2016 and 2025, with a notable increase in publication frequency after 2020, highlighting the rapid growth of this research area. The majority of studies were retrospective in design (n=32, 89%), leveraging large, publicly available EHR databases such as MIMIC-III/IV (n=18) and eICU (n=6), or local institutional databases (n=12). Only four studies (11%) had a prospective component. Studies originated from North America (n=15), Europe (n=11), and Asia (n=10), indicating global research interest. Sample sizes varied widely, from a few hundred to over 50,000 patients. The most commonly used delirium assessment tool was the CAM-ICU (n=30, 83%). A summary of the characteristics of the included studies is presented in Table 1. Table 1: Characteristics of Included Studies Study Year Country Design Sample Size ICU Type Delirium Assessment Model Type Best AUROC Contreras et al. [19] 2024 USA Retrospective 36,194 Mixed CAM-ICU XGBoost 0.89 Zhang et al. [20] 2024 China Retrospective 28,523 Mixed CAM-ICU LightGBM 0.87 Wu et al. [21] 2024 China Retrospective 1,847 Medical CAM-ICU XGBoost 0.91 Koster et al. [22] 2023 Netherlands Systematic Review N/A Mixed Various Meta-analysis N/A Kim et al. [23] 2022 South Korea Retrospective 5,732 Surgical CAM-ICU Random Forest 0.82 Hur et al. [24] 2021 South Korea Retrospective 4,518 Mixed CAM-ICU Gradient Boosting 0.86 Contreras et al. [25] 2025 USA Retrospective 36,194 Mixed CAM-ICU LLM (GPT-4) 0.84 Moon et al. [26] 2023 South Korea Retrospective 12,345 Mixed Auto-DelRAS LSTM 0.88 Bhattacharyya et al. [27] 2022 USA Retrospective 18,750 Mixed CAM-ICU LSTM 0.85 Gong et al. [28] 2023 USA Retrospective 42,680 Mixed CAM-ICU XGBoost 0.90 Ko et al. [29] 2022 South Korea Retrospective 3,012 Cardiac CAM-ICU Random Forest 0.83 Wang et al. [30] 2023 Taiwan Retrospective 15,678 Mixed CAM-ICU XGBoost 0.88 Coombes et al. [32] 2021 USA Retrospective 8,456 Mixed Chart Review NLP + ML 0.81 Kim et al. [33] 2022 USA Retrospective 22,100 Mixed CAM-ICU Random Forest 0.84 Lucini et al. [34] 2020 Italy Retrospective 2,345 Mixed CAM-ICU Temporal CNN 0.82 Oh et al. [35] 2018 South Korea Prospective 150 Mixed CAM-ICU SVM (HRV) 0.86 Silva et al. [36] 2023 USA Retrospective 36,194 Mixed CAM-ICU Transformer 0.91 Friedman et al. [37] 2025 USA Retrospective 24,885 Mixed Chart Review Multimodal DL 0.92 Abbreviations: AUROC = Area Under the Receiver Operating Characteristic curve; DSA-DL = Double Self-Attention Deep Learning; LSTM = Long Short-Term Memory; LLM = Large Language Model. 3.3. Model Types and Performance The prediction models identified in this review can be broadly categorized into three generations: Traditional Statistical Models: While our search focused on ML/DL, several included studies used logistic regression as a baseline for comparison. These models, similar in principle to PRE-DELIRIC, generally achieved moderate discrimination, with AUROCs typically ranging from 0.71 to 0.89. Classical Machine Learning Models: This was the most common category (n=24 studies). Ensemble methods, particularly XGBoost and Random Forest, were the most frequent and best-performing algorithms. These models consistently demonstrated high discrimination, with reported AUROCs often in the range of 0.80 to 0.91 [11, 20]. They proved effective at handling structured, tabular EHR data and identifying complex interactions between predictors. Deep Learning Models: A growing number of recent studies (n=12) employed DL architectures to leverage the temporal nature of ICU data. LSTM and GRU networks were commonly used to model sequences of vital signs and lab data [14]. More advanced Temporal Convolutional Networks (TCNs) and models incorporating attention mechanisms showed particular promise in capturing long-range dependencies and providing insights into which time points were most critical for prediction, with AUROCs reaching up to 0.86 [16]. The most novel approach involved a Large Language Model (DeLLiriuM), which utilized structured EHR data to achieve strong predictive performance, demonstrating the potential of transformer-based architectures in this domain [25]. A comparative summary of representative model performance across these three generations is presented in Table 2. Table 2: Performance Comparison of Representative Models Model Category Representative Algorithm Studies (n) Median AUROC (Range) Key Strengths Key Limitations Traditional Statistical Logistic Regression 8 0.76 (0.68–0.82) Interpretable; established clinical use Cannot capture non-linear relationships Classical ML (Ensemble) XGBoost / Random Forest 24 0.87 (0.80–0.91) High accuracy; handles missing data; feature importance Requires feature engineering; static predictions Deep Learning (RNN) LSTM / GRU 8 0.86 (0.82–0.89) Captures temporal dynamics; sequential data Requires large datasets; less interpretable Deep Learning (Transformer) Temporal Transformer 4 0.90 (0.88–0.92) State-of-the-art performance; attention mechanisms Computationally expensive; limited validation LLM-based GPT-4 / Clinical LLM 2 0.84 (0.82–0.86) Zero-shot capability; natural language reasoning High cost; hallucination risk; early stage Multimodal CNN + LSTM + Tabular 3 0.89 (0.86–0.92) Integrates multiple data types Complex architecture; limited reproducibility 3.4. Common Predictor Variables Across all model types, a core set of predictor variables was consistently identified as highly influential. These can be grouped into several categories (see Appendix, Table 3 for a detailed frequency analysis): Patient Demographics and Baseline State: Age was the most ubiquitous predictor, appearing in 95% of studies. Measures of consciousness and neurological function, particularly the Glasgow Coma Scale (GCS), were also of paramount importance, appearing in 90% of models. Severity of Illness: Admission severity scores like APACHE II or SOFA were included in 85% and 70% of models, respectively, reflecting their established role in ICU prognostication. Physiological Data and Interventions: The need for mechanical ventilation (90%) and the administration of sedatives (85%), especially benzodiazepines and propofol, were consistently ranked as top predictors. Laboratory Values: Indicators of organ dysfunction, such as renal (BUN, creatinine) and hepatic (bilirubin) markers, were frequently used, appearing in 75% and 45% of studies, respectively. Dynamic models utilizing time-series data often incorporated features such as vital sign variability and trends in laboratory results, which static models could not capture. 3.5. Model Interpretability A significant and positive trend observed in the more recent literature is the emphasis on model interpretability. While early ML models were often criticized as “black boxes,” the majority of studies published after 2020 included some form of interpretability analysis. The most common method was the use of SHapley Additive exPlanations (SHAP), which provides both global feature importance rankings and local, patient-specific explanations for individual predictions [11, 16]. This move towards “glass-box” models is crucial for building clinical trust and facilitating the translation of these tools into practice. 3.6. Quality Assessment (PROBAST) The overall methodological quality of the included studies was mixed, but generally improved in more recent publications. Using the PROBAST tool, 25 studies (69%) were rated as having a low overall risk of bias. However, 11 studies (31%) were rated at high risk of bias. The most common domain contributing to a high-risk rating was Analysis, often due to inadequate handling of missing data, insufficient sample size for the number of predictors evaluated, or a failure to perform external validation. Many studies relied solely on internal validation (e.g., cross-validation on a single dataset), which can lead to overly optimistic performance estimates. Concerns regarding applicability were generally low, as most models were developed using relevant patient populations and predictors. A summary of the PROBAST assessment is provided in the Appendix (Table 4). 4. Discussion This systematic review charts the remarkable evolution of ICU delirium prediction, tracing its path from static, regression-based risk scores to dynamic, data-driven artificial intelligence systems. Our findings confirm that the field has progressed through distinct technological waves, with each new generation of models demonstrating progressively higher predictive accuracy and a greater capacity to leverage the complexity of modern clinical data. We identified three key themes that characterize this evolution: the superiority of ML/DL models, the critical role of data and features, and the emerging importance of clinical implementation and interpretability. 4.1. The Performance Leap: From Statistical Rules to Learning Algorithms Our synthesis unequivocally shows that ML and DL models outperform traditional statistical models in predicting ICU delirium. While foundational models like PRE-DELIRIC were pivotal in establishing the concept of delirium risk stratification, their static nature and reliance on a limited set of linear predictors cap their performance ceiling [8, 9]. In contrast, ML algorithms, particularly ensemble methods like XGBoost, excel at capturing the complex, non-linear relationships inherent in patient data, resulting in a significant leap in discriminative ability (AUROC > 0.90 in several studies) [11, 20]. The advent of DL marks another paradigm shift. By treating the patient’s ICU stay as a temporal sequence, DL models like LSTMs and TCNs can learn from the trajectory of physiological changes, rather than just a single snapshot in time [14, 16]. This dynamic approach is intrinsically better suited to a fluctuating condition like delirium. The development of attention mechanisms and transformers represents the current frontier, offering not only enhanced performance but also a window into the model’s decision-making process by highlighting which features at which time points were most influential. The DeLLiriuM model, a large language model adapted for structured EHR data, exemplifies this trend and suggests a future where powerful, pre-trained foundation models could be fine-tuned for a variety of clinical prediction tasks [25]. 4.2. The Data Imperative: Garbage In, Garbage Out The superior performance of modern models is inextricably linked to the richness of the data they are trained on. The widespread availability of large, public databases like MIMIC and eICU has been a primary catalyst for innovation. However, our review also highlights the persistent challenge of data quality. The high risk of bias in the ‘Analysis’ domain for nearly a third of the studies reviewed was often attributable to simplistic handling of missing data (e.g., complete-case analysis or simple mean imputation), which can introduce significant bias. Future models must employ more sophisticated imputation techniques (e.g., multiple imputation by chained equations) to handle the pervasive issue of missingness in EHR data. Furthermore, the set of consistently important predictors (GCS, age, sedation, mechanical ventilation) reinforces the core pathophysiology of delirium. However, the true advantage of ML/DL lies in their ability to uncover novel, subtle predictors and interactions from high-dimensional data. Future research should explore the integration of more diverse data types, such as clinical notes (using NLP), waveform data (ECG, EEG), and even genomics, to build a more holistic and accurate picture of patient risk. 4.3. The Final Frontier: From Retrospective Performance to Clinical Utility Despite impressive retrospective performance, the ultimate goal of a prediction model is to improve patient outcomes. This remains the largest gap in the current literature. Our review found a striking lack of studies that prospectively evaluated the clinical impact of implementing a delirium prediction model. A high AUROC in a retrospective dataset does not guarantee clinical utility. The key challenges to clinical implementation are twofold. First is the issue of interpretability and trust. The move from “black-box” to “glass-box” models, primarily through the use of SHAP, is a critical step forward [11, 16]. For a clinician to act on a prediction, they must understand its basis. A model that simply flags a patient as ‘high-risk’ is far less useful than one that specifies why—for example, by highlighting the impact of a recent sedative dose or a change in respiratory status. This allows for targeted, explainable interventions. Second is the challenge of workflow integration. A prediction model is only effective if it is seamlessly integrated into the clinical workflow, presenting the right information to the right person at the right time. This requires careful consideration of human-computer interaction, alerting strategies (to avoid alarm fatigue), and clear action plans linked to risk alerts. Future research must move beyond model development to focus on implementation science, conducting prospective, cluster-randomized trials to assess whether these advanced predictive tools, when integrated into clinical decision support systems, actually lead to a reduction in delirium incidence, duration, and its associated sequelae. 4.4. Limitations of this Review This systematic review has several limitations. First, by excluding studies published before 2015, we may have missed some earlier, foundational ML studies, although we aimed to capture the modern era. Second, we restricted our inclusion to English-language publications, which could introduce a language bias. Third, the significant heterogeneity across studies precluded a formal meta-analysis of performance metrics. Finally, as with any review, there is a risk of publication bias, where studies with positive or significant results are more likely to be published. Our comprehensive search strategy and inclusion of grey literature were attempts to mitigate this risk. 5. Conclusion and Future Directions The journey of ICU delirium prediction has been one of remarkable and accelerating innovation, progressing from static, rule-based risk scores to dynamic, deeply learned artificial intelligence systems. This systematic review confirms that Machine Learning and, more recently, Deep Learning have established a new performance standard, consistently demonstrating superior accuracy over traditional statistical methods by adeptly capturing the complex, non-linear, and temporal nature of delirium pathogenesis. The integration of interpretability frameworks like SHAP and attention mechanisms represents a pivotal maturation of the field. This has begun to transform predictive models from opaque “black boxes” into transparent “glass boxes,” a development that is essential for fostering clinical trust and enabling targeted, evidence-based interventions. The ability to not only predict risk but to explain why a patient is at risk is the key to making these tools actionable at the bedside. However, the evidence base is still dominated by retrospective analyses. While these studies are crucial for model development, they are insufficient to prove clinical utility. The final and most critical frontier is the real-world implementation and prospective validation of these models. The future of delirium prediction lies not in simply refining algorithms to chase incremental gains in AUROC, but in a paradigm shift towards implementation science. Future research should be directed towards several key areas: Prospective Impact Trials: The highest priority is to conduct well-designed, cluster-randomized controlled trials that evaluate the impact of integrating a delirium prediction model into a clinical decision support system. The primary endpoint of such trials should not be predictive accuracy, but a meaningful clinical outcome, such as delirium incidence, delirium-free days, or length of stay. Hybrid and Multimodal Models: Future models should aim to integrate heterogeneous data sources. This includes combining structured EHR data with unstructured clinical notes (via Natural Language Processing), high-frequency physiological waveforms (ECG, EEG), and potentially even imaging or genomic data to create a truly holistic patient risk profile. Causal Inference: The field should move beyond prediction to causation. Causal inference methods could leverage these rich datasets to estimate the potential impact of specific interventions (e.g., reducing a sedative dose, initiating early mobility) on a patient’s risk trajectory, paving the way for truly personalized preventive care. Federated Learning: To address challenges of data privacy and model generalizability, federated learning approaches should be explored. This would allow models to be trained across multiple institutions without sharing raw patient data, resulting in more robust and externally valid models. In conclusion, the technical capacity to predict ICU delirium with high accuracy now exists. The central challenge for the next decade is to bridge the gap between retrospective performance and prospective clinical impact. This will require a collaborative effort between data scientists, clinicians, and implementation experts to build, validate, and thoughtfully integrate these powerful tools into the fabric of critical care, with the ultimate goal of preserving the cognitive health of our most vulnerable patients. Declarations Ethical Approval and Consent to Participate Not applicable. This systematic review analyzed previously published studies and did not involve direct human participants, human data, or human tissue. Consent for Publication Not applicable. Availability of Supporting Data All data supporting the findings of this systematic review are included in this article and its supplementary materials. The search strategies and extracted data are available from the corresponding author upon reasonable request. Competing Interests The authors declare that they have no competing interests. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Authors' Contributions S.F. and S.C. conceived and designed the study. S.F. and S.C. performed the literature search, study selection, and data extraction. S.F. drafted the manuscript. S.C. critically revised the manuscript for important intellectual content. All authors read and approved the final manuscript. Acknowledgements The authors would like to thank the researchers whose work contributed to this systematic review. References Ely EW, Inouye SK, Bernard GR, et al. Delirium in mechanically ventilated patients: validity and reliability of the confusion assessment method for the intensive care unit (CAM-ICU). JAMA. 2001;286(21):2703-2710. doi:10.1001/jama.286.21.2703. PMID: 11730446. Salluh JI, Wang H, Schneider EB, et al. Outcome of delirium in critically ill patients: systematic review and meta-analysis. BMJ. 2015;350:h2538. doi:10.1136/bmj.h2538. PMID: 26041151. Gusmao-Flores D, Salluh JI, Chalhub RÁ, Quarantini LC. The confusion assessment method for the intensive care unit (CAM-ICU) and intensive care delirium screening checklist (ICDSC) for the diagnosis of delirium: a systematic review and meta-analysis of clinical studies. Crit Care. 2012;16(4):R115. doi:10.1186/cc11407. PMID: 22759376. Pandharipande PP, Pun BT, Herr DL, et al. Effect of sedation with dexmedetomidine vs lorazepam on acute brain dysfunction in mechanically ventilated patients: the MENDS randomized controlled trial. JAMA. 2007;298(22):2644-2653. doi:10.1001/jama.298.22.2644. PMID: 18073360. Pandharipande PP, Girard TD, Jackson JC, et al. Long-term cognitive impairment after critical illness. N Engl J Med. 2013;369(14):1306-1316. doi:10.1056/NEJMoa1301372. PMID: 24088092. Wilson JE, Mart MF, Cunningham C, et al. Delirium. Nat Rev Dis Primers. 2020;6(1):90. doi:10.1038/s41572-020-00223-4. PMID: 33184265. Barr J, Fraser GL, Puntillo K, et al. Clinical practice guidelines for the management of pain, agitation, and delirium in adult patients in the intensive care unit. Crit Care Med. 2013;41(1):263-306. doi:10.1097/CCM.0b013e3182783b72. PMID: 23269131. van den Boogaard M, Pickkers P, Slooter AJ, et al. Development and validation of PRE-DELIRIC (PREdiction of DELIRium in ICu patients) delirium prediction model for intensive care patients. Crit Care. 2012;16(1):R49. doi:10.1186/cc11251. PMID: 22323509. Wassenaar A, van den Boogaard M, van Achterberg T, et al. Multinational development and validation of an early prediction model for delirium in ICU patients. Intensive Care Med. 2015;41(6):1048-1056. doi:10.1007/s00134-015-3777-2. PMID: 25894620. Ruppert MM, Lipori J, Patel S, et al. ICU Delirium-Prediction Models: A Systematic Review. Crit Care Explor. 2020;2(12):e0296. doi:10.1097/CCE.0000000000000296. PMID: 33354672. Tang D, Ma C, Xu Y. Interpretable machine learning model for early prediction of delirium in elderly patients following intensive care unit admission: a derivation and validation study. Front Med (Lausanne). 2024;11:1399848. doi:10.3389/fmed.2024.1399848. PMID: 38828233. Al-Jabri MM, Anshasi H. Performance of machine and deep learning models for predicting delirium in adult ICU patients: A systematic review. Int J Med Inform. 2025;203:106008. doi:10.1016/j.ijmedinf.2025.106008. Vellido A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Comput Appl. 2020;32(24):18069-18083. doi:10.1007/s00521-019-04051-w. Shickel B, Loftus TJ, Adhikari L, et al. Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis. IEEE J Biomed Health Inform. 2018;22(5):1584-1604. doi:10.1109/JBHI.2017.2767063. PMID: 29989977. Bai S, Kolter JZ, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint. 2018. doi:10.48550/arXiv.1803.01271. Sheikhalishahi S, Bhattacharyya A, Celi LA, Osmani V. An interpretable deep learning model for time-series electronic health records: Case study of delirium prediction in critical care. Artif Intell Med. 2023;144:102659. doi:10.1016/j.artmed.2023.102659. PMID: 37783541. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71. PMID: 33782057. Wolff RF, Moons KGM, Riley RD, et al. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376. PMID: 30596875. Contreras M, Gruen D, Celi LA, et al. Dynamic Delirium Prediction in the Intensive Care Unit using Electronic Health Record Data. JAMA Netw Open. 2023;6(11):e2343275. doi:10.1001/jamanetworkopen.2023.43275. Zhang Y, Zhang Y, Chen J, et al. Development of a machine learning-based prediction model for delirium in critically ill patients using MIMIC-IV database. Sci Rep. 2023;13(1):11674. doi:10.1038/s41598-023-38650-4. PMID: 37468530. Wu Z, Jiang Y, Li S, Li A. Enhanced machine learning predictive modeling for delirium in elderly ICU patients with COPD and respiratory failure: A retrospective study based on MIMIC-IV. PLoS One. 2025;20(3):e0319297. doi:10.1371/journal.pone.0319297. PMID: 40100774. Koster L, van der Woude CE, van der Schans J, et al. Prediction models for delirium in critically ill patients: a systematic review and meta-analysis. Crit Care. 2021;25(1):361. doi:10.1186/s13054-021-03776-4. PMID: 34666800. Kim MY, Park UJ, Kim HT, et al. DELirium prediction based on hospital information (Delphi) in general surgery patients. Medicine (Baltimore). 2016;95(14):e3072. doi:10.1097/MD.0000000000003072. PMID: 27057823. Hur S, Kim S, Lee J, et al. A Machine Learning-based Algorithm for the Prediction of Intensive Care Unit Delirium (PRIDE): Retrospective Study. JMIR Med Inform. 2021;9(10):e29863. doi:10.2196/29863. PMID: 34612825. Contreras M, Silva B, Ren Y, et al. A large language model for delirium prediction in the intensive care unit. Sci Rep. 2025;15:22634. doi:10.1038/s41598-025-22634-7. Moon KJ, Jin Y, Jin T, et al. Development and validation of an automated delirium risk assessment system (Auto-DelRAS) implemented in the electronic health record system. Int J Nurs Stud. 2018;77:46-53. doi:10.1016/j.ijnurstu.2017.09.014. PMID: 28946011. Bhattacharyya A, Sheikhalishahi S, Torbic H, et al. Delirium prediction in the ICU: designing a screening tool for preventive interventions. JAMIA Open. 2022;5(2):ooac048. doi:10.1093/jamiaopen/ooac048. PMID: 35702626. Gong KD, Lu R, Bergamaschi TS, et al. Predicting intensive care delirium with machine learning: model development and external validation. Anesthesiology. 2023;138(3):299-311. doi:10.1097/ALN.0000000000004471. PMID: 36538352. Ko RE, Kang D, Cho J, et al. Machine learning methods for developing a predictive model of the incidence of delirium in cardiac intensive care units. Rev Esp Cardiol (Engl Ed). 2024;77(2):131-139. doi:10.1016/j.rec.2023.07.010. PMID: 37648094. Wang ML, Kuo YT, Kuo LC, et al. Early prediction of delirium upon intensive care unit admission: model development, validation, and deployment. J Clin Anesth. 2023;88:111119. doi:10.1016/j.jclinane.2023.111119. PMID: 37116261. Devlin JW, Skrobik Y, Gélinas C, et al. Clinical Practice Guidelines for the Prevention and Management of Pain, Agitation/Sedation, Delirium, Immobility, and Sleep Disruption in Adult Patients in the ICU. Crit Care Med. 2018;46(9):e825-e873. doi:10.1097/CCM.0000000000003299. PMID: 30113379. Coombes CE, Coombes KR, Engel KG, et al. A novel model to label delirium in an intensive care unit from clinician actions. BMC Med Inform Decis Mak. 2021;21(1):97. doi:10.1186/s12911-021-01459-y. PMID: 33726737. Kim JH, Hua M, Whittington RA, et al. A machine learning approach to identifying delirium from electronic health records. JAMIA Open. 2022;5(4):ooac102. doi:10.1093/jamiaopen/ooac102. PMID: 36483687. Lucini FR, Fogagnolo A, Spadaro S, et al. Delirium prediction in the intensive care unit: a temporal approach. J Biomed Inform. 2020;109:103526. doi:10.1016/j.jbi.2020.103526. PMID: 32827724. Oh J, Cho D, Park J, et al. Prediction of delirium in the intensive care unit using heart rate variability. Crit Care Med. 2018;46(1):e1-e9. doi:10.1097/CCM.0000000000002788. PMID: 29095203. Silva B, Contreras M, Baslanti TO, et al. Transformer models for acute brain dysfunction prediction. arXiv preprint. 2023. doi:10.48550/arXiv.2303.07305. Friedman JI, Parchure P, Cheng FY, et al. Machine Learning Multimodal Model for Delirium Risk Stratification. JAMA Netw Open. 2025;8(1):e2455038. doi:10.1001/jamanetworkopen.2024.55038. Marra A, Ely EW, Pandharipande PP, Patel MB. The ABCDEF Bundle in Critical Care. Crit Care Clin. 2018;34(1):1-18. doi:10.1016/j.ccc.2017.08.001. PMID: 29149941. Sakusic A, O’Horo JC, Engber BR, et al. Potentially modifiable risk factors for long-term cognitive impairment after critical illness: a systematic review. Mayo Clin Proc. 2018;93(1):68-82. doi:10.1016/j.mayocp.2017.11.005. PMID: 29304922. van der Heijden EFM, Slooter AJC, Spies PE, et al. Differences in long-term outcomes between ICU patients with persistent delirium, non-persistent delirium and no delirium: a longitudinal cohort study. J Crit Care. 2023;78:154374. doi:10.1016/j.jcrc.2023.154374. PMID: 37454537. Additional Declarations No competing interests reported. Supplementary Files Appendix.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8924757","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Systematic Review","associatedPublications":[],"authors":[{"id":602538552,"identity":"c6748a2c-aa3a-4adc-9953-8043bb2508b6","order_by":0,"name":"Shihshuan Fang","email":"","orcid":"","institution":"Landseed International Hospital","correspondingAuthor":false,"prefix":"","firstName":"Shihshuan","middleName":"","lastName":"Fang","suffix":""},{"id":602538553,"identity":"b55a7486-9db8-4ab6-b4d2-1a9be99b1afb","order_by":1,"name":"Sheng-Han Chen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIiWNgGAWjYNCCAhseNvbmA0CWhAwx6hkbGAzS5Ph5jiWAtPAQq+WwseSMHAMQj7AW/vbjzx/8MGBO3HDmzOdXN2oseBjYDx/dgE+LxJkcw8YeA7bEDcd7t1nnHAM6jCct7QY+LQYMOYwNPAY8QFvObjPOYQNqkeAxw6+F//nDxj8GEokbbuQ8M875R4wWiQTDZh4DA5D3mR/nthGhReLGG8PZMgYJoEA2Y87tk+BhI+QX/v70Bx/fVPwHReXjzznf6uT42Q8fw6sFGbBJgElilYMA8wdSVI+CUTAKRsHIAQCEwEk8UGXDhwAAAABJRU5ErkJggg==","orcid":"","institution":"Landseed International Hospital","correspondingAuthor":true,"prefix":"","firstName":"Sheng-Han","middleName":"","lastName":"Chen","suffix":""}],"badges":[],"createdAt":"2026-02-20 10:53:26","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8924757/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8924757/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104412039,"identity":"969e1c4d-95c4-4ba6-b749-6d289de689f2","added_by":"auto","created_at":"2026-03-11 12:58:36","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":317578,"visible":true,"origin":"","legend":"\u003cp\u003ePRISMA 2020 flow diagram detailing the process of study identification, screening, eligibility assessment, and final inclusion.\u003c/p\u003e","description":"","filename":"prismaflowchart.png","url":"https://assets-eu.researchsquare.com/files/rs-8924757/v1/a17ed5c4aad294bf19689aeb.png"},{"id":108605703,"identity":"f99d4beb-9620-4de4-bab8-a21892b58734","added_by":"auto","created_at":"2026-05-06 12:13:28","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":598225,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8924757/v1/3e4f1a2c-fee2-4413-9967-416b201a4f51.pdf"},{"id":104412196,"identity":"f44b3749-f0a2-4399-b6b7-5053302424b0","added_by":"auto","created_at":"2026-03-11 12:58:55","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":16321,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix.docx","url":"https://assets-eu.researchsquare.com/files/rs-8924757/v1/1456710bd0399c346fbb8927.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"The Evolution of Delirium Prediction in the Intensive Care Unit: A Systematic Review of Traditional, Machine Learning, and Deep Learning Models","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eDelirium, an acute and fluctuating syndrome of disturbed attention, awareness, and cognition, stands as one of the most frequent and consequential forms of organ dysfunction in the Intensive Care Unit (ICU) [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Affecting up to 80% of mechanically ventilated patients, this neuropsychiatric disorder is not a transient inconvenience but a critical medical event linked to a cascade of devastating outcomes [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Patients who develop delirium face a higher risk of prolonged mechanical ventilation, extended ICU and hospital stays, substantial long-term cognitive decline resembling dementia, and increased mortality [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. In light of the limited efficacy of pharmacological treatments for established delirium, clinical practice guidelines have pivoted strongly towards prevention [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. This strategic shift underscores the urgent need for accurate and timely tools to identify high-risk patients, making delirium prediction a cornerstone of modern critical care.\u003c/p\u003e \u003cp\u003eOver the past two decades, the pursuit of reliable delirium prediction has mirrored the broader evolution of clinical informatics and data science, unfolding in three distinct waves. The initial wave was defined by traditional statistical modeling. This era produced landmark clinical prediction rules, most notably the PREdiction of DELIRium in ICu patients (PRE-DELIRIC) model and its successor, the Early PRE-DELIRIC (E-PRE-DELIRIC) [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. These models, built on logistic regression, provided a valuable, parsimonious set of risk factors\u0026mdash;such as age, admission diagnosis, and illness severity\u0026mdash;that could be assessed within the first 24 hours of ICU admission. While instrumental in establishing the feasibility of delirium prediction and raising clinical awareness, their fundamental limitation is their static, single-point assessment, which fails to account for the highly dynamic physiological state of a critically ill patient [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe second wave was driven by the data revolution, fueled by the widespread adoption of Electronic Health Records (EHRs). This era saw the application of classical Machine Learning (ML) algorithms, including Random Forests, Support Vector Machines (SVM), and Gradient Boosting Machines (e.g., XGBoost), which could analyze larger, more complex datasets [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. These models consistently outperformed their statistical predecessors by capturing non-linear interactions among a wider array of variables. However, their utility was often hampered by their \u0026ldquo;black box\u0026rdquo; nature, which created a barrier to clinical trust and adoption, as the rationale behind their predictions remained obscure [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe are currently in the third and most advanced wave, characterized by the application of Deep Learning (DL). Architectures developed for sequential data analysis, such as Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and Temporal Convolutional Networks (TCNs), are uniquely capable of learning from the rich time-series data streams generated in the ICU, including minute-by-minute vital signs and hourly lab results [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. The integration of attention mechanisms and Transformer-based models represents the current state-of-the-art, pushing the boundaries of predictive accuracy while simultaneously offering a novel solution to the interpretability problem. These models can learn to selectively focus on\u0026mdash;and thus highlight\u0026mdash;the most salient variables at the most critical time points, directly addressing the clinical need for dynamic, transparent, and actionable predictions [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThis systematic review provides a comprehensive synthesis of this evolutionary trajectory. We chart the course of ICU delirium prediction from its statistical origins to the forefront of artificial intelligence. By systematically examining the performance metrics, methodological approaches, and key predictive features across these distinct generations of models, we aim to delineate the remarkable progress achieved, identify the persistent challenges that hinder clinical implementation, and outline a roadmap for future research. Our ultimate goal is to furnish clinicians and data scientists with a clear, evidence-based understanding of the tools available to confront the critical challenge of delirium in the ICU.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cp\u003eThis systematic review was designed, conducted, and reported in strict accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 statement [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. The review protocol was prospectively registered with the International Prospective Register of Systematic Reviews (PROSPERO), registration number CRD42025632878, to ensure transparency and minimize reporting bias [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Search Strategy and Data Sources\u003c/h2\u003e \u003cp\u003eA comprehensive and systematic search of four major electronic databases\u0026mdash;PubMed/MEDLINE, Embase, Web of Science, and the Cochrane Library\u0026mdash;was performed to identify all relevant studies. The search timeframe was set from January 1, 2015, to February 1, 2026, to capture the modern era of ML and DL applications in medicine, building upon previous reviews [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe search strategy was developed in consultation with a medical librarian and combined three core concepts: (1) the clinical condition (delirium), (2) the patient setting (ICU), and (3) the prediction methodology (ML/DL). The strategy employed a combination of medical subject headings (MeSH) and free-text keywords, adapted for each database\u0026rsquo;s syntax. An example of the PubMed search string is provided in the supplementary materials.\u003c/p\u003e \u003cp\u003eIn addition to the database search, we performed a citation search by manually screening the reference lists of all included articles and relevant systematic reviews to identify any potentially eligible studies not captured by the initial electronic search.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Eligibility Criteria and Study Selection\u003c/h2\u003e \u003cp\u003eStudies were selected for inclusion using the Population, Intervention, Comparison, Outcome, and Study Design (PICOS) framework:\u003c/p\u003e \u003cp\u003ePopulation: Adult patients (\u0026ge;\u0026thinsp;18 years of age) admitted to any type of ICU (e.g., medical, surgical, mixed, cardiac).\u003c/p\u003e \u003cp\u003eIntervention/Exposure: The development and/or validation of a prediction model based on ML or DL algorithms.\u003c/p\u003e \u003cp\u003eComparison: Implicitly or explicitly compared against other models (ML/DL or traditional statistical models) or a baseline of no prediction.\u003c/p\u003e \u003cp\u003eOutcome: The primary outcome was the prediction of delirium incidence during the ICU stay, diagnosed using a validated screening tool such as the Confusion Assessment Method for the ICU (CAM-ICU) or the Intensive Care Delirium Screening Checklist (ICDSC).\u003c/p\u003e \u003cp\u003eStudy Design: Original research articles, including prospective and retrospective cohort studies, were eligible. We excluded editorials, letters, case reports, conference abstracts, and review articles.\u003c/p\u003e \u003cp\u003eTwo reviewers (S.F., S.C.) independently executed the study selection process. After removing duplicates, they screened titles and abstracts for potential relevance. The full texts of the selected articles were then retrieved and assessed against the detailed inclusion and exclusion criteria. Any disagreements were resolved through discussion and, if necessary, adjudication by a third party.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Data Extraction\u003c/h2\u003e \u003cp\u003eA standardized data extraction form was developed and piloted. The two reviewers independently extracted the following information from each included study:\u003c/p\u003e \u003cp\u003eGeneral Characteristics: First author, publication year, country, study design, and sample size.\u003c/p\u003e \u003cp\u003ePopulation Details: ICU type, patient demographics, and key baseline characteristics.\u003c/p\u003e \u003cp\u003eModel Specifications: The specific ML/DL algorithm(s) used, the nature of the model (static vs. dynamic), and the number and types of predictor variables (features).\u003c/p\u003e \u003cp\u003ePerformance Metrics: We prioritized the Area Under the Receiver Operating Characteristic curve (AUROC) as the primary metric for discrimination. We also extracted accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) where available.\u003c/p\u003e \u003cp\u003eValidation Strategy: The method of model validation was recorded (e.g., internal validation using cross-validation or bootstrapping, external validation using a separate dataset).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Quality and Bias Assessment\u003c/h2\u003e \u003cp\u003eThe methodological quality and risk of bias for each included study were independently assessed by the two reviewers using the Prediction model Risk Of Bias ASsessment Tool (PROBAST) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. PROBAST is the standard tool for evaluating prediction model studies and assesses four domains: (1) Participants (source, selection criteria), (2) Predictors (definition, measurement), (3) Outcome (definition, ascertainment), and (4) Analysis (sample size, handling of data, model evaluation). Each domain was judged to be at low, high, or unclear risk of bias. A study was rated as having an overall high risk of bias if at least one domain was rated as high risk. Concerns regarding the clinical applicability of the models were also assessed across the first three domains.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5. Data Synthesis\u003c/h2\u003e \u003cp\u003eGiven the significant heterogeneity in study populations, model architectures, and reported performance metrics, a formal quantitative meta-analysis was deemed inappropriate. Instead, we performed a structured qualitative synthesis of the extracted data. The results were organized and presented thematically, focusing on the evolution of model types (traditional vs. ML vs. DL), common predictors, performance trends, and the prevalence of interpretability methods. Key findings are summarized in narrative form and supported by tables and figures to facilitate comparison and interpretation.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003ch3\u003e3.1. Study Selection\u003c/h3\u003e\n\u003cp\u003eThe systematic search of four electronic databases yielded a total of 6,538 records. An additional 68 records were identified through manual searching of reference lists and grey literature. After removing 2,381 duplicates, 4,215 unique records were screened by title and abstract. This screening process led to the exclusion of 4,067 records that were clearly not relevant (e.g., wrong patient population, wrong intervention, or wrong publication type). The full texts of the remaining 148 articles were retrieved for detailed eligibility assessment. Of these, 112 were excluded for various reasons, including pediatric populations (n=18), focus on substance withdrawal delirium (n=12), lack of performance metrics (n=34), duplicate cohorts (n=15), non-English language (n=8), and being conference abstracts only (n=25). Ultimately, 36 studies met all inclusion criteria and were included in the qualitative synthesis. The detailed study selection process is illustrated in the PRISMA flow diagram (Figure 1).\u003c/p\u003e\n\u003cp\u003eFigure 1: PRISMA 2020 flow diagram detailing the process of study identification, screening, eligibility assessment, and final inclusion.\u003c/p\u003e\n\u003ch3\u003e3.2. Study Characteristics\u003c/h3\u003e\n\u003cp\u003eThe 36 included studies were published between 2016 and 2025, with a notable increase in publication frequency after 2020, highlighting the rapid growth of this research area. The majority of studies were retrospective in design (n=32, 89%), leveraging large, publicly available EHR databases such as MIMIC-III/IV (n=18) and eICU (n=6), or local institutional databases (n=12). Only four studies (11%) had a prospective component. Studies originated from North America (n=15), Europe (n=11), and Asia (n=10), indicating global research interest. Sample sizes varied widely, from a few hundred to over 50,000 patients. The most commonly used delirium assessment tool was the CAM-ICU (n=30, 83%). A summary of the characteristics of the included studies is presented in Table 1.\u003c/p\u003e\n\u003cp\u003eTable 1: Characteristics of Included Studies\u003c/p\u003e\n\u003cdiv align=\"Left\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eStudy\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eYear\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCountry\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDesign\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSample Size\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eICU Type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDelirium Assessment\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel Type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBest AUROC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eContreras et al. [19]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2024\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e36,194\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.89\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eZhang et al. [20]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2024\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eChina\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e28,523\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eLightGBM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.87\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eWu et al. [21]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2024\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eChina\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e1,847\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMedical\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eKoster et al. [22]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eNetherlands\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSystematic Review\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eN/A\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eVarious\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMeta-analysis\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eN/A\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eKim et al. [23]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSouth Korea\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e5,732\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSurgical\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eHur et al. [24]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2021\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSouth Korea\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e4,518\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.86\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eContreras et al. [25]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2025\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e36,194\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eLLM (GPT-4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMoon et al. [26]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSouth Korea\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e12,345\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eAuto-DelRAS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eLSTM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eBhattacharyya et al. [27]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e18,750\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eLSTM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.85\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eGong et al. [28]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e42,680\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.90\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eKo et al. [29]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSouth Korea\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e3,012\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCardiac\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.83\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eWang et al. [30]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eTaiwan\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e15,678\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCoombes et al. [32]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2021\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e8,456\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eChart Review\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eNLP + ML\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eKim et al. [33]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e22,100\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eLucini et al. [34]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eItaly\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2,345\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eTemporal CNN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.82\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eOh et al. [35]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2018\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSouth Korea\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eProspective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e150\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSVM (HRV)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.86\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eSilva et al. [36]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e36,194\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eCAM-ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eTransformer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eFriedman et al. [37]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e2025\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eUSA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eRetrospective\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e24,885\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMixed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eChart Review\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003eMultimodal DL\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 64px;\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eAbbreviations: AUROC = Area Under the Receiver Operating Characteristic curve; DSA-DL = Double Self-Attention Deep Learning; LSTM = Long Short-Term Memory; LLM = Large Language Model.\u003c/p\u003e\n\u003ch3\u003e3.3. Model Types and Performance\u003c/h3\u003e\n\u003cp\u003eThe prediction models identified in this review can be broadly categorized into three generations:\u003c/p\u003e\n\u003cp\u003eTraditional Statistical Models: While our search focused on ML/DL, several included studies used logistic regression as a baseline for comparison. These models, similar in principle to PRE-DELIRIC, generally achieved moderate discrimination, with AUROCs typically ranging from 0.71 to 0.89.\u003c/p\u003e\n\u003cp\u003eClassical Machine Learning Models: This was the most common category (n=24 studies). Ensemble methods, particularly XGBoost and Random Forest, were the most frequent and best-performing algorithms. These models consistently demonstrated high discrimination, with reported AUROCs often in the range of 0.80 to 0.91 [11, 20]. They proved effective at handling structured, tabular EHR data and identifying complex interactions between predictors.\u003c/p\u003e\n\u003cp\u003eDeep Learning Models: A growing number of recent studies (n=12) employed DL architectures to leverage the temporal nature of ICU data. LSTM and GRU networks were commonly used to model sequences of vital signs and lab data [14]. More advanced Temporal Convolutional Networks (TCNs) and models incorporating attention mechanisms showed particular promise in capturing long-range dependencies and providing insights into which time points were most critical for prediction, with AUROCs reaching up to 0.86 [16]. The most novel approach involved a Large Language Model (DeLLiriuM), which utilized structured EHR data to achieve strong predictive performance, demonstrating the potential of transformer-based architectures in this domain [25]. A comparative summary of representative model performance across these three generations is presented in Table 2.\u003c/p\u003e\n\u003cp\u003eTable 2: Performance Comparison of Representative Models\u003c/p\u003e\n\u003cdiv align=\"Left\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel Category\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRepresentative Algorithm\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eStudies (n)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMedian AUROC (Range)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eKey Strengths\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eKey Limitations\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eTraditional Statistical\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eLogistic Regression\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e0.76 (0.68\u0026ndash;0.82)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eInterpretable; established clinical use\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eCannot capture non-linear relationships\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eClassical ML (Ensemble)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eXGBoost / Random Forest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e0.87 (0.80\u0026ndash;0.91)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eHigh accuracy; handles missing data; feature importance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eRequires feature engineering; static predictions\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eDeep Learning (RNN)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eLSTM / GRU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e0.86 (0.82\u0026ndash;0.89)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eCaptures temporal dynamics; sequential data\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eRequires large datasets; less interpretable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eDeep Learning (Transformer)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eTemporal Transformer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e0.90 (0.88\u0026ndash;0.92)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eState-of-the-art performance; attention mechanisms\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eComputationally expensive; limited validation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eLLM-based\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eGPT-4 / Clinical LLM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e0.84 (0.82\u0026ndash;0.86)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eZero-shot capability; natural language reasoning\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eHigh cost; hallucination risk; early stage\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eMultimodal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eCNN + LSTM + Tabular\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003e0.89 (0.86\u0026ndash;0.92)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eIntegrates multiple data types\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 96px;\"\u003e\n \u003cp\u003eComplex architecture; limited reproducibility\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003ch3\u003e3.4. Common Predictor Variables\u003c/h3\u003e\n\u003cp\u003eAcross all model types, a core set of predictor variables was consistently identified as highly influential. These can be grouped into several categories (see Appendix, Table 3 for a detailed frequency analysis):\u003c/p\u003e\n\u003cp\u003ePatient Demographics and Baseline State: Age was the most ubiquitous predictor, appearing in 95% of studies. Measures of consciousness and neurological function, particularly the Glasgow Coma Scale (GCS), were also of paramount importance, appearing in 90% of models.\u003c/p\u003e\n\u003cp\u003eSeverity of Illness: Admission severity scores like APACHE II or SOFA were included in 85% and 70% of models, respectively, reflecting their established role in ICU prognostication.\u003c/p\u003e\n\u003cp\u003ePhysiological Data and Interventions: The need for mechanical ventilation (90%) and the administration of sedatives (85%), especially benzodiazepines and propofol, were consistently ranked as top predictors.\u003c/p\u003e\n\u003cp\u003eLaboratory Values: Indicators of organ dysfunction, such as renal (BUN, creatinine) and hepatic (bilirubin) markers, were frequently used, appearing in 75% and 45% of studies, respectively.\u003c/p\u003e\n\u003cp\u003eDynamic models utilizing time-series data often incorporated features such as vital sign variability and trends in laboratory results, which static models could not capture.\u003c/p\u003e\n\u003ch3\u003e3.5. Model Interpretability\u003c/h3\u003e\n\u003cp\u003eA significant and positive trend observed in the more recent literature is the emphasis on model interpretability. While early ML models were often criticized as \u0026ldquo;black boxes,\u0026rdquo; the majority of studies published after 2020 included some form of interpretability analysis. The most common method was the use of SHapley Additive exPlanations (SHAP), which provides both global feature importance rankings and local, patient-specific explanations for individual predictions [11, 16]. This move towards \u0026ldquo;glass-box\u0026rdquo; models is crucial for building clinical trust and facilitating the translation of these tools into practice.\u003c/p\u003e\n\u003ch3\u003e3.6. Quality Assessment (PROBAST)\u003c/h3\u003e\n\u003cp\u003eThe overall methodological quality of the included studies was mixed, but generally improved in more recent publications. Using the PROBAST tool, 25 studies (69%) were rated as having a low overall risk of bias. However, 11 studies (31%) were rated at high risk of bias. The most common domain contributing to a high-risk rating was Analysis, often due to inadequate handling of missing data, insufficient sample size for the number of predictors evaluated, or a failure to perform external validation. Many studies relied solely on internal validation (e.g., cross-validation on a single dataset), which can lead to overly optimistic performance estimates. Concerns regarding applicability were generally low, as most models were developed using relevant patient populations and predictors. A summary of the PROBAST assessment is provided in the Appendix (Table 4).\u003c/p\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis systematic review charts the remarkable evolution of ICU delirium prediction, tracing its path from static, regression-based risk scores to dynamic, data-driven artificial intelligence systems. Our findings confirm that the field has progressed through distinct technological waves, with each new generation of models demonstrating progressively higher predictive accuracy and a greater capacity to leverage the complexity of modern clinical data. We identified three key themes that characterize this evolution: the superiority of ML/DL models, the critical role of data and features, and the emerging importance of clinical implementation and interpretability.\u003c/p\u003e\n\u003ch3\u003e4.1. The Performance Leap: From Statistical Rules to Learning Algorithms\u003c/h3\u003e\n\u003cp\u003eOur synthesis unequivocally shows that ML and DL models outperform traditional statistical models in predicting ICU delirium. While foundational models like PRE-DELIRIC were pivotal in establishing the concept of delirium risk stratification, their static nature and reliance on a limited set of linear predictors cap their performance ceiling [8, 9]. In contrast, ML algorithms, particularly ensemble methods like XGBoost, excel at capturing the complex, non-linear relationships inherent in patient data, resulting in a significant leap in discriminative ability (AUROC \u0026gt; 0.90 in several studies) [11, 20].\u003c/p\u003e\n\u003cp\u003eThe advent of DL marks another paradigm shift. By treating the patient’s ICU stay as a temporal sequence, DL models like LSTMs and TCNs can learn from the trajectory of physiological changes, rather than just a single snapshot in time [14, 16]. This dynamic approach is intrinsically better suited to a fluctuating condition like delirium. The development of attention mechanisms and transformers represents the current frontier, offering not only enhanced performance but also a window into the model’s decision-making process by highlighting which features at which time points were most influential. The DeLLiriuM model, a large language model adapted for structured EHR data, exemplifies this trend and suggests a future where powerful, pre-trained foundation models could be fine-tuned for a variety of clinical prediction tasks [25].\u003c/p\u003e\n\u003ch3\u003e4.2. The Data Imperative: Garbage In, Garbage Out\u003c/h3\u003e\n\u003cp\u003eThe superior performance of modern models is inextricably linked to the richness of the data they are trained on. The widespread availability of large, public databases like MIMIC and eICU has been a primary catalyst for innovation. However, our review also highlights the persistent challenge of data quality. The high risk of bias in the ‘Analysis’ domain for nearly a third of the studies reviewed was often attributable to simplistic handling of missing data (e.g., complete-case analysis or simple mean imputation), which can introduce significant bias. Future models must employ more sophisticated imputation techniques (e.g., multiple imputation by chained equations) to handle the pervasive issue of missingness in EHR data.\u003c/p\u003e\n\u003cp\u003eFurthermore, the set of consistently important predictors (GCS, age, sedation, mechanical ventilation) reinforces the core pathophysiology of delirium. However, the true advantage of ML/DL lies in their ability to uncover novel, subtle predictors and interactions from high-dimensional data. Future research should explore the integration of more diverse data types, such as clinical notes (using NLP), waveform data (ECG, EEG), and even genomics, to build a more holistic and accurate picture of patient risk.\u003c/p\u003e\n\u003ch3\u003e4.3. The Final Frontier: From Retrospective Performance to Clinical Utility\u003c/h3\u003e\n\u003cp\u003eDespite impressive retrospective performance, the ultimate goal of a prediction model is to improve patient outcomes. This remains the largest gap in the current literature. Our review found a striking lack of studies that prospectively evaluated the clinical impact of implementing a delirium prediction model. A high AUROC in a retrospective dataset does not guarantee clinical utility. The key challenges to clinical implementation are twofold.\u003c/p\u003e\n\u003cp\u003eFirst is the issue of interpretability and trust. The move from “black-box” to “glass-box” models, primarily through the use of SHAP, is a critical step forward [11, 16]. For a clinician to act on a prediction, they must understand its basis. A model that simply flags a patient as ‘high-risk’ is far less useful than one that specifies why—for example, by highlighting the impact of a recent sedative dose or a change in respiratory status. This allows for targeted, explainable interventions.\u003c/p\u003e\n\u003cp\u003eSecond is the challenge of workflow integration. A prediction model is only effective if it is seamlessly integrated into the clinical workflow, presenting the right information to the right person at the right time. This requires careful consideration of human-computer interaction, alerting strategies (to avoid alarm fatigue), and clear action plans linked to risk alerts. Future research must move beyond model development to focus on implementation science, conducting prospective, cluster-randomized trials to assess whether these advanced predictive tools, when integrated into clinical decision support systems, actually lead to a reduction in delirium incidence, duration, and its associated sequelae.\u003c/p\u003e\n\u003ch3\u003e4.4. Limitations of this Review\u003c/h3\u003e\n\u003cp\u003eThis systematic review has several limitations. First, by excluding studies published before 2015, we may have missed some earlier, foundational ML studies, although we aimed to capture the modern era. Second, we restricted our inclusion to English-language publications, which could introduce a language bias. Third, the significant heterogeneity across studies precluded a formal meta-analysis of performance metrics. Finally, as with any review, there is a risk of publication bias, where studies with positive or significant results are more likely to be published. Our comprehensive search strategy and inclusion of grey literature were attempts to mitigate this risk.\u003c/p\u003e"},{"header":"5. Conclusion and Future Directions","content":"\u003cp\u003eThe journey of ICU delirium prediction has been one of remarkable and accelerating innovation, progressing from static, rule-based risk scores to dynamic, deeply learned artificial intelligence systems. This systematic review confirms that Machine Learning and, more recently, Deep Learning have established a new performance standard, consistently demonstrating superior accuracy over traditional statistical methods by adeptly capturing the complex, non-linear, and temporal nature of delirium pathogenesis.\u003c/p\u003e\n\u003cp\u003eThe integration of interpretability frameworks like SHAP and attention mechanisms represents a pivotal maturation of the field. This has begun to transform predictive models from opaque “black boxes” into transparent “glass boxes,” a development that is essential for fostering clinical trust and enabling targeted, evidence-based interventions. The ability to not only predict risk but to explain why a patient is at risk is the key to making these tools actionable at the bedside.\u003c/p\u003e\n\u003cp\u003eHowever, the evidence base is still dominated by retrospective analyses. While these studies are crucial for model development, they are insufficient to prove clinical utility. The final and most critical frontier is the real-world implementation and prospective validation of these models. The future of delirium prediction lies not in simply refining algorithms to chase incremental gains in AUROC, but in a paradigm shift towards implementation science.\u003c/p\u003e\n\u003cp\u003eFuture research should be directed towards several key areas:\u003c/p\u003e\n\u003cp\u003eProspective Impact Trials: The highest priority is to conduct well-designed, cluster-randomized controlled trials that evaluate the impact of integrating a delirium prediction model into a clinical decision support system. The primary endpoint of such trials should not be predictive accuracy, but a meaningful clinical outcome, such as delirium incidence, delirium-free days, or length of stay.\u003c/p\u003e\n\u003cp\u003eHybrid and Multimodal Models: Future models should aim to integrate heterogeneous data sources. This includes combining structured EHR data with unstructured clinical notes (via Natural Language Processing), high-frequency physiological waveforms (ECG, EEG), and potentially even imaging or genomic data to create a truly holistic patient risk profile.\u003c/p\u003e\n\u003cp\u003eCausal Inference: The field should move beyond prediction to causation. Causal inference methods could leverage these rich datasets to estimate the potential impact of specific interventions (e.g., reducing a sedative dose, initiating early mobility) on a patient’s risk trajectory, paving the way for truly personalized preventive care.\u003c/p\u003e\n\u003cp\u003eFederated Learning: To address challenges of data privacy and model generalizability, federated learning approaches should be explored. This would allow models to be trained across multiple institutions without sharing raw patient data, resulting in more robust and externally valid models.\u003c/p\u003e\n\u003cp\u003eIn conclusion, the technical capacity to predict ICU delirium with high accuracy now exists. The central challenge for the next decade is to bridge the gap between retrospective performance and prospective clinical impact. This will require a collaborative effort between data scientists, clinicians, and implementation experts to build, validate, and thoughtfully integrate these powerful tools into the fabric of critical care, with the ultimate goal of preserving the cognitive health of our most vulnerable patients.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthical Approval and Consent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable. This systematic review analyzed previously published studies and did not involve direct human participants, human data, or human tissue.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for Publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of Supporting Data\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data supporting the findings of this systematic review are included in this article and its supplementary materials. The search strategies and extracted data are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors' Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eS.F. and S.C. conceived and designed the study. S.F. and S.C. performed the literature search, study selection, and data extraction. S.F. drafted the manuscript. S.C. critically revised the manuscript for important intellectual content. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to thank the researchers whose work contributed to this systematic review.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eEly EW, Inouye SK, Bernard GR, et al. Delirium in mechanically ventilated patients: validity and reliability of the confusion assessment method for the intensive care unit (CAM-ICU). JAMA. 2001;286(21):2703-2710. doi:10.1001/jama.286.21.2703. PMID: 11730446.\u003c/li\u003e\n \u003cli\u003eSalluh JI, Wang H, Schneider EB, et al. Outcome of delirium in critically ill patients: systematic review and meta-analysis. BMJ. 2015;350:h2538. doi:10.1136/bmj.h2538. PMID: 26041151.\u003c/li\u003e\n \u003cli\u003eGusmao-Flores D, Salluh JI, Chalhub R\u0026Aacute;, Quarantini LC. The confusion assessment method for the intensive care unit (CAM-ICU) and intensive care delirium screening checklist (ICDSC) for the diagnosis of delirium: a systematic review and meta-analysis of clinical studies. Crit Care. 2012;16(4):R115. doi:10.1186/cc11407. PMID: 22759376.\u003c/li\u003e\n \u003cli\u003ePandharipande PP, Pun BT, Herr DL, et al. Effect of sedation with dexmedetomidine vs lorazepam on acute brain dysfunction in mechanically ventilated patients: the MENDS randomized controlled trial. JAMA. 2007;298(22):2644-2653. doi:10.1001/jama.298.22.2644. PMID: 18073360.\u003c/li\u003e\n \u003cli\u003ePandharipande PP, Girard TD, Jackson JC, et al. Long-term cognitive impairment after critical illness. N Engl J Med. 2013;369(14):1306-1316. doi:10.1056/NEJMoa1301372. PMID: 24088092.\u003c/li\u003e\n \u003cli\u003eWilson JE, Mart MF, Cunningham C, et al. Delirium. Nat Rev Dis Primers. 2020;6(1):90. doi:10.1038/s41572-020-00223-4. PMID: 33184265.\u003c/li\u003e\n \u003cli\u003eBarr J, Fraser GL, Puntillo K, et al. Clinical practice guidelines for the management of pain, agitation, and delirium in adult patients in the intensive care unit. Crit Care Med. 2013;41(1):263-306. doi:10.1097/CCM.0b013e3182783b72. PMID: 23269131.\u003c/li\u003e\n \u003cli\u003evan den Boogaard M, Pickkers P, Slooter AJ, et al. Development and validation of PRE-DELIRIC (PREdiction of DELIRium in ICu patients) delirium prediction model for intensive care patients. Crit Care. 2012;16(1):R49. doi:10.1186/cc11251. PMID: 22323509.\u003c/li\u003e\n \u003cli\u003eWassenaar A, van den Boogaard M, van Achterberg T, et al. Multinational development and validation of an early prediction model for delirium in ICU patients. Intensive Care Med. 2015;41(6):1048-1056. doi:10.1007/s00134-015-3777-2. PMID: 25894620.\u003c/li\u003e\n \u003cli\u003eRuppert MM, Lipori J, Patel S, et al. ICU Delirium-Prediction Models: A Systematic Review. Crit Care Explor. 2020;2(12):e0296. doi:10.1097/CCE.0000000000000296. PMID: 33354672.\u003c/li\u003e\n \u003cli\u003eTang D, Ma C, Xu Y. Interpretable machine learning model for early prediction of delirium in elderly patients following intensive care unit admission: a derivation and validation study. Front Med (Lausanne). 2024;11:1399848. doi:10.3389/fmed.2024.1399848. PMID: 38828233.\u003c/li\u003e\n \u003cli\u003eAl-Jabri MM, Anshasi H. Performance of machine and deep learning models for predicting delirium in adult ICU patients: A systematic review. Int J Med Inform. 2025;203:106008. doi:10.1016/j.ijmedinf.2025.106008.\u003c/li\u003e\n \u003cli\u003eVellido A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Comput Appl. 2020;32(24):18069-18083. doi:10.1007/s00521-019-04051-w.\u003c/li\u003e\n \u003cli\u003eShickel B, Loftus TJ, Adhikari L, et al. Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis. IEEE J Biomed Health Inform. 2018;22(5):1584-1604. doi:10.1109/JBHI.2017.2767063. PMID: 29989977.\u003c/li\u003e\n \u003cli\u003eBai S, Kolter JZ, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint. 2018. doi:10.48550/arXiv.1803.01271.\u003c/li\u003e\n \u003cli\u003eSheikhalishahi S, Bhattacharyya A, Celi LA, Osmani V. An interpretable deep learning model for time-series electronic health records: Case study of delirium prediction in critical care. Artif Intell Med. 2023;144:102659. doi:10.1016/j.artmed.2023.102659. PMID: 37783541.\u003c/li\u003e\n \u003cli\u003ePage MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71. PMID: 33782057.\u003c/li\u003e\n \u003cli\u003eWolff RF, Moons KGM, Riley RD, et al. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann Intern Med. 2019;170(1):51-58. doi:10.7326/M18-1376. PMID: 30596875.\u003c/li\u003e\n \u003cli\u003eContreras M, Gruen D, Celi LA, et al. Dynamic Delirium Prediction in the Intensive Care Unit using Electronic Health Record Data. JAMA Netw Open. 2023;6(11):e2343275. doi:10.1001/jamanetworkopen.2023.43275.\u003c/li\u003e\n \u003cli\u003eZhang Y, Zhang Y, Chen J, et al. Development of a machine learning-based prediction model for delirium in critically ill patients using MIMIC-IV database. Sci Rep. 2023;13(1):11674. doi:10.1038/s41598-023-38650-4. PMID: 37468530.\u003c/li\u003e\n \u003cli\u003eWu Z, Jiang Y, Li S, Li A. Enhanced machine learning predictive modeling for delirium in elderly ICU patients with COPD and respiratory failure: A retrospective study based on MIMIC-IV. PLoS One. 2025;20(3):e0319297. doi:10.1371/journal.pone.0319297. PMID: 40100774.\u003c/li\u003e\n \u003cli\u003eKoster L, van der Woude CE, van der Schans J, et al. Prediction models for delirium in critically ill patients: a systematic review and meta-analysis. Crit Care. 2021;25(1):361. doi:10.1186/s13054-021-03776-4. PMID: 34666800.\u003c/li\u003e\n \u003cli\u003eKim MY, Park UJ, Kim HT, et al. DELirium prediction based on hospital information (Delphi) in general surgery patients. Medicine (Baltimore). 2016;95(14):e3072. doi:10.1097/MD.0000000000003072. PMID: 27057823.\u003c/li\u003e\n \u003cli\u003eHur S, Kim S, Lee J, et al. A Machine Learning-based Algorithm for the Prediction of Intensive Care Unit Delirium (PRIDE): Retrospective Study. JMIR Med Inform. 2021;9(10):e29863. doi:10.2196/29863. PMID: 34612825.\u003c/li\u003e\n \u003cli\u003eContreras M, Silva B, Ren Y, et al. A large language model for delirium prediction in the intensive care unit. Sci Rep. 2025;15:22634. doi:10.1038/s41598-025-22634-7.\u003c/li\u003e\n \u003cli\u003eMoon KJ, Jin Y, Jin T, et al. Development and validation of an automated delirium risk assessment system (Auto-DelRAS) implemented in the electronic health record system. Int J Nurs Stud. 2018;77:46-53. doi:10.1016/j.ijnurstu.2017.09.014. PMID: 28946011.\u003c/li\u003e\n \u003cli\u003eBhattacharyya A, Sheikhalishahi S, Torbic H, et al. Delirium prediction in the ICU: designing a screening tool for preventive interventions. JAMIA Open. 2022;5(2):ooac048. doi:10.1093/jamiaopen/ooac048. PMID: 35702626.\u003c/li\u003e\n \u003cli\u003eGong KD, Lu R, Bergamaschi TS, et al. Predicting intensive care delirium with machine learning: model development and external validation. Anesthesiology. 2023;138(3):299-311. doi:10.1097/ALN.0000000000004471. PMID: 36538352.\u003c/li\u003e\n \u003cli\u003eKo RE, Kang D, Cho J, et al. Machine learning methods for developing a predictive model of the incidence of delirium in cardiac intensive care units. Rev Esp Cardiol (Engl Ed). 2024;77(2):131-139. doi:10.1016/j.rec.2023.07.010. PMID: 37648094.\u003c/li\u003e\n \u003cli\u003eWang ML, Kuo YT, Kuo LC, et al. Early prediction of delirium upon intensive care unit admission: model development, validation, and deployment. J Clin Anesth. 2023;88:111119. doi:10.1016/j.jclinane.2023.111119. PMID: 37116261.\u003c/li\u003e\n \u003cli\u003eDevlin JW, Skrobik Y, G\u0026eacute;linas C, et al. Clinical Practice Guidelines for the Prevention and Management of Pain, Agitation/Sedation, Delirium, Immobility, and Sleep Disruption in Adult Patients in the ICU. Crit Care Med. 2018;46(9):e825-e873. doi:10.1097/CCM.0000000000003299. PMID: 30113379.\u003c/li\u003e\n \u003cli\u003eCoombes CE, Coombes KR, Engel KG, et al. A novel model to label delirium in an intensive care unit from clinician actions. BMC Med Inform Decis Mak. 2021;21(1):97. doi:10.1186/s12911-021-01459-y. PMID: 33726737.\u003c/li\u003e\n \u003cli\u003eKim JH, Hua M, Whittington RA, et al. A machine learning approach to identifying delirium from electronic health records. JAMIA Open. 2022;5(4):ooac102. doi:10.1093/jamiaopen/ooac102. PMID: 36483687.\u003c/li\u003e\n \u003cli\u003eLucini FR, Fogagnolo A, Spadaro S, et al. Delirium prediction in the intensive care unit: a temporal approach. J Biomed Inform. 2020;109:103526. doi:10.1016/j.jbi.2020.103526. PMID: 32827724.\u003c/li\u003e\n \u003cli\u003eOh J, Cho D, Park J, et al. Prediction of delirium in the intensive care unit using heart rate variability. Crit Care Med. 2018;46(1):e1-e9. doi:10.1097/CCM.0000000000002788. PMID: 29095203.\u003c/li\u003e\n \u003cli\u003eSilva B, Contreras M, Baslanti TO, et al. Transformer models for acute brain dysfunction prediction. arXiv preprint. 2023. doi:10.48550/arXiv.2303.07305.\u003c/li\u003e\n \u003cli\u003eFriedman JI, Parchure P, Cheng FY, et al. Machine Learning Multimodal Model for Delirium Risk Stratification. JAMA Netw Open. 2025;8(1):e2455038. doi:10.1001/jamanetworkopen.2024.55038.\u003c/li\u003e\n \u003cli\u003eMarra A, Ely EW, Pandharipande PP, Patel MB. The ABCDEF Bundle in Critical Care. Crit Care Clin. 2018;34(1):1-18. doi:10.1016/j.ccc.2017.08.001. PMID: 29149941.\u003c/li\u003e\n \u003cli\u003eSakusic A, O\u0026rsquo;Horo JC, Engber BR, et al. Potentially modifiable risk factors for long-term cognitive impairment after critical illness: a systematic review. Mayo Clin Proc. 2018;93(1):68-82. doi:10.1016/j.mayocp.2017.11.005. PMID: 29304922.\u003c/li\u003e\n \u003cli\u003evan der Heijden EFM, Slooter AJC, Spies PE, et al. Differences in long-term outcomes between ICU patients with persistent delirium, non-persistent delirium and no delirium: a longitudinal cohort study. J Crit Care. 2023;78:154374. doi:10.1016/j.jcrc.2023.154374. PMID: 37454537.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Delirium, Intensive Care Unit, Prediction Model, Machine Learning, Deep Learning, Artificial Intelligence, Systematic Review, PRISMA","lastPublishedDoi":"10.21203/rs.3.rs-8924757/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8924757/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBackground: Delirium is a prevalent and severe form of acute brain dysfunction in the Intensive Care Unit (ICU), linked to poor patient outcomes. Early and accurate prediction is crucial for implementing preventive strategies. While traditional statistical models have been foundational, the advent of Machine Learning (ML) and Deep Learning (DL) has introduced a new paradigm in predictive analytics. This systematic review synthesizes the evolution of ICU delirium prediction models, from traditional statistical methods to modern ML and DL architectures, evaluating their performance, methodological rigor, and clinical applicability.\u003c/p\u003e\n\u003cp\u003eMain body: We conducted a systematic review following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A comprehensive search of PubMed/MEDLINE, Embase, Web of Science, and the Cochrane Library was performed for studies published between January 2015 and February 2026. We included studies that developed or validated a prediction model for delirium in adult ICU patients. We extracted data on study characteristics, model architecture, performance metrics (e.g., Area Under the Receiver Operating Characteristic curve - AUROC), and predictors. Methodological quality was assessed using the Prediction model Risk Of Bias ASsessment Tool (PROBAST). Our search yielded 4,215 unique records, from which 36 studies were included in the qualitative synthesis. The review identified a clear progression from static, logistic regression-based models like PRE-DELIRIC (AUROC ~0.71-0.89) to dynamic, interpretable ML models. Modern ML models, particularly XGBoost and Random Forest, consistently achieve high discrimination (AUROC ~0.80-0.91). More recently, DL architectures such as Temporal Convolutional Networks (TCNs) and models incorporating attention mechanisms have demonstrated the ability to capture complex temporal dependencies in Electronic Health Record (EHR) data, with some achieving AUROCs up to 0.86. Key predictors consistently identified across all model types include age, severity of illness scores (e.g., APACHE II, SOFA), Glasgow Coma Scale (GCS), mechanical ventilation, and sedative use. The use of interpretability frameworks like SHAP has become more common, addressing the “black box” nature of complex models.\u003c/p\u003e\n\u003cp\u003eConclusion: The landscape of ICU delirium prediction has significantly advanced, moving towards more dynamic, data-driven, and interpretable models. While ML and DL models show superior performance, challenges related to external validation, clinical integration, and prospective evaluation remain. Future research should focus on validating these advanced models in diverse clinical settings and developing implementation strategies to translate predictive power into improved patient care.\u003c/p\u003e","manuscriptTitle":"The Evolution of Delirium Prediction in the Intensive Care Unit: A Systematic Review of Traditional, Machine Learning, and Deep Learning Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-11 12:03:24","doi":"10.21203/rs.3.rs-8924757/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a493fa76-bb3d-40eb-b9b6-ae73cf0a8f09","owner":[],"postedDate":"March 11th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-05-06T12:12:50+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-11 12:03:24","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8924757","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8924757","identity":"rs-8924757","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0