Application of Foundation Models in Emergency and Critical Care: A Scoping Review

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-06, 2026-06-24 · read from full text

This PRISMA-ScR–guided scoping review mapped 49 peer-reviewed original studies (through March 26, 2025) that applied foundation models in emergency and critical care, capturing model types, data modalities, and clinical task domains (including outcome prediction, diagnosis, information extraction, text generation, and treatment recommendation). Most studies focused on language models, with less exploration of multimodal architectures, and they used diverse inputs spanning free text, tabular EHR data, time-series signals, and imaging. The review found that existing evidence was largely retrospective, with minimal external validation or prospective testing, and that many models were trained on internet text potentially misaligned with medical reasoning, raising safety/ethical concerns. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Emergency and critical care (ECC) settings demand rapid decision-making with emergent conditions, dynamic patient trajectories, and diagnostic uncertainty. Foundation models (FMs), large neural networks pretrained on extensive datasets using self-supervised learning, show promise for diverse clinical tasks in these high-acuity environments. However, no prior work has comprehensively reviewed FM applications in ECC or the barriers limiting their implementation. Following PRISMA-ScR guidelines, we identified 49 eligible studies. Most focused on language models, with comparatively limited exploration of multimodal architectures. FMs utilized diverse data modalities, including free text, tabular records, time-series signals, and imaging, supporting tasks such as outcome prediction, diagnosis, information extraction, and text generation. Despite promising applications such as triage, discharge instructions, and disease diagnosis, current evidence is predominantly retrospective, with minimal external validation or prospective testing. Also, many FMs were trained on internet text that may be misaligned with medical reasoning, introducing safety and ethical risks, and highlighting the need for clinically supervised deployment. FMs have yet to demonstrate benefit in the ECC context; none of the included studies had real-world model deployment or improvements in clinical outcomes. Future research should prioritize the development of multimodal FMs with multicenter, temporally robust validation and prospective trials that emphasize safety, equity, and clinician trust.
Full text 255,175 characters · extracted from preprint-html · click to expand
Application of Foundation Models in Emergency and Critical Care: A Scoping Review | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Application of Foundation Models in Emergency and Critical Care: A Scoping Review Yanqing Kong, Xinnie Mai, Andrew Meyer, Yanan Fang, Zidu Xu, Zhixing Song, and 14 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8338830/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Emergency and critical care (ECC) settings demand rapid decision-making with emergent conditions, dynamic patient trajectories, and diagnostic uncertainty. Foundation models (FMs), large neural networks pretrained on extensive datasets using self-supervised learning, show promise for diverse clinical tasks in these high-acuity environments. However, no prior work has comprehensively reviewed FM applications in ECC or the barriers limiting their implementation. Following PRISMA-ScR guidelines, we identified 49 eligible studies. Most focused on language models, with comparatively limited exploration of multimodal architectures. FMs utilized diverse data modalities, including free text, tabular records, time-series signals, and imaging, supporting tasks such as outcome prediction, diagnosis, information extraction, and text generation. Despite promising applications such as triage, discharge instructions, and disease diagnosis, current evidence is predominantly retrospective, with minimal external validation or prospective testing. Also, many FMs were trained on internet text that may be misaligned with medical reasoning, introducing safety and ethical risks, and highlighting the need for clinically supervised deployment. FMs have yet to demonstrate benefit in the ECC context; none of the included studies had real-world model deployment or improvements in clinical outcomes. Future research should prioritize the development of multimodal FMs with multicenter, temporally robust validation and prospective trials that emphasize safety, equity, and clinician trust. Biological sciences/Computational biology and bioinformatics Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Emergency and critical care Emergency department Intensive care unit Foundation models Large language models Multimodal models Clinical decision support Scoping review Figures Figure 1 Figure 2 Figure 3 Introduction In emergency and critical care (ECC) settings such as the emergency departments (ED) and intensive care unit (ICU), patients often present with life-threatening conditions requiring rapid interventions 1 – 3 . Clinicians must integrate diverse data streams 4 – 6 including medical history, imaging, laboratory results, and continuous monitoring under severe time pressure. While both ED and ICU demand urgent responses, their contexts differ: emergency care emphasizes triage and immediate intervention with sparse, heterogeneous data 7 , 8 , whereas critical care involves prolonged management, stabilization, and recovery based on higher-dimensional, longitudinal inputs 9 , 10 , 11 – 13 . Escalating data volume, heterogeneity, and real-time variability have intensified clinicians’ workload and decision-making demands 14 , 15 , underscoring the need for effective tools to support data integration and timely action 16 . Artificial intelligence (AI) has long been applied in ECC to assist emergency physicians and intensivists in processing complex data and enhancing diagnostic reasoning 17 – 19 . Earlier generations of machine learning (ML) models, which include logistic regression 20 , 21 , decision trees 22 , 23 , support vector machines 24 , random forests 25 , and extreme gradient boosting 26 – 29 were typically trained on narrow datasets 30 and required manual feature engineering 31 , 32 . Deep learning (DL) methods such as convolutional neural networks (CNN) 33 – 36 , along with short-term memory networks (LSTM) 37 – 40 , and generative adversarial networks (GAN) 41 improved the ability to learn directly from multi-source clinical data such as imaging, waveforms, and continuous vital signs. However, they still require large task-specific datasets and extensive validation before clinical adoption. These models have been applied in ECC to predict triage acuity, in-hospital mortality, sepsis, and ICU readmission 42 – 45 . Foundation models (FMs) represent the next major shift in this progression. Unlike prior DL models trained for single tasks, FMs are pretrained at scale on diverse datasets using self-supervised learning and can be adapted across a wide range of downstream clinical applications, enabling broader generalization and reuse in application. Consequently, FMs represent a paradigm shift: large, general-purpose neural networks trained on massive, diverse datasets using self-supervised learning, enabling adaptation to a wide range of downstream clinical tasks. These include language models such as Generative Pretrained Transformers (GPT) 46 , Bidirectional Encoder Representations from Transformers (BERT) 47 , and Large Language Model Meta AI (LLaMA) 48 ; vision models such as Vision Transformers (ViT) 49 ; and multimodal models such as Contrastive Language–Image Pretraining (CLIP) 50 and Bootstrapped Language-Image Pretraining (BLIP) 51 . FMs have achieved state-of-the-art performance in natural language processing (NLP), clinical prediction, medical imaging, and multimodal reasoning. In the ECC context, FMs offer unique advantages in handling both unstructured texts (e.g., discharge summaries, triage records) and structured data (e.g., laboratory results, vital signs, diagnostic codes, medical images). Different FM types align with various tasks: language models for narrative text, tabular- or sequence-adapted models for structured EHR data (e.g., labs, vitals, codes) (Appendix eTable 3); vision models for imaging (Appendix eTable 5); audio-focused models (e.g., Wav2Vec) for clinical dialogues (Appendix eTable 5); and multimodal architectures for integrating across data types (Appendix eTable 4). The choice of FM, therefore, depends on both the data modality and the clinical objective; for instance, outcome prediction from admission notes might use a language model, stroke detection from imaging might use a vision model, and risk assessment from physiologic signals and text might use a multimodal model. This task-oriented framing helps clinicians and researchers match FM selection to the specific data types and decision needs of ECC practice. Despite growing interest in FMs, most existing reviews remain focused on conventional AI models 52 – 54 and seldom evaluate how FM capabilities such as few-shot learning, contextual reasoning, multimodal integration, and self-supervised adaptation align with ECC’s time-critical demands. In addition, current reviews do not assess deployment readiness or safety of FM use in ECC, and rarely examine how to address the epistemic mismatch between FM capabilities and clinical reasoning in ECC 55 . A synthesis of the emerging literature is therefore needed to clarify current FM applications, gaps, and promising directions. This review aims to map the landscape of FM applications in ECC, including the models used, data modalities leveraged, and clinical tasks addressed, while consolidating reported limitations to outline priorities for future research and responsible implementation. At the same time, whether FMs are appropriately suited tools for ECC or are simply a solution in search of a problem remains uncertain. We therefore adopt a problem-first, safety-anchored framing to assess not only performance but also alignment with ECC’s operational and clinical realities. Methods Study Design This scoping review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) guidelines (Appendix eTable 2). The protocol was developed in advance and applied consistently across study selection, data extraction, and analysis. A comprehensive search was conducted in five electronic databases: PubMed, Embase, Scopus, Web of Science, and CINAHL. The search included articles published up to March 26, 2025. The strategy was designed to identify studies applying FMs in ECC settings, with full details in the checklist (Appendix eTable 1). Study Selection Studies were eligible if they met the following criteria: (1) applied FMs, defined as large-scale models pre-trained on diverse datasets using self-supervised learning and adapted to downstream tasks via fine-tuning or other transfer-learning methods; (2) focused on emergency or critical care; (3) were original research published in peer-reviewed journals; and (4) had full text available in English. Two independent reviewers (YK and YF) screened titles and abstracts, followed by full-text review using the predefined inclusion and exclusion criteria. Disagreements were resolved through discussion or, if needed, by consulting a third reviewer (FX), an expert in medical AI. Screening consistency was measured using Cohen’s kappa coefficient (κ). Data Extraction and Analysis Data extraction was performed independently by three reviewers (YK, XM, YF) using a standardized form, with discrepancies resolved by consensus and adjudicated by a fourth reviewer (XF) when needed. For each study, we extracted bibliographic details (year, authors, journal), FM architectures and base models (Appendix eTable 2), sample size and population characteristics, data modalities, clinical tasks, ECC setting (Emergency or Critical Care), and medical specialty (e.g., pediatric, respiratory, cardiovascular, radiology). These variables were selected a priori to capture the technical design and clinical context of FM applications in ECC. To ensure consistency, data modalities were defined in four categories: (1) free text, including triage notes, admission histories, radiology reports, and discharge summaries; (2) tabular structured data, such as laboratory values, vital signs, and administrative codes; (3) time-series data, including electrocardiogram waveforms and continuously monitored vital signs; and (4) imaging, such as radiographs, computed tomography (CT) scan, and magnetic resonance imaging (MRI) scan. Clinical applications were grouped into five predefined task domains: (1) outcome prediction (e.g., mortality, readmission), (2) information extraction from unstructured or multimodal data, (3) text generation including summaries and reports, (4) diagnosis through binary, multi-label, or differential classification, and (5) treatment recommendation, such as planning investigations or suggesting therapies. All modality and task summaries were computed at the aggregated ECC level without ED-ICU stratification to reflect overall literature trends. This framework provided a structured, consistent approach to characterizing how FMs have been applied across ECC. Results Study Characteristics Our search retrieved 1,443 studies. After removing duplicates (n = 321), 1,122 titles and abstracts were screened, of which 1,026 were excluded for: irrelevance to emergency or critical care (n = 407), not applying FMs (n = 243), not a peer-reviewed research article (n = 373), or non-English language (n = 3). Subsequently, 96 studies underwent full-text screening, then 47 were excluded for not focusing on emergency or critical care (n = 38), not involving FMs (n = 4), or not being peer-reviewed research articles (n = 5). 49 studies were ultimately included (Table 1 ). The selection process is shown in Fig. 1 . Screening consistency was excellent (inter-rater reliability: κ = 0.897 for title and abstract, κ = 0.812 for full text). Table 1 Characteristics of included studies applying foundation models in emergency and critical care. Author Year Foundation Model Name Data Modality Datasets Clinical Task Clinical Settings Medical Specialties Chen et al. 56 2021 EDisease, BERT Multimodal (Tabular Data + Free Text) 1.04M ED visits from NTUH; 305K from National Hospital Ambulatory ED data Information Extraction Critical Care Clinical Informatics Barash et al. 57 2023 ChatGPT-4 Free Text 40 real ED clinical cases (clinical notes and imaging) from 8 acute pathologies Treatment Recommendation Emergency Radiology Bushuven et al. 58 2023 ChatGPT Free Text 22 case vignettes (2 BLS and the 20 core PALS scenarios) Diagnosis Emergency Pediatric Fraser et al. 59 2023 ChatGPT-3.5, ChatGPT-4.0 Tabular Data 40 ED patient symptom entries Diagnosis Emergency Clinical Informatics Henriksson et al. 60 2023 Clinical KB-BERT Multimodal (Free Text + Tabular Data) Multi-center COVID-19 patient cohort Outcome Prediction Emergency Respiratory Huang et al. 61 2023 PABLO: Pretrained and Adapted BERT for Longitudinal Outcomes Tabular Data California and Florida Statewide EHR Databases Outcome Prediction Emergency Pediatric Savage et al. 62 2023 BioMed-RoBERTa Free Text MIMIC-III Information Extraction Critical Care Clinical Informatics Aityan et al. 63 2024 multiple LLMs (ChatGPT, Claude, Gemini) Multimodal (Free Text + Tabular Data + Time-Series + Imaging) 600 cases (300 sepsis and 300 non-sepsis) from the University archive, Google Scholar, and PubMed Diagnosis Emergency Clinical Informatics Akhondi-Asl et al. 64 2024 BioGPT-Large, LLaMa-7B (Fine-tuned), LLaMa-65B Free Text 1,916,538 notes from 32,454 PICU patients for training; 130 admission notes for evaluation Diagnosis Critical Care Pediatric Alessandri-Bonetti et al. 65 2024 ChatGPT-3.5, ChatGPT-4, Bard Free Text 50 questions with five multiple-choice answers from the ABLS exam Text Generation Emergency Burn Care Amacher et al. 66 2024 ChatGPT-4 Tabular Data Adult cardiac arrest patients' data at University Hospital Basel between 2012 and 2022 Outcome Prediction Critical Care Cardiovascular Haim et al. 67 2024 GPT-4 Free Text 100 consecutive adult ED patient records Outcome Prediction Emergency Clinical Informatics Chung et al. 68 2024 GPT-4 Turbo Free Text task-specific datasets constructed from 2 years of retrospective EHR data collected Outcome Prediction Critical Care Perioperative Medicine Ferri et al. 69 2024 DistilBERT Free Text 1.98M emergency medical call incidents Outcome Prediction Emergency Clinical Informatics Gimeno et al. 70 2024 GPT-4 Free Text 120 discharge instructions Text Generation Emergency Pediatric Glicksberg et al. 71 2024 GPT-4, Bio-Clinical-BERT Multimodal (Free Text + Tabular Data) 864,089 ED encounters from 7 hospitals within NYC Outcome Prediction Emergency Clinical Informatics Guo et al. 72 2024 FMSM (Foundation Model Stanford Medicine) Tabular Data Stanford Medicine EHR, SickKids EHR, MIMIC-IV Outcome Prediction Critical Care Clinical Informatics Huang et al. 73 2024 GPT-4 Free Text 5 synthetic ED encounter notes created by emergency physicians Text Generation Emergency Clinical Informatics Khaldi et al. 74 2024 ChatGPT-4o Free Text 20 common questions of patients about tracheotomy Text Generation Critical Care Otolaryngology Le Guellec et al. 75 2024 Vicuna 13B Free Text 2398 emergency brain MRI reports Information Extraction Emergency Radiology Lee et al. 76 2024 KLUE-BERT, KLUE-RoBERTa, KorBERT, KoBERT Free Text simulated clinician–patient conversations within 6 primary symptom scenarios in emergency triage rooms Information Extraction Emergency Clinical Informatics Levin et al. 77 2024 ChatGPT-4 Free Text Six neonatal ICU clinical case scenarios Treatment Recommendation Critical Care Neonatal Lin et al. 78 2024 multimodal model (combineTabNet and MacBERT) Multimodal (Free Text + Tabular Data) NTUH Retrospective Data Set (745,441 ED visits, 2009–2015); NTUH Prospective Data Set (901 ED visits, May 2020–Feb 2022) Outcome Prediction Emergency Clinical Informatics Masanneck et al. 79 2024 GPT-3.5, GPT-4, LLaMA 3 70B, Gemini 1.5, Mixtral 8x7b Free Text 124 independent emergency cases from the interdisciplinary ED of University Hospital Düsseldorf, Germany Outcome Prediction Emergency Clinical Informatics McCoy and Perlis 80 2024 GPT-4 Turbo, GPT-4-1106-preview Free Text 3059 individuals with a median age of 16 years (interquartile range, 13–18 years) Information Extraction Emergency Pediatric Rothchild et al. 82 2024 ChatGPT-4o Free Text 10 clinical vignettes on common facial trauma presentations Diagnosis Emergency Otolaryngology Saner et al. 83 2024 ChatGPT 4.0 Plus, Bard Free Text 10 simulated ICU patient cases, with each SOFA score calculated Information Extraction Critical Care Respiratory Scquizzato et al. 84 2024 ChatGPT Free Text 40 layperson questions about cardiac arrest and CPR co-created with the Sudden Cardiac Arrest UK community Text Generation Emergency Cardiovascular Seo et al. 85 2024 HyperCLOVA X Free Text 33 ED initial records were generated by 52 participants during the Healthcare Prompt-a-thon event Text Generation Emergency Clinical Informatics Sezgin et al. 86 2024 T5-small, T5-base, PEGASUS-PubMed, BART-Large-CNN Free Text 100 referral conversations among ED clinicians Text Generation Emergency Clinical Informatics Tang et al. 87 2024 GPT-4 Free Text 1,000 de-identified ED imaging reports from 20 most common study types at Stanford Health Care (2020–2023) Information Extraction Emergency Radiology Urquhart et al. 88 2024 GPT-4 API, ChatGPT, LLaMA 2 Free Text ICU admission notes from 11 episodes in 9 patients at Galway University Hospital, Ireland Text Generation Critical Care Clinical Informatics Wali et al. 89 2024 BioMistral 7B Multimodal (Tabular Data + Time-Series) The heart attack dataset from the IEEE Dataport Text Generation Emergency Cardiovascular Wang et al. 90 2024 GPT-3.5, GPT-4 Free Text Electronic Medical Records from the ED of Maoming People’s Hospital Diagnosis Emergency Neurology Yang et al. 91 2024 GPT-3.5, GPT-4, Claude 2, LLaMA 2-7b and 2-13b Free Text 219 MCQs from critical care pharmacotherapy courses at two U.S. colleges of pharmacy Text Generation Critical Care Pharmacotherapy Yau et al. 92 2024 ChatGPT-3.5, Bard Free Text 10 emergency medicine questions Text Generation Emergency Clinical Informatics Amacher et al. 93 2025 ChatGPT-4o Tabular Data 760 consecutive adult ICU patients with status epilepticus from University Hospital Basel (2005–2022) Outcome Prediction Critical Care Neurology Arslan et al. 94 2025 ChatGPT Plus (GPT-4), Copilot Pro (GPT-4) Free Text 468 ED patient cases from a large urban academic hospital Outcome Prediction Emergency Clinical Informatics Balta et al. 95 2025 ChatGPT-3.5, ChatGPT-4.0 Free Text 50 clinical critical care questions across five categories, representative of critical care medicine Treatment Recommendation Critical Care Clinical Informatics Broad et al. 96 2025 GPT-4o Free Text ESO Research Data Collaborative Diagnosis Emergency Pediatric Feng et al. 97 2025 GPT-4-turbo, GPT-3.5-turbo, Jurassic-2 Free Text 490 full-text EHR notes from 125 patients with prior life-threatening arrhythmias Diagnosis Critical Care Cardiovascular Levra et al. 98 2025 Multilingual BERT Free Text 30,320 EMRs from Humanitas Research Hospital Diagnosis Emergency Cardiovascular Ho et al. 99 2025 ChatGPT-3.5, ChatGPT-4.0, T5, LLaMA 2, Mistral-Large, Claude-3 Opus Free Text 70 pediatric emergency clinical vignettes adapted from ESI Outcome Prediction Emergency Clinical Informatics Miller et al. 100 2025 ChatGPT-4 Free Text 104 paramedic PCRs from cloud-based database (EMCE, NMETC) Diagnosis Emergency Clinical Informatics Pathak et al. 101 2025 RespBERT Free Text Radiology notes from Emory University and Grady Memorial Hospital Diagnosis Critical Care Respiratory Shekhar et al. 102 2025 ChatGPT-4oMini Free Text Allegheny County Dispatch Dataset (greater than 1,000,000 ambulance requests from 2015 to 2020) Outcome Prediction Emergency Prehospital Care Williams and Erstad 103 2025 ChatGPT-4, LLaMA 3.1 405B Free Text 43 medication-related PICO questions drawn from 6 CPGs (2023–2024) Treatment Recommendation Critical Care Pharmacotherapy Yang et al. 16 2025 ICU-GPT Tabular Data MIMIC-III; MIMIC-IV; MIMIC-IV-ED; MIMIC-IV-Note; eICU-CRD Information Extraction Critical Care Clinical Informatics Data Modality Adoption of Foundation Models in Emergency and Critical Care Studies (a) Prevalence of base data modalities used across included studies. Bars show the number of studies that used each modality, and the dot matrix below indicates single (a single filled dot) and multimodal (multiple filled dots in the same row) combinations of modalities used within the same study. (b) Modality task mapping. Each cell indicates the number of studies that use a given data modality for a specific clinical task. (c) Associations between data modalities, ECC settings, and medical specialties. The thickness of each flow reflects the number of modality-to-setting-to-topic associations rather than the number of unique studies. The “Unspecified” category includes ECC studies without a defined specialty or disease focus. Values reflect ECC-wide aggregation; ED- and ICU-specific modality distributions are not shown in this primary analysis. Studies using multiple modalities in Figs. 2 b and 2 c are counted once within each modality category (Free Text, Tabular, Imaging, Time-Series). A range of FM architectures have been applied in ECC, most commonly GPT family (35/49, 71.43%), BERT-based models, and the LLaMA series, each suited for different tasks. BERT, an encoder-only model for text understanding, has been used to identify ARDS from radiology notes 101 . It has also been applied to predict ICU admissions or in-hospital mortality 56 , showing strong performance even with limited data. GPT, a decoder-only generative model, has been tested in prehospital acute stroke screening to help identify lesion locations and responsible vessels, which highlights its reasoning and generative capabilities 90 . LLaMA is also a decoder-only, open-source, and computationally efficient model. It can enable on-site fine-tuning without an external server, making it well-suited for privacy-sensitive settings, as shown by its use in generating PICU differential diagnoses 64 and extracting magnetic resonance imaging (MRI) reports in the ED 75 . These applications demonstrate how different FM families are being adapted for ECC tasks; however, most studies fall short of demonstrating improvements in workflow efficiency, decision-making accuracy, or patient outcomes compared with standard practice (Appendix eTable 2). Building on these architecture-specific applications, most studies applied FMs to free text, which was the predominant data modality (85.7%, 42/49), supporting a wide range of applications (Fig. 2 ). Free text was input to support tasks such as outcome prediction, diagnosis, and text generation. For example, one study evaluated ChatGPT-4 as a conversational agent for analyzing ED admission notes to recommend radiology referrals and imaging selection 57 , whereas LLaMA-7B and BioGPT-Large were fine-tuned on PICU admission notes to generate differential diagnoses 64 . These applications highlight the versatility of FMs in leveraging unstructured narratives to inform high-acuity decision-making. However, timeliness and completeness can be limited if free text is used alone without complementary structured data like tabular EHR records. Tabular data was the second most common modality (24.49%, 12/49), most often used for outcome prediction (Fig. 2 ). For example, one study developed the Pretrained and Adapted BERT for Longitudinal Outcomes (PABLO) model to analyze structured EHR data, including demographics, diagnostic codes, and procedures, to predict non-accidental trauma (injury resulting from intentional harm or neglect for children and adolescents) 61 . Another study applied a transformer-based language model to sequences of diagnoses and laboratory results to estimate in-hospital mortality and readmission risk 72 . These studies illustrate how structured records, when modeled longitudinally, can support early risk identification and monitoring in ECC. However, essential time-series or imaging data is not well captured in tabular form, which points to the need for multimodal modeling. Time-series (4.08%, 2/49) and imaging data (2.04%, 1/49) were incorporated only into multimodal applications. For example, high-frequency physiologic signals, such as electrocardiogram traces and imaging inputs, were combined with free-text notes and structured vital signs using models like GPT and Gemini to support diagnosis and risk assessment for sepsis, stroke, and myocardial infarction 63 . Overall, 12.24% of studies (6/49) used explicitly multimodal inputs, most often combining free text with structured EHR data (Fig. 2 ). One study applied Clinical KB-BERT to jointly process clinical notes and structured records to predict 30-day mortality and 14-day readmission for COVID-19 patients, achieving superior performance compared with single-modality models 60 . These findings suggest that incorporating multiple data modalities can improve accuracy and interpretability, but multimodal FM development in ECC remains constrained by limited data accessibility, heterogeneity, and integration challenges. The figure integrates three elements: (1) major FM categories and representative model series, (2) the data modalities commonly used in ECC, and (3) the clinical tasks these models support. FM families span language, multimodal, vision, and audio architectures, which operate on free-text documentation, structured EHR data, time-series signals, and imaging. FM-enabled clinical applications include natural language processing tasks, such as information extraction and text generation, as well as outcome prediction, diagnostic reasoning, and treatment recommendation. Together, the figure depicts how FM types, data modalities, and clinical functions align within ECC workflows. Clinical Applications of Foundation Models in Emergency and Critical Care FMs have been explored across five main applications in ECC: outcome prediction, diagnosis, text generation, information extraction, and treatment recommendation (Fig. 3 ). The emphasis of these applications varies by clinical context and the data modality. In emergency care, studies emphasize rapid triage and diagnosis with limited, heterogeneous data. In contrast, in critical care, the focus shifts to continuous monitoring, dynamic risk stratification, and management of longitudinal, high-dimensional data streams. The following subsections highlight how these domains have been studied, including both the opportunities and the limitations of current FM applications in ECC. Outcome prediction was the most common FM application in ECC (28.6%, 14/49). These prediction categories mainly include mortality 60 , 66 , 93 , hospital resource utilization 68 , 71 , triage and acuity stratification 67 , 69 , 78 , 79 , 94 , 99 , 102 , disease-specific risks and complications 61 , and laboratory abnormalities 72 . In emergency care, applications emphasized rapid risk assessment. For example, the BERT-derived PABLO model used longitudinal EHR data, including demographics, diagnoses, and procedures, to predict non-accidental trauma in children, outperforming traditional ML methods 61 . Other studies used synthetic scenarios rather than real-time triage data. For instance, one study evaluated GPT-4 using pediatric vignettes and reported higher accuracy than LLaMA-2 in predicting ESI scores. However, the researchers acknowledged that vignette-based inputs deviate from actual clinical workflows 99 . In critical care, research focuses more on continuous risk stratification and longitudinal data monitoring. For instance, ChatGPT-4 predicted mortality and neurological outcomes after cardiac arrest with performance comparable to validated post-arrest scoring systems 66 . Diagnosis, including single-label multi-class disease diagnosis 58 , 82 , 100 , differential diagnosis 59 , 63 , 64 , and binary disease diagnosis 81 , 90 , 96 – 98 , 101 , was the second most common FM application in ECC (24.5%, 12/49). In emergency medicine, GPT achieved 93.9% diagnostic accuracy when tested on pediatric prehospital free-text descriptions provided by laypersons 58 . Another study used GPT to evaluate patient-entered symptom data, reporting variation in model performance 59 . In critical care, a fine-tuned LLaMa outperformed BioGPT and LLaMa for generating differential diagnoses from PICU admission notes 64 . Similarly, GPT classified arrhythmia recurrence from post-ablation EHR notes with 91.4% accuracy, exceeding both SapBERT (66.6%) and a rule-based algorithm (82.6%) 97 . FMs have also been applied to text generation (22.4%, 11/49), primarily focusing on automated clinical question answering 65 , 84 , 89 , 91 , 92 , discharge instruction generation 70 , 73 , 74 , initial clinical record drafting 85 , summarization of clinical dialogues 86 , and abstraction of patient records 88 in ECC contexts where efficient, precise communication is essential. For example, GPT-4 was evaluated for generating ED discharge instructions from synthetic free-text clinical notes, with better performance in the return precautions section compared to standard instructions 73 . Another study tested ChatGPT’s ability to answer common post-resuscitation questions from cardiac arrest survivors, relatives, and lay rescuers, finding that its responses were largely accurate and comprehensive. However, cardiopulmonary resuscitation (CPR)-specific answers were weaker 84 . These applications suggest potential for FMs to support documentation and patient education, though their role in direct clinical guidance remains unproven. FMs were applied to information extraction in 16.3% of studies (8/49), including patient-level disease concept embeddings 56 , contraindication identification 62 , symptom element extraction 75 , 76 , severity quantification 80 , 83 , 87 , and structured EHR extraction 16 from clinician notes and radiology reports to support efficient retrieval in high-intensity settings. For example, BioMed-RoBERTa extracted bleeding status from ICU physician notes, enabling more targeted best-practice alerts (14.8% improvement in applicability) 62 . While most models were BERT-based, newer models like Vicuna-13B demonstrate superior flexibility, effectively extracting diverse information, such as symptom presence and causal inferences, from emergency brain MRI reports 75 . Treatment recommendations represented another type of application of FMs (8.2%, 4/49), in which models suggested diagnostic investigation planning 57 and treatment planning based on diagnostic findings 77 , 95 , 103 . These systems integrated diverse clinical data, including physiological status, laboratory results, and patient history, to guide diagnostic uncertainty. One study used ChatGPT-4 to analyze retrospective ED admission records and recommend radiology examinations for acute conditions, achieving 95% concordance with the American College of Radiology Appropriateness Criteria (ACR AC) 57 . Overall, these studies illustrate the breadth of FM applications in ECC, while remaining exploratory and lacking real-world clinical integration. Discussion The growing interest in FMs for ECC reflects their potential to support high-stakes decisions by synthesizing diverse data sources under severe time constraints. Because errors in these environments carry serious consequences, model evaluation must extend beyond accuracy to include safety, fairness, reliability, and oversight. Despite their capabilities, current FMs do not yet fully meet demands in ECC for urgency, contextual nuance, and dependable human supervision. The literature highlights both the breadth of promising applications and the early, exploratory nature of existing work, providing a foundation for evaluating clinical utility and directing future development. Within this landscape, multimodal integration emerges as both a key strength and a central technical challenge, shaping the broader discussion of FM opportunities and limitations in ECC. A significant strength of FMs in ECC is their ability to integrate multimodal data 56 , 60 , 63 , 71 , 78 , 89 , including free text, structured records, time series, and imaging, to enable real-time interpretation of heterogeneous information. Incorporating temporal and spatial signals from continuous monitoring and radiologic imaging can enhance early detection of deterioration and support timely, clinically grounded decisions 104 . Despite their potential, multimodal FMs remain underused in ECC. Models like ViT 105 for radiographic interpretation, CLIP 106 – 108 for image-text alignment, BLIP 109 , 110 for vision-language report generation, and Wav2Vec 2.0 111,112 for acoustic features extraction, remain largely unexplored and warrant greater adaptation to ECC needs. Real-time multimodal deployment also faces operational challenges: asynchronous data streams, latency, noise, and imperfect alignment can degrade performance and interpretability 8 , 113 , 114 . Clinical viability depends on pairing advances in multimodal architectures with robust, automated pipelines for data harmonization, synchronization, and quality assurance, enabling dependable and supervised integration into ECC workflows 115 – 117 . Another advantage of FMs, particularly language models, is their potential to reduce documentation burden in ECC. They can automate the processing of clinical text, such as ED notes 73 , chief complaints 70 , and radiology reports 75 , 101 , freeing clinician time and improving efficiency. Models like GPT-4 have produced accurate and consistent documentation outputs 70 , 73 , while RespBERT 101 and Vicuna 75 have extracted structured information from free text with minimal tuning. However, current applications remain dependent on retrospective notes, such as triage summaries and discharge instructions, which often lack real-time relevance 56 , 57 , 61 , 68 , 73 , 77 , 99 . More systematically, incompleteness in documentation limits downstream reasoning and introduces bias 57 , 60 , 61 , 71 , 76 , 78 , 85 , 93 , 95 . Deciding what to record and omit reflects subjective judgment and resource constraints, a challenge well documented in other ML approaches that rely on imputation. These gaps in completeness and semantic consistency 56 , 66 , 96 , 97 , 101 , 103 hinder causal inference and model reliability 75 , 80 . Future FMs should therefore minimize temporal lags, treat missingness as potentially informative, and improve alignment across documentation to ensure both accuracy and fairness. Beyond data integration and documentation, FMs are poised to transform communication, coordination, and comprehension across both professional and lay interfaces in ECC. Within clinical teams, evaluations of digital scribe systems for summarizing ED consultation calls show that large language models can transcribe and structure complex dialogues in real time, supporting consistent handoffs and reducing information loss in high-acuity workflows 85 . At the lay interface, FMs show comparable promise in addressing gaps in health literacy and language accessibility. Studies of GPT-4-generated discharge instructions in multilingual pediatric emergency settings found that models can produce readable, linguistically concordant summaries of care, empowering caregivers to follow post-discharge recommendations 16 , 70 . Complementary work comparing GPT-based and standard explanations suggests that FMs can improve comprehension and satisfaction when outputs are contextualized and reviewed for accuracy and reproducibility 74 , 84 . These systems demonstrate early success but remain limited by ambient noise, privacy constraints, and semantic drift, emphasizing the need for context-aware fine-tuning. Yet disparities in completeness between English outputs and those in other languages, and the difficulty in achieving simplified reading levels, reveal that fluency does not ensure communicative equity. Collectively, these advances position FMs as communication intermediaries, bridging clinicians, patients, and data. Their responsible ECC deployment will require governance that provides transparency, interpretability, and empathy 73 . Limits and Challenges Despite the increasing integration of FMs into ECC, significant challenges persist. A major limitation concerns the quality and generalizability of data used for training and validation. Many studies have relied on simulated or synthetic datasets 95 , 103 , 58 , 74 , 77 based on idealized clinical cases 70 , 92 , which may overestimate effectiveness compared to real-world scenarios 77 . Small sample sizes 57 , 59 , 62 , 76 , 77 , 88 , 95 , 103 , single-center cohorts 56 , 93 , 94 , 97 , 98 , 100 , and omission of key patient information (e.g., demographics 70 , symptom trajectories 93 , diagnostic uncertainty 77 , 91 , 94 , 97 ) further reduce generalizability and introduce selection bias 64 . Temporal or demographic sampling biases, such as restrictions to specific seasons 94 or age groups 58 , also limit the applicability of the results across diverse ECC populations 59 , 61 , 70 , 72 , 80 , 97 , 98 . In addition, stringent privacy and data protection regulations (e.g., the Health Insurance Portability and Accountability Act [HIPAA] 62 , 80 , 97 , and the General Data Protection Regulation [GDPR] 88 ) limit access to real-world data, thereby restricting the scale and diversity of datasets required for robust FM development 57 , 64 , 71 , 72 . Addressing these limitations requires treating FMs as socio-technical systems, pairing advances such as federated validation, real-time integration, and factual-consistency checks with ethical safeguards and implementation planning. FM performance was generally reported using task-specific metrics, but inconsistent reporting hindered meaningful comparison across studies. Beyond data issues, evaluation methodologies often lacked rigor. Some studies relied on retrospective designs without external validation 93 or prospective testing 65 , both of which are critical for demonstrating real-world utility 64 , 71 , 98 . Reproducibility was further limited by single-run evaluations, absent version control, and reliance on subjective reference standards, all of which hinder fair comparison with existing methods 94 , 56 , 58 , 77 , 96 , . These barriers stem not only from technical gaps but also from structural challenges: patient unpredictability, dynamic care pathways, and intense time pressure complicate prospective validation in ECC, while regulatory uncertainty and institutional hesitancy further slow translation. Potential solutions include federated multi-institutional validation, standardized benchmarking pipelines with version tracking, and pragmatic trials incorporating clinician oversight. These challenges are therefore not only technical but structural, requiring strategies that extend beyond simply expanding datasets. Accordingly, evaluation should consider workflow impact, communication, team dynamics, and clinician trust, not accuracy alone. Persistent challenges in FM research arise not only from generic barriers in data quality and evaluation but from fundamental misalignments between model design and clinical reasoning. Language models, mainly trained on static internet text, capture linguistic patterns rather than the real-time, causal, and context-dependent logic that ECC requires 67 , 101 , 118 . This gap between statistical pattern-matching and the causal, contextual reasoning used by clinicians constitutes an epistemic mismatch. It helps explain why superficially strong test-set performance may not translate into safe decision support at the bedside. The absence of streaming vitals, dynamic documentation, and interprofessional communication in most training corpora leaves FMs linguistically fluent but clinically brittle. In acute settings, they may generate context-blind or unsafe outputs, including misprioritized triage, missed deterioration, or inappropriate treatment advice that undermine clinician judgment and amplify bias 59 , 79 , 82 , 94 , 95 , 103 . By contrast, low-risk and verifiable assistive uses, such as structured note generation, ICU documentation summarization, or extraction of key findings, may improve efficiency when outputs remain easily reviewable 70 , 85 , 88 . Autonomous diagnostic or treatment decisions, however, remain ethically indefensible until models demonstrate contextual reasoning, transparency, and validated safety. In addition, developing and deploying FMs in ECC entails substantial computational, financial, and environmental burdens 119 – 121 : training and tuning require extensive energy and water resources, and safe clinical deployment demands costly infrastructure, monitoring, and validation workflows 122 . Recent analyses show AI’s growing carbon and water footprint in healthcare, highlighting the need for cost and sustainability assessments before large-scale clinical adoption 121 , 123 . Ethical concerns also emerged across the reviewed studies. Limited external validation and single-center sampling not only underdetermine reproducibility but also raise equity concerns, as underrepresented populations may be disproportionately affected. Few studies conducted subgroup analyses, leaving open the risk that FMs could exacerbate existing disparities in ECC. In addition, model-level risks such as hallucinations 65 , 88 , 95 , prompt sensitivity 58 , 59 , 70 , 88 , 91 , and the black-box nature of many architectures undermine reliability and accountability in high-stakes settings 77 , 94 . Without interpretable mechanisms or traceable reasoning pathways, clinicians cannot verify whether outputs are grounded in valid medical knowledge 71 , which erodes confidence and complicates regulatory oversight. These issues resonate with broader debates in emergency AI, where even minor errors can have severe consequences 124 , and align with reviews of generative AI that emphasize transparency, accountability, and fairness as essential safeguards 125 . Addressing these gaps through fairness auditing, interpretable modeling, hallucination detection, and clear governance frameworks will be critical to ensure safe and equitable FM deployment in ECC. Beyond these challenges, the introduction of FMs into ECC also reshapes long-standing professional identities rooted in clinical autonomy, teamwork, and shared accountability. Studies show that clinicians report AI systems as encroaching on their expertise or disrupting traditional hierarchies of judgment and responsibility 126 , 127 . As FMs begin to support documentation, triage, or communication tasks, clinicians risk becoming supervisors of algorithmic output rather than active decision-makers 128 . This shift can undermine trust, situational awareness, and engagement during critical events. As demonstrated by Bienefeld et al., ensuring effective human-AI collaboration will require clear role delineation, workflow integration, and training that preserve clinician judgment and accountability in life-critical settings 129 . Future Directions Future research on FMs in ECC should integrate methodological rigor with ethical and regulatory considerations to ensure safe, trustworthy deployment 130 , 131 . Building multicenter and demographically diverse longitudinal datasets will improve generalizability and reduce bias 132 – 134 . Beyond subgroup fairness evaluations, validation must align with emerging governance frameworks such as the Food and Drug Administration (FDA)’s Predetermined Change Control Plans (PCCP) 135 for managing model updates, Centers for Medicare & Medicaid Services (CMS)’s requirements for human review of automated determinations 136 , and National Institute of Standards and Technology (NIST)’s AI Risk Management Framework 137 for post-deployment monitoring. Reconciling these frameworks is critical, as continuous learning models inherently challenge regulatory expectations for stability and traceability 138 . Evaluations should therefore quantify not only accuracy but also impacts on clinician trust, workflow integration, and patient safety in life-critical settings 139 – 141 . Embedding automated fact-checking, uncertainty estimation, and interpretable reasoning can support traceability and compliance 142 – 144 . Real-time interoperability with EHR and monitoring systems must preserve clinician oversight and accountability 145 . Sustained progress will depend on cross-disciplinary collaboration to harmonize oversight mechanisms while balancing innovation with safety, transparency, and equity 146 . Conclusion This scoping review emphasizes the growing importance of FMs in ECC, a field defined by clinical urgency, high complexity, and the need for rapid information synthesis. While language models dominate current applications, emerging work on multimodal architectures signals a move toward more comprehensive, context-aware support tools. Most current studies remain constrained by small or single-center datasets, retrospective designs, and limited external validation, highlighting the early and fragmented nature of the evidence base. Despite accelerating research activity, no FM has yet demonstrated verified clinical deployment or measurable impact on patient outcomes. Future work must prioritize rigorous validation, cross-system interoperability, and consistent, safety-oriented evaluation frameworks, while also addressing fairness, transparency, and reproducibility. Progress along these dimensions will be essential for realizing clinically meaningful, trustworthy, and equitable FM applications in life-critical ECC settings. Declarations Funding: L.C. is funded by the National Institute of Health through DS-I Africa U54 TW012043-01 and Bridge2AI OT2OD032701, the National Science Foundation through ITEST #2148451, a grant of the Boston-Korea Innovative Research Project (RS-2024-00403047) and a grant of the Korea Health Technology R&D Project (RS-2024-00439677) through the Korea Health Industry Development Institute (KHIDI) as funded by the Ministry of Health & Welfare, Republic of Korea. Author Contributions: F.X. conceptualized the study and led the work. Y.K. conducted a literature search. Y.K., Y.F., and F.X. screened the titles, abstracts, and full texts. Y.K., X.M., Y.F., and F.X. conducted data extraction. Y.K. drafted the initial manuscript. Y.K. and X.M. performed data synthesis and incorporated revisions based on feedback. Y.K., X.M., A.M., Z.X., Z.S., C.G., D.W., J.H., Q.W., A.H., R.Z., M.L., G.S., E.L., X.X., M.P., N.L., L.C., & F.X. revised the manuscript. F.X. supervised the study. All authors read and approved the final version of the manuscript. Competing Interests: Michael Puskarich has served on scientific advisory boards for Inflammatix LLC, Cvtovale LLC, and Opticyte LLC, companies working in the sepsis diagnosis and prognosis space. References Choi, A. et al. Development of a machine learning-based clinical decision support system to predict clinical deterioration in patients visiting the emergency department. Sci. Rep. 13, 8561 (2023). Kwon, J.-M. et al. Validation of deep-learning-based triage and acuity score using a large national dataset. PLoS One 13, e0205836 (2018). Chang, H., Yu, J. Y., Yoon, S., Kim, T. & Cha, W. C. Machine learning-based suggestion for critical interventions in the management of potentially severe conditioned patients in emergency department triage. Sci. Rep. 12, 10537 (2022). Clifton, D. A. et al. A large-scale clinical validation of an integrated monitoring system in the emergency department. IEEE J. Biomed. Health Inform. 17, 835–842 (2013). Sanchez-Pinto, L. N., Luo, Y. & Churpek, M. M. Big data and data science in critical care. Chest 154, 1239–1248 (2018). Liu, N. et al. Leveraging large-scale electronic health records and interpretable machine learning for clinical decision making at the emergency department: Protocol for system development and validation. JMIR Res. Protoc. 11, e34201 (2022). Chen, C.-H. et al. Emergency department disposition prediction using a deep neural network with integrated clinical narratives and structured data. Int. J. Med. Inform. 139, 104146 (2020). King, Z. et al. Machine learning for real-time aggregated prediction of hospital admission for emergency patients. NPJ Digit. Med. 5, 104 (2022). Candel, B. G. J. et al. The effect of treatment and clinical course during Emergency Department stay on severity scoring and predicted mortality risk in Intensive Care patients. Crit. Care 26, 112 (2022). Gunnerson, K. J. et al. Association of an emergency department-based intensive care unit with survival and inpatient intensive care unit admissions. JAMA Netw. Open 2, e197584 (2019). Kang, D.-Y. et al. Artificial intelligence algorithm to predict the need for critical care in prehospital emergency medical services. Scand. J. Trauma Resusc. Emerg. Med. 28, 17 (2020). Gravesteijn, B. Y., Steyerberg, E. W. & Lingsma, H. F. Modern learning from big data in critical care: Primum non nocere. Neurocrit. Care 37, 174–184 (2022). Deasy, J., Liò, P. & Ercole, A. Dynamic survival prediction in intensive care units from heterogeneous time series without the need for variable selection or curation. Sci. Rep. 10, 22129 (2020). Hu, Y. et al. Use of real-time information to predict future arrivals in the emergency department. Ann. Emerg. Med. 81, 728–737 (2023). Yoo, J. et al. A real-time autonomous dashboard for the emergency department: 5-year case study. JMIR MHealth UHealth 6, e10666 (2018). Yang, Z. et al. Large language model-based critical care big data deployment and extraction: Descriptive analysis. JMIR Med. Inform. 13, e63216 (2025). Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C. & Faisal, A. A. The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nat. Med. 24, 1716–1720 (2018). Kachman, M. M., Brennan, I., Oskvarek, J. J., Waseem, T. & Pines, J. M. How artificial intelligence could transform emergency care. Am. J. Emerg. Med. 81, 40–46 (2024). Goh, E. et al. Physician clinical decision modification and bias assessment in a randomized controlled trial of AI assistance. Commun. Med. (Lond.) 5, 59 (2025). Ataman, M. G. & Sarıyer, G. Predicting waiting and treatment times in emergency departments using ordinal logistic regression models. Am. J. Emerg. Med. 46, 45–50 (2021). Wu, T., Wei, Y., Wu, J., Yi, B. & Li, H. Logistic regression technique is comparable to complex machine learning algorithms in predicting cognitive impairment related to post intensive care syndrome. Sci. Rep. 13, 2485 (2023). Subudhi, S. et al. Comparing machine learning algorithms for predicting ICU admission and mortality in COVID-19. NPJ Digit. Med. 4, 87 (2021). Hinson, J. S. et al. Multisite implementation of a workflow-integrated machine learning system to optimize COVID-19 hospital admission decisions. NPJ Digit. Med. 5, 94 (2022). Marafino, B. J., Davies, J. M., Bardach, N. S., Dean, M. L. & Dudley, R. A. N-gram support vector machines for scalable procedure and diagnosis classification, with applications to clinical free text data from the intensive care unit. J. Am. Med. Inform. Assoc. 21, 871–875 (2014). Verdaasdonk, M. J. A. & M. de Carvalho, R. From predictions to recommendations: Tackling bottlenecks and overstaying in the Emergency Room through a sequence of Random Forests. Healthc. Anal. (N. Y.) 2, 100040 (2022). Zheng, L., Xue, Y.-J., Yuan, Z.-N. & Xing, X.-Z. Explainable SHAP-XGBoost models for pressure injuries among patients requiring with mechanical ventilation in intensive care unit. Sci. Rep. 15, 9878 (2025). Chen, Z., Li, T., Guo, S., Zeng, D. & Wang, K. Machine learning-based in-hospital mortality risk prediction tool for intensive care unit patients with heart failure. Front. Cardiovasc. Med. 10, 1119699 (2023). Liu, M., Guo, C. & Guo, S. An explainable knowledge distillation method with XGBoost for ICU mortality prediction. Comput. Biol. Med. 152, 106466 (2023). Yun, H., Choi, J. & Park, J. H. Prediction of critical care outcome for adult patients presenting to emergency department using initial triage information: An XGBoost algorithm analysis. JMIR Med. Inform. 9, e30770 (2021). Duckworth, C. et al. Using explainable machine learning to characterise data drift and detect emergent health risks for emergency department admissions during COVID-19. Sci. Rep. 11, 23017 (2021). Alghatani, K., Ammar, N., Rezgui, A. & Shaban-Nejad, A. Predicting Intensive Care unit length of stay and mortality using patient vital signs: Machine learning model development and validation. JMIR Med. Inform. 9, e21347 (2021). Ivanov, O. et al. Improving ED Emergency Severity Index acuity assignment using machine learning and clinical natural language processing. J. Emerg. Nurs. 47, 265–278.e7 (2021). Yao, L.-H., Leung, K.-C., Tsai, C.-L., Huang, C.-H. & Fu, L.-C. A novel deep learning-based system for triage in the emergency department using electronic medical records: Retrospective cohort study. J. Med. Internet Res. 23, e27008 (2021). Park, J. J. et al. Convolutional-neural-network-based diagnosis of appendicitis via CT scans in patients with acute abdominal pain presenting in the emergency department. Sci. Rep. 10, 9556 (2020). Arntfield, R. et al. Development of a convolutional neural network to differentiate among the etiology of similar appearing pathological B lines on lung ultrasound: a deep learning study. BMJ Open 11, e045120 (2021). Le, S. et al. Convolutional neural network model for Intensive Care unit acute kidney injury prediction. Kidney Int. Rep. 6, 1289–1298 (2021). Kessler, S. et al. Predicting readmission to the cardiovascular intensive care unit using recurrent neural networks. Digit. Health 9, 20552076221149529 (2023). Gandin, I., Scagnetto, A., Romani, S. & Barbati, G. Interpretability of time-series deep learning models: A study in cardiovascular patients admitted to Intensive care unit. J. Biomed. Inform. 121, 103876 (2021). Thorsen-Meyer, H.-C. et al. Discrete-time survival analysis in the critically ill: a deep learning approach using heterogeneous data. NPJ Digit. Med. 5, 142 (2022). Hyland, S. L. et al. Early prediction of circulatory failure in the intensive care unit using machine learning. Nat. Med. 26, 364–373 (2020). Li, J., Cairns, B. J., Li, J. & Zhu, T. Generating synthetic mixed-type longitudinal electronic health records for artificial intelligent applications. NPJ Digit. Med. 6, 98 (2023). Rahmatinejad, Z. et al. A comparative study of explainable ensemble learning and logistic regression for predicting in-hospital mortality in the emergency department. Sci. Rep. 14, 3406 (2024). Elhazmi, A. et al. Machine learning decision tree algorithm role for predicting mortality in critically ill adult COVID-19 patients admitted to the ICU. J. Infect. Public Health 15, 826–834 (2022). Goto, T., Camargo, C. A., Jr, Faridi, M. K., Freishtat, R. J. & Hasegawa, K. Machine learning-based prediction of clinical outcomes for children during emergency department triage. JAMA Netw. Open 2, e186937 (2019). Liu, Z., Shu, W., Li, T., Zhang, X. & Chong, W. Interpretable machine learning for predicting sepsis risk in emergency triage patients. Sci. Rep. 15, 887 (2025). Brown, T. B. et al. Language Models are Few-Shot Learners. https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf . Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of deep bidirectional Transformers for language understanding. arXiv [cs.CL] (2018). Touvron, H. et al. LLaMA: Open and efficient foundation language models. arXiv [cs.CL] (2023). Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv [cs.CV] (2020). Radford, A. et al. Learning transferable visual models from natural language supervision. arXiv [cs.CV] (2021). Li, J., Li, D., Xiong, C. & Hoi, S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv [cs.CV] (2022). Naemi, A. et al. Machine learning techniques for mortality prediction in emergency departments: a systematic review. BMJ Open 11, e052663 (2021). Piliuk, K. & Tomforde, S. Artificial intelligence in emergency medicine. A systematic literature review. Int. J. Med. Inform. 180, 105274 (2023). Yang, Z., Cui, X. & Song, Z. Predicting sepsis onset in ICU using machine learning models: a systematic review and meta-analysis. BMC Infect. Dis. 23, 635 (2023). Preiksaitis, C. et al. The role of large language models in transforming emergency medicine: Scoping review. JMIR Med. Inform. 12, e53787 (2024). Chen, Y.-P., Lo, Y.-H., Lai, F. & Huang, C.-H. Disease concept-embedding based on the self-supervised method for medical information extraction from electronic health records and disease retrieval: Algorithm development and validation study. J. Med. Internet Res. 23, e25113 (2021). Barash, Y., Klang, E., Konen, E. & Sorin, V. ChatGPT-4 assistance in optimizing emergency department radiology referrals and imaging selection. J. Am. Coll. Radiol. 20, 998–1003 (2023). Bushuven, S. et al. ‘ChatGPT, can you help me save my child’s life?’ - diagnostic accuracy and supportive capabilities to lay rescuers by ChatGPT in prehospital basic life support and paediatric advanced life support cases - an in-silico analysis. J. Med. Syst. 47, 123 (2023). Fraser, H. et al. Comparison of diagnostic and triage accuracy of Ada Health and WebMD symptom checkers, ChatGPT, and physicians for patients in an emergency department: Clinical data analysis study. JMIR MHealth UHealth 11, e49995 (2023). Henriksson, A., Pawar, Y., Hedberg, P. & Nauclér, P. Multimodal fine-tuning of clinical language models for predicting COVID-19 outcomes. Artif. Intell. Med. 146, 102695 (2023). Huang, D., Cogill, S., Hsia, R. Y., Yang, S. & Kim, D. Development and external validation of a pretrained deep learning model for the prediction of non-accidental trauma. NPJ Digit. Med. 6, 131 (2023). Savage, T., Wang, J. & Shieh, L. A large language model screening tool to target patients for best Practice Alerts: Development and validation. JMIR Med. Inform. 11, e49886 (2023). Aityan, S. K. et al. Integrated AI medical emergency diagnostics advising system. Electronics (Basel) 13, 4389 (2024). Akhondi-Asl, A. et al. Comparing the quality of domain-specific versus general language models for artificial intelligence-generated differential diagnoses in PICU patients. Pediatr. Crit. Care Med. 25, e273–e282 (2024). Alessandri-Bonetti, M., Liu, H. Y., Donovan, J. M., Ziembicki, J. A. & Egro, F. M. A comparative analysis of ChatGPT, ChatGPT-4, and Google Bard performances at the Advanced Burn Life Support exam. J. Burn Care Res. 45, 945–948 (2024). Amacher, S. A. et al. Prediction of outcomes after cardiac arrest by a generative artificial intelligence model. Resusc. Plus 18, 100587 (2024). Haim, G. B. et al. Evaluating large language model-assisted emergency triage: A comparison of acuity assessments by GPT-4 and medical experts. J. Clin. Nurs. (2024) doi: 10.1111/jocn.17490 . Chung, P. et al. Large language model capabilities in perioperative risk prediction and prognostication. JAMA Surg. 159, 928–937 (2024). Ferri, P. et al. Deep continual learning for medical call incidents text classification under the presence of dataset shifts. Comput. Biol. Med. 175, 108548 (2024). Gimeno, A., Krause, K., D’Souza, S. & Walsh, C. G. Completeness and readability of GPT-4-generated multilingual discharge instructions in the pediatric emergency department. JAMIA Open 7, ooae050 (2024). Glicksberg, B. S. et al. Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency room. J. Am. Med. Inform. Assoc. 31, 1921–1928 (2024). Guo, L. L. et al. A multi-center study on the adaptability of a shared foundation model for electronic health records. NPJ Digit. Med. 7, 171 (2024). Huang, T. et al. Patient-representing population’s perceptions of GPT-generated versus standard emergency department discharge instructions: Randomized blind survey assessment. J. Med. Internet Res. 26, e60336 (2024). Khaldi, A. et al. Accuracy of ChatGPT responses on tracheotomy for patient education. Eur. Arch. Otorhinolaryngol. 281, 6167–6172 (2024). Le Guellec, B. et al. Performance of an Open-Source large Language Model in extracting information from free-text radiology reports. Radiol. Artif. Intell. 6, e230364 (2024). Lee, S. et al. Deep learning-based natural language processing for detecting medical symptoms and histories in emergency patient triage. Am. J. Emerg. Med. 77, 29–38 (2024). Levin, C., Kagan, T., Rosen, S. & Saban, M. An evaluation of the capabilities of language models and nurses in providing neonatal clinical decision support. Int. J. Nurs. Stud. 155, 104771 (2024). Lin, Y.-T., Deng, Y.-X., Tsai, C.-L., Huang, C.-H. & Fu, L.-C. Interpretable deep learning system for identifying critical patients through the prediction of triage level, hospitalization, and length of stay: Prospective study. JMIR Med. Inform. 12, e48862 (2024). Masanneck, L. et al. Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: Comparative study. J. Med. Internet Res. 26, e53297 (2024). McCoy, T. H., Jr & Perlis, R. H. Dimensional measures of psychopathology in children and adolescents using large language models. Biol. Psychiatry 96, 940–947 (2024). Oliveira, L. L., Jiang, X., Babu, A. N., Karajagi, P. & Daneshkhah, A. Effective Natural Language Processing algorithms for early alerts of gout flares from chief complaints. Forecasting 6, 224–238 (2024). Rothchild, E., Baker, C., Smith, I. T., Tanna, N. & Ricci, J. A. Evaluating the utility of ChatGPT in diagnosing and managing maxillofacial trauma. J. Craniofac. Surg. 36, 1237–1241 (2025). Saner, F. H., Saner, Y. M., Abufarhaneh, E., Broering, D. C. & Raptis, D. A. Comparative analysis of artificial intelligence (AI) languages in predicting Sequential Organ Failure Assessment (SOFA) scores. Cureus 16, e59662 (2024). Scquizzato, T. et al. Testing ChatGPT ability to answer laypeople questions about cardiac arrest and cardiopulmonary resuscitation. Resuscitation 194, 110077 (2024). Seo, J. et al. Evaluation framework of large language models in medical documentation: Development and usability study. J. Med. Internet Res. 26, e58329 (2024). Sezgin, E., Sirrianni, J. W. & Kranz, K. Evaluation of a digital scribe: Conversation summarization for Emergency Department consultation calls. Appl. Clin. Inform. 15, 600–611 (2024). Tang, K., Reisler, J., Rose, C., Suffoletto, B. & Kim, D. GPT-4 accurately classifies the clinical actionability of emergency department imaging reports. J. Med. Artif. Intell. 7, 36–36 (2024). Urquhart, E. et al. A pilot feasibility study comparing large language models in extracting key information from ICU patient text records from an Irish population. Intensive Care Med. Exp. 12, 71 (2024). Wali, T., Bolatbekov, A., Maimaitijiang, E., Salman, D. & Mamatjan, Y. A novel recommender framework with chatbot to stratify heart attack risk. Discov Med (Cham) 1, 161 (2024). Wang, X. et al. Performance of ChatGPT on prehospital acute ischemic stroke and large vessel occlusion (LVO) stroke screening. Digit. Health 10, 20552076241297127 (2024). Yang, H. et al. Evaluating accuracy and reproducibility of large language model performance on critical care assessments in pharmacy education. Front. Artif. Intell. 7, 1514896 (2024). Yau, J. Y.-S. et al. Accuracy of prospective assessments of 4 large language model chatbot responses to patient questions about emergency care: Experimental comparative study. J. Med. Internet Res. 26, e60291 (2024). Amacher, S. A. et al. Can the large language model ChatGPT-4omni predict outcomes in adult patients with status epilepticus? Epilepsia 66, 674–685 (2025). Arslan, B., Nuhoglu, C., Satici, M. O. & Altinbilek, E. Evaluating LLM-based generative AI tools in emergency triage: A comparative study of ChatGPT Plus, Copilot Pro, and triage nurses. Am. J. Emerg. Med. 89, 174–181 (2025). Balta, K. Y., Javidan, A. P., Walser, E., Arntfield, R. & Prager, R. Evaluating the appropriateness, consistency, and readability of ChatGPT in critical care recommendations. J. Intensive Care Med. 40, 184–190 (2025). Broad, A. et al. Factors associated with abusive head trauma in young children presenting to Emergency Medical Services using a large language model. Prehosp. Emerg. Care 29, 227–237 (2025). Feng, R. et al. Engineering of generative artificial intelligence and natural language processing models to accurately identify arrhythmia recurrence. Circ. Arrhythm. Electrophysiol. 18, e013023 (2025). Levra, A. G. et al. A large language model-based clinical decision support system for syncope recognition in the emergency department: A framework for clinical workflow integration. Eur. J. Intern. Med. 131, 113–120 (2025). Ho, B. et al. Evaluation of generative artificial intelligence models in predicting pediatric Emergency Severity Index levels. Pediatr. Emerg. Care 41, 251–255 (2025). Miller, E. D. et al. Accuracy of commercial large language model (ChatGPT) to predict the diagnosis for prehospital patients suitable for ambulance transport decisions: Diagnostic accuracy study. Prehosp. Emerg. Care 29, 238–242 (2025). Pathak, A., Marshall, C., Davis, C., Yang, P. & Kamaleswaran, R. RespBERT: A multi-site validation of a Natural Language Processing algorithm, of radiology notes to identify acute respiratory distress syndrome (ARDS). IEEE J. Biomed. Health Inform. 29, 1455–1463 (2025). Shekhar, A. C. et al. Use of a large language model (LLM) for ambulance dispatch and triage. Am. J. Emerg. Med. 89, 27–29 (2025). Williams, B. & Erstad, B. L. Analysis of responses from artificial intelligence programs to medication-related questions derived from critical care guidelines. Am. J. Health. Syst. Pharm. (2025) doi: 10.1093/ajhp/zxaf075 . Khader, F. et al. Medical transformer for multimodal survival prediction in intensive care: integration of imaging and non-imaging data. Sci. Rep. 13, 10666 (2023). Park, S. et al. Multi-task vision transformer using low-level chest X-ray feature corpus for COVID-19 diagnosis and severity quantification. Med. Image Anal. 75, 102299 (2022). Lian, C., Zhou, H.-Y., Liang, D., Qin, J. & Wang, L. Efficient medical vision-language alignment through adapting masked vision models. IEEE Trans. Med. Imaging PP, 1–1 (2025). Yu, X., Zhang, L., Wu, Z. & Zhu, D. Core-periphery multi-modality feature alignment for zero-shot medical image analysis. IEEE Trans. Med. Imaging PP, (2024). Wang, P., Zhang, H. & Yuan, Y. MCPL: Multi-modal Collaborative Prompt Learning for medical vision-language model. IEEE Trans. Med. Imaging 43, 4224–4235 (2024). Li, Y. et al. Visual analytics for efficient image exploration and user-guided image captioning. IEEE Trans. Vis. Comput. Graph. 30, 2875–2887 (2024). Ji, J., Hou, Y., Chen, X., Pan, Y. & Xiang, Y. Vision-language model for generating textual descriptions from clinical images: Model development and validation study. JMIR Form. Res. 8, e32690 (2024). Zhang, X., Zhang, X., Chen, W., Li, C. & Yu, C. Improving speech depression detection using transfer learning with wav2vec 2.0 in low-resource environments. Sci. Rep. 14, 9543 (2024). Klempir, O., Skryjova, A., Tichopad, A. & Krupicka, R. Ranking pre-trained speech embeddings in Parkinson’s disease detection: Does Wav2Vec 2.0 outperform its 1.0 version across speech modes and languages? Comput. Struct. Biotechnol. J. 27, 2584–2601 (2025). Ardic, N. & Dinc, R. Emerging trends in multi-modal artificial intelligence for clinical decision support: A narrative review. Health Informatics J. 31, 14604582251366141 (2025). Zhou, Y. et al. A contrastive learning approach for ICU false arrhythmia alarm reduction. Sci. Rep. 12, 4689 (2022). Hao, Y. et al. Multimodal integration in healthcare: Development with applications in disease management (preprint). J. Med. Internet Res. 27, e76557 (2025). Barrit, S. et al. Intracranial multimodal monitoring in neurocritical care (Neurocore-iMMM): an open, decentralized consensus. Crit. Care 28, 427 (2024). Subasri, V. et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Netw. Open 8, e2513685 (2025). Biesheuvel, L. A. et al. Large language models in critical care. J. Intensive Med. 5, 113–118 (2025). El Arab, R. A. & Al Moosa, O. A. Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare. NPJ Digit. Med. 8, 548 (2025). Luccioni, A. S., Jernite, Y. & Strubell, E. Power hungry processing: Watts driving the cost of AI deployment? arXiv [cs.LG] (2023). Katirai, A. The environmental costs of artificial intelligence for healthcare. Asian Bioeth. Rev. 16, 527–538 (2024). Khanna, N. N. et al. Economics of artificial Intelligence in healthcare: Diagnosis vs. Treatment. Healthcare (Basel) 10, 2493 (2022). Ueda, D. et al. Climate change and artificial intelligence in healthcare: Review and recommendations towards a sustainable future. Diagn. Interv. Imaging 105, 453–459 (2024). Nord-Bronzyk, A. et al. Assessing risk in implementing new artificial intelligence triage tools-how much risk is reasonable in an already risky world? Asian Bioeth. Rev. 17, 187–205 (2025). Ning, Y. et al. Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist. Lancet Digit. Health 6, e848–e856 (2024). Henry, K. E. et al. Human-machine teaming is key to AI adoption: clinicians’ experiences with a deployed machine learning system. NPJ Digit. Med. 5, 97 (2022). Ling Kuo, R. Y. et al. Stakeholder perspectives towards diagnostic artificial intelligence: a co-produced qualitative evidence synthesis. EClinicalMedicine 71, 102555 (2024). Smith, H., Downer, J. & Ives, J. Clinicians and AI use: where is the professional guidance? J. Med. Ethics 50, 437–441 (2024). Bienefeld, N., Keller, E. & Grote, G. Human-AI teaming in critical care: A comparative analysis of data scientists’ and clinicians' perspectives on AI augmentation and automation. J. Med. Internet Res. 26, e50130 (2024). Moons, K. G. M. et al. PROBAST + AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 388, e082505 (2025). Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. https://www.who.int/publications/i/item/9789240084759 (2025). Collins, G. S. et al. TRIPOD + AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385, e078378 (2024). Arora, A. et al. The value of standards for health datasets in artificial intelligence-based applications. Nat. Med. 29, 2929–2938 (2023). Rockenschaub, P. et al. The impact of multi-institution datasets on the generalizability of machine learning prediction models in the ICU. Crit. Care Med. 52, 1710–1721 (2024). Food and Drug Administration Staff. https://www.fda.gov/media/180978/download (2024). Thomas, K. S. et al. Prior authorization and utilization management for post-acute home health in Medicare Advantage: the motivations, players, processes, unique challenges, and impacts on patient care. Health Aff. Sch. 3, qxaf020 (2025). Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0): AI RMF (1.0) . https://doi.org/10.6028/NIST.AI.100-1 (2023) doi:10.6028/nist.ai.100-1. Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 25, 44–56 (2019). Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28, 924–933 (2022). Liu, X. et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat. Med. 26, 1364–1374 (2020). Cruz Rivera, S. et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat. Med. 26, 1351–1363 (2020). Yang, R. et al. Retrieval-augmented generation for generative artificial intelligence in health care. Npj Health Syst. 2, 1–5 (2025). Gargari, O. K. & Habibi, G. Enhancing medical AI with retrieval-augmented generation: A mini narrative review. Digit. Health 11, 20552076251337177 (2025). Sadeghi, Z. et al. A review of Explainable Artificial Intelligence in healthcare. Comput. Electr. Eng. 118, 109370 (2024). Foldy, S. et al. Public Health FHIR® Playbook. https://www.cdc.gov/data-interoperability/media/pdfs/PHFIC_Public-Health-FHIR-Playbook.pdf (2023). Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles. U.S. Food and Drug Administration https://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles (2024). Additional Declarations Competing interest reported. Michael Puskarich has served on scientific advisory boards for Inflammatix LLC, Cvtovale LLC, and Opticyte LLC, companies working in the sepsis diagnosis and prognosis space. Supplementary Files Appendix.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8338830","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":572650711,"identity":"f3b2233d-9d91-480d-b1ad-490855a4aea1","order_by":0,"name":"Yanqing Kong","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Yanqing","middleName":"","lastName":"Kong","suffix":""},{"id":572650715,"identity":"e8b9e251-2478-487b-aff2-f0fce0cc9de9","order_by":1,"name":"Xinnie Mai","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Xinnie","middleName":"","lastName":"Mai","suffix":""},{"id":572650717,"identity":"690c9b80-54de-49a0-b691-bc9ec8df9a04","order_by":2,"name":"Andrew Meyer","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Andrew","middleName":"","lastName":"Meyer","suffix":""},{"id":572650718,"identity":"5f8751fb-08ca-4d3b-b0e4-ba6a4dcc3c98","order_by":3,"name":"Yanan Fang","email":"","orcid":"","institution":"University of Michigan","correspondingAuthor":false,"prefix":"","firstName":"Yanan","middleName":"","lastName":"Fang","suffix":""},{"id":572650719,"identity":"b36187f0-e462-4c19-a462-669aa5745b1c","order_by":4,"name":"Zidu Xu","email":"","orcid":"","institution":"Harvard Medical School","correspondingAuthor":false,"prefix":"","firstName":"Zidu","middleName":"","lastName":"Xu","suffix":""},{"id":572650721,"identity":"08081b58-d678-40c6-8288-c43746c9ec6b","order_by":5,"name":"Zhixing Song","email":"","orcid":"","institution":"University of Pittsburgh Medical Center","correspondingAuthor":false,"prefix":"","firstName":"Zhixing","middleName":"","lastName":"Song","suffix":""},{"id":572650723,"identity":"69a185a3-cee8-4529-a561-372c6a012e56","order_by":6,"name":"Changlin Gong","email":"","orcid":"","institution":"Jacobi Medical Center, Albert Einstein College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Changlin","middleName":"","lastName":"Gong","suffix":""},{"id":572650726,"identity":"5dc5b8d2-1e92-4fbb-a2f7-2e7eca8669e9","order_by":7,"name":"David A. Wacker","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"A.","lastName":"Wacker","suffix":""},{"id":572650727,"identity":"fce68a3d-7c9c-4ddc-bc25-dae03af55b66","order_by":8,"name":"Julia Heneghan","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Julia","middleName":"","lastName":"Heneghan","suffix":""},{"id":572650728,"identity":"36836e7f-61e6-4b07-95f0-888d91279ab9","order_by":9,"name":"Qianwen Wang","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Qianwen","middleName":"","lastName":"Wang","suffix":""},{"id":572650729,"identity":"5f2bcdae-9999-4377-8329-16deb1bbc525","order_by":10,"name":"Andrew Fu Wah Ho","email":"","orcid":"","institution":"Singapore General Hospital","correspondingAuthor":false,"prefix":"","firstName":"Andrew","middleName":"Fu Wah","lastName":"Ho","suffix":""},{"id":572650730,"identity":"bd75eb35-7466-49bb-aae7-d75153eb612b","order_by":11,"name":"Rui Zhang","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Rui","middleName":"","lastName":"Zhang","suffix":""},{"id":572650731,"identity":"d52ef24e-999a-4c47-b23d-5761de17bd5a","order_by":12,"name":"Mingquan Lin","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Mingquan","middleName":"","lastName":"Lin","suffix":""},{"id":572650732,"identity":"0a28661e-4ec3-48fb-a017-0297fb2a139d","order_by":13,"name":"Geetha Saarunya","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Geetha","middleName":"","lastName":"Saarunya","suffix":""},{"id":572650733,"identity":"6430ad0c-cd8c-4abd-a4f5-aef7ba8cd931","order_by":14,"name":"Xuhai Xu","email":"","orcid":"","institution":"Columbia University","correspondingAuthor":false,"prefix":"","firstName":"Xuhai","middleName":"","lastName":"Xu","suffix":""},{"id":572650734,"identity":"91d43cb7-da0c-4f6f-996d-f41bc07f0fc5","order_by":15,"name":"Elizabeth Lusczek","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Elizabeth","middleName":"","lastName":"Lusczek","suffix":""},{"id":572650735,"identity":"4058e131-ab8f-48a4-a6b6-892941dfd3dc","order_by":16,"name":"Michael A. Puskarich","email":"","orcid":"","institution":"University of Minnesota","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"A.","lastName":"Puskarich","suffix":""},{"id":572650736,"identity":"8e3d5849-60c8-4ca7-9531-3f3debee6437","order_by":17,"name":"Nan Liu","email":"","orcid":"","institution":"Duke-NUS Medical School","correspondingAuthor":false,"prefix":"","firstName":"Nan","middleName":"","lastName":"Liu","suffix":""},{"id":572650738,"identity":"9aee7f77-7383-4c04-8716-a160affafc09","order_by":18,"name":"Leo Anthony Celi","email":"","orcid":"","institution":"Massachusetts Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"Leo","middleName":"Anthony","lastName":"Celi","suffix":""},{"id":572650740,"identity":"3c785a7f-03e7-484e-8f75-2ddd8dee7786","order_by":19,"name":"Feng Xie","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAw0lEQVRIiWNgGAWjYDACZgaGAwk8NnL8cC4RWhgffJBJM5ZsIFoLUJXhDJtDiRsOEKvF4DjzMWmenAOJm29kpz1gqLBObCCkRbKZLU2a58wd4203crcbMJxJJ6yFn5nHTJq355ksUMs2Cca2w4S1sDHzf5Pm/XeYcfMMkJZ/RGgB2gL0Ps9hxQ0SIC0NRGgB+sXwwQeeNGOJM2+3SSQcSzcmqMXg/OEHkKhsB9ryocZalqAWVJBAmvJRMApGwSgYBbgAAEz0Pon/pHanAAAAAElFTkSuQmCC","orcid":"","institution":"University of Minnesota","correspondingAuthor":true,"prefix":"","firstName":"Feng","middleName":"","lastName":"Xie","suffix":""}],"badges":[],"createdAt":"2025-12-11 16:38:38","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8338830/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8338830/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100399950,"identity":"caa41bd6-7b56-45c0-969d-f80dfa44d91d","added_by":"auto","created_at":"2026-01-16 11:57:48","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4793613,"visible":true,"origin":"","legend":"","description":"","filename":"ApplicationofFoundationModelsinEmergencyandCriticalCareAScopingReview1.docx","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/67e3d04622aa654bfab4c5f9.docx"},{"id":100400276,"identity":"1321a713-8e02-4c8d-adc5-e3546d4b38e2","added_by":"auto","created_at":"2026-01-16 11:58:02","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18395,"visible":true,"origin":"","legend":"","description":"","filename":"b28b3642b8254b1b8a72aa3e94bf90d4.json","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/96b6b6d0afada573e8b294ea.json"},{"id":100400082,"identity":"ba2a6d5c-6dc9-482e-85e1-dac43038de4e","added_by":"auto","created_at":"2026-01-16 11:57:52","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":233136,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/ac251882c4c2ec1361ae5a23.pdf"},{"id":100399372,"identity":"6f8a6ee7-c132-45eb-9cec-cbf5e2f0ff26","added_by":"auto","created_at":"2026-01-16 11:56:49","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":283788,"visible":true,"origin":"","legend":"","description":"","filename":"b28b3642b8254b1b8a72aa3e94bf90d41enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/dcd3dc6ff79231998ec43a3f.xml"},{"id":100400138,"identity":"bda56b5f-9dfe-4e11-a36b-c500a84e69e1","added_by":"auto","created_at":"2026-01-16 11:57:55","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":203027,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/6e1554df866c898bc1bf0bf4.png"},{"id":100399939,"identity":"a8c9eda4-d9a7-4160-88a1-be93fff5c98f","added_by":"auto","created_at":"2026-01-16 11:57:47","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":300717,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/b967ab3f1c3fbfc9c691bad9.jpeg"},{"id":100399703,"identity":"49b9b743-113b-4f52-85bc-3f6700c7447b","added_by":"auto","created_at":"2026-01-16 11:57:32","extension":"jpeg","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":352831,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/be220c8bf136715af89942b4.jpeg"},{"id":100399628,"identity":"31e5ce35-db10-4be7-8e60-741508c316aa","added_by":"auto","created_at":"2026-01-16 11:57:21","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":594993,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/8c41bd9a57ef6935908d3e9c.png"},{"id":100399377,"identity":"5bda8402-fc89-4abf-a8e2-6f74da9c2e22","added_by":"auto","created_at":"2026-01-16 11:56:49","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":68869,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/10ebb12bd1e9c27ffbb45c5a.png"},{"id":100399760,"identity":"336bf6ea-c481-4ce5-84e1-b8bfd356bdc3","added_by":"auto","created_at":"2026-01-16 11:57:34","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":75494,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/d719e19ba883aba90cd6e023.png"},{"id":100399630,"identity":"f6fc1694-ac97-48a3-b04b-eb252234c605","added_by":"auto","created_at":"2026-01-16 11:57:21","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":65093,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/de548ff88228a062f7385370.png"},{"id":100400117,"identity":"0d5c17f5-139e-4be5-a6f4-9d317b59d20f","added_by":"auto","created_at":"2026-01-16 11:57:54","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":135411,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/d77d272fb761b77aa5bc8b7b.png"},{"id":100400062,"identity":"83e98e31-8e72-4016-801c-71534c850c2b","added_by":"auto","created_at":"2026-01-16 11:57:51","extension":"xml","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":282480,"visible":true,"origin":"","legend":"","description":"","filename":"b28b3642b8254b1b8a72aa3e94bf90d41structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/201bc78fc6c24f1acf809cb1.xml"},{"id":100400039,"identity":"25fbb94e-c954-4e0d-a15d-0ffcae46eb6a","added_by":"auto","created_at":"2026-01-16 11:57:49","extension":"html","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":307411,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/1b47ee2307b1f03dee2a9008.html"},{"id":100400128,"identity":"447b98f2-0648-4ebc-80e5-d5855f3a9a72","added_by":"auto","created_at":"2026-01-16 11:57:54","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":203027,"visible":true,"origin":"","legend":"\u003cp\u003ePRIMSA-ScR flow diagram of the study selection process for foundation model applications in emergency and critical care.\u003c/p\u003e\n\u003cp\u003eOur search retrieved 1,443 studies. After removing duplicates (n = 321), 1,122 titles and abstracts were screened, of which 1,026 were excluded for: irrelevance to emergency or critical care (n = 407), not applying FMs (n = 243), not a peer-reviewed research article (n = 373), or non-English language (n = 3). Subsequently, 96 studies underwent full-text screening, then 47 were excluded for not focusing on emergency or critical care (n = 38), not involving FMs (n = 4), or not being peer-reviewed research articles (n = 5). 49 studies were ultimately included (Table 1). The selection process is shown in Figure 1. Screening consistency was excellent (inter-rater reliability: κ = 0.897 for title and abstract, κ = 0.812 for full text).\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/4a4632e717fe9f93d18876ca.png"},{"id":100400139,"identity":"84e35c6f-24c9-4508-9275-48cfa4b57037","added_by":"auto","created_at":"2026-01-16 11:57:55","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1206690,"visible":true,"origin":"","legend":"\u003cp\u003eData Modality Analysis of Included Studies on Emergency and Critical Care\u003c/p\u003e\n\u003cp\u003e(a) Prevalence of base data modalities used across included studies. Bars show the number of studies that used each modality, and the dot matrix below indicates single (a single filled dot) and multimodal (multiple filled dots in the same row) combinations of modalities used within the same study.\u003c/p\u003e\n\u003cp\u003e(b) Modality task mapping. Each cell indicates the number of studies that use a given data modality for a specific clinical task.\u003c/p\u003e\n\u003cp\u003e(c) Associations between data modalities, ECC settings, and medical specialties. The thickness of each flow reflects the number of modality-to-setting-to-topic associations rather than the number of unique studies. The “Unspecified” category includes ECC studies without a defined specialty or disease focus. Values reflect ECC-wide aggregation; ED- and ICU-specific modality distributions are not shown in this primary analysis. Studies using multiple modalities in Figures 2b and 2c are counted once within each modality category (Free Text, Tabular, Imaging, Time-Series).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/83661a3a3982f0df4d0b77bd.png"},{"id":100400294,"identity":"f7f8fe9c-b9cd-4d37-9499-31c8779b7a0e","added_by":"auto","created_at":"2026-01-16 11:58:03","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":594993,"visible":true,"origin":"","legend":"\u003cp\u003eFramework and clinical task landscape of foundation model applications in ECC.\u003c/p\u003e\n\u003cp\u003eThe figure integrates three elements: (1) major FM categories and representative model series, (2) the data modalities commonly used in ECC, and (3) the clinical tasks these models support. FM families span language, multimodal, vision, and audio architectures, which operate on free-text documentation, structured EHR data, time-series signals, and imaging. FM-enabled clinical applications include natural language processing tasks, such as information extraction and text generation, as well as outcome prediction, diagnostic reasoning, and treatment recommendation. Together, the figure depicts how FM types, data modalities, and clinical functions align within ECC workflows.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/e92e5f003bdf4550e12866e5.png"},{"id":101869169,"identity":"20e2d3db-c2a8-4bd2-b05e-75c5f0021494","added_by":"auto","created_at":"2026-02-04 12:58:27","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3378769,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/841939db-4113-40d1-89e2-fd07851a20e2.pdf"},{"id":100399609,"identity":"563a05cd-f897-4d37-b046-43e331ff9a09","added_by":"auto","created_at":"2026-01-16 11:57:19","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":233136,"visible":true,"origin":"","legend":"","description":"","filename":"Appendix.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8338830/v1/7bbc0209ab9f272df1c76f45.pdf"}],"financialInterests":"Competing interest reported. Michael Puskarich has served on scientific advisory boards for Inflammatix LLC, Cvtovale LLC, and Opticyte LLC, companies working in the sepsis diagnosis and prognosis space.","formattedTitle":"Application of Foundation Models in Emergency and Critical Care: A Scoping Review","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn emergency and critical care (ECC) settings such as the emergency departments (ED) and intensive care unit (ICU), patients often present with life-threatening conditions requiring rapid interventions\u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Clinicians must integrate diverse data streams\u003csup\u003e\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e including medical history, imaging, laboratory results, and continuous monitoring under severe time pressure. While both ED and ICU demand urgent responses, their contexts differ: emergency care emphasizes triage and immediate intervention with sparse, heterogeneous data\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e, whereas critical care involves prolonged management, stabilization, and recovery based on higher-dimensional, longitudinal inputs\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan additionalcitationids=\"CR12\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Escalating data volume, heterogeneity, and real-time variability have intensified clinicians\u0026rsquo; workload and decision-making demands\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e, underscoring the need for effective tools to support data integration and timely action\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eArtificial intelligence (AI) has long been applied in ECC to assist emergency physicians and intensivists in processing complex data and enhancing diagnostic reasoning\u003csup\u003e\u003cspan additionalcitationids=\"CR18\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Earlier generations of machine learning (ML) models, which include logistic regression\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e,\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e, decision trees\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e,\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e, support vector machines\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e, random forests\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e, and extreme gradient boosting\u003csup\u003e\u003cspan additionalcitationids=\"CR27 CR28\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e were typically trained on narrow datasets\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e and required manual feature engineering\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e,\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. Deep learning (DL) methods such as convolutional neural networks (CNN)\u003csup\u003e\u003cspan additionalcitationids=\"CR34 CR35\" citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e, along with short-term memory networks (LSTM)\u003csup\u003e\u003cspan additionalcitationids=\"CR38 CR39\" citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e, and generative adversarial networks (GAN)\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e improved the ability to learn directly from multi-source clinical data such as imaging, waveforms, and continuous vital signs. However, they still require large task-specific datasets and extensive validation before clinical adoption. These models have been applied in ECC to predict triage acuity, in-hospital mortality, sepsis, and ICU readmission\u003csup\u003e\u003cspan additionalcitationids=\"CR43 CR44\" citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. Foundation models (FMs) represent the next major shift in this progression. Unlike prior DL models trained for single tasks, FMs are pretrained at scale on diverse datasets using self-supervised learning and can be adapted across a wide range of downstream clinical applications, enabling broader generalization and reuse in application.\u003c/p\u003e \u003cp\u003eConsequently, FMs represent a paradigm shift: large, general-purpose neural networks trained on massive, diverse datasets using self-supervised learning, enabling adaptation to a wide range of downstream clinical tasks. These include language models such as Generative Pretrained Transformers (GPT)\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e, Bidirectional Encoder Representations from Transformers (BERT)\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e, and Large Language Model Meta AI (LLaMA)\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e; vision models such as Vision Transformers (ViT)\u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e; and multimodal models such as Contrastive Language\u0026ndash;Image Pretraining (CLIP)\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e and Bootstrapped Language-Image Pretraining (BLIP)\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. FMs have achieved state-of-the-art performance in natural language processing (NLP), clinical prediction, medical imaging, and multimodal reasoning. In the ECC context, FMs offer unique advantages in handling both unstructured texts (e.g., discharge summaries, triage records) and structured data (e.g., laboratory results, vital signs, diagnostic codes, medical images). Different FM types align with various tasks: language models for narrative text, tabular- or sequence-adapted models for structured EHR data (e.g., labs, vitals, codes) (Appendix eTable 3); vision models for imaging (Appendix eTable 5); audio-focused models (e.g., Wav2Vec) for clinical dialogues (Appendix eTable 5); and multimodal architectures for integrating across data types (Appendix eTable 4). The choice of FM, therefore, depends on both the data modality and the clinical objective; for instance, outcome prediction from admission notes might use a language model, stroke detection from imaging might use a vision model, and risk assessment from physiologic signals and text might use a multimodal model. This task-oriented framing helps clinicians and researchers match FM selection to the specific data types and decision needs of ECC practice.\u003c/p\u003e \u003cp\u003eDespite growing interest in FMs, most existing reviews remain focused on conventional AI models\u003csup\u003e\u003cspan additionalcitationids=\"CR53\" citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e and seldom evaluate how FM capabilities such as few-shot learning, contextual reasoning, multimodal integration, and self-supervised adaptation align with ECC\u0026rsquo;s time-critical demands. In addition, current reviews do not assess deployment readiness or safety of FM use in ECC, and rarely examine how to address the epistemic mismatch between FM capabilities and clinical reasoning in ECC\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e. A synthesis of the emerging literature is therefore needed to clarify current FM applications, gaps, and promising directions. This review aims to map the landscape of FM applications in ECC, including the models used, data modalities leveraged, and clinical tasks addressed, while consolidating reported limitations to outline priorities for future research and responsible implementation. At the same time, whether FMs are appropriately suited tools for ECC or are simply a solution in search of a problem remains uncertain. We therefore adopt a problem-first, safety-anchored framing to assess not only performance but also alignment with ECC\u0026rsquo;s operational and clinical realities.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Design\u003c/h2\u003e \u003cp\u003eThis scoping review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) guidelines (Appendix eTable 2). The protocol was developed in advance and applied consistently across study selection, data extraction, and analysis. A comprehensive search was conducted in five electronic databases: PubMed, Embase, Scopus, Web of Science, and CINAHL. The search included articles published up to March 26, 2025. The strategy was designed to identify studies applying FMs in ECC settings, with full details in the checklist (Appendix eTable 1).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eStudy Selection\u003c/h3\u003e\n\u003cp\u003eStudies were eligible if they met the following criteria: (1) applied FMs, defined as large-scale models pre-trained on diverse datasets using self-supervised learning and adapted to downstream tasks via fine-tuning or other transfer-learning methods; (2) focused on emergency or critical care; (3) were original research published in peer-reviewed journals; and (4) had full text available in English. Two independent reviewers (YK and YF) screened titles and abstracts, followed by full-text review using the predefined inclusion and exclusion criteria. Disagreements were resolved through discussion or, if needed, by consulting a third reviewer (FX), an expert in medical AI. Screening consistency was measured using Cohen\u0026rsquo;s kappa coefficient (κ).\u003c/p\u003e\n\u003ch3\u003eData Extraction and Analysis\u003c/h3\u003e\n\u003cp\u003eData extraction was performed independently by three reviewers (YK, XM, YF) using a standardized form, with discrepancies resolved by consensus and adjudicated by a fourth reviewer (XF) when needed. For each study, we extracted bibliographic details (year, authors, journal), FM architectures and base models (Appendix eTable 2), sample size and population characteristics, data modalities, clinical tasks, ECC setting (Emergency or Critical Care), and medical specialty (e.g., pediatric, respiratory, cardiovascular, radiology). These variables were selected a priori to capture the technical design and clinical context of FM applications in ECC.\u003c/p\u003e \u003cp\u003eTo ensure consistency, data modalities were defined in four categories: (1) free text, including triage notes, admission histories, radiology reports, and discharge summaries; (2) tabular structured data, such as laboratory values, vital signs, and administrative codes; (3) time-series data, including electrocardiogram waveforms and continuously monitored vital signs; and (4) imaging, such as radiographs, computed tomography (CT) scan, and magnetic resonance imaging (MRI) scan. Clinical applications were grouped into five predefined task domains: (1) outcome prediction (e.g., mortality, readmission), (2) information extraction from unstructured or multimodal data, (3) text generation including summaries and reports, (4) diagnosis through binary, multi-label, or differential classification, and (5) treatment recommendation, such as planning investigations or suggesting therapies. All modality and task summaries were computed at the aggregated ECC level without ED-ICU stratification to reflect overall literature trends. This framework provided a structured, consistent approach to characterizing how FMs have been applied across ECC.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eStudy Characteristics\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOur search retrieved 1,443 studies. After removing duplicates (n\u0026thinsp;=\u0026thinsp;321), 1,122 titles and abstracts were screened, of which 1,026 were excluded for: irrelevance to emergency or critical care (n\u0026thinsp;=\u0026thinsp;407), not applying FMs (n\u0026thinsp;=\u0026thinsp;243), not a peer-reviewed research article (n\u0026thinsp;=\u0026thinsp;373), or non-English language (n\u0026thinsp;=\u0026thinsp;3). Subsequently, 96 studies underwent full-text screening, then 47 were excluded for not focusing on emergency or critical care (n\u0026thinsp;=\u0026thinsp;38), not involving FMs (n\u0026thinsp;=\u0026thinsp;4), or not being peer-reviewed research articles (n\u0026thinsp;=\u0026thinsp;5). 49 studies were ultimately included (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The selection process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Screening consistency was excellent (inter-rater reliability: κ\u0026thinsp;=\u0026thinsp;0.897 for title and abstract, κ\u0026thinsp;=\u0026thinsp;0.812 for full text).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCharacteristics of included studies applying foundation models in emergency and critical care.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAuthor\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eYear\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFoundation Model Name\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eData Modality\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eClinical Task\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eClinical Settings\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eMedical Specialties\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChen et al. \u003csup\u003e56\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2021\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEDisease, BERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultimodal (Tabular Data\u0026thinsp;+\u0026thinsp;Free Text)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.04M ED visits from NTUH; 305K from National Hospital Ambulatory ED data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBarash et al. \u003csup\u003e57\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e40 real ED clinical cases (clinical notes and imaging) from 8 acute pathologies\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTreatment Recommendation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRadiology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBushuven et al. \u003csup\u003e58\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e22 case vignettes (2 BLS and the 20 core PALS scenarios)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePediatric\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFraser et al. \u003csup\u003e59\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-3.5, ChatGPT-4.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTabular Data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e40 ED patient symptom entries\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHenriksson et al. \u003csup\u003e60\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eClinical KB-BERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultimodal (Free Text\u0026thinsp;+\u0026thinsp;Tabular Data)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMulti-center COVID-19 patient cohort\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRespiratory\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHuang et al. \u003csup\u003e61\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePABLO: Pretrained and Adapted BERT for Longitudinal Outcomes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTabular Data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCalifornia and Florida Statewide EHR Databases\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePediatric\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSavage et al. \u003csup\u003e62\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBioMed-RoBERTa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMIMIC-III\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAityan et al. \u003csup\u003e63\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003emultiple LLMs (ChatGPT, Claude, Gemini)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultimodal (Free Text\u0026thinsp;+\u0026thinsp;Tabular Data\u0026thinsp;+\u0026thinsp;Time-Series\u0026thinsp;+\u0026thinsp;Imaging)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e600 cases (300 sepsis and 300 non-sepsis) from the University archive, Google Scholar, and PubMed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAkhondi-Asl et al. \u003csup\u003e64\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBioGPT-Large, LLaMa-7B (Fine-tuned), LLaMa-65B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1,916,538 notes from 32,454 PICU patients for training; 130 admission notes for evaluation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePediatric\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlessandri-Bonetti et al. \u003csup\u003e65\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-3.5, ChatGPT-4, Bard\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e50 questions with five multiple-choice answers from the ABLS exam\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eBurn Care\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAmacher et al. \u003csup\u003e66\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTabular Data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAdult cardiac arrest patients' data at University Hospital Basel between 2012 and 2022\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCardiovascular\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHaim et al. \u003csup\u003e67\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e100 consecutive adult ED patient records\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChung et al. \u003csup\u003e68\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4 Turbo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003etask-specific datasets constructed from 2 years of retrospective EHR data collected\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePerioperative Medicine\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFerri et al. \u003csup\u003e69\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDistilBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.98M emergency medical call incidents\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGimeno et al. \u003csup\u003e70\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e120 discharge instructions\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePediatric\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGlicksberg et al. \u003csup\u003e71\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4, Bio-Clinical-BERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultimodal (Free Text\u0026thinsp;+\u0026thinsp;Tabular Data)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e864,089 ED encounters from 7 hospitals within NYC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGuo et al. \u003csup\u003e72\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFMSM (Foundation Model Stanford Medicine)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTabular Data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStanford Medicine EHR, SickKids EHR, MIMIC-IV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHuang et al. \u003csup\u003e73\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5 synthetic ED encounter notes created by emergency physicians\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKhaldi et al. \u003csup\u003e74\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e20 common questions of patients about tracheotomy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eOtolaryngology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLe Guellec et al. \u003csup\u003e75\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVicuna 13B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2398 emergency brain MRI reports\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRadiology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLee et al. \u003csup\u003e76\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eKLUE-BERT, KLUE-RoBERTa, KorBERT, KoBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003esimulated clinician\u0026ndash;patient conversations\u003c/p\u003e \u003cp\u003ewithin 6 primary symptom scenarios in emergency triage rooms\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLevin et al. \u003csup\u003e77\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSix neonatal ICU clinical case scenarios\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTreatment Recommendation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eNeonatal\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLin et al. \u003csup\u003e78\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003emultimodal model (combineTabNet and MacBERT)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultimodal (Free Text\u0026thinsp;+\u0026thinsp;Tabular Data)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNTUH Retrospective Data Set (745,441 ED visits, 2009\u0026ndash;2015); NTUH Prospective Data Set (901 ED visits, May 2020\u0026ndash;Feb 2022)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMasanneck et al. \u003csup\u003e79\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-3.5, GPT-4, LLaMA 3 70B, Gemini 1.5, Mixtral 8x7b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e124 independent emergency cases from the interdisciplinary ED of University Hospital D\u0026uuml;sseldorf, Germany\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMcCoy and Perlis \u003csup\u003e\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4 Turbo, GPT-4-1106-preview\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3059 individuals with a median age of 16 years (interquartile range, 13\u0026ndash;18 years)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePediatric\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRothchild et al. \u003csup\u003e82\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10 clinical vignettes on common facial trauma presentations\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eOtolaryngology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSaner et al. \u003csup\u003e83\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT 4.0 Plus, Bard\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10 simulated ICU patient cases, with each SOFA score calculated\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRespiratory\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScquizzato et al. \u003csup\u003e84\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e40 layperson questions about cardiac arrest and CPR co-created with the Sudden Cardiac Arrest UK community\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCardiovascular\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSeo et al. \u003csup\u003e85\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHyperCLOVA X\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e33 ED initial records were generated by 52 participants during the Healthcare Prompt-a-thon event\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSezgin et al. \u003csup\u003e86\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eT5-small, T5-base, PEGASUS-PubMed, BART-Large-CNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e100 referral conversations among ED clinicians\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTang et al. \u003csup\u003e87\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1,000 de-identified ED imaging reports from 20 most common study types at Stanford Health Care (2020\u0026ndash;2023)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRadiology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUrquhart et al. \u003csup\u003e88\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4 API, ChatGPT, LLaMA 2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eICU admission notes from 11 episodes in 9 patients at Galway University Hospital, Ireland\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWali et al. \u003csup\u003e89\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBioMistral 7B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultimodal (Tabular Data\u0026thinsp;+\u0026thinsp;Time-Series)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eThe heart attack dataset from the IEEE Dataport\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCardiovascular\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWang et al. \u003csup\u003e90\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-3.5, GPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eElectronic Medical Records from the ED of Maoming People\u0026rsquo;s Hospital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eNeurology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYang et al. \u003csup\u003e91\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-3.5, GPT-4, Claude 2, LLaMA 2-7b and 2-13b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e219 MCQs from critical care pharmacotherapy courses at two U.S. colleges of pharmacy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePharmacotherapy\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYau et al. \u003csup\u003e92\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2024\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-3.5, Bard\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10 emergency medicine questions\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eText Generation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAmacher et al. \u003csup\u003e93\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTabular Data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e760 consecutive adult ICU patients with status epilepticus from University Hospital Basel (2005\u0026ndash;2022)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eNeurology\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eArslan et al. \u003csup\u003e94\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT Plus (GPT-4), Copilot Pro (GPT-4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e468 ED patient cases from a large urban academic hospital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBalta et al. \u003csup\u003e95\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-3.5, ChatGPT-4.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e50 clinical critical care questions across five categories, representative of critical care medicine\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTreatment Recommendation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBroad et al. \u003csup\u003e96\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4o\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eESO Research Data Collaborative\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePediatric\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFeng et al. \u003csup\u003e97\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPT-4-turbo, GPT-3.5-turbo, Jurassic-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e490 full-text EHR notes from 125 patients with prior life-threatening arrhythmias\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCardiovascular\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLevra et al. \u003csup\u003e98\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMultilingual BERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e30,320 EMRs from Humanitas Research Hospital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCardiovascular\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHo et al. \u003csup\u003e99\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-3.5, ChatGPT-4.0, T5, LLaMA 2, Mistral-Large, Claude-3 Opus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e70 pediatric emergency clinical vignettes adapted from ESI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMiller et al. \u003csup\u003e100\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e104 paramedic PCRs from cloud-based database (EMCE, NMETC)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePathak et al. \u003csup\u003e101\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRespBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRadiology notes from Emory University and Grady Memorial Hospital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDiagnosis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRespiratory\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShekhar et al. \u003csup\u003e102\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4oMini\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAllegheny County Dispatch Dataset (greater than 1,000,000 ambulance requests from 2015 to 2020)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eOutcome Prediction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eEmergency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePrehospital Care\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWilliams and Erstad \u003csup\u003e\u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e103\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChatGPT-4, LLaMA 3.1 405B\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFree Text\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e43 medication-related PICO questions drawn from 6 CPGs (2023\u0026ndash;2024)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTreatment Recommendation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePharmacotherapy\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYang et al. \u003csup\u003e16\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eICU-GPT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTabular Data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMIMIC-III;\u003c/p\u003e \u003cp\u003eMIMIC-IV; MIMIC-IV-ED; MIMIC-IV-Note; eICU-CRD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation Extraction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCritical Care\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eClinical Informatics\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eData Modality Adoption of Foundation Models in Emergency and Critical Care Studies\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e(a) Prevalence of base data modalities used across included studies. Bars show the number of studies that used each modality, and the dot matrix below indicates single (a single filled dot) and multimodal (multiple filled dots in the same row) combinations of modalities used within the same study.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e(b) Modality task mapping. Each cell indicates the number of studies that use a given data modality for a specific clinical task.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e(c) Associations between data modalities, ECC settings, and medical specialties. The thickness of each flow reflects the number of modality-to-setting-to-topic associations rather than the number of unique studies. The \u0026ldquo;Unspecified\u0026rdquo; category includes ECC studies without a defined specialty or disease focus. Values reflect ECC-wide aggregation; ED- and ICU-specific modality distributions are not shown in this primary analysis. Studies using multiple modalities in Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec are counted once within each modality category (Free Text, Tabular, Imaging, Time-Series).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eA range of FM architectures have been applied in ECC, most commonly GPT family (35/49, 71.43%), BERT-based models, and the LLaMA series, each suited for different tasks. BERT, an encoder-only model for text understanding, has been used to identify ARDS from radiology notes\u003csup\u003e\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e\u003c/sup\u003e. It has also been applied to predict ICU admissions or in-hospital mortality\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e, showing strong performance even with limited data. GPT, a decoder-only generative model, has been tested in prehospital acute stroke screening to help identify lesion locations and responsible vessels, which highlights its reasoning and generative capabilities\u003csup\u003e\u003cspan citationid=\"CR90\" class=\"CitationRef\"\u003e90\u003c/span\u003e\u003c/sup\u003e. LLaMA is also a decoder-only, open-source, and computationally efficient model. It can enable on-site fine-tuning without an external server, making it well-suited for privacy-sensitive settings, as shown by its use in generating PICU differential diagnoses\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e and extracting magnetic resonance imaging (MRI) reports in the ED\u003csup\u003e75\u003c/sup\u003e. These applications demonstrate how different FM families are being adapted for ECC tasks; however, most studies fall short of demonstrating improvements in workflow efficiency, decision-making accuracy, or patient outcomes compared with standard practice (Appendix eTable 2).\u003c/p\u003e \u003cp\u003eBuilding on these architecture-specific applications, most studies applied FMs to free text, which was the predominant data modality (85.7%, 42/49), supporting a wide range of applications (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Free text was input to support tasks such as outcome prediction, diagnosis, and text generation. For example, one study evaluated ChatGPT-4 as a conversational agent for analyzing ED admission notes to recommend radiology referrals and imaging selection\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e, whereas LLaMA-7B and BioGPT-Large were fine-tuned on PICU admission notes to generate differential diagnoses\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e. These applications highlight the versatility of FMs in leveraging unstructured narratives to inform high-acuity decision-making. However, timeliness and completeness can be limited if free text is used alone without complementary structured data like tabular EHR records.\u003c/p\u003e \u003cp\u003eTabular data was the second most common modality (24.49%, 12/49), most often used for outcome prediction (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). For example, one study developed the Pretrained and Adapted BERT for Longitudinal Outcomes (PABLO) model to analyze structured EHR data, including demographics, diagnostic codes, and procedures, to predict non-accidental trauma (injury resulting from intentional harm or neglect for children and adolescents)\u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e. Another study applied a transformer-based language model to sequences of diagnoses and laboratory results to estimate in-hospital mortality and readmission risk\u003csup\u003e\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e. These studies illustrate how structured records, when modeled longitudinally, can support early risk identification and monitoring in ECC. However, essential time-series or imaging data is not well captured in tabular form, which points to the need for multimodal modeling.\u003c/p\u003e \u003cp\u003eTime-series (4.08%, 2/49) and imaging data (2.04%, 1/49) were incorporated only into multimodal applications. For example, high-frequency physiologic signals, such as electrocardiogram traces and imaging inputs, were combined with free-text notes and structured vital signs using models like GPT and Gemini to support diagnosis and risk assessment for sepsis, stroke, and myocardial infarction\u003csup\u003e\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e. Overall, 12.24% of studies (6/49) used explicitly multimodal inputs, most often combining free text with structured EHR data (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). One study applied Clinical KB-BERT to jointly process clinical notes and structured records to predict 30-day mortality and 14-day readmission for COVID-19 patients, achieving superior performance compared with single-modality models\u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e. These findings suggest that incorporating multiple data modalities can improve accuracy and interpretability, but multimodal FM development in ECC remains constrained by limited data accessibility, heterogeneity, and integration challenges.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe figure integrates three elements: (1) major FM categories and representative model series, (2) the data modalities commonly used in ECC, and (3) the clinical tasks these models support. FM families span language, multimodal, vision, and audio architectures, which operate on free-text documentation, structured EHR data, time-series signals, and imaging. FM-enabled clinical applications include natural language processing tasks, such as information extraction and text generation, as well as outcome prediction, diagnostic reasoning, and treatment recommendation. Together, the figure depicts how FM types, data modalities, and clinical functions align within ECC workflows.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eClinical Applications of Foundation Models in Emergency and Critical Care\u003c/h3\u003e\n\u003cp\u003eFMs have been explored across five main applications in ECC: outcome prediction, diagnosis, text generation, information extraction, and treatment recommendation (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The emphasis of these applications varies by clinical context and the data modality. In emergency care, studies emphasize rapid triage and diagnosis with limited, heterogeneous data. In contrast, in critical care, the focus shifts to continuous monitoring, dynamic risk stratification, and management of longitudinal, high-dimensional data streams. The following subsections highlight how these domains have been studied, including both the opportunities and the limitations of current FM applications in ECC.\u003c/p\u003e \u003cp\u003eOutcome prediction was the most common FM application in ECC (28.6%, 14/49). These prediction categories mainly include mortality\u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e,\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e,\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e\u003c/sup\u003e, hospital resource utilization\u003csup\u003e\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e,\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e\u003c/sup\u003e, triage and acuity stratification\u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e,\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e,\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e,\u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e79\u003c/span\u003e,\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e,\u003cspan citationid=\"CR99\" class=\"CitationRef\"\u003e99\u003c/span\u003e,\u003cspan citationid=\"CR102\" class=\"CitationRef\"\u003e102\u003c/span\u003e\u003c/sup\u003e, disease-specific risks and complications\u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e, and laboratory abnormalities\u003csup\u003e\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e. In emergency care, applications emphasized rapid risk assessment. For example, the BERT-derived PABLO model used longitudinal EHR data, including demographics, diagnoses, and procedures, to predict non-accidental trauma in children, outperforming traditional ML methods\u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e. Other studies used synthetic scenarios rather than real-time triage data. For instance, one study evaluated GPT-4 using pediatric vignettes and reported higher accuracy than LLaMA-2 in predicting ESI scores. However, the researchers acknowledged that vignette-based inputs deviate from actual clinical workflows\u003csup\u003e\u003cspan citationid=\"CR99\" class=\"CitationRef\"\u003e99\u003c/span\u003e\u003c/sup\u003e. In critical care, research focuses more on continuous risk stratification and longitudinal data monitoring. For instance, ChatGPT-4 predicted mortality and neurological outcomes after cardiac arrest with performance comparable to validated post-arrest scoring systems\u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eDiagnosis, including single-label multi-class disease diagnosis\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e82\u003c/span\u003e,\u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e100\u003c/span\u003e\u003c/sup\u003e, differential diagnosis\u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e,\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e,\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e, and binary disease diagnosis\u003csup\u003e\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e81\u003c/span\u003e,\u003cspan citationid=\"CR90\" class=\"CitationRef\"\u003e90\u003c/span\u003e,\u003cspan additionalcitationids=\"CR97\" citationid=\"CR96\" class=\"CitationRef\"\u003e96\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e,\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e\u003c/sup\u003e, was the second most common FM application in ECC (24.5%, 12/49). In emergency medicine, GPT achieved 93.9% diagnostic accuracy when tested on pediatric prehospital free-text descriptions provided by laypersons\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e. Another study used GPT to evaluate patient-entered symptom data, reporting variation in model performance\u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e. In critical care, a fine-tuned LLaMa outperformed BioGPT and LLaMa for generating differential diagnoses from PICU admission notes\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e. Similarly, GPT classified arrhythmia recurrence from post-ablation EHR notes with 91.4% accuracy, exceeding both SapBERT (66.6%) and a rule-based algorithm (82.6%)\u003csup\u003e97\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eFMs have also been applied to text generation (22.4%, 11/49), primarily focusing on automated clinical question answering\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e,\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e,\u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e89\u003c/span\u003e,\u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e91\u003c/span\u003e,\u003cspan citationid=\"CR92\" class=\"CitationRef\"\u003e92\u003c/span\u003e\u003c/sup\u003e, discharge instruction generation\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e,\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e,\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e\u003c/sup\u003e, initial clinical record drafting\u003csup\u003e\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e\u003c/sup\u003e, summarization of clinical dialogues\u003csup\u003e\u003cspan citationid=\"CR86\" class=\"CitationRef\"\u003e86\u003c/span\u003e\u003c/sup\u003e, and abstraction of patient records\u003csup\u003e\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e\u003c/sup\u003e in ECC contexts where efficient, precise communication is essential. For example, GPT-4 was evaluated for generating ED discharge instructions from synthetic free-text clinical notes, with better performance in the return precautions section compared to standard instructions\u003csup\u003e\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e\u003c/sup\u003e. Another study tested ChatGPT\u0026rsquo;s ability to answer common post-resuscitation questions from cardiac arrest survivors, relatives, and lay rescuers, finding that its responses were largely accurate and comprehensive. However, cardiopulmonary resuscitation (CPR)-specific answers were weaker\u003csup\u003e\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e\u003c/sup\u003e. These applications suggest potential for FMs to support documentation and patient education, though their role in direct clinical guidance remains unproven.\u003c/p\u003e \u003cp\u003eFMs were applied to information extraction in 16.3% of studies (8/49), including patient-level disease concept embeddings\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e, contraindication identification\u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e, symptom element extraction\u003csup\u003e\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e,\u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e\u003c/sup\u003e, severity quantification\u003csup\u003e\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e,\u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e83\u003c/span\u003e,\u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e87\u003c/span\u003e\u003c/sup\u003e, and structured EHR extraction\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e from clinician notes and radiology reports to support efficient retrieval in high-intensity settings. For example, BioMed-RoBERTa extracted bleeding status from ICU physician notes, enabling more targeted best-practice alerts (14.8% improvement in applicability)\u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e. While most models were BERT-based, newer models like Vicuna-13B demonstrate superior flexibility, effectively extracting diverse information, such as symptom presence and causal inferences, from emergency brain MRI reports\u003csup\u003e\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTreatment recommendations represented another type of application of FMs (8.2%, 4/49), in which models suggested diagnostic investigation planning\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e and treatment planning based on diagnostic findings\u003csup\u003e\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e,\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e,\u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e103\u003c/span\u003e\u003c/sup\u003e. These systems integrated diverse clinical data, including physiological status, laboratory results, and patient history, to guide diagnostic uncertainty. One study used ChatGPT-4 to analyze retrospective ED admission records and recommend radiology examinations for acute conditions, achieving 95% concordance with the American College of Radiology Appropriateness Criteria (ACR AC)\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e. Overall, these studies illustrate the breadth of FM applications in ECC, while remaining exploratory and lacking real-world clinical integration.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe growing interest in FMs for ECC reflects their potential to support high-stakes decisions by synthesizing diverse data sources under severe time constraints. Because errors in these environments carry serious consequences, model evaluation must extend beyond accuracy to include safety, fairness, reliability, and oversight. Despite their capabilities, current FMs do not yet fully meet demands in ECC for urgency, contextual nuance, and dependable human supervision. The literature highlights both the breadth of promising applications and the early, exploratory nature of existing work, providing a foundation for evaluating clinical utility and directing future development. Within this landscape, multimodal integration emerges as both a key strength and a central technical challenge, shaping the broader discussion of FM opportunities and limitations in ECC.\u003c/p\u003e \u003cp\u003eA significant strength of FMs in ECC is their ability to integrate multimodal data\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e,\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e,\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e,\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e,\u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e89\u003c/span\u003e\u003c/sup\u003e, including free text, structured records, time series, and imaging, to enable real-time interpretation of heterogeneous information. Incorporating temporal and spatial signals from continuous monitoring and radiologic imaging can enhance early detection of deterioration and support timely, clinically grounded decisions\u003csup\u003e\u003cspan citationid=\"CR104\" class=\"CitationRef\"\u003e104\u003c/span\u003e\u003c/sup\u003e. Despite their potential, multimodal FMs remain underused in ECC. Models like ViT\u003csup\u003e\u003cspan citationid=\"CR105\" class=\"CitationRef\"\u003e105\u003c/span\u003e\u003c/sup\u003e for radiographic interpretation, CLIP\u003csup\u003e\u003cspan additionalcitationids=\"CR107\" citationid=\"CR106\" class=\"CitationRef\"\u003e106\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR108\" class=\"CitationRef\"\u003e108\u003c/span\u003e\u003c/sup\u003e for image-text alignment, BLIP\u003csup\u003e\u003cspan citationid=\"CR109\" class=\"CitationRef\"\u003e109\u003c/span\u003e,\u003cspan citationid=\"CR110\" class=\"CitationRef\"\u003e110\u003c/span\u003e\u003c/sup\u003e for vision-language report generation, and Wav2Vec 2.0\u003csup\u003e111,112\u003c/sup\u003e for acoustic features extraction, remain largely unexplored and warrant greater adaptation to ECC needs. Real-time multimodal deployment also faces operational challenges: asynchronous data streams, latency, noise, and imperfect alignment can degrade performance and interpretability\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR113\" class=\"CitationRef\"\u003e113\u003c/span\u003e,\u003cspan citationid=\"CR114\" class=\"CitationRef\"\u003e114\u003c/span\u003e\u003c/sup\u003e. Clinical viability depends on pairing advances in multimodal architectures with robust, automated pipelines for data harmonization, synchronization, and quality assurance, enabling dependable and supervised integration into ECC workflows\u003csup\u003e\u003cspan additionalcitationids=\"CR116\" citationid=\"CR115\" class=\"CitationRef\"\u003e115\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR117\" class=\"CitationRef\"\u003e117\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAnother advantage of FMs, particularly language models, is their potential to reduce documentation burden in ECC. They can automate the processing of clinical text, such as ED notes\u003csup\u003e\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e\u003c/sup\u003e, chief complaints\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e, and radiology reports\u003csup\u003e\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e,\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e\u003c/sup\u003e, freeing clinician time and improving efficiency. Models like GPT-4 have produced accurate and consistent documentation outputs\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e,\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e\u003c/sup\u003e, while RespBERT\u003csup\u003e\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e\u003c/sup\u003e and Vicuna\u003csup\u003e\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e\u003c/sup\u003e have extracted structured information from free text with minimal tuning. However, current applications remain dependent on retrospective notes, such as triage summaries and discharge instructions, which often lack real-time relevance\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e,\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e,\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e,\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e,\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e,\u003cspan citationid=\"CR99\" class=\"CitationRef\"\u003e99\u003c/span\u003e\u003c/sup\u003e. More systematically, incompleteness in documentation limits downstream reasoning and introduces bias\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e,\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e,\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e,\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e,\u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e,\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e,\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e,\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e,\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e\u003c/sup\u003e. Deciding what to record and omit reflects subjective judgment and resource constraints, a challenge well documented in other ML approaches that rely on imputation. These gaps in completeness and semantic consistency\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e,\u003cspan citationid=\"CR96\" class=\"CitationRef\"\u003e96\u003c/span\u003e,\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e,\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e,\u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e103\u003c/span\u003e\u003c/sup\u003e hinder causal inference and model reliability\u003csup\u003e\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e,\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e\u003c/sup\u003e. Future FMs should therefore minimize temporal lags, treat missingness as potentially informative, and improve alignment across documentation to ensure both accuracy and fairness.\u003c/p\u003e \u003cp\u003eBeyond data integration and documentation, FMs are poised to transform communication, coordination, and comprehension across both professional and lay interfaces in ECC. Within clinical teams, evaluations of digital scribe systems for summarizing ED consultation calls show that large language models can transcribe and structure complex dialogues in real time, supporting consistent handoffs and reducing information loss in high-acuity workflows\u003csup\u003e\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e\u003c/sup\u003e. At the lay interface, FMs show comparable promise in addressing gaps in health literacy and language accessibility. Studies of GPT-4-generated discharge instructions in multilingual pediatric emergency settings found that models can produce readable, linguistically concordant summaries of care, empowering caregivers to follow post-discharge recommendations\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e. Complementary work comparing GPT-based and standard explanations suggests that FMs can improve comprehension and satisfaction when outputs are contextualized and reviewed for accuracy and reproducibility\u003csup\u003e\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e,\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e\u003c/sup\u003e. These systems demonstrate early success but remain limited by ambient noise, privacy constraints, and semantic drift, emphasizing the need for context-aware fine-tuning. Yet disparities in completeness between English outputs and those in other languages, and the difficulty in achieving simplified reading levels, reveal that fluency does not ensure communicative equity. Collectively, these advances position FMs as communication intermediaries, bridging clinicians, patients, and data. Their responsible ECC deployment will require governance that provides transparency, interpretability, and empathy\u003csup\u003e\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eLimits and Challenges\u003c/h2\u003e \u003cp\u003eDespite the increasing integration of FMs into ECC, significant challenges persist. A major limitation concerns the quality and generalizability of data used for training and validation. Many studies have relied on simulated or synthetic datasets\u003csup\u003e\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e,\u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e103\u003c/span\u003e,\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e,\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e\u003c/sup\u003e based on idealized clinical cases\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e,\u003cspan citationid=\"CR92\" class=\"CitationRef\"\u003e92\u003c/span\u003e\u003c/sup\u003e, which may overestimate effectiveness compared to real-world scenarios\u003csup\u003e\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e\u003c/sup\u003e. Small sample sizes\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e,\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e,\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e,\u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e,\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e,\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e,\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e,\u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e103\u003c/span\u003e\u003c/sup\u003e, single-center cohorts\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e,\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e,\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e,\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e,\u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e100\u003c/span\u003e\u003c/sup\u003e, and omission of key patient information (e.g., demographics\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e\u003c/sup\u003e, symptom trajectories\u003csup\u003e\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e\u003c/sup\u003e, diagnostic uncertainty\u003csup\u003e\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e,\u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e91\u003c/span\u003e,\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e,\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e\u003c/sup\u003e) further reduce generalizability and introduce selection bias\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e. Temporal or demographic sampling biases, such as restrictions to specific seasons\u003csup\u003e\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e\u003c/sup\u003e or age groups\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e, also limit the applicability of the results across diverse ECC populations\u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e,\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e,\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e,\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e,\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e,\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e,\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e\u003c/sup\u003e. In addition, stringent privacy and data protection regulations (e.g., the Health Insurance Portability and Accountability Act [HIPAA]\u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e,\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e,\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e\u003c/sup\u003e, and the General Data Protection Regulation [GDPR]\u003csup\u003e\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e\u003c/sup\u003e) limit access to real-world data, thereby restricting the scale and diversity of datasets required for robust FM development\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e,\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e,\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e,\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e. Addressing these limitations requires treating FMs as socio-technical systems, pairing advances such as federated validation, real-time integration, and factual-consistency checks with ethical safeguards and implementation planning.\u003c/p\u003e \u003cp\u003eFM performance was generally reported using task-specific metrics, but inconsistent reporting hindered meaningful comparison across studies. Beyond data issues, evaluation methodologies often lacked rigor. Some studies relied on retrospective designs without external validation\u003csup\u003e\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e\u003c/sup\u003e or prospective testing\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e, both of which are critical for demonstrating real-world utility\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e,\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e,\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e\u003c/sup\u003e. Reproducibility was further limited by single-run evaluations, absent version control, and reliance on subjective reference standards, all of which hinder fair comparison with existing methods\u003csup\u003e\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e,\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e,\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e,\u003cspan citationid=\"CR96\" class=\"CitationRef\"\u003e96\u003c/span\u003e,\u003c/sup\u003e. These barriers stem not only from technical gaps but also from structural challenges: patient unpredictability, dynamic care pathways, and intense time pressure complicate prospective validation in ECC, while regulatory uncertainty and institutional hesitancy further slow translation. Potential solutions include federated multi-institutional validation, standardized benchmarking pipelines with version tracking, and pragmatic trials incorporating clinician oversight. These challenges are therefore not only technical but structural, requiring strategies that extend beyond simply expanding datasets. Accordingly, evaluation should consider workflow impact, communication, team dynamics, and clinician trust, not accuracy alone.\u003c/p\u003e \u003cp\u003ePersistent challenges in FM research arise not only from generic barriers in data quality and evaluation but from fundamental misalignments between model design and clinical reasoning. Language models, mainly trained on static internet text, capture linguistic patterns rather than the real-time, causal, and context-dependent logic that ECC requires\u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e,\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e,\u003cspan citationid=\"CR118\" class=\"CitationRef\"\u003e118\u003c/span\u003e\u003c/sup\u003e. This gap between statistical pattern-matching and the causal, contextual reasoning used by clinicians constitutes an epistemic mismatch. It helps explain why superficially strong test-set performance may not translate into safe decision support at the bedside. The absence of streaming vitals, dynamic documentation, and interprofessional communication in most training corpora leaves FMs linguistically fluent but clinically brittle. In acute settings, they may generate context-blind or unsafe outputs, including misprioritized triage, missed deterioration, or inappropriate treatment advice that undermine clinician judgment and amplify bias\u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e,\u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e79\u003c/span\u003e,\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e82\u003c/span\u003e,\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e,\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e,\u003cspan citationid=\"CR103\" class=\"CitationRef\"\u003e103\u003c/span\u003e\u003c/sup\u003e. By contrast, low-risk and verifiable assistive uses, such as structured note generation, ICU documentation summarization, or extraction of key findings, may improve efficiency when outputs remain easily reviewable\u003csup\u003e\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e,\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e,\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e\u003c/sup\u003e. Autonomous diagnostic or treatment decisions, however, remain ethically indefensible until models demonstrate contextual reasoning, transparency, and validated safety. In addition, developing and deploying FMs in ECC entails substantial computational, financial, and environmental burdens\u003csup\u003e\u003cspan additionalcitationids=\"CR120\" citationid=\"CR119\" class=\"CitationRef\"\u003e119\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR121\" class=\"CitationRef\"\u003e121\u003c/span\u003e\u003c/sup\u003e: training and tuning require extensive energy and water resources, and safe clinical deployment demands costly infrastructure, monitoring, and validation workflows\u003csup\u003e\u003cspan citationid=\"CR122\" class=\"CitationRef\"\u003e122\u003c/span\u003e\u003c/sup\u003e. Recent analyses show AI\u0026rsquo;s growing carbon and water footprint in healthcare, highlighting the need for cost and sustainability assessments before large-scale clinical adoption\u003csup\u003e\u003cspan citationid=\"CR121\" class=\"CitationRef\"\u003e121\u003c/span\u003e,\u003cspan citationid=\"CR123\" class=\"CitationRef\"\u003e123\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eEthical concerns also emerged across the reviewed studies. Limited external validation and single-center sampling not only underdetermine reproducibility but also raise equity concerns, as underrepresented populations may be disproportionately affected. Few studies conducted subgroup analyses, leaving open the risk that FMs could exacerbate existing disparities in ECC. In addition, model-level risks such as hallucinations\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e,\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e,\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e\u003c/sup\u003e, prompt sensitivity\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e,\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e,\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e,\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e,\u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e91\u003c/span\u003e\u003c/sup\u003e, and the black-box nature of many architectures undermine reliability and accountability in high-stakes settings\u003csup\u003e\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e,\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e\u003c/sup\u003e. Without interpretable mechanisms or traceable reasoning pathways, clinicians cannot verify whether outputs are grounded in valid medical knowledge\u003csup\u003e\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e\u003c/sup\u003e, which erodes confidence and complicates regulatory oversight. These issues resonate with broader debates in emergency AI, where even minor errors can have severe consequences \u003csup\u003e\u003cspan citationid=\"CR124\" class=\"CitationRef\"\u003e124\u003c/span\u003e\u003c/sup\u003e, and align with reviews of generative AI that emphasize transparency, accountability, and fairness as essential safeguards \u003csup\u003e\u003cspan citationid=\"CR125\" class=\"CitationRef\"\u003e125\u003c/span\u003e\u003c/sup\u003e. Addressing these gaps through fairness auditing, interpretable modeling, hallucination detection, and clear governance frameworks will be critical to ensure safe and equitable FM deployment in ECC.\u003c/p\u003e \u003cp\u003eBeyond these challenges, the introduction of FMs into ECC also reshapes long-standing professional identities rooted in clinical autonomy, teamwork, and shared accountability. Studies show that clinicians report AI systems as encroaching on their expertise or disrupting traditional hierarchies of judgment and responsibility\u003csup\u003e\u003cspan citationid=\"CR126\" class=\"CitationRef\"\u003e126\u003c/span\u003e,\u003cspan citationid=\"CR127\" class=\"CitationRef\"\u003e127\u003c/span\u003e\u003c/sup\u003e. As FMs begin to support documentation, triage, or communication tasks, clinicians risk becoming supervisors of algorithmic output rather than active decision-makers\u003csup\u003e\u003cspan citationid=\"CR128\" class=\"CitationRef\"\u003e128\u003c/span\u003e\u003c/sup\u003e. This shift can undermine trust, situational awareness, and engagement during critical events. As demonstrated by Bienefeld et al., ensuring effective human-AI collaboration will require clear role delineation, workflow integration, and training that preserve clinician judgment and accountability in life-critical settings\u003csup\u003e\u003cspan citationid=\"CR129\" class=\"CitationRef\"\u003e129\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eFuture Directions\u003c/h2\u003e \u003cp\u003eFuture research on FMs in ECC should integrate methodological rigor with ethical and regulatory considerations to ensure safe, trustworthy deployment\u003csup\u003e\u003cspan citationid=\"CR130\" class=\"CitationRef\"\u003e130\u003c/span\u003e,\u003cspan citationid=\"CR131\" class=\"CitationRef\"\u003e131\u003c/span\u003e\u003c/sup\u003e. Building multicenter and demographically diverse longitudinal datasets will improve generalizability and reduce bias\u003csup\u003e\u003cspan additionalcitationids=\"CR133\" citationid=\"CR132\" class=\"CitationRef\"\u003e132\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR134\" class=\"CitationRef\"\u003e134\u003c/span\u003e\u003c/sup\u003e. Beyond subgroup fairness evaluations, validation must align with emerging governance frameworks such as the Food and Drug Administration (FDA)\u0026rsquo;s Predetermined Change Control Plans (PCCP)\u003csup\u003e\u003cspan citationid=\"CR135\" class=\"CitationRef\"\u003e135\u003c/span\u003e\u003c/sup\u003e for managing model updates, Centers for Medicare \u0026amp; Medicaid Services (CMS)\u0026rsquo;s requirements for human review of automated determinations\u003csup\u003e\u003cspan citationid=\"CR136\" class=\"CitationRef\"\u003e136\u003c/span\u003e\u003c/sup\u003e, and National Institute of Standards and Technology (NIST)\u0026rsquo;s AI Risk Management Framework\u003csup\u003e\u003cspan citationid=\"CR137\" class=\"CitationRef\"\u003e137\u003c/span\u003e\u003c/sup\u003e for post-deployment monitoring. Reconciling these frameworks is critical, as continuous learning models inherently challenge regulatory expectations for stability and traceability\u003csup\u003e\u003cspan citationid=\"CR138\" class=\"CitationRef\"\u003e138\u003c/span\u003e\u003c/sup\u003e. Evaluations should therefore quantify not only accuracy but also impacts on clinician trust, workflow integration, and patient safety in life-critical settings\u003csup\u003e\u003cspan additionalcitationids=\"CR140\" citationid=\"CR139\" class=\"CitationRef\"\u003e139\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR141\" class=\"CitationRef\"\u003e141\u003c/span\u003e\u003c/sup\u003e. Embedding automated fact-checking, uncertainty estimation, and interpretable reasoning can support traceability and compliance\u003csup\u003e\u003cspan additionalcitationids=\"CR143\" citationid=\"CR142\" class=\"CitationRef\"\u003e142\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR144\" class=\"CitationRef\"\u003e144\u003c/span\u003e\u003c/sup\u003e. Real-time interoperability with EHR and monitoring systems must preserve clinician oversight and accountability\u003csup\u003e\u003cspan citationid=\"CR145\" class=\"CitationRef\"\u003e145\u003c/span\u003e\u003c/sup\u003e. Sustained progress will depend on cross-disciplinary collaboration to harmonize oversight mechanisms while balancing innovation with safety, transparency, and equity\u003csup\u003e\u003cspan citationid=\"CR146\" class=\"CitationRef\"\u003e146\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis scoping review emphasizes the growing importance of FMs in ECC, a field defined by clinical urgency, high complexity, and the need for rapid information synthesis. While language models dominate current applications, emerging work on multimodal architectures signals a move toward more comprehensive, context-aware support tools. Most current studies remain constrained by small or single-center datasets, retrospective designs, and limited external validation, highlighting the early and fragmented nature of the evidence base. Despite accelerating research activity, no FM has yet demonstrated verified clinical deployment or measurable impact on patient outcomes. Future work must prioritize rigorous validation, cross-system interoperability, and consistent, safety-oriented evaluation frameworks, while also addressing fairness, transparency, and reproducibility. Progress along these dimensions will be essential for realizing clinically meaningful, trustworthy, and equitable FM applications in life-critical ECC settings.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eL.C. is funded by the National Institute of Health through DS-I Africa U54 TW012043-01 and Bridge2AI OT2OD032701, the National Science Foundation through ITEST #2148451, a grant of the Boston-Korea Innovative Research Project (RS-2024-00403047) and a grant of the Korea Health Technology R\u0026amp;D Project (RS-2024-00439677) through the Korea Health Industry Development Institute (KHIDI) as funded by the Ministry of Health \u0026amp; Welfare, Republic of Korea.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eF.X. conceptualized the study and led the work. Y.K. conducted a literature search. Y.K., Y.F., and F.X. screened the titles, abstracts, and full texts. Y.K., X.M., Y.F., and \u0026nbsp;F.X. conducted data extraction. Y.K. drafted the initial manuscript. Y.K. and X.M. performed data synthesis and incorporated revisions based on feedback. Y.K., X.M., A.M., Z.X., Z.S., C.G., D.W., J.H., Q.W., A.H., R.Z., M.L., G.S., E.L., X.X., M.P., N.L., L.C., \u0026amp; F.X. revised the manuscript. \u0026nbsp;F.X. supervised the study. All authors read and approved the final version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMichael Puskarich has served on scientific advisory boards for Inflammatix LLC, Cvtovale LLC, and Opticyte LLC, companies working in the sepsis diagnosis and prognosis space.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eChoi, A. \u003cem\u003eet al.\u003c/em\u003e Development of a machine learning-based clinical decision support system to predict clinical deterioration in patients visiting the emergency department. \u003cem\u003eSci. Rep.\u003c/em\u003e 13, 8561 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKwon, J.-M. \u003cem\u003eet al.\u003c/em\u003e Validation of deep-learning-based triage and acuity score using a large national dataset. \u003cem\u003ePLoS One\u003c/em\u003e 13, e0205836 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang, H., Yu, J. Y., Yoon, S., Kim, T. \u0026amp; Cha, W. C. Machine learning-based suggestion for critical interventions in the management of potentially severe conditioned patients in emergency department triage. \u003cem\u003eSci. Rep.\u003c/em\u003e 12, 10537 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClifton, D. A. \u003cem\u003eet al.\u003c/em\u003e A large-scale clinical validation of an integrated monitoring system in the emergency department. \u003cem\u003eIEEE J. Biomed. Health Inform.\u003c/em\u003e 17, 835\u0026ndash;842 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSanchez-Pinto, L. N., Luo, Y. \u0026amp; Churpek, M. M. Big data and data science in critical care. \u003cem\u003eChest\u003c/em\u003e 154, 1239\u0026ndash;1248 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, N. \u003cem\u003eet al.\u003c/em\u003e Leveraging large-scale electronic health records and interpretable machine learning for clinical decision making at the emergency department: Protocol for system development and validation. \u003cem\u003eJMIR Res. Protoc.\u003c/em\u003e 11, e34201 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, C.-H. \u003cem\u003eet al.\u003c/em\u003e Emergency department disposition prediction using a deep neural network with integrated clinical narratives and structured data. \u003cem\u003eInt. J. Med. Inform.\u003c/em\u003e 139, 104146 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKing, Z. \u003cem\u003eet al.\u003c/em\u003e Machine learning for real-time aggregated prediction of hospital admission for emergency patients. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 5, 104 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCandel, B. G. J. \u003cem\u003eet al.\u003c/em\u003e The effect of treatment and clinical course during Emergency Department stay on severity scoring and predicted mortality risk in Intensive Care patients. \u003cem\u003eCrit. Care\u003c/em\u003e 26, 112 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGunnerson, K. J. \u003cem\u003eet al.\u003c/em\u003e Association of an emergency department-based intensive care unit with survival and inpatient intensive care unit admissions. \u003cem\u003eJAMA Netw. Open\u003c/em\u003e 2, e197584 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKang, D.-Y. \u003cem\u003eet al.\u003c/em\u003e Artificial intelligence algorithm to predict the need for critical care in prehospital emergency medical services. \u003cem\u003eScand. J. Trauma Resusc. Emerg. Med.\u003c/em\u003e 28, 17 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGravesteijn, B. Y., Steyerberg, E. W. \u0026amp; Lingsma, H. F. Modern learning from big data in critical care: Primum non nocere. \u003cem\u003eNeurocrit. Care\u003c/em\u003e 37, 174\u0026ndash;184 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeasy, J., Li\u0026ograve;, P. \u0026amp; Ercole, A. Dynamic survival prediction in intensive care units from heterogeneous time series without the need for variable selection or curation. \u003cem\u003eSci. Rep.\u003c/em\u003e 10, 22129 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu, Y. \u003cem\u003eet al.\u003c/em\u003e Use of real-time information to predict future arrivals in the emergency department. \u003cem\u003eAnn. Emerg. Med.\u003c/em\u003e 81, 728\u0026ndash;737 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYoo, J. \u003cem\u003eet al.\u003c/em\u003e A real-time autonomous dashboard for the emergency department: 5-year case study. \u003cem\u003eJMIR MHealth UHealth\u003c/em\u003e 6, e10666 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, Z. \u003cem\u003eet al.\u003c/em\u003e Large language model-based critical care big data deployment and extraction: Descriptive analysis. \u003cem\u003eJMIR Med. Inform.\u003c/em\u003e 13, e63216 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKomorowski, M., Celi, L. A., Badawi, O., Gordon, A. C. \u0026amp; Faisal, A. A. The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. \u003cem\u003eNat. Med.\u003c/em\u003e 24, 1716\u0026ndash;1720 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKachman, M. M., Brennan, I., Oskvarek, J. J., Waseem, T. \u0026amp; Pines, J. M. How artificial intelligence could transform emergency care. \u003cem\u003eAm. J. Emerg. Med.\u003c/em\u003e 81, 40\u0026ndash;46 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoh, E. \u003cem\u003eet al.\u003c/em\u003e Physician clinical decision modification and bias assessment in a randomized controlled trial of AI assistance. \u003cem\u003eCommun. Med. (Lond.)\u003c/em\u003e 5, 59 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAtaman, M. G. \u0026amp; Sarıyer, G. Predicting waiting and treatment times in emergency departments using ordinal logistic regression models. \u003cem\u003eAm. J. Emerg. Med.\u003c/em\u003e 46, 45\u0026ndash;50 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, T., Wei, Y., Wu, J., Yi, B. \u0026amp; Li, H. Logistic regression technique is comparable to complex machine learning algorithms in predicting cognitive impairment related to post intensive care syndrome. \u003cem\u003eSci. Rep.\u003c/em\u003e 13, 2485 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubudhi, S. \u003cem\u003eet al.\u003c/em\u003e Comparing machine learning algorithms for predicting ICU admission and mortality in COVID-19. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 4, 87 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHinson, J. S. \u003cem\u003eet al.\u003c/em\u003e Multisite implementation of a workflow-integrated machine learning system to optimize COVID-19 hospital admission decisions. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 5, 94 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarafino, B. J., Davies, J. M., Bardach, N. S., Dean, M. L. \u0026amp; Dudley, R. A. N-gram support vector machines for scalable procedure and diagnosis classification, with applications to clinical free text data from the intensive care unit. \u003cem\u003eJ. Am. Med. Inform. Assoc.\u003c/em\u003e 21, 871\u0026ndash;875 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVerdaasdonk, M. J. A. \u0026amp; M. de Carvalho, R. From predictions to recommendations: Tackling bottlenecks and overstaying in the Emergency Room through a sequence of Random Forests. \u003cem\u003eHealthc. Anal. (N. Y.)\u003c/em\u003e 2, 100040 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng, L., Xue, Y.-J., Yuan, Z.-N. \u0026amp; Xing, X.-Z. Explainable SHAP-XGBoost models for pressure injuries among patients requiring with mechanical ventilation in intensive care unit. \u003cem\u003eSci. Rep.\u003c/em\u003e 15, 9878 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, Z., Li, T., Guo, S., Zeng, D. \u0026amp; Wang, K. Machine learning-based in-hospital mortality risk prediction tool for intensive care unit patients with heart failure. \u003cem\u003eFront. Cardiovasc. Med.\u003c/em\u003e 10, 1119699 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, M., Guo, C. \u0026amp; Guo, S. An explainable knowledge distillation method with XGBoost for ICU mortality prediction. \u003cem\u003eComput. Biol. Med.\u003c/em\u003e 152, 106466 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYun, H., Choi, J. \u0026amp; Park, J. H. Prediction of critical care outcome for adult patients presenting to emergency department using initial triage information: An XGBoost algorithm analysis. \u003cem\u003eJMIR Med. Inform.\u003c/em\u003e 9, e30770 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuckworth, C. \u003cem\u003eet al.\u003c/em\u003e Using explainable machine learning to characterise data drift and detect emergent health risks for emergency department admissions during COVID-19. \u003cem\u003eSci. Rep.\u003c/em\u003e 11, 23017 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlghatani, K., Ammar, N., Rezgui, A. \u0026amp; Shaban-Nejad, A. Predicting Intensive Care unit length of stay and mortality using patient vital signs: Machine learning model development and validation. \u003cem\u003eJMIR Med. Inform.\u003c/em\u003e 9, e21347 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIvanov, O. \u003cem\u003eet al.\u003c/em\u003e Improving ED Emergency Severity Index acuity assignment using machine learning and clinical natural language processing. \u003cem\u003eJ. Emerg. Nurs.\u003c/em\u003e 47, 265\u0026ndash;278.e7 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYao, L.-H., Leung, K.-C., Tsai, C.-L., Huang, C.-H. \u0026amp; Fu, L.-C. A novel deep learning-based system for triage in the emergency department using electronic medical records: Retrospective cohort study. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 23, e27008 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark, J. J. \u003cem\u003eet al.\u003c/em\u003e Convolutional-neural-network-based diagnosis of appendicitis via CT scans in patients with acute abdominal pain presenting in the emergency department. \u003cem\u003eSci. Rep.\u003c/em\u003e 10, 9556 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArntfield, R. \u003cem\u003eet al.\u003c/em\u003e Development of a convolutional neural network to differentiate among the etiology of similar appearing pathological B lines on lung ultrasound: a deep learning study. \u003cem\u003eBMJ Open\u003c/em\u003e 11, e045120 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLe, S. \u003cem\u003eet al.\u003c/em\u003e Convolutional neural network model for Intensive Care unit acute kidney injury prediction. \u003cem\u003eKidney Int. Rep.\u003c/em\u003e 6, 1289\u0026ndash;1298 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKessler, S. \u003cem\u003eet al.\u003c/em\u003e Predicting readmission to the cardiovascular intensive care unit using recurrent neural networks. \u003cem\u003eDigit. Health\u003c/em\u003e 9, 20552076221149529 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGandin, I., Scagnetto, A., Romani, S. \u0026amp; Barbati, G. Interpretability of time-series deep learning models: A study in cardiovascular patients admitted to Intensive care unit. \u003cem\u003eJ. Biomed. Inform.\u003c/em\u003e 121, 103876 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThorsen-Meyer, H.-C. \u003cem\u003eet al.\u003c/em\u003e Discrete-time survival analysis in the critically ill: a deep learning approach using heterogeneous data. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 5, 142 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHyland, S. L. \u003cem\u003eet al.\u003c/em\u003e Early prediction of circulatory failure in the intensive care unit using machine learning. \u003cem\u003eNat. Med.\u003c/em\u003e 26, 364\u0026ndash;373 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, J., Cairns, B. J., Li, J. \u0026amp; Zhu, T. Generating synthetic mixed-type longitudinal electronic health records for artificial intelligent applications. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 6, 98 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRahmatinejad, Z. \u003cem\u003eet al.\u003c/em\u003e A comparative study of explainable ensemble learning and logistic regression for predicting in-hospital mortality in the emergency department. \u003cem\u003eSci. Rep.\u003c/em\u003e 14, 3406 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eElhazmi, A. \u003cem\u003eet al.\u003c/em\u003e Machine learning decision tree algorithm role for predicting mortality in critically ill adult COVID-19 patients admitted to the ICU. \u003cem\u003eJ. Infect. Public Health\u003c/em\u003e 15, 826\u0026ndash;834 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoto, T., Camargo, C. A., Jr, Faridi, M. K., Freishtat, R. J. \u0026amp; Hasegawa, K. Machine learning-based prediction of clinical outcomes for children during emergency department triage. \u003cem\u003eJAMA Netw. Open\u003c/em\u003e 2, e186937 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, Z., Shu, W., Li, T., Zhang, X. \u0026amp; Chong, W. Interpretable machine learning for predicting sepsis risk in emergency triage patients. \u003cem\u003eSci. Rep.\u003c/em\u003e 15, 887 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrown, T. B. \u003cem\u003eet al.\u003c/em\u003e Language Models are Few-Shot Learners. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf\u003c/span\u003e\u003cspan address=\"https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDevlin, J., Chang, M.-W., Lee, K. \u0026amp; Toutanova, K. BERT: Pre-training of deep bidirectional Transformers for language understanding. \u003cem\u003earXiv [cs.CL]\u003c/em\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTouvron, H. \u003cem\u003eet al.\u003c/em\u003e LLaMA: Open and efficient foundation language models. \u003cem\u003earXiv [cs.CL]\u003c/em\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDosovitskiy, A. \u003cem\u003eet al.\u003c/em\u003e An image is worth 16x16 words: Transformers for image recognition at scale. \u003cem\u003earXiv [cs.CV]\u003c/em\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRadford, A. \u003cem\u003eet al.\u003c/em\u003e Learning transferable visual models from natural language supervision. \u003cem\u003earXiv [cs.CV]\u003c/em\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, J., Li, D., Xiong, C. \u0026amp; Hoi, S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. \u003cem\u003earXiv [cs.CV]\u003c/em\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaemi, A. \u003cem\u003eet al.\u003c/em\u003e Machine learning techniques for mortality prediction in emergency departments: a systematic review. \u003cem\u003eBMJ Open\u003c/em\u003e 11, e052663 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePiliuk, K. \u0026amp; Tomforde, S. Artificial intelligence in emergency medicine. A systematic literature review. \u003cem\u003eInt. J. Med. Inform.\u003c/em\u003e 180, 105274 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, Z., Cui, X. \u0026amp; Song, Z. Predicting sepsis onset in ICU using machine learning models: a systematic review and meta-analysis. \u003cem\u003eBMC Infect. Dis.\u003c/em\u003e 23, 635 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePreiksaitis, C. \u003cem\u003eet al.\u003c/em\u003e The role of large language models in transforming emergency medicine: Scoping review. \u003cem\u003eJMIR Med. Inform.\u003c/em\u003e 12, e53787 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, Y.-P., Lo, Y.-H., Lai, F. \u0026amp; Huang, C.-H. Disease concept-embedding based on the self-supervised method for medical information extraction from electronic health records and disease retrieval: Algorithm development and validation study. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 23, e25113 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarash, Y., Klang, E., Konen, E. \u0026amp; Sorin, V. ChatGPT-4 assistance in optimizing emergency department radiology referrals and imaging selection. \u003cem\u003eJ. Am. Coll. Radiol.\u003c/em\u003e 20, 998\u0026ndash;1003 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBushuven, S. \u003cem\u003eet al.\u003c/em\u003e \u0026lsquo;ChatGPT, can you help me save my child\u0026rsquo;s life?\u0026rsquo; - diagnostic accuracy and supportive capabilities to lay rescuers by ChatGPT in prehospital basic life support and paediatric advanced life support cases - an in-silico analysis. \u003cem\u003eJ. Med. Syst.\u003c/em\u003e 47, 123 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFraser, H. \u003cem\u003eet al.\u003c/em\u003e Comparison of diagnostic and triage accuracy of Ada Health and WebMD symptom checkers, ChatGPT, and physicians for patients in an emergency department: Clinical data analysis study. \u003cem\u003eJMIR MHealth UHealth\u003c/em\u003e 11, e49995 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHenriksson, A., Pawar, Y., Hedberg, P. \u0026amp; Naucl\u0026eacute;r, P. Multimodal fine-tuning of clinical language models for predicting COVID-19 outcomes. \u003cem\u003eArtif. Intell. Med.\u003c/em\u003e 146, 102695 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, D., Cogill, S., Hsia, R. Y., Yang, S. \u0026amp; Kim, D. Development and external validation of a pretrained deep learning model for the prediction of non-accidental trauma. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 6, 131 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSavage, T., Wang, J. \u0026amp; Shieh, L. A large language model screening tool to target patients for best Practice Alerts: Development and validation. \u003cem\u003eJMIR Med. Inform.\u003c/em\u003e 11, e49886 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAityan, S. K. \u003cem\u003eet al.\u003c/em\u003e Integrated AI medical emergency diagnostics advising system. \u003cem\u003eElectronics (Basel)\u003c/em\u003e 13, 4389 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAkhondi-Asl, A. \u003cem\u003eet al.\u003c/em\u003e Comparing the quality of domain-specific versus general language models for artificial intelligence-generated differential diagnoses in PICU patients. \u003cem\u003ePediatr. Crit. Care Med.\u003c/em\u003e 25, e273\u0026ndash;e282 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlessandri-Bonetti, M., Liu, H. Y., Donovan, J. M., Ziembicki, J. A. \u0026amp; Egro, F. M. A comparative analysis of ChatGPT, ChatGPT-4, and Google Bard performances at the Advanced Burn Life Support exam. \u003cem\u003eJ. Burn Care Res.\u003c/em\u003e 45, 945\u0026ndash;948 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmacher, S. A. \u003cem\u003eet al.\u003c/em\u003e Prediction of outcomes after cardiac arrest by a generative artificial intelligence model. \u003cem\u003eResusc. Plus\u003c/em\u003e 18, 100587 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaim, G. B. \u003cem\u003eet al.\u003c/em\u003e Evaluating large language model-assisted emergency triage: A comparison of acuity assessments by GPT-4 and medical experts. \u003cem\u003eJ. Clin. Nurs.\u003c/em\u003e (2024) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/jocn.17490\u003c/span\u003e\u003cspan address=\"10.1111/jocn.17490\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChung, P. \u003cem\u003eet al.\u003c/em\u003e Large language model capabilities in perioperative risk prediction and prognostication. \u003cem\u003eJAMA Surg.\u003c/em\u003e 159, 928\u0026ndash;937 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFerri, P. \u003cem\u003eet al.\u003c/em\u003e Deep continual learning for medical call incidents text classification under the presence of dataset shifts. \u003cem\u003eComput. Biol. Med.\u003c/em\u003e 175, 108548 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGimeno, A., Krause, K., D\u0026rsquo;Souza, S. \u0026amp; Walsh, C. G. Completeness and readability of GPT-4-generated multilingual discharge instructions in the pediatric emergency department. \u003cem\u003eJAMIA Open\u003c/em\u003e 7, ooae050 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGlicksberg, B. S. \u003cem\u003eet al.\u003c/em\u003e Evaluating the accuracy of a state-of-the-art large language model for prediction of admissions from the emergency room. \u003cem\u003eJ. Am. Med. Inform. Assoc.\u003c/em\u003e 31, 1921\u0026ndash;1928 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, L. L. \u003cem\u003eet al.\u003c/em\u003e A multi-center study on the adaptability of a shared foundation model for electronic health records. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 7, 171 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, T. \u003cem\u003eet al.\u003c/em\u003e Patient-representing population\u0026rsquo;s perceptions of GPT-generated versus standard emergency department discharge instructions: Randomized blind survey assessment. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 26, e60336 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhaldi, A. \u003cem\u003eet al.\u003c/em\u003e Accuracy of ChatGPT responses on tracheotomy for patient education. \u003cem\u003eEur. Arch. Otorhinolaryngol.\u003c/em\u003e 281, 6167\u0026ndash;6172 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLe Guellec, B. \u003cem\u003eet al.\u003c/em\u003e Performance of an Open-Source large Language Model in extracting information from free-text radiology reports. \u003cem\u003eRadiol. Artif. Intell.\u003c/em\u003e 6, e230364 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, S. \u003cem\u003eet al.\u003c/em\u003e Deep learning-based natural language processing for detecting medical symptoms and histories in emergency patient triage. \u003cem\u003eAm. J. Emerg. Med.\u003c/em\u003e 77, 29\u0026ndash;38 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLevin, C., Kagan, T., Rosen, S. \u0026amp; Saban, M. An evaluation of the capabilities of language models and nurses in providing neonatal clinical decision support. \u003cem\u003eInt. J. Nurs. Stud.\u003c/em\u003e 155, 104771 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, Y.-T., Deng, Y.-X., Tsai, C.-L., Huang, C.-H. \u0026amp; Fu, L.-C. Interpretable deep learning system for identifying critical patients through the prediction of triage level, hospitalization, and length of stay: Prospective study. \u003cem\u003eJMIR Med. Inform.\u003c/em\u003e 12, e48862 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMasanneck, L. \u003cem\u003eet al.\u003c/em\u003e Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: Comparative study. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 26, e53297 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcCoy, T. H., Jr \u0026amp; Perlis, R. H. Dimensional measures of psychopathology in children and adolescents using large language models. \u003cem\u003eBiol. Psychiatry\u003c/em\u003e 96, 940\u0026ndash;947 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOliveira, L. L., Jiang, X., Babu, A. N., Karajagi, P. \u0026amp; Daneshkhah, A. Effective Natural Language Processing algorithms for early alerts of gout flares from chief complaints. \u003cem\u003eForecasting\u003c/em\u003e 6, 224\u0026ndash;238 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRothchild, E., Baker, C., Smith, I. T., Tanna, N. \u0026amp; Ricci, J. A. Evaluating the utility of ChatGPT in diagnosing and managing maxillofacial trauma. \u003cem\u003eJ. Craniofac. Surg.\u003c/em\u003e 36, 1237\u0026ndash;1241 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaner, F. H., Saner, Y. M., Abufarhaneh, E., Broering, D. C. \u0026amp; Raptis, D. A. Comparative analysis of artificial intelligence (AI) languages in predicting Sequential Organ Failure Assessment (SOFA) scores. \u003cem\u003eCureus\u003c/em\u003e 16, e59662 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eScquizzato, T. \u003cem\u003eet al.\u003c/em\u003e Testing ChatGPT ability to answer laypeople questions about cardiac arrest and cardiopulmonary resuscitation. \u003cem\u003eResuscitation\u003c/em\u003e 194, 110077 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeo, J. \u003cem\u003eet al.\u003c/em\u003e Evaluation framework of large language models in medical documentation: Development and usability study. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 26, e58329 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSezgin, E., Sirrianni, J. W. \u0026amp; Kranz, K. Evaluation of a digital scribe: Conversation summarization for Emergency Department consultation calls. \u003cem\u003eAppl. Clin. Inform.\u003c/em\u003e 15, 600\u0026ndash;611 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang, K., Reisler, J., Rose, C., Suffoletto, B. \u0026amp; Kim, D. GPT-4 accurately classifies the clinical actionability of emergency department imaging reports. \u003cem\u003eJ. Med. Artif. Intell.\u003c/em\u003e 7, 36\u0026ndash;36 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUrquhart, E. \u003cem\u003eet al.\u003c/em\u003e A pilot feasibility study comparing large language models in extracting key information from ICU patient text records from an Irish population. \u003cem\u003eIntensive Care Med. Exp.\u003c/em\u003e 12, 71 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWali, T., Bolatbekov, A., Maimaitijiang, E., Salman, D. \u0026amp; Mamatjan, Y. A novel recommender framework with chatbot to stratify heart attack risk. \u003cem\u003eDiscov Med (Cham)\u003c/em\u003e 1, 161 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, X. \u003cem\u003eet al.\u003c/em\u003e Performance of ChatGPT on prehospital acute ischemic stroke and large vessel occlusion (LVO) stroke screening. \u003cem\u003eDigit. Health\u003c/em\u003e 10, 20552076241297127 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, H. \u003cem\u003eet al.\u003c/em\u003e Evaluating accuracy and reproducibility of large language model performance on critical care assessments in pharmacy education. \u003cem\u003eFront. Artif. Intell.\u003c/em\u003e 7, 1514896 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYau, J. Y.-S. \u003cem\u003eet al.\u003c/em\u003e Accuracy of prospective assessments of 4 large language model chatbot responses to patient questions about emergency care: Experimental comparative study. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 26, e60291 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmacher, S. A. \u003cem\u003eet al.\u003c/em\u003e Can the large language model ChatGPT-4omni predict outcomes in adult patients with status epilepticus? \u003cem\u003eEpilepsia\u003c/em\u003e 66, 674\u0026ndash;685 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArslan, B., Nuhoglu, C., Satici, M. O. \u0026amp; Altinbilek, E. Evaluating LLM-based generative AI tools in emergency triage: A comparative study of ChatGPT Plus, Copilot Pro, and triage nurses. \u003cem\u003eAm. J. Emerg. Med.\u003c/em\u003e 89, 174\u0026ndash;181 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBalta, K. Y., Javidan, A. P., Walser, E., Arntfield, R. \u0026amp; Prager, R. Evaluating the appropriateness, consistency, and readability of ChatGPT in critical care recommendations. \u003cem\u003eJ. Intensive Care Med.\u003c/em\u003e 40, 184\u0026ndash;190 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBroad, A. \u003cem\u003eet al.\u003c/em\u003e Factors associated with abusive head trauma in young children presenting to Emergency Medical Services using a large language model. \u003cem\u003ePrehosp. Emerg. Care\u003c/em\u003e 29, 227\u0026ndash;237 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFeng, R. \u003cem\u003eet al.\u003c/em\u003e Engineering of generative artificial intelligence and natural language processing models to accurately identify arrhythmia recurrence. \u003cem\u003eCirc. Arrhythm. Electrophysiol.\u003c/em\u003e 18, e013023 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLevra, A. G. \u003cem\u003eet al.\u003c/em\u003e A large language model-based clinical decision support system for syncope recognition in the emergency department: A framework for clinical workflow integration. \u003cem\u003eEur. J. Intern. Med.\u003c/em\u003e 131, 113\u0026ndash;120 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHo, B. \u003cem\u003eet al.\u003c/em\u003e Evaluation of generative artificial intelligence models in predicting pediatric Emergency Severity Index levels. \u003cem\u003ePediatr. Emerg. Care\u003c/em\u003e 41, 251\u0026ndash;255 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiller, E. D. \u003cem\u003eet al.\u003c/em\u003e Accuracy of commercial large language model (ChatGPT) to predict the diagnosis for prehospital patients suitable for ambulance transport decisions: Diagnostic accuracy study. \u003cem\u003ePrehosp. Emerg. Care\u003c/em\u003e 29, 238\u0026ndash;242 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePathak, A., Marshall, C., Davis, C., Yang, P. \u0026amp; Kamaleswaran, R. RespBERT: A multi-site validation of a Natural Language Processing algorithm, of radiology notes to identify acute respiratory distress syndrome (ARDS). \u003cem\u003eIEEE J. Biomed. Health Inform.\u003c/em\u003e 29, 1455\u0026ndash;1463 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShekhar, A. C. \u003cem\u003eet al.\u003c/em\u003e Use of a large language model (LLM) for ambulance dispatch and triage. \u003cem\u003eAm. J. Emerg. Med.\u003c/em\u003e 89, 27\u0026ndash;29 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilliams, B. \u0026amp; Erstad, B. L. Analysis of responses from artificial intelligence programs to medication-related questions derived from critical care guidelines. \u003cem\u003eAm. J. Health. Syst. Pharm.\u003c/em\u003e (2025) doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/ajhp/zxaf075\u003c/span\u003e\u003cspan address=\"10.1093/ajhp/zxaf075\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhader, F. \u003cem\u003eet al.\u003c/em\u003e Medical transformer for multimodal survival prediction in intensive care: integration of imaging and non-imaging data. \u003cem\u003eSci. Rep.\u003c/em\u003e 13, 10666 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark, S. \u003cem\u003eet al.\u003c/em\u003e Multi-task vision transformer using low-level chest X-ray feature corpus for COVID-19 diagnosis and severity quantification. \u003cem\u003eMed. Image Anal.\u003c/em\u003e 75, 102299 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLian, C., Zhou, H.-Y., Liang, D., Qin, J. \u0026amp; Wang, L. Efficient medical vision-language alignment through adapting masked vision models. \u003cem\u003eIEEE Trans. Med. Imaging\u003c/em\u003e PP, 1\u0026ndash;1 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu, X., Zhang, L., Wu, Z. \u0026amp; Zhu, D. Core-periphery multi-modality feature alignment for zero-shot medical image analysis. \u003cem\u003eIEEE Trans. Med. Imaging\u003c/em\u003e PP, (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, P., Zhang, H. \u0026amp; Yuan, Y. MCPL: Multi-modal Collaborative Prompt Learning for medical vision-language model. \u003cem\u003eIEEE Trans. Med. Imaging\u003c/em\u003e 43, 4224\u0026ndash;4235 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, Y. \u003cem\u003eet al.\u003c/em\u003e Visual analytics for efficient image exploration and user-guided image captioning. \u003cem\u003eIEEE Trans. Vis. Comput. Graph.\u003c/em\u003e 30, 2875\u0026ndash;2887 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi, J., Hou, Y., Chen, X., Pan, Y. \u0026amp; Xiang, Y. Vision-language model for generating textual descriptions from clinical images: Model development and validation study. \u003cem\u003eJMIR Form. Res.\u003c/em\u003e 8, e32690 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, X., Zhang, X., Chen, W., Li, C. \u0026amp; Yu, C. Improving speech depression detection using transfer learning with wav2vec 2.0 in low-resource environments. \u003cem\u003eSci. Rep.\u003c/em\u003e 14, 9543 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKlempir, O., Skryjova, A., Tichopad, A. \u0026amp; Krupicka, R. Ranking pre-trained speech embeddings in Parkinson\u0026rsquo;s disease detection: Does Wav2Vec 2.0 outperform its 1.0 version across speech modes and languages? \u003cem\u003eComput. Struct. Biotechnol. J.\u003c/em\u003e 27, 2584\u0026ndash;2601 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArdic, N. \u0026amp; Dinc, R. Emerging trends in multi-modal artificial intelligence for clinical decision support: A narrative review. \u003cem\u003eHealth Informatics J.\u003c/em\u003e 31, 14604582251366141 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, Y. \u003cem\u003eet al.\u003c/em\u003e A contrastive learning approach for ICU false arrhythmia alarm reduction. \u003cem\u003eSci. Rep.\u003c/em\u003e 12, 4689 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHao, Y. \u003cem\u003eet al.\u003c/em\u003e Multimodal integration in healthcare: Development with applications in disease management (preprint). \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 27, e76557 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarrit, S. \u003cem\u003eet al.\u003c/em\u003e Intracranial multimodal monitoring in neurocritical care (Neurocore-iMMM): an open, decentralized consensus. \u003cem\u003eCrit. Care\u003c/em\u003e 28, 427 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubasri, V. \u003cem\u003eet al.\u003c/em\u003e Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. \u003cem\u003eJAMA Netw. Open\u003c/em\u003e 8, e2513685 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBiesheuvel, L. A. \u003cem\u003eet al.\u003c/em\u003e Large language models in critical care. \u003cem\u003eJ. Intensive Med.\u003c/em\u003e 5, 113\u0026ndash;118 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEl Arab, R. A. \u0026amp; Al Moosa, O. A. Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 8, 548 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuccioni, A. S., Jernite, Y. \u0026amp; Strubell, E. Power hungry processing: Watts driving the cost of AI deployment? \u003cem\u003earXiv [cs.LG]\u003c/em\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatirai, A. The environmental costs of artificial intelligence for healthcare. \u003cem\u003eAsian Bioeth. Rev.\u003c/em\u003e 16, 527\u0026ndash;538 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhanna, N. N. \u003cem\u003eet al.\u003c/em\u003e Economics of artificial Intelligence in healthcare: Diagnosis vs. Treatment. \u003cem\u003eHealthcare (Basel)\u003c/em\u003e 10, 2493 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUeda, D. \u003cem\u003eet al.\u003c/em\u003e Climate change and artificial intelligence in healthcare: Review and recommendations towards a sustainable future. \u003cem\u003eDiagn. Interv. Imaging\u003c/em\u003e 105, 453\u0026ndash;459 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNord-Bronzyk, A. \u003cem\u003eet al.\u003c/em\u003e Assessing risk in implementing new artificial intelligence triage tools-how much risk is reasonable in an already risky world? \u003cem\u003eAsian Bioeth. Rev.\u003c/em\u003e 17, 187\u0026ndash;205 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNing, Y. \u003cem\u003eet al.\u003c/em\u003e Generative artificial intelligence and ethical considerations in health care: a scoping review and ethics checklist. \u003cem\u003eLancet Digit. Health\u003c/em\u003e 6, e848\u0026ndash;e856 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHenry, K. E. \u003cem\u003eet al.\u003c/em\u003e Human-machine teaming is key to AI adoption: clinicians\u0026rsquo; experiences with a deployed machine learning system. \u003cem\u003eNPJ Digit. Med.\u003c/em\u003e 5, 97 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLing Kuo, R. Y. \u003cem\u003eet al.\u003c/em\u003e Stakeholder perspectives towards diagnostic artificial intelligence: a co-produced qualitative evidence synthesis. \u003cem\u003eEClinicalMedicine\u003c/em\u003e 71, 102555 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith, H., Downer, J. \u0026amp; Ives, J. Clinicians and AI use: where is the professional guidance? \u003cem\u003eJ. Med. Ethics\u003c/em\u003e 50, 437\u0026ndash;441 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBienefeld, N., Keller, E. \u0026amp; Grote, G. Human-AI teaming in critical care: A comparative analysis of data scientists\u0026rsquo; and clinicians' perspectives on AI augmentation and automation. \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e 26, e50130 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoons, K. G. M. \u003cem\u003eet al.\u003c/em\u003e PROBAST\u0026thinsp;+\u0026thinsp;AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. \u003cem\u003eBMJ\u003c/em\u003e 388, e082505 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEthics and governance of artificial intelligence for health: Guidance on large multi-modal models. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.who.int/publications/i/item/9789240084759\u003c/span\u003e\u003cspan address=\"https://www.who.int/publications/i/item/9789240084759\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollins, G. S. \u003cem\u003eet al.\u003c/em\u003e TRIPOD\u0026thinsp;+\u0026thinsp;AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. \u003cem\u003eBMJ\u003c/em\u003e 385, e078378 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArora, A. \u003cem\u003eet al.\u003c/em\u003e The value of standards for health datasets in artificial intelligence-based applications. \u003cem\u003eNat. Med.\u003c/em\u003e 29, 2929\u0026ndash;2938 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRockenschaub, P. \u003cem\u003eet al.\u003c/em\u003e The impact of multi-institution datasets on the generalizability of machine learning prediction models in the ICU. \u003cem\u003eCrit. Care Med.\u003c/em\u003e 52, 1710\u0026ndash;1721 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFood and Drug Administration Staff. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.fda.gov/media/180978/download\u003c/span\u003e\u003cspan address=\"https://www.fda.gov/media/180978/download\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomas, K. S. \u003cem\u003eet al.\u003c/em\u003e Prior authorization and utilization management for post-acute home health in Medicare Advantage: the motivations, players, processes, unique challenges, and impacts on patient care. \u003cem\u003eHealth Aff. Sch.\u003c/em\u003e 3, qxaf020 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTabassi, E. \u003cem\u003eArtificial Intelligence Risk Management Framework (AI RMF 1.0): AI RMF (1.0)\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.6028/NIST.AI.100-1\u003c/span\u003e\u003cspan address=\"10.6028/NIST.AI.100-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023) doi:10.6028/nist.ai.100-1.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTopol, E. J. High-performance medicine: the convergence of human and artificial intelligence. \u003cem\u003eNat. Med.\u003c/em\u003e 25, 44\u0026ndash;56 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVasey, B. \u003cem\u003eet al.\u003c/em\u003e Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. \u003cem\u003eNat. Med.\u003c/em\u003e 28, 924\u0026ndash;933 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, X. \u003cem\u003eet al.\u003c/em\u003e Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. \u003cem\u003eNat. Med.\u003c/em\u003e 26, 1364\u0026ndash;1374 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCruz Rivera, S. \u003cem\u003eet al.\u003c/em\u003e Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. \u003cem\u003eNat. Med.\u003c/em\u003e 26, 1351\u0026ndash;1363 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, R. \u003cem\u003eet al.\u003c/em\u003e Retrieval-augmented generation for generative artificial intelligence in health care. \u003cem\u003eNpj Health Syst.\u003c/em\u003e 2, 1\u0026ndash;5 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGargari, O. K. \u0026amp; Habibi, G. Enhancing medical AI with retrieval-augmented generation: A mini narrative review. \u003cem\u003eDigit. Health\u003c/em\u003e 11, 20552076251337177 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSadeghi, Z. \u003cem\u003eet al.\u003c/em\u003e A review of Explainable Artificial Intelligence in healthcare. \u003cem\u003eComput. Electr. Eng.\u003c/em\u003e 118, 109370 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFoldy, S. \u003cem\u003eet al.\u003c/em\u003e Public Health FHIR\u0026reg; Playbook. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cdc.gov/data-interoperability/media/pdfs/PHFIC_Public-Health-FHIR-Playbook.pdf\u003c/span\u003e\u003cspan address=\"https://www.cdc.gov/data-interoperability/media/pdfs/PHFIC_Public-Health-FHIR-Playbook.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTransparency for Machine Learning-Enabled Medical Devices: Guiding Principles. \u003cem\u003eU.S. Food and Drug Administration\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles\u003c/span\u003e\u003cspan address=\"https://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Emergency and critical care, Emergency department, Intensive care unit, Foundation models, Large language models, Multimodal models, Clinical decision support, Scoping review","lastPublishedDoi":"10.21203/rs.3.rs-8338830/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8338830/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eEmergency and critical care (ECC) settings demand rapid decision-making with emergent conditions, dynamic patient trajectories, and diagnostic uncertainty. Foundation models (FMs), large neural networks pretrained on extensive datasets using self-supervised learning, show promise for diverse clinical tasks in these high-acuity environments. However, no prior work has comprehensively reviewed FM applications in ECC or the barriers limiting their implementation. Following PRISMA-ScR guidelines, we identified 49 eligible studies. Most focused on language models, with comparatively limited exploration of multimodal architectures. FMs utilized diverse data modalities, including free text, tabular records, time-series signals, and imaging, supporting tasks such as outcome prediction, diagnosis, information extraction, and text generation. Despite promising applications such as triage, discharge instructions, and disease diagnosis, current evidence is predominantly retrospective, with minimal external validation or prospective testing. Also, many FMs were trained on internet text that may be misaligned with medical reasoning, introducing safety and ethical risks, and highlighting the need for clinically supervised deployment. FMs have yet to demonstrate benefit in the ECC context; none of the included studies had real-world model deployment or improvements in clinical outcomes. Future research should prioritize the development of multimodal FMs with multicenter, temporally robust validation and prospective trials that emphasize safety, equity, and clinician trust.\u003c/p\u003e","manuscriptTitle":"Application of Foundation Models in Emergency and Critical Care: A Scoping Review","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-16 08:50:42","doi":"10.21203/rs.3.rs-8338830/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2e2172a5-fcb8-475f-9d0c-69063f254366","owner":[],"postedDate":"January 16th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":60934086,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":60934087,"name":"Health sciences/Health care"},{"id":60934088,"name":"Physical sciences/Mathematics and computing"},{"id":60934089,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-02-04T12:56:53+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-16 08:50:42","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8338830","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8338830","identity":"rs-8338830","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00