Enhancing Predictive Modeling for Respiratory Support with LLM-Driven Guideline Adherence

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Optimal respiratory support selection between high-flow nasal cannula (HFNC) and noninvasive ventilation (NIV) for intensive care units (ICU) patients at risk of invasive mechanical ventilation (IMV) remains unclear, particularly in cases not represented in prior clinical trials. We previously developed RepFlow-CFR, a deep counterfactual model estimating individualized treatment effects (ITE) of HFNC versus NIV. However, interpretability and guideline alignment remain challenges for clinical adoption. This study describes the development and integration of a clinical guideline-driven LLM to enhance deep counterfactual model recommendations for NIV versus HFNC in patients at high-risk for invasive mechanical ventilation. Methods We enhanced RepFlow-CFR by incorporating a large language model (LLM, Claude 3.5 Sonnet) to enforce clinical guideline adherence and generate explainable treatment recommendations. The LLM was configured in a HIPAA-compliant AWS environment and prompted using structured patient data, clinical notes, and formal guideline criteria. Recommendations from RepFlow-CFR and LLM were compared to actual treatment decisions to assess concordance. We evaluated IMV and mortality/hospice rates across concordant and discordant groups. Additionally, we conducted a structured chart review of 20 cases to assess the clinical validity and safety of LLM-driven recommendations. Results Among 1,261 ICU encounters, treatments concordant with LLM-enhanced recommendations were associated with significantly lower IMV rates (e.g., 24.47% when concordant versus 52.94% when discordant with the HFNC recommendation, corresponding to a 97.33% relative risk increase when discordant) and reduced odds of mortality or hospice discharge (odds ratio = 0.670, p = 0.046). In the chart review, 95% of LLM recommendations aligned with clinical guidelines, and physicians agreed with 65% of final recommendations. Errors were noted in 11/20 cases, with most deemed low or moderate risk; only 2 were rated as potentially causing severe harm. Conclusions Integrating LLMs for guideline enforcement improves the interpretability and clinical alignment of counterfactual models in respiratory support decision-making. This hybrid framework not only enhances concordance with real-world practice but may also improve patient outcomes. Future work will refine contraindication detection and expand validation to prospective clinical trials.
Full text 153,883 characters · extracted from preprint-html · click to expand
Enhancing Predictive Modeling for Respiratory Support with LLM-Driven Guideline Adherence | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Enhancing Predictive Modeling for Respiratory Support with LLM-Driven Guideline Adherence Xiaolei Lu, Michael Miller, Alex K. Pearce, Preeti Gupta, Thaidan T. Pham, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7230335/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Nov, 2025 Read the published version in Critical Care → Version 1 posted 10 You are reading this latest preprint version Abstract Background Optimal respiratory support selection between high-flow nasal cannula (HFNC) and noninvasive ventilation (NIV) for intensive care units (ICU) patients at risk of invasive mechanical ventilation (IMV) remains unclear, particularly in cases not represented in prior clinical trials. We previously developed RepFlow-CFR, a deep counterfactual model estimating individualized treatment effects (ITE) of HFNC versus NIV. However, interpretability and guideline alignment remain challenges for clinical adoption. This study describes the development and integration of a clinical guideline-driven LLM to enhance deep counterfactual model recommendations for NIV versus HFNC in patients at high-risk for invasive mechanical ventilation. Methods We enhanced RepFlow-CFR by incorporating a large language model (LLM, Claude 3.5 Sonnet) to enforce clinical guideline adherence and generate explainable treatment recommendations. The LLM was configured in a HIPAA-compliant AWS environment and prompted using structured patient data, clinical notes, and formal guideline criteria. Recommendations from RepFlow-CFR and LLM were compared to actual treatment decisions to assess concordance. We evaluated IMV and mortality/hospice rates across concordant and discordant groups. Additionally, we conducted a structured chart review of 20 cases to assess the clinical validity and safety of LLM-driven recommendations. Results Among 1,261 ICU encounters, treatments concordant with LLM-enhanced recommendations were associated with significantly lower IMV rates (e.g., 24.47% when concordant versus 52.94% when discordant with the HFNC recommendation, corresponding to a 97.33% relative risk increase when discordant) and reduced odds of mortality or hospice discharge (odds ratio = 0.670, p = 0.046). In the chart review, 95% of LLM recommendations aligned with clinical guidelines, and physicians agreed with 65% of final recommendations. Errors were noted in 11/20 cases, with most deemed low or moderate risk; only 2 were rated as potentially causing severe harm. Conclusions Integrating LLMs for guideline enforcement improves the interpretability and clinical alignment of counterfactual models in respiratory support decision-making. This hybrid framework not only enhances concordance with real-world practice but may also improve patient outcomes. Future work will refine contraindication detection and expand validation to prospective clinical trials. Causal inference large language models individualized treatment effect high-flow nasal cannula noninvasive ventilation guideline adherence Figures Figure 1 1. Background Acute respiratory failure (ARF) is common in critically ill patients, affecting about half of all intensive care units (ICU) admissions either upon arrival or during their stay. Among patients not requiring urgent intubation and invasive mechanical ventilation (IMV), high-flow nasal cannula (HFNC) and noninvasive ventilation (NIV) are commonly used respiratory support therapies. The choice of initial respiratory support therapy can impact important patient outcomes like need for IMV or mortality, however, the optimal choice based on prior randomized controlled trials (RCTs. Patients can have multiple indications for either HFNC or NIV at the same time, confusing treatment decisions, and many patients with ARF in the real-world who receive HFNC and NIV would not fit the inclusion and exclusion criteria of prior RCTs comparing the two modalities. There is also the possibility of heterogeneity of treatment effect, in which baseline patient characteristics influence response to therapy. The uncertainty introduced by these features into the decision to initially treat a patient with ARF with either NIV or HFNC highlights the need for data-driven precision medicine techniques to assist with choice of NIV or HFNC. Machine learning techniques can be employed to estimate the individualized treatment effect (ITE) of high-flow nasal cannula (HFNC) versus noninvasive ventilation (NIV) for each patient 7 . Recent advancements in causal inference and counterfactual modeling have produced a diverse toolkit for ITE estimation in observational data. These tools include tree-based methods (e.g., Causal Forests 12 ), representation learning approaches (e.g., TARNet 13 ), meta-learners (e.g., X-learner 14 ), and balancing techniques such as Bayesian Additive Regression Trees 15 . Building upon this foundation, we previously developed and validated RepFlow-CFR, a deep counterfactual representation model tailored to the critical care setting. RepFlow-CFR is designed to estimate patient-specific outcomes under both HFNC and NIV by learning balanced representations that adjust for confounding and capture latent patient heterogeneity. The model combines shared representation learning, normalizing flows for outcome modeling, and a second-stage adjustment for unmeasured confounding, enabling it to generate robust, individualized predictions for response to NIV versus HFNC to decrease risk of invasive mechanical ventilation. When applied to ICU patients identified as high-risk for respiratory failure, patients receiving treatments concordant with RepFlow-CFR recommendations showed lower rates of IMV and mortality or hospice discharge, highlighting the model’s ability to identify beneficial therapy pathways in complex ICU populations. While RepFlow-CFR demonstrated strong performance in generating individualized treatment recommendations and identifying beneficial therapy pathways, it shares a common limitation with many deep learning models, namely, a lack of transparency in its decision-making process. This “black box” nature makes it difficult for clinicians to interpret or trust the model’s output, especially when recommendations diverge from established clinical guidelines. Moreover, RepFlow-CFR is not inherently constrained to follow these guidelines, which may lead to discordance between data-driven recommendations and evidence-based best practices. To address these challenges, we introduce a large language model (LLM) into the decision pipeline to serve as a guideline-aware, explainable reasoning layer. By leveraging structured patient data, clinical notes, and formal guideline criteria, the LLM can validate or refine the model’s recommendations, ensuring that treatment suggestions align with current standards of care while also providing interpretable justifications. In this study, we aim to refine treatment recommendations from a deep counterfactual model by integrating LLM to enforce adherence to clinical guidelines and enhance interpretability. We further evaluate the impact of this LLM-augmented framework on concordance with real-world treatment decisions and patient outcomes, with additional validation through structured expert chart review. We hypothesized that integrating a large language model with a deep counterfactual model to refine treatment recommendations would enhance interpretability, result in treatment suggestions that better align with established guideline driven standard of care. 2. Methods 2.1 RepFlow-CFR Prediction Data Sources. We used de-identified structured Electronic Health Record (EHR) data from ICU patients at UC San Diego Health (UCSD) between January 1, 2016, and December 31, 2023, focusing on encounters where either HFNC or NIV was administered as the first respiratory support following Vent.io-predicted 16 high-risk timepoint (Vent.io T0). Figure 1 illustrates the cohort derivation process. Extracted variables included 50 vital signs and laboratory measurements (resampled into hourly bins), 6 demographic features, 12 SIRS/SOFA criteria, 12 medication categories, and 62 comorbidities. Additional derived features included baseline values, local trends, and time since last measurement (TSLM). Missing values were forward-filled up to 24 hours or mean-imputed (Supplementary Section 1). RepFlow-CFR Model. The RepFlow-CFR model is a deep counterfactual inference framework that adjusts for confounding using both observed and inferred latent variables. It consists of three stages (Supplementary Section 2 for the model architecture and mathematical formulation): Stage 0 applies counterfactual regression (CFR) 13 to learn shared representations balancing measured confounders across treatment groups using an integral probability metric (e.g., Wasserstein distance); Stage 1 uses a conditional normalizing flow (CNF) 17 to model the outcome distribution given the representation and treatments; Stage 2 introduces a second CNF to adjust for unmeasured confounding by transforming the treatment-dependent latent variable into an interventional distribution. During inference, RepFlow-CFR generates potential outcomes under both HFNC and NIV by sampling from the learned distributions. The ITE is defined as the difference in predicted probabilities of IMV under NIV versus HFNC. Based on the ITE, the model provides treatment recommendations: NIV preferred (ITE 0.001), or Indifferent (ITE between − 0.001 and 0.001). 2.2 LLM-Driven Guideline Enforcement 2.2.1 Claude Sonnet Configuration and Deployment We utilized Claude 3.5 Sonnet 19 , a large language model with a 200,000-token context window, optimized for high-accuracy reasoning and structured clinical analysis. The model was deployed within a HIPAA-compliant Amazon Web Services (AWS) environment using an EC2 instance configured with authenticated access to Amazon Bedrock. Inference was performed via the Bedrock Converse API, which supports structured JSON output for reliable parsing. To minimize variability and ensure consistent, deterministic responses, the model operated with a temperature setting of 0.1. Claude Sonnet was used to evaluate whether RepFlow-CFR recommendations were aligned with clinical guidelines and to independently generate treatment recommendations based on structured patient data, clinical notes and formal guideline criteria. 2.2.2 Clinical Guidelines Criteria We summarize the recommended indications for NIV and HFNC based on established clinical guidelines. Specifically, we refer to the ERS/ATS 2017 Guidelines for NIV 11 and the ERS 2022 Guidelines for HFNC 10 . The guideline criteria are outlined in Table 1 below, highlighting when each therapy is recommended, conditionally recommended, or not advised. Table 1 Clinical Indications for NIV and HFNC According to ERS/ATS Guidelines. NIV – ERS/ATS 2017 Guidelines Recommend NIV if any of the following are true: HFNC – ERS 2022 Guidelines Recommend HFNC if any of the following are true: – COPD exacerbation with acute or acute-on-chronic respiratory acidosis (pH ≤ 7.35, PaCO₂ >45 mmHg) – De novo (or acute) hypoxemic respiratory failure (e.g., AHRF, ARDS, pneumonia, COVID-19) – Acute respiratory failure due to cardiogenic pulmonary edema, including acute (or “flash”) pulmonary edema – Post-operative post-extubation in high-risk patients: HFNC is preferred over conventional oxygen but NIV is also an acceptable prophylactic treatment for high-risk patients – Neuromuscular disease or obesity hypoventilation syndrome (OHS) with acute or acute-on-chronic respiratory acidosis (pH ≤ 7.35, PaCO₂ >45 mmHg) – The patient is intolerant to NIV and moderate to severe acute respiratory failure is present – Use NIV prophylactically post-extubation in high-risk patients (e.g., those with COPD or CHF) but not in low-risk patients or in patients with established post-extubation respiratory failure – Moderate to severe acute respiratory failure is present and there is not an indication for treatment with NIV – Immunocompromised patients with mild-to-moderate acute respiratory failure (conditional recommendation) Do not recommend HFNC as first-line if: – The patient has acute hypercapnic respiratory failure (e.g., COPD with acidosis) unless NIV is contraindicated or not tolerated – There is clear guideline indication for NIV – Post-operative ARF (Conditional recommendation, moderate certainty of evidence) – Chest trauma patients with ARF (Conditional recommendation, moderate certainty of evidence) 2.2.3 LLM-Assisted Guideline Concordance Assessment Inputs to LLM To support explainable and guideline-aligned treatment recommendations, we provided LLM with a structured input set that includes clinical guideline knowledge, model predictions, and patient-specific data. These inputs allowed the LLM to generate context-aware treatment rationales and assess alignment between clinical recommendations and actual care. We summarize the input components in Table 2 below. Table 2 Inputs Provided to the LLM. Input Type Description Summarized Clinical Guidelines Guideline-based indications and contraindications for NIV and HFNC (from Section 2.2.1 ), enabling the LLM to reason using ERS/ATS 2017 and ERS 2022 recommendations. Vent.io T0 Timestamp indicating when the patient was first identified as high-risk for respiratory failure by the Vent.io early warning model. RepFlow-CFR Model Output at T0 Includes (1) the recommended treatment modality (NIV, HFNC, or Indifferent), and (2) the model rationale based on EHR-derived features. The rationale includes the top 50 most influential SHAP-ranked features impacting the decision. Recent Clinical Parameters (pre-T0) Most recent values prior to T0 for: pH, PaCO₂, respiratory rate, SpO₂, and FiO₂. Relevant Clinical Notes (≤ 72h pre-T0) Free-text notes extracted from the EHR within 72 hours before T0, including: ED Provider Notes, ED MD Progress Notes, Consults, ED OBS HPI, Progress Notes, H&P, Event/Update entries, ED OBS Progress Notes, and Radiology Reports. Prompting Strategy : The LLM was prompted to: (1) evaluate the concordance between the RepFlow-CFR model’s recommendation and guideline-based recommendations derived from patient data; (2) independently provide a final recommendation including NIV, HFNC, or Indifferent (if either option is acceptable); and (3) generate a rationale by explicitly citing relevant clinical guideline statements to support the recommendation. (Supplementary Section 3 for prompt details) Output of the LLM The LLM was prompted to return a structured JSON object as listed in Table 3 . Table 3 Structured output returned by the LLM. Key Value / Format “NIV_recommendation” { “recommendation”: “Yes” or “No” or “Either”, “confidence”: “high” / “medium” / “low”, “explanation”: “Brief rationale including guideline Reference” } “HFNC_recommendation” { “recommendation”: “Yes” or “No” or “Either”, “confidence”: “high” / “medium” / “low”, “explanation”: “Brief rationale including guideline Reference” } “Model_alignment” { “alignment”: true / false, “explanation”: “Was the RepFlow-CFR model aligned with guideline-based decision?” } 2.3 Concordance Analysis We assessed the concordance between model-generated treatment recommendations and actual clinical decisions at two levels. First, RepFlow-CFR Concordance was defined as agreement between the treatment modality recommended by the RepFlow-CFR model: NIV, HFNC, or Indifferent, and the treatment the patient actually received immediately following the Vent.io T0 timepoint. Second, LLM-Enhanced Concordance was defined as agreement between the final treatment recommendation generated by the LLM, which integrated clinical guidelines and patient-specific data, and the actual treatment administered. For both types of recommendations, we compared key clinical outcomes between concordant and discordant groups, specifically focusing on the incidence of IMV and a composite endpoint of in-hospital mortality or discharge to hospice (Mortality/Hospice). 2.4 Chart Review Validation To assess the clinical validity and safety of the LLM-enhanced recommendation framework, we conducted a retrospective chart review of 20 patient cases. These cases were selected to represent a variety of clinical scenarios, including variations in disease severity, instances where the LLM and the RepFlow-CFR model generated discordant treatment recommendations, and cases with diverse real-world treatment pathways. This purposive sampling strategy was designed to challenge the LLM's interpretive and decision-making capacity across a spectrum of relevant clinical contexts. Each case was independently reviewed by three board-certified critical care physicians. Reviewers evaluated the LLM's treatment recommendation for congruence with established clinical guidelines, specifically the ERS/ATS 2017 guideline for NIV and the ERS 2022 guideline for HFNC. In addition to assessing whether the LLM’s final recommendation aligned with guideline-based care, the reviewers also examined the explanation provided by the LLM, with particular attention to the presence of factual errors, clinically significant omissions, or hallucinated statements. To evaluate the LLM’s comprehension, knowledge retrieval, and reasoning ability, a structured evaluation framework was used, adapted from prior work in validating medical LLM outputs 18 . Given the known tendency of LLMs to mix correct and incorrect content, the review framework was designed to explicitly capture both accurate and inaccurate components of the model's response. To assess the potential for harm, the reviewers used the harm classification framework from the Agency for Healthcare Research and Quality (AHRQ) Common Formats. The structured evaluation criteria used by physician reviewers are summarized in Table 4 . These include assessments of recommendation congruence, explanation accuracy, potential for harm, clinical agreement, and LLM-specific reasoning and retrieval performance. Each criterion was accompanied by categorical evaluation options and a space for free-text reviewer comments. Table 4 Structured Evaluation Criteria for LLM Chart Review. Domain Evaluation Criteria Options / Scale Reviewer comments Recommendation Congruence Is recommendation congruent with guideline? Yes / No Explanation Accuracy Is there incorrect content in the explanation? Yes / No Is the incorrect explanation clinically significant? Yes / No Is the explanation missing important clinical content? Yes / No Error Harm Assessment Harm Likelihood for LLM error Low / Medium / High / N.A. N.A. was put in for the cases where there was no LLM error Harm Extent from LLM error None/no harm, Mild/Moderate, Severe/death, NA N.A. was put in for the cases where there was no LLM error Clinical Judgment Overall MD agreement with LLM recommendation Agree / Disagree / Partially Agree Second MD agreement (if any) Agree / Disagree / Partially Agree LLM Comprehension & Reasoning Is comprehension of the question correct? Yes / No Is evidence retrieval correct? Yes / No Is reasoning for LLM recommendation correct? Yes / No / Partially Any incorrect comprehension from LLM? Yes / No Any incorrect or missing clinical info (retrieval error or hallucination)? Yes / No Did the LLM recommendation contain incorrect rationale? Yes / No 3. Results 3.1 Patient Characteristics, Interventions, and Outcomes by RepFlow-CFR and LLM-Enhanced Recommendation Our study cohort included 1,261 ICU encounters, stratified by treatment recommendations generated by the RepFlow-CFR and LLM-enhanced models. Based on RepFlow-CFR recommendations, 859 patients were classified as NIV-preferred, 279 as HFNC-preferred, and 123 as indifferent. The LLM-enhanced recommendations reassigned patients into 759 NIV-preferred, 205 HFNC-preferred, and 297 indifferent. Across both models, baseline characteristics were generally balanced. The mean age across groups ranged from 60 to 63 years, and the proportion of male patients varied between 52.7% and 63.4%. Comorbidity burden, as measured by Charlson Comorbidity Index, was comparable across groups (median: 1.0–3.0), and the distribution of specific comorbidities such as congestive heart failure and COPD showed no significant differences. Similarly, SOFA scores at Vent.io T0 were consistent across recommendations. Post-T0 interventions, including the use of steroids, antibiotics, vasopressors, and diuretics, did not significantly differ across groups. Notably, among patients recommended HFNC by the LLM-enhanced model, 91.7% actually received HFNC, compared to 70.3% in the corresponding RepFlow-CFR group, indicating improved alignment between recommendation and practice. Clinical outcomes, including IMV, mortality, and hospice enrollment, remained similar across all recommendation groups. IMV occurred in approximately 23–29% of patients, and mortality ranged from 24.4–41.8%, with no statistically significant differences observed. These findings suggest that both models stratify patients into demographically and clinically comparable groups, with the LLM-enhanced model demonstrating improved adherence to recommended respiratory support strategies. Table 5 Baseline characteristics of patients by RepFlow-CFR and LLM-enhanced recommendations. Variable RepFlow-CFR recommendation LLM-enhanced recommendation NIV preferred HFNC preferred Indifferent P b value NIV preferred HFNC preferred Indifferent P value Characteristic Encounters, N 859 279 123 - 759 205 297 - Age(years), mean (SD) 62(16.7) 62(16.0) 63(16.0) .580 63(16.4) 60(18.0) 60(15.5) .580 Gender, N (%) Male 512(59.6) 147(52.7) 78(63.4) - 442(58.2) 118(57.6) 177(59.6) - Organ dysfunction Charlson Comorbidity Index, Median (IQR) 2.0(1.0–5.0) 2.0(1.0–4.0) 3.0(1.0–5.0) .361 3.0(1.0–5.0) 1.0(1.0–3.0) 2.0(2.0–6.0) .361 Congestive Heart Failure Component, N (%) 208(24.2) 73(26.2) 31(25.2) .800 261(34.4) 18(8.8) 33(11.1) .800 Chronic Obstructive Pulmonary Disease, N (%) 162(18.9) 52(18.6) 23(18.7) .996 176(23.2) 28(13.7) 33(11.1) .996 SOFA score (at Vent.io T0), Median (IQR) 1.0(1.0–3.0) 1.0(0.0–4.0) 1.0(0.0–3.0) .512 2.0(0.0-3.5) 0.0(0.0–2.0) 1.0(0.0–3.0) .512 Interventions after Vent.io T0 a N (%) Steroids administration 289(33.6) 80(28.7) 42(34.1) .129 225(29.6) 65(31.7) 121(40.7) .129 Antibiotics administration 717(83.5) 225(80.6) 106(86.2) .154 627(82.6) 154(75.1) 267(89.9) .154 Vasopressors administration 263(30.6) 90(32.3) 31(25.2) .306 276(36.4) 39(19.0) 69(23.2) .306 Diuretics administration 512(59.6) 159(57.0) 78(63.4) .424 462(60.9) 118(57.6) 169(56.9) .424 Initial respiratory support after Vent.io T0 (N %) HFNC 634(73.8) 196(70.3) 93(75.6) .414 485(63.9) 188(91.7) 250(84.2) - Outcomes (N %) IMV 245(28.5) 65(23.3) 29(23.6) .159 197(26.0) 55(26.8) 87(29.3) .159 Mortality 292(34.0) 93(33.3) 42(34.1) .977 253(33.3) 50(24.4) 122(41.8) .977 Hospice 9(1.0) 2(0.7) 0(0.0) .480 8(1.1) 2(1.0) 1(0.3) .480 a We focused on interventions administered from Vent.io T0 to one hour before ICU discharge for the control group (those not intubated), and from Vent.io T0 to one hour before the time of intubation for the positive group (those who require intubation). b Testing for the difference P value were χ2 for categorical variables and Kruskal-Wallis rank-sum test for continuous variables. 3.2 Concordance Analysis To evaluate the clinical utility of model-guided recommendations, we analyzed the concordance between predicted treatment recommendations and the actual treatments received by patients, stratified by outcome. We compared both the baseline RepFlow-CFR model and the LLM-enhanced framework across two major clinical outcomes: the need for IMV and Mortality & Hospice. Tables below present the rate of IMV and mortality/hospice in patients who received treatments that were concordant versus discordant with the model recommendations. Relative reduction (or increase) in outcome rates is reported to capture the direction and magnitude of benefit or risk associated with following model-predicted guidance. Across both models, patients who received treatments concordant with model recommendations had lower IMV rates. The LLM-enhanced model demonstrated greater relative reductions in IMV rates in the concordant groups, particularly under HFNC recommendations, where discordant treatment was associated with a 97% relative increase in IMV risk. For the mortality and hospice outcome, the LLM-enhanced model again showed larger differences between concordant and discordant groups. Under the LLM framework, NIV-concordant cases showed a 14.03% relative reduction in mortality/hospice, compared to only 6.83% under the baseline model. We further assessed the association between concordant treatment and outcomes using multivariable logistic regression, controlling for potential confounders including age, gender, CCI, SOFA score, and Vent.io risk score. The results are summarized in Table 8 . Concordance with the LLM-enhanced recommendation was significantly associated with reduced risk of both IMV and mortality/hospice. In particular, HFNC concordance under the LLM framework was associated with a significantly lower odds of mortality or hospice discharge (OR = 0.670, p = 0.046), suggesting a clinically meaningful improvement in outcome alignment when recommendations followed the LLM-augmented guidance. Table 6 IMV Rates Stratified by Model Recommendation and Concordance. Model Recommendation Total IMV Concordant IMV Discordant IMV Relative reduction if concordant Relative increase if discordant RepFlow-CFR NIV 28.28 22.73 30.10 19.64 6.44 HFNC 27.78 25.00 33.75 10.00 21.50 LLM-enhanced NIV 25.96 21.17 28.65 18.44 10.42 HFNC 26.83 24.47 52.94 8.80 97.33 Table 7 Mortality & Hospice Rates Stratified by Recommendation and Concordance. Model Recommendation Total Mortality & Hospice Concordant Mortality & Hospice Discordant Mortality & Hospice Relative reduction if concordant Relative increase if discordant RepFlow-CFR NIV 36.59 34.09 37.41 6.83 2.24 HFNC 29.76 30.23 28.75 -1.58 -3.40 LLM-enhanced NIV 34.39 29.56 37.11 14.03 7.93 HFNC 25.37 25.00 29.41 1.44 15.95 Table 8 Multivariable Logistic Regression Results (Odds Ratios and p-values) for Predicting the Need for IMV and Mortality & Hospice Across Methods and Sites. Outcome Model NIV concordance HFNC concordance Age Gender CCI score SOFA score Vent.io score IMV RepFlow-CFR 0.678 p = 0.032 0.729 p = 0.108 0.984 p = 0.000 1.021 p = 0.877 0.913 p = 0.000 1.063 p = 0.051 0.836 p = 0.284 LLM-enhanced 0.664 p = 0.023 0.679 p = 0.057 0.983 p = 0.000 1.068 p = 0.662 0.904 p = 0.001 1.053 p = 0.145 0.946 p = 0.766 Mortality & Hospice RepFlow-CFR 0.822 p = 0.249 0.702 p = 0.066 1.015 p = 0.000 0.732 p = 0.018 0.959 p = 0.070 1.312 p = 0.000 0.908 p = 0.547 LLM-enhanced 0.716 p = 0.048 0.670 p = 0.046 1.019 p = 0.000 0.685 p = 0.009 0.929 p = 0.005 1.272 p = 0.000 0.851 p = 0.377 3.3 Chart review findings Based on selected chart review by three critical care physicians, 95% of LLM recommendations were consistent with the ERS/ATS 2017 guideline for NIV and the ERS 2022 guideline for HFNC. Despite high congruence with the guidelines, physicians overall agreed with the LLM recommendation for only 65% of cases. In the free form comment section provided to reviewers, reasons for disagreement with the LLM included incorrect content in the explanation which was identified in 3 cases (15%). The incorrect content was deemed clinically significant in all 3 cases by critical care physician reviewers. Reviewers also identified missing clinically important information in 6 cases (30%). Of the 11 cases with errors identified in LLM accuracy- either missing important clinical content or incorrect content in the explanation, likelihood of potential harm from these errors was determined to be low in 64% (7/11), medium in 27% (3/11), and high in 9% (1/11) of cases. Potential extent of harm from these errors was determined to be severe/death in 2 cases, mild/moderate in 5 cases and, no harm in 4 cases. LLM comprehension was excellent, exhibiting correct question comprehension in 100% of reviewed cases. In one case there was both evidence of correct and incorrect question comprehension. The LLM had evidence of correct evidence retrieval and reasoning in 19/20 cases. However, in 6/20 cases the LLM also showed evidence of incorrect retrieval and in 4/20 cases showed evidence of incorrect rationale. Table 9 Summary of Chart Review Results for LLM Recommendations (n = 20). Domain Evaluation Criteria n % Recommendation Congruence Is recommendation congruent with the guidelines? 19 95% Explanation Accuracy Is there incorrect content in the explanation? 3 15% Is the incorrect explanation clinically significant? 3 15% Is the explanation missing important clinical content? 6 30% Error Harm Assessment Harm likelihood for LLM error = Low/Med/High/N.A. Low:7; Med:3; High: 1; N.A.:9 – Harm extent from LLM error = None/no harm,Mild/Moderate,Severe/death,N.A. No harm:4; Mild/Mod:5; Severe/death:2; N.A.:9 – Clinical Judgment Overall MD agreement with LLM recommendation Yes: 13;No:7; Partially: 0 65% First MD agreement Yes: 16;No: 4; Partially: 1 80% Second MD agreement Yes: 14;No: 5; Partially: 1 70% Third MD agreement Yes: 15;No: 5; Partially: 0 75% LLM Comprehension & Reasoning Is comprehension of the question correct? 20 100% Is evidence retrieval correct? 19 95% Is reasoning for LLM recommendation correct? Yes:19; No:0; Partially: 0 - Any incorrect comprehension from LLM? 1 5% Any incorrect or missing clinical info (retrieval error or hallucination)? 6 30% Did the LLM recommendation contain incorrect rationale? 4 20% 4. Discussion In this study, we propose and evaluate a novel framework that combines a deep counterfactual inference model (RepFlow-CFR) with LLM to generate individualized, guideline-aligned treatment recommendations for patients with ARF. Our findings show that LLM-guided reinforcement of clinical guidelines improves alignment between model recommendations and real-world decisions, and that concordance with model recommendations is associated with improved patient outcomes, including reduced rates of IMV and mortality or hospice discharge. Impact of LLM-Guided Recommendations on Concordance and Outcomes By embedding clinical guidelines into the recommendation pipeline, the LLM-enhanced model yielded significantly higher concordance between suggested treatments and actual clinical actions. More importantly, patients who received treatments concordant with LLM recommendations experienced better outcomes. For instance, HFNC-concordant cases had markedly lower IMV rates compared to discordant ones, with a 97% relative increase in IMV risk when recommendations were not followed. These results demonstrate the clinical utility of guideline-informed recommendations and suggest that integrating LLMs into decision support systems can operationalize evidence-based care more effectively. Chart Review and Clinical Validity Structured chart review by critical care physicians confirmed that LLM-generated recommendations were consistent with established guidelines in 95% of cases. However, full clinical agreement with these recommendations was only observed in 65% of cases, reflecting a critical limitation of guideline-only approaches. Reviewers noted that guidelines alone do not fully capture patient complexity and cannot account for every clinical variable. Specific patient conditions such as right ventricular failure, hematemesis, or altered mental status were cited as potential contraindications to NIV that were not recognized by the LLM despite technically aligning with guideline criteria. Furthermore, several cases involved clinical ambiguity where either HFNC or NIV could be considered acceptable. Additionally, critical care physician reviewers did not agree on the same management approach to all cases, likely reflecting variability in clinical practice patterns. This factor may also pose challenges to implementing LLM-generating recommendations. Importantly, the current implementation did not emphasize training the LLM to identify contraindications to NIV or HFNC, nor did it incorporate complex patient-specific modifiers beyond guideline definitions. However, this represents a promising area for future development. Such contraindication awareness can be modularly integrated into a decision support tool to better reflect real-world clinical reasoning and improve the robustness of LLM outputs. Limitations Several limitations merit consideration. First, although LLM-guided recommendations improved interpretability and clinical alignment, hallucination and reasoning errors still occurred. In the chart review, 30% of cases involved explanation inaccuracies or omissions, including a small number (15%) with clinically significant issues. Second, while guidelines are essential, they are inherently limited in scope and may not generalize well to complex patients with multiple comorbidities. Third, the retrospective design and single-center dataset limit the external validity of our findings. Finally, the chart review sample size (n = 20) was relatively small, and although it included diverse clinical presentations, larger validation cohorts are needed. Clinical Implications and Future Work Our study demonstrates that LLMs can play a powerful role in bridging the gap between black-box machine learning models and interpretable, guideline-adherent clinical decision-making. Future work should focus on enhancing the LLM's ability to detect contraindications, reason about overlapping clinical conditions, and adapt to emerging guidelines. Integrating clinician feedback and deploying this framework prospectively will be essential for developing trustworthy, adaptive AI tools in high-acuity environments such as the ICU. Conclusions We developed and validated a hybrid framework that enhances deep counterfactual inference with LLM-based guideline enforcement to support individualized respiratory support decisions for ICU patients. The LLM-enhanced model improved treatment concordance and was associated with better patient outcomes, including lower rates of IMV and mortality/hospice discharge. While the LLM achieved high guideline adherence, discrepancies with physician judgment highlighted the need to better account for real-world clinical complexity and contraindications. Our findings support the potential of combining explainable AI with evidence-based medicine to build interpretable, high-impact decision support tools in critical care. Further refinement and prospective validation are needed to safely translate this approach into routine clinical practice. Abbreviations Abbreviation Full Term ARF Acute Respiratory Failure HFNC High-Flow Nasal Cannula NIV Noninvasive Ventilation IMV Invasive Mechanical Ventilation ICU Intensive Care Unit LLM Large Language Model ITE Individualized Treatment Effect CFR Counterfactual Regression CNF Conditional Normalizing Flow SOFA Sequential Organ Failure Assessment SIRS Systemic Inflammatory Response Syndrome CCI Charlson Comorbidity Index TSLM Time Since Last Measurement EHR Electronic Health Record AWS Amazon Web Services ED Emergency Department H&P History and Physical HPI History of Present Illness AHRQ Agency for Healthcare Research and Quality RCT Randomized Controlled Trial OR Odds Ratio Declarations Ethics approval and consent to participate Ethics approval was obtained from the University of California San Diego Institutional Review Board (UC San Diego IRB Protocol \#800258 (“VentNet: A Real-Time Multimodal Data Integration Model for Prediction of Respiratory Failure in Patients with COVID-19”) ). Consent for publication Not applicable Availability of data and materials The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request. Competing interests S.N., A.B., and A.M. are co-founders of a UCSD start-up, Clairyon Inc. (formerly Healcisio), a digital health company formed in compliance with UCSD conflict of interest policies. A.M. reports additional income from Eli Lilly, Livanova, Zoll and Powell Mansfield. ResMed provided a philanthropic donation to UCSD. The remaining authors declare no competing interests. Funding This work was supported by the National Heart, Lung, and Blood Institute (R01HL157985) and the National Library of Medicine (R01LM013998, T15LM011271). Authors' contributions Xiaolei Lu, Michael Miller, Alex Pearce, Atul Malhotra and Shamim Nemati were involved in the original conception and design of the work. Xiaolei Lu and Shamim Nemati developed the network architectures, conducted the experiments and analyzed the data. Michael Miller, Alex Pearce, Preeti Gupta, Thaidan T. Pham and Atul Malhotra provided clinical expertise and assisted with interpretation of the results. Xiaolei Lu, Michael Miller, Alex Pearce, Atul Malhotra and Shamim Nemati wrote the initial draft of the manuscript. All authors contributed feedback and approved the final manuscript. Shamim Nemati is the guarantor for the paper. Acknowledgements Not applicable References Arunachala S, Parthasarathi A, Basavaraj CK, et al. The use of high-flow nasal cannula and non-invasive mechanical ventilation in the management of COVID-19 patients: A Prospective Study. Viruses. 2023;15(9):1879. Doshi P, Whittle JS, Bublewicz M, et al. High-velocity nasal insufflation in the treatment of respiratory failure: a randomized clinical trial. Ann Emerg Med. 2018;72(1):73–83. e5. Dugan KC, Hall JB, Patel BK. High-flow nasal oxygen—the pendulum continues to swing in the assessment of critical care technology. JAMA. 2018;320(20):2083–4. Ferreyro BL, Angriman F, Munshi L, et al. Noninvasive oxygenation strategies in adult patients with acute respiratory failure: a protocol for a systematic review and network meta-analysis. Syst Reviews. 2020;9:1–6. Francio F, Weigert RM, Mattei EDB, et al. High-flow nasal oxygen vs noninvasive ventilation in patients with acute respiratory failure: the renovate randomized clinical trial. JAMA. 2025;333(10):875–90. Grieco DL, Menga LS, Cesarano M, et al. Effect of helmet noninvasive ventilation vs high-flow nasal oxygen on days free of respiratory support in patients with COVID-19 and moderate to severe hypoxemic respiratory failure: the HENIVOT randomized clinical trial. JAMA. 2021;325(17):1731–43. Linck EJG, Goligher EC, Semler MW, et al. Toward precision in critical care research: Methods for observational and interventional studies. Crit Care Med. 2024;52(9):1439–50. Munroe ES, Prevalska I, Hyer M, et al. High-Flow Nasal Cannula Versus Noninvasive Ventilation as Initial Treatment in Acute Hypoxia: A Propensity Score-Matched Study. Crit Care Explorations. 2024;6(5):e1092. Nair PR, Haritha D, Behera S, et al. Comparison of high-flow nasal cannula and noninvasive ventilation in acute hypoxemic respiratory failure due to severe COVID-19 pneumonia. Respir Care. 2021;66(12):1824–30. Oczkowski S, Ergan B, Bos L et al. ERS clinical practice guidelines: high-flow nasal cannula in acute respiratory failure. Eur Respir J 2022;59(4). Rochwerg B, Brochard L, Elliott MW et al. Official ERS/ATS clinical practice guidelines: noninvasive ventilation for acute respiratory failure. Eur Respir J 2017;50(2). Athey S, Wager S. Estimating treatment effects with causal forests: An application. Observational Stud. 2019;5(2):37–51. Shalit U, Johansson FD, Sontag D. Estimating individual treatment effect: generalization bounds and algorithms. PMLR: 2017. Künzel SR, Sekhon JS, Bickel PJ et al. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences. 2019;116(10):4156–4165. Chipman HA, George EI, McCulloch RE. BART: Bayesian additive regression trees. 2010. Lam JY, Lu X, Shashikumar SP et al. Development, deployment, and continuous monitoring of a machine learning model to predict respiratory failure in critically ill patients. JAMIA open 2024;7(4). Winkler C, Worrall D, Hoogeboom E et al. Learning likelihoods with conditional normalizing flows. arXiv preprint arXiv:191200042 2019. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–80. Anthropic. (2024). Claude 3.5 Sonnet [Large language model]. https://www.anthropic.com/news/claude-3-5-sonnet Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterials.docx Cite Share Download PDF Status: Published Journal Publication published 14 Nov, 2025 Read the published version in Critical Care → Version 1 posted Editorial decision: Revision requested 16 Sep, 2025 Reviews received at journal 24 Aug, 2025 Reviewers agreed at journal 09 Aug, 2025 Reviews received at journal 08 Aug, 2025 Reviewers agreed at journal 07 Aug, 2025 Reviewers agreed at journal 07 Aug, 2025 Reviewers invited by journal 07 Aug, 2025 Editor assigned by journal 29 Jul, 2025 Submission checks completed at journal 29 Jul, 2025 First submitted to journal 28 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7230335","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":497981817,"identity":"144ebf2b-5183-44d9-a8af-6982468a5142","order_by":0,"name":"Xiaolei Lu","email":"","orcid":"","institution":"University of California, San Diego","correspondingAuthor":false,"prefix":"","firstName":"Xiaolei","middleName":"","lastName":"Lu","suffix":""},{"id":497981818,"identity":"7ed8caea-9d52-4155-8ccb-998f8b8dafe2","order_by":1,"name":"Michael Miller","email":"","orcid":"","institution":"University of California, San Diego","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"","lastName":"Miller","suffix":""},{"id":497981819,"identity":"d53ff454-02cd-4d11-ac02-89a393e461c4","order_by":2,"name":"Alex K. Pearce","email":"","orcid":"","institution":"University of California, San Diego","correspondingAuthor":false,"prefix":"","firstName":"Alex","middleName":"K.","lastName":"Pearce","suffix":""},{"id":497981820,"identity":"d8a244b6-3c7e-4374-85c9-8ee66677dd67","order_by":3,"name":"Preeti Gupta","email":"","orcid":"","institution":"University of California, San Diego","correspondingAuthor":false,"prefix":"","firstName":"Preeti","middleName":"","lastName":"Gupta","suffix":""},{"id":497981821,"identity":"c53873ae-54e0-418c-9db7-352a9091e788","order_by":4,"name":"Thaidan T. Pham","email":"","orcid":"","institution":"University of California, San Diego","correspondingAuthor":false,"prefix":"","firstName":"Thaidan","middleName":"T.","lastName":"Pham","suffix":""},{"id":497981822,"identity":"a1bced14-71ab-4533-9985-6f7795739405","order_by":5,"name":"Atul Malhotra","email":"","orcid":"","institution":"University of California, San Diego","correspondingAuthor":false,"prefix":"","firstName":"Atul","middleName":"","lastName":"Malhotra","suffix":""},{"id":497981823,"identity":"cfc484e2-24fd-46dd-b76a-8a47f4fd154a","order_by":6,"name":"Shamim Nemati","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA90lEQVRIiWNgGAWjYPACCTkDBgbGA4wNYB6QzUZIR4KEMVAZw4GDJGhhSNxAtBb52c1HN/z8YZG+nb35weGPOxgS+/sPb2D4UHYYpxaDO8fSbvYkSOTu7DlmcODgGYbEGTfSChhnnMOjRSLH7AYPUMuGGwlALW1AF0rwGDDztuHWIj8j/9vNPwkS6Qb3n3+AaOE/Y8D8F48Whhs5bLeBtiQY3OCB2sKQY8DMiEeLwY00s9syaRKGG87kFBw42yZhDPLLwZ5z6Xgclvzs5hubOnmD48c3Pqhss5EFhtjGBz/KrHE7DA1IgMkDRKsfBaNgFIyCUYAVAABVa2JeBwG3PAAAAABJRU5ErkJggg==","orcid":"","institution":"University of California, San Diego","correspondingAuthor":true,"prefix":"","firstName":"Shamim","middleName":"","lastName":"Nemati","suffix":""}],"badges":[],"createdAt":"2025-07-28 06:23:21","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7230335/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7230335/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s13054-025-05739-3","type":"published","date":"2025-11-14T15:57:21+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":88894885,"identity":"d20eaaa0-02b7-4b36-807b-ce9809b146e1","added_by":"auto","created_at":"2025-08-12 13:04:14","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":197067,"visible":true,"origin":"","legend":"\u003cp\u003eCohort selection flowchart for early HFNC/NIV analysis.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7230335/v1/f2eef2490e476e8c38b2fd2c.png"},{"id":96104980,"identity":"80a4a820-b6c5-4e91-b878-fb0ff52d3d4c","added_by":"auto","created_at":"2025-11-17 16:05:52","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1776211,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7230335/v1/7974040e-7d01-4c3d-81f5-afa6287ba0c5.pdf"},{"id":88894892,"identity":"09db8394-9618-400f-b462-c5cead69b2ea","added_by":"auto","created_at":"2025-08-12 13:04:14","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":328173,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterials.docx","url":"https://assets-eu.researchsquare.com/files/rs-7230335/v1/830e683e9a425e4609ea5142.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Enhancing Predictive Modeling for Respiratory Support with LLM-Driven Guideline Adherence","fulltext":[{"header":"1. Background","content":"\u003cp\u003eAcute respiratory failure (ARF) is common in critically ill patients, affecting about half of all intensive care units (ICU) admissions either upon arrival or during their stay. Among patients not requiring urgent intubation and invasive mechanical ventilation (IMV), high-flow nasal cannula (HFNC) and noninvasive ventilation (NIV) are commonly used respiratory support therapies. The choice of initial respiratory support therapy can impact important patient outcomes like need for IMV or mortality, however, the optimal choice based on prior randomized controlled trials (RCTs. Patients can have multiple indications for either HFNC or NIV at the same time, confusing treatment decisions, and many patients with ARF in the real-world who receive HFNC and NIV would not fit the inclusion and exclusion criteria of prior RCTs comparing the two modalities. There is also the possibility of heterogeneity of treatment effect, in which baseline patient characteristics influence response to therapy. The uncertainty introduced by these features into the decision to initially treat a patient with ARF with either NIV or HFNC highlights the need for data-driven precision medicine techniques to assist with choice of NIV or HFNC.\u003c/p\u003e\u003cp\u003eMachine learning techniques can be employed to estimate the individualized treatment effect (ITE) of high-flow nasal cannula (HFNC) versus noninvasive ventilation (NIV) for each patient \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Recent advancements in causal inference and counterfactual modeling have produced a diverse toolkit for ITE estimation in observational data. These tools include tree-based methods (e.g., Causal Forests\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e ), representation learning approaches (e.g., TARNet\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e), meta-learners (e.g., X-learner\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e), and balancing techniques such as Bayesian Additive Regression Trees\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. Building upon this foundation, we previously developed and validated RepFlow-CFR, a deep counterfactual representation model tailored to the critical care setting. RepFlow-CFR is designed to estimate patient-specific outcomes under both HFNC and NIV by learning balanced representations that adjust for confounding and capture latent patient heterogeneity. The model combines shared representation learning, normalizing flows for outcome modeling, and a second-stage adjustment for unmeasured confounding, enabling it to generate robust, individualized predictions for response to NIV versus HFNC to decrease risk of invasive mechanical ventilation. When applied to ICU patients identified as high-risk for respiratory failure, patients receiving treatments concordant with RepFlow-CFR recommendations showed lower rates of IMV and mortality or hospice discharge, highlighting the model\u0026rsquo;s ability to identify beneficial therapy pathways in complex ICU populations.\u003c/p\u003e\u003cp\u003eWhile RepFlow-CFR demonstrated strong performance in generating individualized treatment recommendations and identifying beneficial therapy pathways, it shares a common limitation with many deep learning models, namely, a lack of transparency in its decision-making process. This \u0026ldquo;black box\u0026rdquo; nature makes it difficult for clinicians to interpret or trust the model\u0026rsquo;s output, especially when recommendations diverge from established clinical guidelines. Moreover, RepFlow-CFR is not inherently constrained to follow these guidelines, which may lead to discordance between data-driven recommendations and evidence-based best practices. To address these challenges, we introduce a large language model (LLM) into the decision pipeline to serve as a guideline-aware, explainable reasoning layer. By leveraging structured patient data, clinical notes, and formal guideline criteria, the LLM can validate or refine the model\u0026rsquo;s recommendations, ensuring that treatment suggestions align with current standards of care while also providing interpretable justifications. In this study, we aim to refine treatment recommendations from a deep counterfactual model by integrating LLM to enforce adherence to clinical guidelines and enhance interpretability. We further evaluate the impact of this LLM-augmented framework on concordance with real-world treatment decisions and patient outcomes, with additional validation through structured expert chart review. We hypothesized that integrating a large language model with a deep counterfactual model to refine treatment recommendations would enhance interpretability, result in treatment suggestions that better align with established guideline driven standard of care.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1 RepFlow-CFR Prediction\u003c/h2\u003e\u003cp\u003e\u003cb\u003eData Sources.\u003c/b\u003e We used de-identified structured Electronic Health Record (EHR) data from ICU patients at UC San Diego Health (UCSD) between January 1, 2016, and December 31, 2023, focusing on encounters where either HFNC or NIV was administered as the first respiratory support following Vent.io-predicted\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e high-risk timepoint (Vent.io T0). Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates the cohort derivation process. Extracted variables included 50 vital signs and laboratory measurements (resampled into hourly bins), 6 demographic features, 12 SIRS/SOFA criteria, 12 medication categories, and 62 comorbidities. Additional derived features included baseline values, local trends, and time since last measurement (TSLM). Missing values were forward-filled up to 24 hours or mean-imputed (Supplementary Section 1).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eRepFlow-CFR Model.\u003c/b\u003e The RepFlow-CFR model is a deep counterfactual inference framework that adjusts for confounding using both observed and inferred latent variables. It consists of three stages (Supplementary Section 2 for the model architecture and mathematical formulation): \u003cb\u003eStage 0\u003c/b\u003e applies counterfactual regression (CFR)\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e to learn shared representations balancing measured confounders across treatment groups using an integral probability metric (e.g., Wasserstein distance); \u003cb\u003eStage 1\u003c/b\u003e uses a conditional normalizing flow (CNF)\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e to model the outcome distribution given the representation and treatments; \u003cb\u003eStage 2\u003c/b\u003e introduces a second CNF to adjust for unmeasured confounding by transforming the treatment-dependent latent variable into an interventional distribution. During inference, RepFlow-CFR generates potential outcomes under both HFNC and NIV by sampling from the learned distributions. The ITE is defined as the difference in predicted probabilities of IMV under NIV versus HFNC. Based on the ITE, the model provides treatment recommendations: NIV preferred (ITE \u0026lt; -0.001), HFNC preferred (ITE\u0026thinsp;\u0026gt;\u0026thinsp;0.001), or Indifferent (ITE between \u0026minus;\u0026thinsp;0.001 and 0.001).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2 LLM-Driven Guideline Enforcement\u003c/h2\u003e\u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\u003ch2\u003e2.2.1 Claude Sonnet Configuration and Deployment\u003c/h2\u003e\u003cp\u003eWe utilized Claude 3.5 Sonnet\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, a large language model with a 200,000-token context window, optimized for high-accuracy reasoning and structured clinical analysis. The model was deployed within a HIPAA-compliant Amazon Web Services (AWS) environment using an EC2 instance configured with authenticated access to Amazon Bedrock. Inference was performed via the Bedrock Converse API, which supports structured JSON output for reliable parsing. To minimize variability and ensure consistent, deterministic responses, the model operated with a temperature setting of 0.1. Claude Sonnet was used to evaluate whether RepFlow-CFR recommendations were aligned with clinical guidelines and to independently generate treatment recommendations based on structured patient data, clinical notes and formal guideline criteria.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section3\"\u003e\u003ch2\u003e2.2.2 Clinical Guidelines Criteria\u003c/h2\u003e\u003cp\u003eWe summarize the recommended indications for NIV and HFNC based on established clinical guidelines. Specifically, we refer to the ERS/ATS 2017 Guidelines for NIV\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e and the ERS 2022 Guidelines for HFNC\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. The guideline criteria are outlined in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e below, highlighting when each therapy is recommended, conditionally recommended, or not advised.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eClinical Indications for NIV and HFNC According to ERS/ATS Guidelines.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNIV \u0026ndash; ERS/ATS 2017 Guidelines\u003c/p\u003e\u003cp\u003eRecommend NIV if any of the following are true:\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHFNC \u0026ndash; ERS 2022 Guidelines\u003c/p\u003e\u003cp\u003eRecommend HFNC if any of the following are true:\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; COPD exacerbation with acute or acute-on-chronic respiratory acidosis (pH\u0026thinsp;\u0026le;\u0026thinsp;7.35, PaCO₂ \u0026gt;45 mmHg)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026ndash; De novo (or acute) hypoxemic respiratory failure (e.g., AHRF, ARDS, pneumonia, COVID-19)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; Acute respiratory failure due to cardiogenic pulmonary edema, including acute (or \u0026ldquo;flash\u0026rdquo;) pulmonary edema\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026ndash; Post-operative post-extubation in high-risk patients: HFNC is preferred over conventional oxygen but NIV is also an acceptable prophylactic treatment for high-risk patients\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; Neuromuscular disease or obesity hypoventilation syndrome (OHS) with acute or acute-on-chronic respiratory acidosis (pH\u0026thinsp;\u0026le;\u0026thinsp;7.35, PaCO₂ \u0026gt;45 mmHg)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026ndash; The patient is intolerant to NIV and moderate to severe acute respiratory failure is present\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; Use NIV prophylactically post-extubation in high-risk patients (e.g., those with COPD or CHF) but not in low-risk patients or in patients with established post-extubation respiratory failure\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u0026ndash; Moderate to severe acute respiratory failure is present and there is not an indication for treatment with NIV\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; Immunocompromised patients with mild-to-moderate acute respiratory failure (conditional recommendation)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eDo not recommend HFNC as first-line if:\u003c/p\u003e\u003cp\u003e\u0026ndash; The patient has acute hypercapnic respiratory failure (e.g., COPD with acidosis) unless NIV is contraindicated or not tolerated\u003c/p\u003e\u003cp\u003e\u0026ndash; There is clear guideline indication for NIV\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; Post-operative ARF (Conditional recommendation, moderate certainty of evidence)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u0026ndash; Chest trauma patients with ARF (Conditional recommendation, moderate certainty of evidence)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section3\"\u003e\u003ch2\u003e2.2.3 LLM-Assisted Guideline Concordance Assessment\u003c/h2\u003e\u003cp\u003e\u003cstrong\u003eInputs to LLM\u003c/strong\u003e\u003cp\u003eTo support explainable and guideline-aligned treatment recommendations, we provided LLM with a structured input set that includes clinical guideline knowledge, model predictions, and patient-specific data. These inputs allowed the LLM to generate context-aware treatment rationales and assess alignment between clinical recommendations and actual care. We summarize the input components in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e below.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eInputs Provided to the LLM.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eInput Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSummarized Clinical Guidelines\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGuideline-based indications and contraindications for NIV and HFNC (from Section \u003cspan refid=\"Sec5\" class=\"InternalRef\"\u003e2.2.1\u003c/span\u003e), enabling the LLM to reason using ERS/ATS 2017 and ERS 2022 recommendations.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eVent.io T0\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTimestamp indicating when the patient was first identified as high-risk for respiratory failure by the Vent.io early warning model.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRepFlow-CFR Model Output at T0\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIncludes (1) the recommended treatment modality (NIV, HFNC, or Indifferent), and (2) the model rationale based on EHR-derived features. The rationale includes the top 50 most influential SHAP-ranked features impacting the decision.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRecent Clinical Parameters (pre-T0)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMost recent values prior to T0 for: pH, PaCO₂, respiratory rate, SpO₂, and FiO₂.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRelevant Clinical Notes (\u0026le;\u0026thinsp;72h pre-T0)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFree-text notes extracted from the EHR within 72 hours before T0, including: ED Provider Notes, ED MD Progress Notes, Consults, ED OBS HPI, Progress Notes, H\u0026amp;P, Event/Update entries, ED OBS Progress Notes, and Radiology Reports.\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003ePrompting Strategy\u003c/b\u003e: The LLM was prompted to: (1) evaluate the concordance between the RepFlow-CFR model\u0026rsquo;s recommendation and guideline-based recommendations derived from patient data; (2) independently provide a final recommendation including NIV, HFNC, or Indifferent (if either option is acceptable); and (3) generate a rationale by explicitly citing relevant clinical guideline statements to support the recommendation. (Supplementary Section 3 for prompt details)\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eOutput of the LLM\u003c/strong\u003e\u003cp\u003eThe LLM was prompted to return a structured JSON object as listed in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eStructured output returned by the LLM.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eKey\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eValue / Format\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e\u0026ldquo;NIV_recommendation\u0026rdquo;\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e{\u003c/p\u003e\u003cp\u003e\u0026ldquo;recommendation\u0026rdquo;: \u0026ldquo;Yes\u0026rdquo; or \u0026ldquo;No\u0026rdquo; or \u0026ldquo;Either\u0026rdquo;,\u003c/p\u003e\u003cp\u003e\u0026ldquo;confidence\u0026rdquo;: \u0026ldquo;high\u0026rdquo; / \u0026ldquo;medium\u0026rdquo; / \u0026ldquo;low\u0026rdquo;,\u003c/p\u003e\u003cp\u003e\u0026ldquo;explanation\u0026rdquo;: \u0026ldquo;Brief rationale including guideline\u003c/p\u003e\u003cp\u003eReference\u0026rdquo;\u003c/p\u003e\u003cp\u003e}\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e\u0026ldquo;HFNC_recommendation\u0026rdquo;\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e{\u003c/p\u003e\u003cp\u003e\u0026ldquo;recommendation\u0026rdquo;: \u0026ldquo;Yes\u0026rdquo; or \u0026ldquo;No\u0026rdquo; or \u0026ldquo;Either\u0026rdquo;,\u003c/p\u003e\u003cp\u003e\u0026ldquo;confidence\u0026rdquo;: \u0026ldquo;high\u0026rdquo; / \u0026ldquo;medium\u0026rdquo; / \u0026ldquo;low\u0026rdquo;,\u003c/p\u003e\u003cp\u003e\u0026ldquo;explanation\u0026rdquo;: \u0026ldquo;Brief rationale including guideline\u003c/p\u003e\u003cp\u003eReference\u0026rdquo;\u003c/p\u003e\u003cp\u003e}\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003e\u0026ldquo;Model_alignment\u0026rdquo;\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e{\u003c/p\u003e\u003cp\u003e\u0026ldquo;alignment\u0026rdquo;: true / false,\u003c/p\u003e\u003cp\u003e\u0026ldquo;explanation\u0026rdquo;: \u0026ldquo;Was the RepFlow-CFR model\u003c/p\u003e\u003cp\u003ealigned with guideline-based decision?\u0026rdquo;\u003c/p\u003e\u003cp\u003e}\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e2.3 Concordance Analysis\u003c/h2\u003e\u003cp\u003eWe assessed the concordance between model-generated treatment recommendations and actual clinical decisions at two levels. First, RepFlow-CFR Concordance was defined as agreement between the treatment modality recommended by the RepFlow-CFR model: NIV, HFNC, or Indifferent, and the treatment the patient actually received immediately following the Vent.io T0 timepoint. Second, LLM-Enhanced Concordance was defined as agreement between the final treatment recommendation generated by the LLM, which integrated clinical guidelines and patient-specific data, and the actual treatment administered. For both types of recommendations, we compared key clinical outcomes between concordant and discordant groups, specifically focusing on the incidence of IMV and a composite endpoint of in-hospital mortality or discharge to hospice (Mortality/Hospice).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e2.4 Chart Review Validation\u003c/h2\u003e\u003cp\u003eTo assess the clinical validity and safety of the LLM-enhanced recommendation framework, we conducted a retrospective chart review of 20 patient cases. These cases were selected to represent a variety of clinical scenarios, including variations in disease severity, instances where the LLM and the RepFlow-CFR model generated discordant treatment recommendations, and cases with diverse real-world treatment pathways. This purposive sampling strategy was designed to challenge the LLM's interpretive and decision-making capacity across a spectrum of relevant clinical contexts.\u003c/p\u003e\u003cp\u003eEach case was independently reviewed by three board-certified critical care physicians. Reviewers evaluated the LLM's treatment recommendation for congruence with established clinical guidelines, specifically the ERS/ATS 2017 guideline for NIV and the ERS 2022 guideline for HFNC. In addition to assessing whether the LLM\u0026rsquo;s final recommendation aligned with guideline-based care, the reviewers also examined the explanation provided by the LLM, with particular attention to the presence of factual errors, clinically significant omissions, or hallucinated statements.\u003c/p\u003e\u003cp\u003eTo evaluate the LLM\u0026rsquo;s comprehension, knowledge retrieval, and reasoning ability, a structured evaluation framework was used, adapted from prior work in validating medical LLM outputs\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. Given the known tendency of LLMs to mix correct and incorrect content, the review framework was designed to explicitly capture both accurate and inaccurate components of the model's response. To assess the potential for harm, the reviewers used the harm classification framework from the Agency for Healthcare Research and Quality (AHRQ) Common Formats.\u003c/p\u003e\u003cp\u003eThe structured evaluation criteria used by physician reviewers are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. These include assessments of recommendation congruence, explanation accuracy, potential for harm, clinical agreement, and LLM-specific reasoning and retrieval performance. Each criterion was accompanied by categorical evaluation options and a space for free-text reviewer comments.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eStructured Evaluation Criteria for LLM Chart Review.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDomain\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEvaluation Criteria\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eOptions / Scale\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eReviewer comments\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRecommendation Congruence\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs recommendation congruent with guideline?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003e\u003cb\u003eExplanation Accuracy\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs there incorrect content in the explanation?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs the incorrect explanation clinically significant?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs the explanation missing important clinical content?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eError Harm Assessment\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHarm Likelihood for LLM error\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLow / Medium / High / N.A.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eN.A. was put in for the cases where there was no LLM error\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHarm Extent from LLM error\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNone/no harm, Mild/Moderate, Severe/death, NA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eN.A. was put in for the cases where there was no LLM error\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eClinical Judgment\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eOverall MD agreement with LLM recommendation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAgree / Disagree / Partially Agree\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSecond MD agreement (if any)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAgree / Disagree / Partially Agree\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e\u003cp\u003e\u003cb\u003eLLM Comprehension \u0026amp; Reasoning\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs comprehension of the question correct?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs evidence retrieval correct?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs reasoning for LLM recommendation correct?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No / Partially\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAny incorrect comprehension from LLM?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAny incorrect or missing clinical info (retrieval error or hallucination)?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDid the LLM recommendation contain incorrect rationale?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes / No\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.1 Patient Characteristics, Interventions, and Outcomes by RepFlow-CFR and LLM-Enhanced Recommendation\u003c/h2\u003e\u003cp\u003eOur study cohort included 1,261 ICU encounters, stratified by treatment recommendations generated by the RepFlow-CFR and LLM-enhanced models. Based on RepFlow-CFR recommendations, 859 patients were classified as NIV-preferred, 279 as HFNC-preferred, and 123 as indifferent. The LLM-enhanced recommendations reassigned patients into 759 NIV-preferred, 205 HFNC-preferred, and 297 indifferent. Across both models, baseline characteristics were generally balanced. The mean age across groups ranged from 60 to 63 years, and the proportion of male patients varied between 52.7% and 63.4%. Comorbidity burden, as measured by Charlson Comorbidity Index, was comparable across groups (median: 1.0\u0026ndash;3.0), and the distribution of specific comorbidities such as congestive heart failure and COPD showed no significant differences. Similarly, SOFA scores at Vent.io T0 were consistent across recommendations.\u003c/p\u003e\u003cp\u003ePost-T0 interventions, including the use of steroids, antibiotics, vasopressors, and diuretics, did not significantly differ across groups. Notably, among patients recommended HFNC by the LLM-enhanced model, 91.7% actually received HFNC, compared to 70.3% in the corresponding RepFlow-CFR group, indicating improved alignment between recommendation and practice. Clinical outcomes, including IMV, mortality, and hospice enrollment, remained similar across all recommendation groups. IMV occurred in approximately 23\u0026ndash;29% of patients, and mortality ranged from 24.4\u0026ndash;41.8%, with no statistically significant differences observed. These findings suggest that both models stratify patients into demographically and clinically comparable groups, with the LLM-enhanced model demonstrating improved adherence to recommended respiratory support strategies.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eBaseline characteristics of patients by RepFlow-CFR and LLM-enhanced recommendations.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eVariable\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"4\" nameend=\"c5\" namest=\"c2\"\u003e\u003cp\u003eRepFlow-CFR recommendation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"4\" nameend=\"c9\" namest=\"c6\"\u003e\u003cp\u003eLLM-enhanced recommendation\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNIV preferred\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHFNC\u003c/p\u003e\u003cp\u003epreferred\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eIndifferent\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u003csup\u003e\u003cem\u003eb\u003c/em\u003e\u003c/sup\u003e value\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eNIV\u003c/p\u003e\u003cp\u003epreferred\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eHFNC\u003c/p\u003e\u003cp\u003epreferred\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eIndifferent\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u003c/p\u003e\u003cp\u003evalue\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003eCharacteristic\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEncounters, N\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e859\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e279\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e123\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e759\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e205\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e297\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAge(years), mean (SD)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e62(16.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e62(16.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e63(16.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.580\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e63(16.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e60(18.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e60(15.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.580\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGender, N (%)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e512(59.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e147(52.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e78(63.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e442(58.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e118(57.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e177(59.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003eOrgan dysfunction\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eCharlson Comorbidity Index, Median (IQR)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e2.0(1.0\u0026ndash;5.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2.0(1.0\u0026ndash;4.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3.0(1.0\u0026ndash;5.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.361\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3.0(1.0\u0026ndash;5.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e1.0(1.0\u0026ndash;3.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e2.0(2.0\u0026ndash;6.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.361\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCongestive Heart Failure Component, N (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e208(24.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e73(26.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e31(25.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.800\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e261(34.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e18(8.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e33(11.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.800\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChronic Obstructive Pulmonary Disease, N (%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e162(18.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e52(18.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e23(18.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.996\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e176(23.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e28(13.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e33(11.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.996\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSOFA score (at Vent.io T0), Median (IQR)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1.0(1.0\u0026ndash;3.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1.0(0.0\u0026ndash;4.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.0(0.0\u0026ndash;3.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.512\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2.0(0.0-3.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.0(0.0\u0026ndash;2.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e1.0(0.0\u0026ndash;3.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.512\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003eInterventions after Vent.io T0\u003c/b\u003e\u003csup\u003e\u003cb\u003ea\u003c/b\u003e\u003c/sup\u003e \u003cb\u003eN (%)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSteroids administration\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e289(33.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e80(28.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e42(34.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.129\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e225(29.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e65(31.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e121(40.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.129\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAntibiotics administration\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e717(83.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e225(80.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e106(86.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.154\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e627(82.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e154(75.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e267(89.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.154\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eVasopressors administration\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e263(30.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e90(32.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e31(25.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.306\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e276(36.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e39(19.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e69(23.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.306\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDiuretics administration\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e512(59.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e159(57.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e78(63.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.424\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e462(60.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e118(57.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e169(56.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.424\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003eInitial respiratory support after Vent.io T0 (N %)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHFNC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e634(73.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e196(70.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e93(75.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.414\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e485(63.9)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e188(91.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e250(84.2)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003e\u003cb\u003eOutcomes (N %)\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eIMV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e245(28.5)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e65(23.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e29(23.6)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.159\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e197(26.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e55(26.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e87(29.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.159\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMortality\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e292(34.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e93(33.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e42(34.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.977\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e253(33.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e50(24.4)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e122(41.8)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.977\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHospice\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e9(1.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2(0.7)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0(0.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e.480\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e8(1.1)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2(1.0)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e1(0.3)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e.480\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"8\" nameend=\"c8\" namest=\"c1\"\u003e\u003cp\u003e\u003csup\u003ea\u003c/sup\u003e We focused on interventions administered from Vent.io T0 to one hour before ICU discharge for the control group (those not intubated), and from Vent.io T0 to one hour before the time of intubation for the positive group (those who require intubation).\u003c/p\u003e\u003cp\u003e\u003csup\u003eb\u003c/sup\u003e Testing for the difference \u003cem\u003eP\u003c/em\u003e value were χ2 for categorical variables and Kruskal-Wallis rank-sum test for continuous variables.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Concordance Analysis\u003c/h2\u003e\u003cp\u003eTo evaluate the clinical utility of model-guided recommendations, we analyzed the concordance between predicted treatment recommendations and the actual treatments received by patients, stratified by outcome. We compared both the baseline RepFlow-CFR model and the LLM-enhanced framework across two major clinical outcomes: the need for IMV and Mortality \u0026amp; Hospice.\u003c/p\u003e\u003cp\u003eTables below present the rate of IMV and mortality/hospice in patients who received treatments that were concordant versus discordant with the model recommendations. Relative reduction (or increase) in outcome rates is reported to capture the direction and magnitude of benefit or risk associated with following model-predicted guidance. Across both models, patients who received treatments concordant with model recommendations had lower IMV rates. The LLM-enhanced model demonstrated greater relative reductions in IMV rates in the concordant groups, particularly under HFNC recommendations, where discordant treatment was associated with a 97% relative increase in IMV risk. For the mortality and hospice outcome, the LLM-enhanced model again showed larger differences between concordant and discordant groups. Under the LLM framework, NIV-concordant cases showed a 14.03% relative reduction in mortality/hospice, compared to only 6.83% under the baseline model.\u003c/p\u003e\u003cp\u003eWe further assessed the association between concordant treatment and outcomes using multivariable logistic regression, controlling for potential confounders including age, gender, CCI, SOFA score, and Vent.io risk score. The results are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. Concordance with the LLM-enhanced recommendation was significantly associated with reduced risk of both IMV and mortality/hospice. In particular, HFNC concordance under the LLM framework was associated with a significantly lower odds of mortality or hospice discharge (OR\u0026thinsp;=\u0026thinsp;0.670, p\u0026thinsp;=\u0026thinsp;0.046), suggesting a clinically meaningful improvement in outcome alignment when recommendations followed the LLM-augmented guidance.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eIMV Rates Stratified by Model Recommendation and Concordance.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRecommendation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTotal\u003c/p\u003e\u003cp\u003eIMV\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eConcordant\u003c/p\u003e\u003cp\u003eIMV\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiscordant\u003c/p\u003e\u003cp\u003eIMV\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRelative reduction if concordant\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eRelative increase if discordant\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eRepFlow-CFR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eNIV\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e28.28\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e22.73\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e30.10\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e19.64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e6.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eHFNC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e27.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e25.00\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e33.75\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e10.00\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e21.50\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eLLM-enhanced\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eNIV\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e25.96\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e21.17\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e28.65\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e18.44\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e10.42\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eHFNC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e26.83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e24.47\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e52.94\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e8.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e97.33\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eMortality \u0026amp; Hospice Rates Stratified by Recommendation and Concordance.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRecommendation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTotal\u003c/p\u003e\u003cp\u003eMortality \u0026amp; Hospice\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eConcordant\u003c/p\u003e\u003cp\u003eMortality \u0026amp; Hospice\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eDiscordant\u003c/p\u003e\u003cp\u003eMortality \u0026amp; Hospice\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRelative reduction if concordant\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eRelative increase if discordant\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eRepFlow-CFR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eNIV\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e36.59\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e34.09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e37.41\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e6.83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e2.24\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eHFNC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e29.76\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e30.23\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e28.75\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e-1.58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e-3.40\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eLLM-enhanced\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eNIV\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e34.39\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e29.56\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e37.11\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e14.03\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e7.93\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eHFNC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e25.37\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e25.00\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e29.41\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.44\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e15.95\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eMultivariable Logistic Regression Results (Odds Ratios and p-values) for Predicting the Need for IMV and Mortality \u0026amp; Hospice Across Methods and Sites.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"9\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eOutcome\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNIV\u003c/p\u003e\u003cp\u003econcordance\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eHFNC\u003c/p\u003e\u003cp\u003econcordance\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eGender\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eCCI score\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eSOFA score\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c9\"\u003e\u003cp\u003eVent.io\u003c/p\u003e\u003cp\u003escore\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eIMV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRepFlow-CFR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.678\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.032\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.729\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.108\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.984\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.021\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.877\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.913\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e1.063\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.051\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.836\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.284\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM-enhanced\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.664\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.023\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.679\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.057\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.983\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.068\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.662\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.904\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e1.053\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.145\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.946\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.766\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eMortality \u0026amp; Hospice\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRepFlow-CFR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.822\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.249\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.702\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.066\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.015\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.732\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.018\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.959\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.070\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e1.312\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.908\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.547\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLLM-enhanced\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.716\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.048\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.670\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.046\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.019\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.685\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.009\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.929\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.005\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e1.272\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.000\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.851\u003c/p\u003e\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.377\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Chart review findings\u003c/h2\u003e\u003cp\u003eBased on selected chart review by three critical care physicians, 95% of LLM recommendations were consistent with the ERS/ATS 2017 guideline for NIV and the ERS 2022 guideline for HFNC. Despite high congruence with the guidelines, physicians overall agreed with the LLM recommendation for only 65% of cases. In the free form comment section provided to reviewers, reasons for disagreement with the LLM included incorrect content in the explanation which was identified in 3 cases (15%). The incorrect content was deemed clinically significant in all 3 cases by critical care physician reviewers. Reviewers also identified missing clinically important information in 6 cases (30%). Of the 11 cases with errors identified in LLM accuracy- either missing important clinical content or incorrect content in the explanation, likelihood of potential harm from these errors was determined to be low in 64% (7/11), medium in 27% (3/11), and high in 9% (1/11) of cases. Potential extent of harm from these errors was determined to be severe/death in 2 cases, mild/moderate in 5 cases and, no harm in 4 cases.\u003c/p\u003e\u003cp\u003eLLM comprehension was excellent, exhibiting correct question comprehension in 100% of reviewed cases. In one case there was both evidence of correct and incorrect question comprehension. The LLM had evidence of correct evidence retrieval and reasoning in 19/20 cases. However, in 6/20 cases the LLM also showed evidence of incorrect retrieval and in 4/20 cases showed evidence of incorrect rationale.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSummary of Chart Review Results for LLM Recommendations (n\u0026thinsp;=\u0026thinsp;20).\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDomain\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEvaluation Criteria\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003en\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003e%\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRecommendation Congruence\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs recommendation congruent with the guidelines?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e19\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e95%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003e\u003cb\u003eExplanation Accuracy\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs there incorrect content in the explanation?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e15%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs the incorrect explanation clinically significant?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e15%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs the explanation missing important clinical content?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e30%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eError Harm Assessment\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHarm likelihood for LLM error\u0026thinsp;=\u0026thinsp;Low/Med/High/N.A.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLow:7; Med:3; High: 1; N.A.:9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u0026ndash;\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHarm extent from LLM error\u0026thinsp;=\u0026thinsp;None/no harm,Mild/Moderate,Severe/death,N.A.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNo harm:4; Mild/Mod:5;\u003c/p\u003e\u003cp\u003eSevere/death:2;\u003c/p\u003e\u003cp\u003eN.A.:9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u0026ndash;\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003e\u003cb\u003eClinical Judgment\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eOverall MD agreement with LLM recommendation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes: 13;No:7; Partially: 0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e65%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFirst MD agreement\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes: 16;No: 4; Partially: 1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e80%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSecond MD agreement\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes: 14;No: 5; Partially: 1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e70%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eThird MD agreement\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes: 15;No: 5; Partially: 0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e75%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e\u003cp\u003e\u003cb\u003eLLM Comprehension \u0026amp; Reasoning\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs comprehension of the question correct?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e20\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e100%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs evidence retrieval correct?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e19\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e95%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eIs reasoning for LLM recommendation correct?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eYes:19; No:0; Partially: 0\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e-\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAny incorrect comprehension from LLM?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e5%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAny incorrect or missing clinical info (retrieval error or hallucination)?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e30%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDid the LLM recommendation contain incorrect rationale?\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e20%\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eIn this study, we propose and evaluate a novel framework that combines a deep counterfactual inference model (RepFlow-CFR) with LLM to generate individualized, guideline-aligned treatment recommendations for patients with ARF. Our findings show that LLM-guided reinforcement of clinical guidelines improves alignment between model recommendations and real-world decisions, and that concordance with model recommendations is associated with improved patient outcomes, including reduced rates of IMV and mortality or hospice discharge.\u003c/p\u003e\u003cp\u003e\u003cb\u003eImpact of LLM-Guided Recommendations on Concordance and Outcomes\u003c/b\u003e\u003c/p\u003e\u003cp\u003eBy embedding clinical guidelines into the recommendation pipeline, the LLM-enhanced model yielded significantly higher concordance between suggested treatments and actual clinical actions. More importantly, patients who received treatments concordant with LLM recommendations experienced better outcomes. For instance, HFNC-concordant cases had markedly lower IMV rates compared to discordant ones, with a 97% relative increase in IMV risk when recommendations were not followed. These results demonstrate the clinical utility of guideline-informed recommendations and suggest that integrating LLMs into decision support systems can operationalize evidence-based care more effectively.\u003c/p\u003e\u003cp\u003e\u003cb\u003eChart Review and Clinical Validity\u003c/b\u003e\u003c/p\u003e\u003cp\u003eStructured chart review by critical care physicians confirmed that LLM-generated recommendations were consistent with established guidelines in 95% of cases. However, full clinical agreement with these recommendations was only observed in 65% of cases, reflecting a critical limitation of guideline-only approaches. Reviewers noted that guidelines alone do not fully capture patient complexity and cannot account for every clinical variable. Specific patient conditions such as right ventricular failure, hematemesis, or altered mental status were cited as potential contraindications to NIV that were not recognized by the LLM despite technically aligning with guideline criteria. Furthermore, several cases involved clinical ambiguity where either HFNC or NIV could be considered acceptable. Additionally, critical care physician reviewers did not agree on the same management approach to all cases, likely reflecting variability in clinical practice patterns. This factor may also pose challenges to implementing LLM-generating recommendations.\u003c/p\u003e\u003cp\u003eImportantly, the current implementation did not emphasize training the LLM to identify contraindications to NIV or HFNC, nor did it incorporate complex patient-specific modifiers beyond guideline definitions. However, this represents a promising area for future development. Such contraindication awareness can be modularly integrated into a decision support tool to better reflect real-world clinical reasoning and improve the robustness of LLM outputs.\u003c/p\u003e\u003cp\u003e\u003cb\u003eLimitations\u003c/b\u003e\u003c/p\u003e\u003cp\u003eSeveral limitations merit consideration. First, although LLM-guided recommendations improved interpretability and clinical alignment, hallucination and reasoning errors still occurred. In the chart review, 30% of cases involved explanation inaccuracies or omissions, including a small number (15%) with clinically significant issues. Second, while guidelines are essential, they are inherently limited in scope and may not generalize well to complex patients with multiple comorbidities. Third, the retrospective design and single-center dataset limit the external validity of our findings. Finally, the chart review sample size (n\u0026thinsp;=\u0026thinsp;20) was relatively small, and although it included diverse clinical presentations, larger validation cohorts are needed.\u003c/p\u003e\u003cp\u003e\u003cb\u003eClinical Implications and Future Work\u003c/b\u003e\u003c/p\u003e\u003cp\u003eOur study demonstrates that LLMs can play a powerful role in bridging the gap between black-box machine learning models and interpretable, guideline-adherent clinical decision-making. Future work should focus on enhancing the LLM's ability to detect contraindications, reason about overlapping clinical conditions, and adapt to emerging guidelines. Integrating clinician feedback and deploying this framework prospectively will be essential for developing trustworthy, adaptive AI tools in high-acuity environments such as the ICU.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eWe developed and validated a hybrid framework that enhances deep counterfactual inference with LLM-based guideline enforcement to support individualized respiratory support decisions for ICU patients. The LLM-enhanced model improved treatment concordance and was associated with better patient outcomes, including lower rates of IMV and mortality/hospice discharge. While the LLM achieved high guideline adherence, discrepancies with physician judgment highlighted the need to better account for real-world clinical complexity and contraindications. Our findings support the potential of combining explainable AI with evidence-based medicine to build interpretable, high-impact decision support tools in critical care. Further refinement and prospective validation are needed to safely translate this approach into routine clinical practice.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eAbbreviation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003e\u003cstrong\u003eFull Term\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eARF\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAcute Respiratory Failure\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHFNC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHigh-Flow Nasal Cannula\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eNIV\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eNoninvasive Ventilation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eIMV\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eInvasive Mechanical Ventilation\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eIntensive Care Unit\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eLLM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eLarge Language Model\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eITE\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eIndividualized Treatment Effect\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCFR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCounterfactual Regression\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCNF\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eConditional Normalizing Flow\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eSOFA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSequential Organ Failure Assessment\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eSIRS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eSystemic Inflammatory Response Syndrome\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eCCI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eCharlson Comorbidity Index\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eTSLM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eTime Since Last Measurement\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eEHR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eElectronic Health Record\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eAWS\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAmazon Web Services\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eED\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eEmergency Department\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eH\u0026amp;P\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHistory and Physical\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eHPI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eHistory of Present Illness\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eAHRQ\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eAgency for Healthcare Research and Quality\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eRCT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eRandomized Controlled Trial\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\n \u003cp\u003eOR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd\u003e\n \u003cp\u003eOdds Ratio\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"},{"header":"Declarations","content":"\u003cul\u003e\n \u003cli\u003eEthics approval and consent to participate\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eEthics approval was obtained from the University of California San Diego Institutional Review Board (UC San Diego IRB Protocol \\#800258 (“VentNet: A Real-Time Multimodal Data Integration Model for Prediction of Respiratory Failure in Patients with COVID-19”) ).\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eConsent for publication\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eAvailability of data and materials\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThe datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eCompeting interests\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eS.N., A.B., and A.M. are co-founders of a UCSD start-up, Clairyon Inc. (formerly Healcisio), a digital health company formed in compliance with UCSD conflict of interest policies. A.M. reports additional income from Eli Lilly, Livanova, Zoll and Powell Mansfield. ResMed provided a philanthropic donation to UCSD. The remaining authors declare no competing interests.\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eFunding\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThis work was supported by the National Heart, Lung, and Blood Institute (R01HL157985) and the National Library of Medicine (R01LM013998, T15LM011271).\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eAuthors' contributions\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eXiaolei Lu, Michael Miller, Alex Pearce, Atul Malhotra and Shamim Nemati were involved in the original conception and design of the work. Xiaolei Lu and Shamim Nemati developed the network architectures, conducted the experiments and analyzed the data. Michael Miller, Alex Pearce, Preeti Gupta, Thaidan T. Pham \u0026nbsp;and\u0026nbsp;Atul Malhotra provided clinical expertise and assisted with interpretation of the results. Xiaolei Lu, Michael Miller, Alex Pearce, Atul Malhotra and Shamim Nemati wrote the initial draft of the manuscript. All authors contributed feedback and approved the final manuscript. Shamim Nemati is the guarantor for the paper.\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eAcknowledgements\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eArunachala S, Parthasarathi A, Basavaraj CK, et al. The use of high-flow nasal cannula and non-invasive mechanical ventilation in the management of COVID-19 patients: A Prospective Study. Viruses. 2023;15(9):1879.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDoshi P, Whittle JS, Bublewicz M, et al. High-velocity nasal insufflation in the treatment of respiratory failure: a randomized clinical trial. Ann Emerg Med. 2018;72(1):73\u0026ndash;83. e5.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDugan KC, Hall JB, Patel BK. High-flow nasal oxygen\u0026mdash;the pendulum continues to swing in the assessment of critical care technology. JAMA. 2018;320(20):2083\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFerreyro BL, Angriman F, Munshi L, et al. Noninvasive oxygenation strategies in adult patients with acute respiratory failure: a protocol for a systematic review and network meta-analysis. Syst Reviews. 2020;9:1\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFrancio F, Weigert RM, Mattei EDB, et al. High-flow nasal oxygen vs noninvasive ventilation in patients with acute respiratory failure: the renovate randomized clinical trial. JAMA. 2025;333(10):875\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGrieco DL, Menga LS, Cesarano M, et al. Effect of helmet noninvasive ventilation vs high-flow nasal oxygen on days free of respiratory support in patients with COVID-19 and moderate to severe hypoxemic respiratory failure: the HENIVOT randomized clinical trial. JAMA. 2021;325(17):1731\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLinck EJG, Goligher EC, Semler MW, et al. Toward precision in critical care research: Methods for observational and interventional studies. Crit Care Med. 2024;52(9):1439\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMunroe ES, Prevalska I, Hyer M, et al. High-Flow Nasal Cannula Versus Noninvasive Ventilation as Initial Treatment in Acute Hypoxia: A Propensity Score-Matched Study. Crit Care Explorations. 2024;6(5):e1092.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNair PR, Haritha D, Behera S, et al. Comparison of high-flow nasal cannula and noninvasive ventilation in acute hypoxemic respiratory failure due to severe COVID-19 pneumonia. Respir Care. 2021;66(12):1824\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOczkowski S, Ergan B, Bos L et al. ERS clinical practice guidelines: high-flow nasal cannula in acute respiratory failure. Eur Respir J 2022;59(4).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRochwerg B, Brochard L, Elliott MW et al. Official ERS/ATS clinical practice guidelines: noninvasive ventilation for acute respiratory failure. Eur Respir J 2017;50(2).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAthey S, Wager S. Estimating treatment effects with causal forests: An application. Observational Stud. 2019;5(2):37\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eShalit U, Johansson FD, Sontag D. Estimating individual treatment effect: generalization bounds and algorithms. PMLR: 2017.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eK\u0026uuml;nzel SR, Sekhon JS, Bickel PJ et al. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences. 2019;116(10):4156\u0026ndash;4165.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChipman HA, George EI, McCulloch RE. BART: Bayesian additive regression trees. 2010.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLam JY, Lu X, Shashikumar SP et al. Development, deployment, and continuous monitoring of a machine learning model to predict respiratory failure in critically ill patients. JAMIA open 2024;7(4).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWinkler C, Worrall D, Hoogeboom E et al. Learning likelihoods with conditional normalizing flows. arXiv preprint arXiv:191200042 2019.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSinghal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAnthropic. (2024). Claude 3.5 Sonnet [Large language model]. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.anthropic.com/news/claude-3-5-sonnet\u003c/span\u003e\u003cspan address=\"https://www.anthropic.com/news/claude-3-5-sonnet\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"critical-care","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"cric","sideBox":"Learn more about [Critical Care](http://ccforum.biomedcentral.com/)","snPcode":"13054","submissionUrl":"https://submission.nature.com/new-submission/13054/3","title":"Critical Care","twitterHandle":"@Crit_Care","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Causal inference, large language models, individualized treatment effect, high-flow nasal cannula, noninvasive ventilation, guideline adherence","lastPublishedDoi":"10.21203/rs.3.rs-7230335/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7230335/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003cbr\u003e\n Optimal respiratory support selection between high-flow nasal cannula (HFNC) and noninvasive ventilation (NIV) for intensive care units (ICU) patients at risk of invasive mechanical ventilation (IMV) remains unclear, particularly in cases not represented in prior clinical trials. We previously developed RepFlow-CFR, a deep counterfactual model estimating individualized treatment effects (ITE) of HFNC versus NIV. However, interpretability and guideline alignment remain challenges for clinical adoption. This study describes the development and integration of a clinical guideline-driven LLM to enhance deep counterfactual model recommendations for NIV versus HFNC in patients at high-risk for invasive mechanical ventilation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003cbr\u003e\n We enhanced RepFlow-CFR by incorporating a large language model (LLM, Claude 3.5 Sonnet) to enforce clinical guideline adherence and generate explainable treatment recommendations. The LLM was configured in a HIPAA-compliant AWS environment and prompted using structured patient data, clinical notes, and formal guideline criteria. Recommendations from RepFlow-CFR and LLM were compared to actual treatment decisions to assess concordance. We evaluated IMV and mortality/hospice rates across concordant and discordant groups. Additionally, we conducted a structured chart review of 20 cases to assess the clinical validity and safety of LLM-driven recommendations.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003cbr\u003e\n Among 1,261 ICU encounters, treatments concordant with LLM-enhanced recommendations were associated with significantly lower IMV rates (e.g., 24.47% when concordant versus 52.94% when discordant with the HFNC recommendation, corresponding to a 97.33% relative risk increase when discordant) and reduced odds of mortality or hospice discharge (odds ratio = 0.670, p = 0.046). In the chart review, 95% of LLM recommendations aligned with clinical guidelines, and physicians agreed with 65% of final recommendations. Errors were noted in 11/20 cases, with most deemed low or moderate risk; only 2 were rated as potentially causing severe harm.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003cbr\u003e\n Integrating LLMs for guideline enforcement improves the interpretability and clinical alignment of counterfactual models in respiratory support decision-making. This hybrid framework not only enhances concordance with real-world practice but may also improve patient outcomes. Future work will refine contraindication detection and expand validation to prospective clinical trials.\u003c/p\u003e","manuscriptTitle":"Enhancing Predictive Modeling for Respiratory Support with LLM-Driven Guideline Adherence","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-12 13:04:09","doi":"10.21203/rs.3.rs-7230335/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-09-16T12:29:06+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-24T13:22:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"26184218956579967778158407641447295975","date":"2025-08-09T07:04:50+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-08T12:24:48+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"885826839053438631943971356915903160","date":"2025-08-07T05:42:13+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"93790401443223923327189977986010683633","date":"2025-08-07T05:10:29+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-08-07T05:03:45+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-07-30T01:58:37+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-07-30T00:36:07+00:00","index":"","fulltext":""},{"type":"submitted","content":"Critical Care","date":"2025-07-28T06:12:52+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"critical-care","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"cric","sideBox":"Learn more about [Critical Care](http://ccforum.biomedcentral.com/)","snPcode":"13054","submissionUrl":"https://submission.nature.com/new-submission/13054/3","title":"Critical Care","twitterHandle":"@Crit_Care","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e9617da7-4304-48d2-a7e3-0ceccdff5d30","owner":[],"postedDate":"August 12th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-11-17T16:00:11+00:00","versionOfRecord":{"articleIdentity":"rs-7230335","link":"https://doi.org/10.1186/s13054-025-05739-3","journal":{"identity":"critical-care","isVorOnly":false,"title":"Critical Care"},"publishedOn":"2025-11-14 15:57:21","publishedOnDateReadable":"November 14th, 2025"},"versionCreatedAt":"2025-08-12 13:04:09","video":"","vorDoi":"10.1186/s13054-025-05739-3","vorDoiUrl":"https://doi.org/10.1186/s13054-025-05739-3","workflowStages":[]},"version":"v1","identity":"rs-7230335","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7230335","identity":"rs-7230335","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00