Dynamic Landmark-Based Prediction of Sepsis Using Interpretable and Balanced Machine Learning Models in Respiratory-Supported Critically ill Patients | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Dynamic Landmark-Based Prediction of Sepsis Using Interpretable and Balanced Machine Learning Models in Respiratory-Supported Critically ill Patients Ayao Sangenis Assogba, Jennifer H. Gladius, Komi Selassi Gayi, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8737800/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Early recognition of sepsis in critically ill patients remains challenging due to dynamic physiological changes and nonspecific clinical presentation. Most prediction models rely on static or continuously updated data streams without explicitly accounting for evolving risk over clinically meaningful time intervals. The study aimed to develop and evaluate a landmark-based dynamic machine learning framework to predict sepsis within a 6-hour horizon among respiratory-supported intensive care unit (ICU) patients. Methods This is a secondary analysis using data from the MIMIC-IV database. Adult intensive care unit patients receiving respiratory support were evaluated at four landmarks (6, 12, 18, and 24 hours). At each point, sepsis-free patients were used to predict sepsis onset within the next 6 hours. Models included logistic regression, random forest, and XGBoost. The patient–level train–test splitting and group cross-validation prevented information leakage. Performance was assessed using discrimination, classification metrics, and calibration. A balanced ensemble approach addressed class imbalance in sensitivity analysis, and interpretability was examined using permutation importance and regression effect estimates. Results A total of 41,871, 39,912, 36,472, and 31,367 patients were included at the 6, 12, 18, and 24-hour landmarks, respectively. Sepsis incidence declined from 1.48% to 0.37% across time points. Model performance varied, with the 18-hour landmark showing the best balance between discrimination and clinically meaningful operating characteristics. Logistic regression achieved the highest discrimination in the primary analysis (AUROC = 0.78), while random forest performed best in sensitivity analyses (AUROC = 0.77). Both consistently identified the 18-hour landmark as optimal, indicating that temporal risk structure outweighed algorithm choice. Calibration was checked overall but showed overestimation at higher predicted risks. Key predictors reflected respiratory, hemodynamic, neurological, and comorbidity factors. Conclusions Landmark-based dynamic modelling provides a clinically interpretable and temporally informed strategy for early sepsis prediction in respiratory-supported intensive care unit patients. The consistent identification of the 18-hour window as the most informative prediction point suggests that intermediate ICU time frames may offer the best balance between timeliness and predictive stability. Further work should focus on recalibration, threshold optimization, and external validation before clinical implementation. Sepsis Machine learning Landmark modelling ICU Dynamic prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Sepsis is a life-threatening organ dysfunction resulting from a disordered host response to an infectious agent and is recognized as a significant global health concern. According to the Global Burden of Disease estimates, nearly 49 million cases of sepsis and 11 million deaths attributable to sepsis occur annually, which accounts for approximately one-fifth of all deaths globally [ 1 ]. In addition to being a major cause of morbidity and mortality in the intensive care unit (ICU), sepsis also contributes to extended lengths of hospitalization, elevated healthcare costs, and increased risk of chronic complications [ 2 ]. Therefore, early recognition and timely initiation of antibiotics and supportive therapy are essential, as each hour of delay in treatment has been associated with a significant increase in mortality [ 3 ]. Although important advances have been made in critical care, the diverse and often non-specific presentation of sepsis in critically ill patients continues to make early detection of sepsis a persistent challenge for clinicians. Traditionally, clinicians primarily relied upon severity scores (such as the sequential organ failure assessment (SOFA) and the quick SOFA (qSOFA)) to assist in assessing the risk of patients. While these scoring systems can provide valuable prognostic information, they are limited because they are based solely on static physiological thresholds and may be insensitive to the early stages of sepsis [ 4 ]. Furthermore, when using scoring systems (such as SOFA) for both defining outcomes and developing predictive models, they can result in circular reasoning and limit the discovery of new predictors independent of the established criteria [ 5 ]. During the last decade, improvements in electronic health records (EHRs) and computational resources have enabled the development of machine learning (ML) approaches for the prediction of sepsis. ML approaches can analyze large amounts of complex and multi-dimensional clinical data and identify nonlinear associations between risk factors and outcomes. Multiple studies have demonstrated strong discriminative performance of ML models in ICU populations, with many studies reporting area under the receiver operating characteristic curve (AUROC) values of greater than 0.80 [ 6 , 7 ]. However, systematic reviews have emphasized several serious shortcomings of ML models for the prediction of sepsis: the performance of ML models varies significantly between studies; the calibration of ML models is commonly poor; and prospective or external validation of ML models is uncommon [ 6 , 8 ]. There are a number of significant methodological limitations to the majority of the previous literature concerning sepsis prediction models. First, many previously published sepsis prediction models utilized only static, or admission time data, thus failing to capture the changing physiological states of ICU patients over time [ 6 , 9 ]. Second, the vast majority of sepsis prediction models were developed in general ICU populations and there has been very limited focus on those patients that require respiratory support and are at a significantly higher risk of developing sepsis due to the need for invasive medical devices, compromised immune systems and increased susceptibility to nosocomial infections. Thirdly, the use of temporal modeling techniques to predict a patient’s risk of developing sepsis based upon new data collected over time is a novel area of research, but one that may be able to offer clinicians with more clinically relevant information regarding a patient’s risk of developing sepsis than static models. A landmark analysis offers a clinically intuitive approach to predictively determining a patient’s risk of developing sepsis, by defining a series of predefined time points (landmarks) at which a patient’s risk will be estimated based upon all the data available up to that landmark [ 10 ]. Although the structural representation of each landmark model is static, by making predictions at each landmark, patient risk profiles can be continuously updated and therefore reflect changes in both physiological status and clinical context. By doing so, this strategy maintains temporal validity, prevents information leakage and aligns very well with how clinicians assess risk throughout the ICU care process [ 11 ]. In order to address the aforementioned limitations, our study applies a landmark-based dynamic prediction strategy for predicting short-term sepsis risk in critically ill patients who receive mechanical ventilation (respiratory) support. The risk estimations were made at 6, 12, 18 and 24 hours post-ICU admission using common statistical and machine learning models; logistic regression, random forest and gradient boosting. Logistic regression is included as an interpretable reference model, while ensemble-based methods are considered to examine whether nonlinear modeling strategies may offer advantages in the presence of complex clinical interactions and class imbalance. To understand differences in predictive discrimination, calibration, and clinical utility across the course of treatment, model performance is independently assessed at each landmark. The primary objective of this study was to develop and evaluate a dynamic landmark-based machine learning framework for early sepsis risk prediction in critically ill patients requiring respiratory support and to determine how the predictive performance, calibration, and clinical interpretability of the model evolved over the first 24 hours of critical care. By integrating landmark analysis with machine learning and rigorous evaluation strategies, our work seeks to advance the development of transparent, temporally valid, and clinically meaningful decision-support tools for early sepsis detection in critical care. Methods Study Design and Data Source A secondary data analysis was conducted using the publicly available, de-identified Medical Information Mart for Intensive Care IV (MIMIC-IV) database [ 12 ]. The MIMIC-IV is a large public database containing highly detailed, time-stamped information regarding adult patients admitted to ICUs, which includes physiological measurements, clinical interventions and patient outcomes. All analyses were performed in accordance with ethical standards for secondary analysis of de-identified data; institutional review board approval and informed consent were not required. Study Population and Cohort Construction The study population consisted of adult patients (> 18 years) who were admitted to an ICU and received some form of respiratory support, including mechanical ventilation (invasive or non-invasive), or high-flow oxygen therapy for at least one of the first 24 hours post-ICU admission. In order to generate longitudinal data sufficient to create landmark-specific feature windows, only patients with such data were included. To construct a valid at-risk cohort and prevent temporal leakage, patients were excluded from a given landmark and all subsequent landmarks if sepsis occurred prior to that landmark time. Data censoring was performed at either ICU discharge or at the time of a patient's death, and no landmark observations were created after these events. Landmark observations with insufficient data within the predefined feature window were excluded. The final analytic database was structured in a long format; with each row representing a single patient-landmark observation associated with a unique prediction time point (see Fig. 1). Patients were permitted to contribute observations to multiple landmarks provided they were still alive in the ICU and sepsis-free at each corresponding landmark. The overall study design, cohort selection process, landmark-based dataset construction, and prediction modeling framework are illustrated in Fig. 2. Landmark Analysis Framework A landmark analysis refers to the practice of designating a time point occurring during the follow-up period (known as the landmark time ) and analyzing only those subjects who have survived until the landmark time [ 13 ]. A comprehensive overview of the landmark analysis method and its use has been provided by Dafni [ 14 ]. A landmarking strategy was employed to enable dynamic risk prediction at clinically meaningful time points following ICU admission. Our landmarks were prespecified at 6, 12, 18, and 24 hours after ICU admission. Each landmark constituted an independent prediction task with its own risk set, including only patients who were alive in the ICU and free of sepsis at that time. Although each landmark model was static in form, allowing patients to contribute repeated observations across landmarks enabled implicit modeling of patient trajectories through repeated risk re-estimation over time. This approach preserved temporal validity, avoided information leakage, and aligned with routine clinical reassessment practices in the ICU. Temporal Alignment of Predictors and Outcomes For each landmark, predictors were extracted from a fixed 6-hour feature window immediately preceding the landmark time: [ t landmark – 6, t landmark ) The outcome was defined as the occurrence of sepsis within the subsequent 6-hour prediction window: ( t landmark , t landmark + 6] This strict temporal alignment ensured that all predictors were observed prior to outcome assessment and eliminated forward-looking bias. Outcome Definition The primary outcome, sepsis_next_6h , was defined as the occurrence of sepsis within 6 hours following each landmark, based on Sepsis-3 criteria implemented in the dataset [ 2 ]. Patients with any evidence of sepsis prior to a given landmark were excluded from that landmark’s risk set to ensure that only incident sepsis events were modelled. At each landmark, control observations were defined as patient–landmark instances in which sepsis did not occur within the subsequent 6-hour prediction window. The composite organ dysfunction scores used to define sepsis (e.g., SOFA) were explicitly excluded from the predictor set to avoid circularity and artificial inflation of model performance [ 5 ]. Predictor Engineering and Selection Physiological Variables Predictors were selected a priori based on clinical relevance, interpretability, and availability prior to each landmark. Only information observed within the pre-landmark feature window was used. Physiological instability was summarized using clinically meaningful extreme values, which better reflect acute deterioration in critically ill patients than averages. Selected variables included: Maximum heart rate Maximum respiratory rate Minimum oxygen saturation Maximum temperature Minimum Glasgow Coma Scale (GCS) score Blood Pressure Representation Non-invasive systolic and diastolic blood pressure measurements were used to derive mean arterial pressure (MAP) using the standard formula: MAP = (SBP + 2 × DBP) / 3 [ 15 ]. MAP values were summarized using their mean within each landmark feature window and retained as continuous variables due to their relevance to tissue perfusion and hemodynamic stability. Explicit missingness indicators were incorporated to capture clinically informative measurement patterns related to illness severity and care processes. Organ Support Variables Variables representing markers of organ support included vasopressor use, continuous renal replacement therapy (CRRT), invasive mechanical ventilation, non-invasive ventilation, and high-flow oxygen therapy. These variables were represented as proportions of exposure time within each feature window. Organ support variables were interpreted as indicators of increasing illness severity rather than as predictive factors of sepsis. Static Covariates Patient characteristics collected at baseline included age at ICU admission, sex, race, and the Elixhauser–van Walraven comorbidity index, which measures comorbidity burden. Initial consideration was given to including body mass index (BMI) but it was ultimately removed from the final models because of excessive missingness. Missing Data Handling Exploratory Missingness Assessment Patterns of missing data were explored in total and stratified by landmark time (as shown in Fig. S1 of Additional file 1). Most physiological variables had relatively small amounts of missing data (less than 10%), while some of the non-invasive blood pressure variables were moderately missing (about 20%) and BMI had large amounts of missing data (more than 50%). All organ support variables were complete by construction. Row-Level Exclusions Prior to imputation, the following exclusions were applied: Observations in which the outcome or landmark time were missing Observations in which all of the core physiological variables were missing Observations in which greater than 50% of the selected predictors were missing Imputation Strategy Landmark-aware imputation strategy was employed in order to maintain temporal relationships and avoid leaking information. Time-varying physiological variables were imputed by the median value for the landmark time in which they were measured. There was no imputation of organ support variables, categorical variables, or the outcome. Binary missingness indicators were also created for GCS and non-invasive blood pressure variables. Events-Per-Variable (EPV) Assessment Before beginning model development, the number of sepsis events at each landmark was examined relative to the number of candidate predictors in order to evaluate the number of events per variable (EPV). EPV was sufficient to justify the selected model specifications and guided the model complexity and regularization to reduce the risk of overfitting. Model Development and Evaluation To ensure that model development is robust and maintains clinical realism, a standard two-stage training and evaluation paradigm was defined as the main analytical methodology. In addition to this, a balanced ensemble methodology was developed and employed as a sensitivity analysis to determine the impact of different methods of managing class imbalance on model performance. Both methodologies were applied uniformly to each landmark time point (6, 12, 18, and 24 hours) and each candidate model. Modeling Approaches Approach 1: Standard Two-Stage Training and Evaluation In the primary analysis, a conventional two-stage machine learning pipeline was implemented. For each landmark and model type, data were split at the patient level into training and independent test sets to prevent information leakage across repeated landmark observations. Hyperparameters were optimized using cross-validation (5 folds) within the training set, and final performance was evaluated on the held-out test set. Class imbalance was handled using class-weighted loss functions, preserving the natural incidence of sepsis and maintaining clinically meaningful calibration. This approach reflects real-world deployment where models are trained on imbalanced data and applied to unseen patients. Approach 2: Balanced Ensemble Strategy (Sensitivity Analysis) To assess robustness to alternative imbalance handling, a sensitivity analysis using a balanced ensemble strategy adapted from prior early-prediction studies [ 16 ] was conducted. For each landmark (6, 12, 18, and 24 hours), non-sepsis observations were partitioned into multiple mutually exclusive subgroups via stratified random sampling. The number of subgroups was predetermined based upon the degree of class imbalance at each landmark; the number of subgroups increased as the number of sepsis events decreased at later landmarks: 6-hour: 20 subgroups 12-hour: 30 subgroups 18-hour: 40 subgroups 24-hour: 50 subgroups Non-sepsis cases from each subgroup were paired with the total number of sepsis cases at the corresponding landmark window to create a balanced training dataset (as shown in Fig. S2 of Additional file 2). For each balanced dataset, an independent model was trained using the same predictors and preprocessing as in the primary approach. The ensemble predictions were averaged across all models for each landmark time point. This strategy allowed evaluation of whether conclusions regarding optimal landmark timing and comparative model performance were sensitive to subsampling-based imbalance correction. Candidate Models Three modeling algorithms were independently evaluated at each landmark time using both methodologies: Logistic Regression Random Forest XGBoost Gradient Boosting All models were trained using the same feature sets and temporal feature definitions. Performance Evaluation Performance of a model was evaluated on test data from each landmark and each method of analysis. Discriminative ability was evaluated using both AUROC and AUPRC. Additionally, clinical utility was evaluated through calculation of sensitivity, specificity and accuracy. Metrics were reported separately by landmark, model type, and modeling approach. Model Selection Strategy For each landmark, models were compared within each analytical approach. The best performing model was selected based upon discrimination and clinical utility metrics, with an emphasis on AUROC, AUPRC, sensitivity and specificity. Robustness of results across primary and sensitivity analyses were evaluated to assess agreement and consistency of results. Role of the Dual-Approach Framework The dual-strategy approach provided the opportunity to evaluate the predictive stability of the models under different class imbalance assumptions. The consistent results obtained across the different approaches provided additional confidence that the results can be generalized, while inconsistencies provide caution when interpreting the results. Model Interpretability Interpretability was evaluated for the best-performing model at the most clinically informative landmark. For tree-based models, permutation feature importance was used to quantify the impact of each predictor on model performance. For logistic regression, associations were expressed as odds ratios, enhancing transparency in identifying key predictors of sepsis risk. Software All analyses were conducted using Python (version 3.11). Data processing and preprocessing were performed using pandas and NumPy. Machine-learning modeling and evaluation were implemented using Scikit-learn, XGBoost, and Joblib. Model interpretability analyses were conducted using SHAP and permutation importance methods. Visualizations were generated using Matplotlib and Seaborn. Results Patient characteristics A total of 50,920 adult ICU admissions requiring respiratory support were initially identified from the MIMIC-IV v2.2 database. After applying the predefined inclusion and exclusion criteria, 41,871 patients were retained in the 6-hour landmark risk set. As landmark-specific eligibility criteria were sequentially applied, cohort sizes decreased to 39,912, 36,472, and 31,367 patients at the 12, 18, and 24 hour landmarks, respectively. The Table 1 summarizes the demographic characteristics, physiological measurements, comorbidity burden, and treatment-related features of patients included at each landmark time point. All characteristics are reported descriptively. Table 1 Baseline characteristics of the study cohort at each landmark time Variable Landmark 6 h (n = 41,871) Landmark 12 h (n = 39,912) Landmark 18 h (n = 36,472) Landmark 24 h (n = 31,367) Age 65 (53–77) 65 (53–77) 65 (53–77) 66 (53–77) Heart rate, max 90 (79.5–104.5) 89 (78–102) 89 (78–102) 90 (79–103) SpO₂, min 95.5 (93–98) 95 (93–97) 95 (93–97) 95 (93–97) Temperature, max 36.9 (36.6–37.2) 36.9 (36.7–37.2) 36.9 (36.7–37.2) 36.9 (36.7–37.2) Respiratory rate, max 22 (19–26) 22 (19–26) 22 (19–26) 23 (20–27) MAP, non-invasive 83.7 (77.2–91.1) 81.9 (76.0–88.2) 82.2 (76.6–88.5) 82.1 (76.1–88.8) Sex Female 19,030 (45.4) 18,157 (45.5) 16,543 (45.4) 14,088 (44.9) Male 22,841 (54.6) 21,755 (54.5) 19,929 (54.6) 17,279 (55.1) Race/ethnicity Asian 1,266 (3.0) 1,210 (3.0) 1,111 (3.0) 947 (3.0) Black 4,159 (9.9) 3,974 (10.0) 3,555 (9.7) 3,011 (9.6) Hispanic/Latino 1,646 (3.9) 1,585 (4.0) 1,435 (3.9) 1,242 (4.0) Native American 75 (0.2) 73 (0.2) 69 (0.2) 61 (0.2) Other/Unknown 6,527 (15.6) 6,138 (15.4) 5,640 (15.5) 4,968 (15.8) Pacific Islander 66 (0.2) 62 (0.2) 54 (0.1) 44 (0.1) White 28,132 (67.2) 26,870 (67.3) 24,608 (67.5) 21,094 (67.2) Vasopressor use No 36,064 (86.1) 34,014 (85.2) 31,427 (86.2) 27,241 (86.8) Yes 5,807 (13.9) 5,898 (14.8) 5,045 (13.8) 4,126 (13.2) Continuous renal replacement therapy No 41,835 (99.9) 39,813 (99.8) 36,329 (99.6) 31,184 (99.4) Yes 36 (0.1) 99 (0.2) 143 (0.4) 183 (0.6) Invasive ventilation No 33,220 (79.3) 31,935 (80.0) 30,370 (83.3) 26,475 (84.4) Yes 8,651 (20.7) 7,977 (20.0) 6,102 (16.7) 4,892 (15.6) Non-invasive ventilation No 41,583 (99.3) 39,661 (99.4) 36,259 (99.4) 31,182 (99.4) Yes 288 (0.7) 251 (0.6) 213 (0.6) 185 (0.6) High-flow oxygen No 41,705 (99.6) 39,713 (99.5) 36,273 (99.5) 31,156 (99.3) Yes 166 (0.4) 199 (0.5) 199 (0.5) 211 (0.7) Glasgow Coma Scale Severe (cat_1) 1,381 (3.3) 787 (2.0) 550 (1.5) 518 (1.7) Moderate (cat_2) 1,907 (4.6) 1,554 (3.9) 1,366 (3.7) 1,219 (3.9) Mild (cat_3) 38,583 (92.1) 37,571 (94.1) 34,556 (94.7) 29,630 (94.5) Elixhauser comorbidity Low (cat_1) 16,159 (38.6) 15,279 (38.3) 13,693 (37.5) 11,352 (36.2) Moderate (cat_2) 16,208 (38.7) 15,516 (38.9) 14,263 (39.1) 12,404 (39.5) High (cat_3) 9,504 (22.7) 9,117 (22.8) 8,516 (23.3) 7,611 (24.3) Across the landmark risk sets, the incidence of sepsis decreased progressively with increasing landmark time. Among patients included in the 6-hour landmark cohort, 621 sepsis events were observed among 41,871 patients, corresponding to an incidence of 1.48%. At subsequent landmarks, sepsis incidence declined to 0.55% (220 of 39,912) at 12 hours, 0.45% (164 of 36,472) at 18 hours, and 0.37% (117 of 31,367) at 24 hours (Table 2). This reduction reflects both attrition of higher-risk patients over time and the exclusion of patients who developed sepsis prior to each landmark. Table 2 Incidence of Sepsis Across Landmark Windows ICU landmark (hours) Patients at risk, n Sepsis events, n (%) EPV 6 41,871 621 (1.48) 36.5 12 39,912 220 (0.55) 12.9 18 36,472 164 (0.45) 9.6 24 31,367 117 (0.37) 6.9 Training and evaluation by Landmark The model performance was evaluated separately at each landmark time using a standard two-stage training and evaluation framework. Cross-validated discrimination metrics were first estimated within the training data, followed by final performance assessment on an independent held-out test set. Two stage classic approach Cross-validated model performance Cross-validated discrimination varied across landmark times and modeling approaches (Table 3). Logistic regression demonstrated more balanced sensitivity and specificity across landmarks, whereas ensemble models consistently favoured high specificity at the expense of sensitivity. Table 3 Cross-validated discrimination performance across ICU landmark times in the primary approach hr model AUROC_ mean AUROC_ sd AUPRC_ mean Sensitivity_ mean Specificity_ mean n_patients n_events 6 Logistic 0.72 0.02 0.06 0.55 0.76 31403 465 6 RandomForest 0.73 0.04 0.10 0.31 0.96 31403 465 6 XGBoost 0.71 0.04 0.08 0.00 1.00 31403 465 12 Logistic 0.69 0.01 0.02 0.64 0.66 29920 174 12 RandomForest 0.64 0.03 0.01 0.05 0.98 29920 174 12 XGBoost 0.61 0.03 0.01 0.00 1.00 29920 174 18 Logistic 0.65 0.06 0.01 0.52 0.68 27386 121 18 RandomForest 0.67 0.07 0.01 0.01 0.99 27386 121 18 XGBoost 0.66 0.06 0.01 0.00 1.00 27386 121 24 Logistic 0.67 0.07 0.01 0.56 0.71 23604 90 24 RandomForest 0.63 0.04 0.01 0.00 0.99 23604 90 24 XGBoost 0.59 0.07 0.01 0.00 1.00 23604 90 Final test set performance The final test set evaluation confirmed these trends (Table 4). At the 18-hour landmark, logistic regression achieved the highest overall discrimination (Fig. 3), with improved sensitivity and comparable specificity relative to other landmarks. Table 4 Final test set performance of selected sepsis prediction models across ICU landmark times hr model AUROC AUPRC Accuracy Sensitivity Specificity n_patients n_events 6 RF 0.72 0.07 0.95 0.30 0.95 10468 156 12 Logistic 0.67 0.01 0.66 0.63 0.66 9992 46 18 Logistic 0.78 0.02 0.68 0.77 0.68 9086 43 24 Logistic 0.70 0.01 0.71 0.70 0.718 7763 27 Balanced ensemble models Sensitivity analyses using balanced ensemble models supported the robustness of the primary findings. Across landmarks, the 18-hour window consistently demonstrated the most favourable predictive performance, with the random forest ensemble achieving the best balance between discrimination (AUROC 0.77) and clinically relevant sensitivity while preserving acceptable specificity compared to others (Table 5). Although the optimal algorithm varied across landmarks in the sensitivity analysis, the relative ranking of landmark windows and the identification of 18 hours as the optimal prediction time remained unchanged, confirming the stability of the primary conclusions under alternative class-imbalance handling strategies. Table 5 Model performance using the balanced ensemble approach (sensitivity analysis) hr Model AUROC AUPRC Accuracy Sensitivity Specificity Precision Recall F1 6 LR 0.73 0.05 0.95 0.26 0.96 0.10 0.26 0.14 12 RF 0.65 0.01 0.72 0.48 0.72 0.01 0.48 0.02 18 RF 0.77 0.02 0.74 0.65 0.74 0.01 0.65 0.02 24 LR 0.67 0.01 0.98 0 0.99 0 0 0 Calibration Performance Calibration was assessed for the best-performing model at each landmark using calibration plots and Brier scores. At the 18-hour landmarking the primary approach, the logistic regression model showed reasonable agreement between predicted and observed sepsis risk (Brier score = 0.213), with mild overprediction at higher risk levels and good alignment in the low-risk range where most patients clustered. Other landmarks showed similar patterns, with greater instability at later times due to fewer events. Sensitivity analysis of balanced ensemble models indicated improved discrimination at some landmarks but more variable calibration, reflecting resampling and aggregation across sub-models (as shown in Fig. S3 of Additional file 3). Overall, findings support using the 18-hour model for risk stratification while emphasizing caution in interpreting absolute risk estimates in highly imbalanced settings. Model Interpretability To enhance clinical interpretability of the selected models, feature importance analyses were conducted for the best-performing models identified at the 18-hour landmark in both the primary and sensitivity analyses. For the primary analysis, interpretability of the logistic regression model was assessed using a forest plot of adjusted odds ratios with 95% confidence intervals (Fig. 4.). This analysis highlighted physiologically plausible predictors of near-term sepsis, including markers of respiratory compromise, hemodynamic instability, and neurological status, with consistent directions of effect across covariates. The magnitude and direction of associations were clinically coherent, supporting the transparency and interpretability of the linear model. In the sensitivity analysis, permutation-based feature importance was derived for the random forest model selected at the 18-hour landmark (Fig. 5.). This approach identified a similar set of high-impact features, particularly vital sign extremes and indicators of organ support, suggesting concordance between linear and non-linear modeling approaches despite algorithmic difference. Together, these findings demonstrate that the predictive signal driving model performance was clinically meaningful and robust across modeling strategies. Interpretability analyses therefore reinforced the validity of the 18-hour landmark as a clinically actionable time point and supported the reliability of the identified predictors for early sepsis risk stratification. Discussion Principal findings In this secondary data analysis of respiratory-supported ICU patients, dynamic machine-learning models were developed and evaluated to predict sepsis within a 6-hour horizon using a landmarking framework. Several important findings emerged. First, the incidence of sepsis declined across successive landmark windows, reflecting both clinical evolution and exclusion of earlier events. Second, model performance varied by landmark time, with the 18-hour landmark consistently demonstrating the most favourable balance between discrimination and clinically relevant operating characteristics. Third, although the best performing algorithms were different for each strategy (logistic regression in the primary analysis and random forest in the sensitivity analysis), the fact that the two studies converged upon the same landmark time point suggests that the predictive performance of the models was based on the temporal risk structure of sepsis development and not the specific classification method used. Finally, the calibration was performed and the key features were clinically interpretable and reflected respiratory, hemodynamic and neurological dysfunction that is consistent with the well-established pathophysiology of sepsis [ 6 , 7 , 17 ]. Comparison with existing literature Early identification of sepsis in critically ill patients remains challenging due to heterogeneous clinical presentation and evolving physiology [ 6 , 18 ]. Research has shown that machine learning can provide effective methods for predicting sepsis in the ICU population, however, most of these studies utilize either static baseline data or continuous real-time monitoring without explicit temporal risk stratification [ 19 , 20 ]. Fleuren et al. in their systematic review and meta-analysis reported an average AUROC range from 0.68 to 0.99 for all ICU-based sepsis models [ 6 ] but they also found variability in calibration among the reviewed models as well as inconsistency in the reporting of the F1 measure, often leading to inflated impressions of clinical utility. In contrast to previous studies reported by Fleuren et al. showing near-perfect discrimination, our model achieved a more moderate AUROC of 0.78 when evaluated under real-world class imbalance. This finding suggests that discrimination alone may be insufficient to determine the clinical utility of a predictive model, particularly in settings where sepsis incidence is low and false alarms carry significant consequences. Beyond overall discrimination in our analysis we found respiratory compromise, hemodynamic instability, temperature and neurological dysfunction as key contributors to predicting sepsis. These findings are consistent with prior interpretable machine learning studies by Nemati et al. [ 21 ] and Tang et al. [ 22 ], which highlighted oxygen saturation, respiratory rate, and temperature among the most influential predictors. The alignment of our model’s important features with established physiological markers of sepsis supports the biological plausibility and clinical relevance of our approach. The inclusion of Elixhauser comorbidity [ 23 ] as a strong contributor further supports the relevance of chronic disease burden in predisposing critically ill patients to infection-related organ dysfunction. This convergence across studies strengthens the face validity and clinical credibility of our model. Similar landmarking strategies have shown promise in other critical care contexts, including early infection detection and mortality risk modeling, supporting the relevance of temporally updated prediction frameworks [ 24 , 25 ]. Notably, our findings align with prior work showing improved predictive stability when models account for evolving patient trajectories rather than single-time-point measurements [ 25 , 26 ]. Clinical implications The identification of the 18-hour landmark as the optimal prediction window has potential clinical relevance. At this time point, models achieved a favourable trade-off between discrimination, sensitivity, and specificity, suggesting that risk stratification may be most informative once initial stabilization and early therapeutic interventions have occurred [ 19 , 27 ]. Furthermore, the consistency of this landmark across primary and sensitivity analyses underscores the robustness of this temporal signal. Interpretability analyses highlighted clinically plausible predictors, including markers of respiratory compromise, presence of comorbidities, hemodynamic instability, and altered neurological status, reinforcing the face validity of the models and supporting their potential role in clinical decision support [ 17 , 28 ]. Robustness and sensitivity analyses In addition to the issues caused by large amounts of imbalanced data for predicting sepsis, a balanced ensemble modeling method was used as a sensitivity test, given the large degree of class imbalance [ 16 , 29 , 30 ] as a sensitivity analysis. Although different models were selected at each landmark, the results from all three landmarks showed generally similar clinical performance characteristics and discrimination relative to the primary analysis. Additionally, both methods found the 18-hour landmark to be the best time to predict sepsis, this further emphasizes the consistency of the temporal trend seen in the previous findings. These findings suggest that the predictive signal was not an artifact of sampling strategy or algorithm choice, but rather reflects underlying clinical risk dynamics. Our research is strong from both clinical and methodological perspectives. First, we used a very large, fully characterized critical care data set, which had high resolution, clinically relevant data [ 31 ]. The use of a dynamic landmarking model was able to capture the changing nature of sepsis risk in real-time, as opposed to the static predictions provided by many other models, and it aligned with recommendations regarding the application of time-sensitive machine learning in the field of critical care [ 32 – 33 ]. To minimize the risk of information leakage and model over-fitting, we employed two methods of internal validation (patient level splitting and group based cross-validation), which are common among predictive modeling applications [ 32 – 34 ]. Additionally, we applied two types of interpretability analysis (forest plot and permutation feature importance) to increase the clarity of how clinical factors contribute to an individual's sepsis risk and to build clinician confidence through increasing their understanding of how these models’ function [ 32 – 35 ]. Several limitations should be acknowledged. The retrospective design limits causal inference and may introduce selection bias. Consistent with systematic reviews in this field, many machine learning sepsis prediction models including ours are developed on single databases, which restricts generalizability and highlights the need for independent validation across diverse settings [ 32 , 36 ]. Although we employed standard techniques to mitigate class imbalance, the low incidence of sepsis constrained precision and positive predictive value, reflecting a common challenge in imbalanced clinical prediction tasks [ 32 , 34 ]. Additionally, certain potentially informative biomarkers (e.g., lactate, inflammatory markers) were unavailable or inconsistently measured in our dataset, which may have limited predictive performance compared with models incorporating broader laboratory data [ 32 , 37 ]. Finally, integration into real-time clinical workflows and prospective evaluation remain necessary before practical deployment. Future work should prioritize external validation in multicenter and multinational cohorts to assess robustness and transportability, as underscored in recent meta-analyses [ 36 ]. Efforts to improve probability calibration and threshold optimization for clinically actionable decision thresholds are essential before implementation in decision-support systems [ 32 , 34 ]. Exploring adaptive updating strategies that incorporate additional longitudinal features and real-time data streams may further enhance predictive performance and clinical utility. Conclusion A dynamic landmark-based machine learning approach to predict early sepsis in respiratory supported ICU patients was developed from routinely collected EHR data. The ability to update risk at clinically relevant times resulted in better capture of a patient's trajectory over time, and demonstrated that the 18 hour landmark provided the best trade-off between discrimination and clinical applicability. Model interpretability demonstrated reliance on physiological and logical indicators of respiratory, hemodynamic, and neurological dysfunction to support clinical relevance. Although good discrimination was achieved in our model, low precision due to a large degree of class imbalance in the real world indicates that additional recalibration of probabilities, optimization of thresholds, and additional validation of the models externally is necessary before implementation as a clinical tool. In general, dynamic landmark-based prediction appears to be a promising and clinically appropriate strategy for sepsis risk stratification in critically ill patients. Abbreviations AI Artificial Intelligence AUC Area Under the Curve AUROC Area Under the Receiver Operating Characteristic Curve AUPRC Area Under the Precision-Recall Curve BIDMC Beth Israel Deaconess Medical Center CI Confidence Interval CV Cross-Validation CITI Collaborative Institutional Training Initiative CRRT Continuous Renal Replacement Therapy EHR Electronic Health Record EPV Events Per Variable ICU Intensive Care Unit LR Logistic Regression ML Machine Learning MIMIC IV-Medical Information Mart for Intensive Care IV RF Random Forest ROC Receiver Operating Characteristic qSOFA Quick Sequential Organ Failure Assessment SHAP SHapley Additive exPlanations SOFA Sequential Organ Failure Assessment SpO₂ Peripheral Capillary Oxygen Saturation XGB Extreme Gradient Boosting (XGBoost) Declarations Funding The authors received no specific funding for this work. Ethics approval and consent to participate The MIMIC-IV database is a publicly available, de-identified critical care database. Access was granted after completion of the required data use training and credentialing process. The use of MIMIC-IV data is approved by the institutional review boards of the Massachusetts Institute of Technology and Beth Israel Deaconess Medical Center, and informed consent was waived due to the de-identified nature of the dataset. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Author Contribution ASA conceptualized the study, conducted the analysis, and drafted the manuscript. JHG supervised the project, contributed to the study design, and critically revised the manuscript. KSG, ST, YK, and RST assisted with data processing, statistical interpretation, and manuscript review. RD contributed to study design, methodology, and manuscript revision. All authors read and approved the final manuscript. Acknowledgement The authors gratefully acknowledge the support provided by the SRM School of Public Health, Faculty of Medicine and Health Sciences, SRM Institute of Science and Technology (SRMIST), Kattankulathur. We also thank the PhysioNet team for maintaining open access to the MIMIC-IV database, which made this research possible. Data Availability The data that support the findings of this study are available from the MIMIC-IV database (PhysioNet) and require credentialed access and completion of a data use agreement.The full analysis pipeline, including cohort extraction, landmark dataset construction, feature engineering, model development, and evaluation scripts, is publicly available on GitHub at:[https://github.com/Sangenis11/sepsis-landmark-prediction](https:/github.com/Sangenis11/sepsis-landmark-prediction) .The repository contains all code necessary to reproduce the analyses, excluding patient-level data in accordance with the MIMIC data use agreement. References Rudd KE, Johnson SC, Agesa KM, et al. Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study. Lancet. 2020;395(10219):200–11. Singer M, Deutschman CS, Seymour CW, et al. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA. 2016;315(8):801–10. Seymour CW, Gesten F, Prescott HC, et al. Time to treatment and mortality during mandated emergency care for sepsis. N Engl J Med. 2017;376(23):2235–44. Raith EP, Udy AA, Bailey M, et al. Prognostic accuracy of SOFA score, SIRS criteria, and qSOFA score for in-hospital mortality among adults with suspected infection admitted to the intensive care unit. JAMA. 2017;317(3):290–300. Lindner HA, Schamoni S, Kirschning T, et al. Ground truth labels challenge the validity of sepsis consensus definitions in critical illness. Crit Care. 2022;26:350. Fleuren LM, Klausch TLT, Zwager CL, et al. Machine learning for the prediction of sepsis: a systematic review and meta-analysis of diagnostic test accuracy. Intensive Care Med. 2020;46(3):383–400. Moor M, Rieck B, Horn M, et al. Early prediction of sepsis in the ICU using machine learning: a systematic review. Front Med (Lausanne). 2021;8:607952. Zhang Z, Luo L, Song D, et al. Early sepsis mortality prediction model based on interpretable machine learning: development and external validation. BMC Med Inf Decis Mak. 2024;24(1):46. Komorowski M, Celi LA, Badawi O, et al. The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nat Med. 2018;24(11):1716–20. Barrett JK, Sweeting MJ, Wood AM. Dynamic risk prediction for cardiovascular disease: an illustration using the ARIC Study. In: Rao ASR, Pyne S, Rao CR, editors. Handbook of Statistics. Volume 36. Amsterdam: Elsevier; 2017. pp. 47–65. Rizopoulos D, Molenberghs G, Lesaffre EMEH. Dynamic predictions with time-dependent covariates in survival analysis using joint modeling and landmarking. Biom J. 2017;59(6):1261–76. Moukheiber M, Moukheiber L, Moukheiber D, Hao S, Celi LA, Lee H. A temporal dataset for respiratory support in critically ill patients (version 1.1.0). PhysioNet. 2025. Available from: https://doi.org/10.13026/wewp-sj67 Anderson JR, Cain KC, Gelber RD. Analysis of survival by tumor response. J Clin Oncol. 1983;1:710–9. Dafni U. Landmark analysis at the 25-year landmark point. Circ Cardiovasc Qual Outcomes. 2011;4:363–71. DeMers D, Wachs D. Physiology, mean arterial pressure. In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing; 2025. Available from: https://www.ncbi.nlm.nih.gov/books/NBK538226/ Liang Y, Zhu C, Tian C, et al. Early prediction of ventilator-associated pneumonia in critical care patients: a machine learning model. BMC Pulm Med. 2022;22:250. Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1:206–15. Shashikumar SP, Stanley MD, Sadiq I, et al. Early sepsis detection in critical care patients using multiscale physiological signals. J Electrocardiol. 2017;50(6):739–43. Henry KE, Hager DN, Pronovost PJ, Saria S. A targeted real-time early warning score (TREWScore) for septic shock. Sci Transl Med. 2015;7(299):299ra122. Desautels T, Calvert J, Hoffman J, et al. Prediction of sepsis in the ICU with minimal electronic health record data. JMIR Med Inf. 2016;4(3):e28. Nemati S, Holder A, Razmi F, et al. An interpretable machine learning model for accurate prediction of sepsis in the ICU. Crit Care Med. 2018;46(4):547–53. Tang J, Li J, Luo X, et al. Interpretable machine learning-based prediction of 28-day mortality in ICU sepsis patients. Front Public Health. 2024;12:1349928. Elixhauser A, Steiner C, Harris DR, Coffey RM. Comorbidity measures for use with administrative data. Med Care. 1998;36(1):8–27. van Houwelingen HC. Dynamic prediction by landmarking in event history analysis. Scand J Stat. 2007;34(1):70–85. Varkila MRJ, Lancia G, van Smeden M, et al. Early detection of ICU-acquired infections using high-frequency electronic health record data. BMC Med Inf Decis Mak. 2025;25:273. https://doi.org/10.1186/s12911-025-03031-6 . van Houwelingen HC, Putter H. Dynamic Prediction in Clinical Survival Analysis. Boca Raton: CRC; 2012. Sendak MP, D’Arcy J, Kashyap S, et al. A path for translation of machine learning products into healthcare delivery. NPJ Digit Med. 2020;3:1–7. Kawamoto K, Houlihan CA, Balas EA, Lobach DF. Improving clinical practice using clinical decision support systems. BMJ. 2005;330:765. Saito T, Rehmsmeier M. The precision–recall plot is more informative than ROC plots for imbalanced datasets. PLoS ONE. 2015;10(3):e0118432. He H, Garcia EA. Learning from imbalanced data. IEEE Trans Knowl Data Eng. 2009;21(9):1263–84. Johnson AEW, Pollard TJ, Shen L, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10:1. Zubair M, Din I, Sarwar N, Elov B, Makhmudov S, Trabelsi Z. Revolutionizing sepsis diagnosis using machine learning and deep learning models: a systematic literature review. BMC Infect Dis . 2025;25(1):1396. Published 2025 Oct 23. 10.1186/s12879-025-11423-2 Zhang M, Zhong M, Cheng Y, Zhang T. Intelligent Prediction Platform for Sepsis Risk Based on Real-Time Dynamic Temporal Features: Design Study. JMIR Med Inf. 2025;13. 10.2196/74940 . https://medinform.jmir.org/2025/1/e74940 . e74940, URL. Zhang SZ, Ding HY, Shen YM, et al. Harness machine learning for multiple prognoses prediction in sepsis patients: evidence from the MIMIC-IV database. BMC Med Inf Decis Mak. 2025;25:152. https://doi.org/10.1186/s12911-025-02976-y . Hu C, Li L, Huang W, et al. Interpretable Machine Learning for Early Prediction of Prognosis in Sepsis: A Discovery and Validation Study. Infect Dis Ther. 2022;11(3):1117–32. 10.1007/s40121-022-00628-6 . Yadgarov MY, Landoni G, Berikashvili LB, Polyakov PA, Kadantseva KK, Smirnova AV et al. Early detection of sepsis using machine learning algorithms: a systematic review and network meta-analysis. Front Med (Lausanne). 2024;11:1491358. 10.3389/fmed.2024.1491358 . Available from: https://www.frontiersin.org/journals/medicine/articles/10.3389/fmed.2024.1491358 Liu Z, Shu W, Li T, Zhang X, Chong W. Interpretable machine learning for predicting sepsis risk in emergency triage patients. Sci Rep. 2025;15(1):887. 10.1038/s41598-025-85121-z . Published 2025 Jan 6. Additional Declarations No competing interests reported. Supplementary Files missingnessheatmap.pdf balancingensembleapproach.pdf CalibrationCurveLM18Logistic.png CalibrationRandomForestLandmark18.png Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8737800","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":597453075,"identity":"d32ef983-e27a-405b-adb6-104dbcfeb56c","order_by":0,"name":"Ayao Sangenis Assogba","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIie2QsQrCMBRFUwqCILgGivoLlQ4upf5KQ6EuDk7SURDSzVEq+BF+wpMHupR2LXQRBCeHgB+gD7qJpLo55EDCHe4hlzBmMPwlNoBKfArWijV3EzR0xDHL46YM3yk9D7sSmwysvc/66SlES5bTyRbXD8X8wQFsedEpPI8AF0Ut9rWQHFjsHcBKXe0zEIeYLeuQO0LSMBSkSK4zRuXdxV6nmJKyVsCe7YpbzUmRYGWOWNEwaFfG1S2kT44EKZLnbuTtsEUZlhEqlQQ0bHZVSRIMNuf0plXed9Kxf+gbDAaD4TMv2IFXJ/CCweoAAAAASUVORK5CYII=","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":true,"prefix":"","firstName":"Ayao","middleName":"Sangenis","lastName":"Assogba","suffix":""},{"id":597453076,"identity":"5203050f-9e70-40b9-9002-8fbf7537742f","order_by":1,"name":"Jennifer H. Gladius","email":"","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Jennifer","middleName":"H.","lastName":"Gladius","suffix":""},{"id":597453077,"identity":"6d1e931c-02b8-421b-bbc9-ccbb0cd25d7b","order_by":2,"name":"Komi Selassi Gayi","email":"","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Komi","middleName":"Selassi","lastName":"Gayi","suffix":""},{"id":597453078,"identity":"3ec1d2be-bcbc-49a6-a04d-f27bddbc1d6f","order_by":3,"name":"Samadou Tchakondo","email":"","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Samadou","middleName":"","lastName":"Tchakondo","suffix":""},{"id":597453079,"identity":"c6f51d94-21dd-40b2-aecd-e0da835a74ca","order_by":4,"name":"Yendouname Kandjoni","email":"","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Yendouname","middleName":"","lastName":"Kandjoni","suffix":""},{"id":597453081,"identity":"267625e8-de22-4f0f-951d-1c92dc0980c5","order_by":5,"name":"Richard Sagacity Tugbeh","email":"","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Richard","middleName":"Sagacity","lastName":"Tugbeh","suffix":""},{"id":597453083,"identity":"abf379f0-6d21-4531-a2fe-40de9625bdff","order_by":6,"name":"Rachana Das","email":"","orcid":"","institution":"SRM Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Rachana","middleName":"","lastName":"Das","suffix":""}],"badges":[],"createdAt":"2026-01-30 06:24:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8737800/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8737800/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105354843,"identity":"d1a17c59-c885-4e99-b74a-082419d3496c","added_by":"auto","created_at":"2026-03-25 06:34:44","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":102479,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLandmark dataset construction process and final long-format dataset\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/30c68b6e08ee9b548724954b.png"},{"id":105354849,"identity":"7113edbf-78b3-4da1-8c3e-288bd362ad5f","added_by":"auto","created_at":"2026-03-25 06:34:45","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":44942,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe flowchart and framework of the prediction models\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/052602f00ab339ecdf598a31.png"},{"id":105565557,"identity":"7d0a0c04-42a6-4aa9-b9fa-ad40898218d2","added_by":"auto","created_at":"2026-03-27 12:53:35","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":219016,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFinal test ROC curves across ICU landmark times\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/cead7b161ac82729ba3037ac.png"},{"id":105565485,"identity":"6eb0e149-1c0e-44d8-a0db-e6c6fc006573","added_by":"auto","created_at":"2026-03-27 12:53:23","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":355199,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eForest plot of Adjusted Odd Ratios from the logistic regression model at the 18 hours: Primary approach\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/5c90722dffc63a3539c5ed8f.png"},{"id":105354844,"identity":"23f56b26-8c68-4270-adbc-0992b7376ff0","added_by":"auto","created_at":"2026-03-25 06:34:44","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":50985,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePermutation features importance from the random forest model at the 18 hours: Sensitivity analysis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/8ac3815398dcdd2456b491ab.png"},{"id":105570035,"identity":"a030d91f-228e-4f87-a2c5-3b40a56f60cb","added_by":"auto","created_at":"2026-03-27 13:14:10","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2496511,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/767ae761-060d-4f68-96d5-be314c60827e.pdf"},{"id":105354842,"identity":"e39afbb7-3c83-4e2b-ad67-11eeea42e8c9","added_by":"auto","created_at":"2026-03-25 06:34:44","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":24286,"visible":true,"origin":"","legend":"","description":"","filename":"missingnessheatmap.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/ce506cbb880bf558eebc9a65.pdf"},{"id":105354846,"identity":"ddf831b8-8020-482c-86b3-d79fe7e9b0bc","added_by":"auto","created_at":"2026-03-25 06:34:44","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":49196,"visible":true,"origin":"","legend":"","description":"","filename":"balancingensembleapproach.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/beb46deb39ebf20d423898f4.pdf"},{"id":105354848,"identity":"dac7f13c-651f-469f-ac43-15bad1ccdc0f","added_by":"auto","created_at":"2026-03-25 06:34:45","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":126363,"visible":true,"origin":"","legend":"","description":"","filename":"CalibrationCurveLM18Logistic.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/2b9c84b1b4c59971422085e3.png"},{"id":105354850,"identity":"2d3fb19a-e9d8-4dc1-befd-a44e5dde5826","added_by":"auto","created_at":"2026-03-25 06:34:45","extension":"png","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":116087,"visible":true,"origin":"","legend":"","description":"","filename":"CalibrationRandomForestLandmark18.png","url":"https://assets-eu.researchsquare.com/files/rs-8737800/v1/a30ea8c1227ca2ab706e9180.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eDynamic Landmark-Based Prediction of Sepsis Using Interpretable and Balanced Machine Learning Models in Respiratory-Supported Critically ill Patients\u003c/p\u003e","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSepsis is a life-threatening organ dysfunction resulting from a disordered host response to an infectious agent and is recognized as a significant global health concern. According to the Global Burden of Disease estimates, nearly 49\u0026nbsp;million cases of sepsis and 11\u0026nbsp;million deaths attributable to sepsis occur annually, which accounts for approximately one-fifth of all deaths globally [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. In addition to being a major cause of morbidity and mortality in the intensive care unit (ICU), sepsis also contributes to extended lengths of hospitalization, elevated healthcare costs, and increased risk of chronic complications [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Therefore, early recognition and timely initiation of antibiotics and supportive therapy are essential, as each hour of delay in treatment has been associated with a significant increase in mortality [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Although important advances have been made in critical care, the diverse and often non-specific presentation of sepsis in critically ill patients continues to make early detection of sepsis a persistent challenge for clinicians.\u003c/p\u003e \u003cp\u003eTraditionally, clinicians primarily relied upon severity scores (such as the sequential organ failure assessment (SOFA) and the quick SOFA (qSOFA)) to assist in assessing the risk of patients. While these scoring systems can provide valuable prognostic information, they are limited because they are based solely on static physiological thresholds and may be insensitive to the early stages of sepsis [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Furthermore, when using scoring systems (such as SOFA) for both defining outcomes and developing predictive models, they can result in circular reasoning and limit the discovery of new predictors independent of the established criteria [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. During the last decade, improvements in electronic health records (EHRs) and computational resources have enabled the development of machine learning (ML) approaches for the prediction of sepsis. ML approaches can analyze large amounts of complex and multi-dimensional clinical data and identify nonlinear associations between risk factors and outcomes. Multiple studies have demonstrated strong discriminative performance of ML models in ICU populations, with many studies reporting area under the receiver operating characteristic curve (AUROC) values of greater than 0.80 [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. However, systematic reviews have emphasized several serious shortcomings of ML models for the prediction of sepsis: the performance of ML models varies significantly between studies; the calibration of ML models is commonly poor; and prospective or external validation of ML models is uncommon [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThere are a number of significant methodological limitations to the majority of the previous literature concerning sepsis prediction models. First, many previously published sepsis prediction models utilized only static, or admission time data, thus failing to capture the changing physiological states of ICU patients over time [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Second, the vast majority of sepsis prediction models were developed in general ICU populations and there has been very limited focus on those patients that require respiratory support and are at a significantly higher risk of developing sepsis due to the need for invasive medical devices, compromised immune systems and increased susceptibility to nosocomial infections. Thirdly, the use of temporal modeling techniques to predict a patient\u0026rsquo;s risk of developing sepsis based upon new data collected over time is a novel area of research, but one that may be able to offer clinicians with more clinically relevant information regarding a patient\u0026rsquo;s risk of developing sepsis than static models.\u003c/p\u003e \u003cp\u003eA landmark analysis offers a clinically intuitive approach to predictively determining a patient\u0026rsquo;s risk of developing sepsis, by defining a series of predefined time points (landmarks) at which a patient\u0026rsquo;s risk will be estimated based upon all the data available up to that landmark [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Although the structural representation of each landmark model is static, by making predictions at each landmark, patient risk profiles can be continuously updated and therefore reflect changes in both physiological status and clinical context. By doing so, this strategy maintains temporal validity, prevents information leakage and aligns very well with how clinicians assess risk throughout the ICU care process [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn order to address the aforementioned limitations, our study applies a landmark-based dynamic prediction strategy for predicting short-term sepsis risk in critically ill patients who receive mechanical ventilation (respiratory) support. The risk estimations were made at 6, 12, 18 and 24 hours post-ICU admission using common statistical and machine learning models; logistic regression, random forest and gradient boosting. Logistic regression is included as an interpretable reference model, while ensemble-based methods are considered to examine whether nonlinear modeling strategies may offer advantages in the presence of complex clinical interactions and class imbalance. To understand differences in predictive discrimination, calibration, and clinical utility across the course of treatment, model performance is independently assessed at each landmark.\u003c/p\u003e \u003cp\u003eThe primary objective of this study was to develop and evaluate a dynamic landmark-based machine learning framework for early sepsis risk prediction in critically ill patients requiring respiratory support and to determine how the predictive performance, calibration, and clinical interpretability of the model evolved over the first 24 hours of critical care.\u003c/p\u003e \u003cp\u003eBy integrating landmark analysis with machine learning and rigorous evaluation strategies, our work seeks to advance the development of transparent, temporally valid, and clinically meaningful decision-support tools for early sepsis detection in critical care.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Design and Data Source\u003c/h2\u003e \u003cp\u003eA secondary data analysis was conducted using the publicly available, de-identified Medical Information Mart for Intensive Care IV (MIMIC-IV) database [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The MIMIC-IV is a large public database containing highly detailed, time-stamped information regarding adult patients admitted to ICUs, which includes physiological measurements, clinical interventions and patient outcomes. All analyses were performed in accordance with ethical standards for secondary analysis of de-identified data; institutional review board approval and informed consent were not required.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eStudy Population and Cohort Construction\u003c/h3\u003e\n\u003cp\u003eThe study population consisted of adult patients (\u0026gt;\u0026thinsp;18 years) who were admitted to an ICU and received some form of respiratory support, including mechanical ventilation (invasive or non-invasive), or high-flow oxygen therapy for at least one of the first 24 hours post-ICU admission. In order to generate longitudinal data sufficient to create landmark-specific feature windows, only patients with such data were included.\u003c/p\u003e \u003cp\u003eTo construct a valid at-risk cohort and prevent temporal leakage, patients were excluded from a given landmark and all subsequent landmarks if sepsis occurred prior to that landmark time. Data censoring was performed at either ICU discharge or at the time of a patient's death, and no landmark observations were created after these events. Landmark observations with insufficient data within the predefined feature window were excluded.\u003c/p\u003e \u003cp\u003eThe final analytic database was structured in a long format; with each row representing a single patient-landmark observation associated with a unique prediction time point (see Fig.\u0026nbsp;1). Patients were permitted to contribute observations to multiple landmarks provided they were still alive in the ICU and sepsis-free at each corresponding landmark.\u003c/p\u003e \u003cp\u003eThe overall study design, cohort selection process, landmark-based dataset construction, and prediction modeling framework are illustrated in Fig.\u0026nbsp;2.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eLandmark Analysis Framework\u003c/h3\u003e\n\u003cp\u003eA landmark analysis refers to the practice of designating a time point occurring during the follow-up period (known as the \u003cem\u003elandmark time\u003c/em\u003e) and analyzing only those subjects who have survived until the landmark time [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. A comprehensive overview of the landmark analysis method and its use has been provided by Dafni [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. A landmarking strategy was employed to enable dynamic risk prediction at clinically meaningful time points following ICU admission. Our landmarks were prespecified at 6, 12, 18, and 24 hours after ICU admission. Each landmark constituted an independent prediction task with its own risk set, including only patients who were alive in the ICU and free of sepsis at that time.\u003c/p\u003e \u003cp\u003eAlthough each landmark model was static in form, allowing patients to contribute repeated observations across landmarks enabled implicit modeling of patient trajectories through repeated risk re-estimation over time. This approach preserved temporal validity, avoided information leakage, and aligned with routine clinical reassessment practices in the ICU.\u003c/p\u003e\n\u003ch3\u003eTemporal Alignment of Predictors and Outcomes\u003c/h3\u003e\n\u003cp\u003eFor each landmark, predictors were extracted from a fixed 6-hour feature window immediately preceding the landmark time:\u003c/p\u003e \u003cp\u003e[\u003cem\u003et\u003c/em\u003e\u003csub\u003elandmark\u003c/sub\u003e \u0026ndash; 6, \u003cem\u003et\u003c/em\u003e\u003csub\u003elandmark\u003c/sub\u003e)\u003c/p\u003e \u003cp\u003eThe outcome was defined as the occurrence of sepsis within the subsequent 6-hour prediction window:\u003c/p\u003e \u003cp\u003e(\u003cem\u003et\u003c/em\u003e\u003csub\u003elandmark\u003c/sub\u003e, \u003cem\u003et\u003c/em\u003e\u003csub\u003elandmark\u003c/sub\u003e + 6]\u003c/p\u003e \u003cp\u003eThis strict temporal alignment ensured that all predictors were observed prior to outcome assessment and eliminated forward-looking bias.\u003c/p\u003e\n\u003ch3\u003eOutcome Definition\u003c/h3\u003e\n\u003cp\u003eThe primary outcome, \u003cem\u003esepsis_next_6h\u003c/em\u003e, was defined as the occurrence of sepsis within 6 hours following each landmark, based on Sepsis-3 criteria implemented in the dataset [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Patients with any evidence of sepsis prior to a given landmark were excluded from that landmark\u0026rsquo;s risk set to ensure that only incident sepsis events were modelled.\u003c/p\u003e \u003cp\u003eAt each landmark, control observations were defined as patient\u0026ndash;landmark instances in which sepsis did not occur within the subsequent 6-hour prediction window. The composite organ dysfunction scores used to define sepsis (e.g., SOFA) were explicitly excluded from the predictor set to avoid circularity and artificial inflation of model performance [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003ePredictor Engineering and Selection\u003c/h2\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003ePhysiological Variables\u003c/h2\u003e \u003cp\u003ePredictors were selected a priori based on clinical relevance, interpretability, and availability prior to each landmark. Only information observed within the pre-landmark feature window was used. Physiological instability was summarized using clinically meaningful extreme values, which better reflect acute deterioration in critically ill patients than averages. Selected variables included:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eMaximum heart rate\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMaximum respiratory rate\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMinimum oxygen saturation\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMaximum temperature\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMinimum Glasgow Coma Scale (GCS) score\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e\n\u003ch3\u003eBlood Pressure Representation\u003c/h3\u003e\n\u003cp\u003eNon-invasive systolic and diastolic blood pressure measurements were used to derive mean arterial pressure (MAP) using the standard formula: MAP = (SBP\u0026thinsp;+\u0026thinsp;2 \u0026times; DBP) / 3 [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMAP values were summarized using their mean within each landmark feature window and retained as continuous variables due to their relevance to tissue perfusion and hemodynamic stability. Explicit missingness indicators were incorporated to capture clinically informative measurement patterns related to illness severity and care processes.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eOrgan Support Variables\u003c/h2\u003e \u003cp\u003eVariables representing markers of organ support included vasopressor use, continuous renal replacement therapy (CRRT), invasive mechanical ventilation, non-invasive ventilation, and high-flow oxygen therapy. These variables were represented as proportions of exposure time within each feature window. Organ support variables were interpreted as indicators of increasing illness severity rather than as predictive factors of sepsis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eStatic Covariates\u003c/h2\u003e \u003cp\u003ePatient characteristics collected at baseline included age at ICU admission, sex, race, and the Elixhauser\u0026ndash;van Walraven comorbidity index, which measures comorbidity burden. Initial consideration was given to including body mass index (BMI) but it was ultimately removed from the final models because of excessive missingness.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eMissing Data Handling\u003c/h2\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003eExploratory Missingness Assessment\u003c/h2\u003e \u003cp\u003ePatterns of missing data were explored in total and stratified by landmark time (as shown in Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e of Additional file 1). Most physiological variables had relatively small amounts of missing data (less than 10%), while some of the non-invasive blood pressure variables were moderately missing (about 20%) and BMI had large amounts of missing data (more than 50%). All organ support variables were complete by construction.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eRow-Level Exclusions\u003c/h2\u003e \u003cp\u003ePrior to imputation, the following exclusions were applied:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eObservations in which the outcome or landmark time were missing\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eObservations in which all of the core physiological variables were missing\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eObservations in which greater than 50% of the selected predictors were missing\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eImputation Strategy\u003c/h2\u003e \u003cp\u003eLandmark-aware imputation strategy was employed in order to maintain temporal relationships and avoid leaking information. Time-varying physiological variables were imputed by the median value for the landmark time in which they were measured. There was no imputation of organ support variables, categorical variables, or the outcome. Binary missingness indicators were also created for GCS and non-invasive blood pressure variables.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eEvents-Per-Variable (EPV) Assessment\u003c/h2\u003e \u003cp\u003eBefore beginning model development, the number of sepsis events at each landmark was examined relative to the number of candidate predictors in order to evaluate the number of events per variable (EPV). EPV was sufficient to justify the selected model specifications and guided the model complexity and regularization to reduce the risk of overfitting.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eModel Development and Evaluation\u003c/h2\u003e \u003cp\u003eTo ensure that model development is robust and maintains clinical realism, a standard two-stage training and evaluation paradigm was defined as the main analytical methodology. In addition to this, a balanced ensemble methodology was developed and employed as a sensitivity analysis to determine the impact of different methods of managing class imbalance on model performance. Both methodologies were applied uniformly to each landmark time point (6, 12, 18, and 24 hours) and each candidate model.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eModeling Approaches\u003c/h2\u003e \u003cdiv id=\"Sec20\" class=\"Section3\"\u003e \u003ch2\u003eApproach 1: Standard Two-Stage Training and Evaluation\u003c/h2\u003e \u003cp\u003eIn the primary analysis, a conventional two-stage machine learning pipeline was implemented. For each landmark and model type, data were split at the patient level into training and independent test sets to prevent information leakage across repeated landmark observations. Hyperparameters were optimized using cross-validation (5 folds) within the training set, and final performance was evaluated on the held-out test set. Class imbalance was handled using class-weighted loss functions, preserving the natural incidence of sepsis and maintaining clinically meaningful calibration. This approach reflects real-world deployment where models are trained on imbalanced data and applied to unseen patients.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eApproach 2: Balanced Ensemble Strategy (Sensitivity Analysis)\u003c/h2\u003e \u003cp\u003eTo assess robustness to alternative imbalance handling, a sensitivity analysis using a balanced ensemble strategy adapted from prior early-prediction studies [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] was conducted. For each landmark (6, 12, 18, and 24 hours), non-sepsis observations were partitioned into multiple mutually exclusive subgroups via stratified random sampling. The number of subgroups was predetermined based upon the degree of class imbalance at each landmark; the number of subgroups increased as the number of sepsis events decreased at later landmarks:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e6-hour: 20 subgroups\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e12-hour: 30 subgroups\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e18-hour: 40 subgroups\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e24-hour: 50 subgroups\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eNon-sepsis cases from each subgroup were paired with the total number of sepsis cases at the corresponding landmark window to create a balanced training dataset (as shown in Fig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e of Additional file 2). For each balanced dataset, an independent model was trained using the same predictors and preprocessing as in the primary approach. The ensemble predictions were averaged across all models for each landmark time point. This strategy allowed evaluation of whether conclusions regarding optimal landmark timing and comparative model performance were sensitive to subsampling-based imbalance correction.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eCandidate Models\u003c/h2\u003e \u003cp\u003eThree modeling algorithms were independently evaluated at each landmark time using both methodologies:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eLogistic Regression\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eXGBoost Gradient Boosting\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eAll models were trained using the same feature sets and temporal feature definitions.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003ePerformance Evaluation\u003c/h2\u003e \u003cp\u003ePerformance of a model was evaluated on test data from each landmark and each method of analysis. Discriminative ability was evaluated using both AUROC and AUPRC. Additionally, clinical utility was evaluated through calculation of sensitivity, specificity and accuracy. Metrics were reported separately by landmark, model type, and modeling approach.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eModel Selection Strategy\u003c/h2\u003e \u003cp\u003eFor each landmark, models were compared within each analytical approach. The best performing model was selected based upon discrimination and clinical utility metrics, with an emphasis on AUROC, AUPRC, sensitivity and specificity. Robustness of results across primary and sensitivity analyses were evaluated to assess agreement and consistency of results.\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003eRole of the Dual-Approach Framework\u003c/h2\u003e \u003cp\u003eThe dual-strategy approach provided the opportunity to evaluate the predictive stability of the models under different class imbalance assumptions. The consistent results obtained across the different approaches provided additional confidence that the results can be generalized, while inconsistencies provide caution when interpreting the results.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003eModel Interpretability\u003c/h2\u003e \u003cp\u003eInterpretability was evaluated for the best-performing model at the most clinically informative landmark. For tree-based models, permutation feature importance was used to quantify the impact of each predictor on model performance. For logistic regression, associations were expressed as odds ratios, enhancing transparency in identifying key predictors of sepsis risk.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section3\"\u003e \u003ch2\u003eSoftware\u003c/h2\u003e \u003cp\u003eAll analyses were conducted using Python (version 3.11). Data processing and preprocessing were performed using pandas and NumPy. Machine-learning modeling and evaluation were implemented using Scikit-learn, XGBoost, and Joblib. Model interpretability analyses were conducted using SHAP and permutation importance methods. Visualizations were generated using Matplotlib and Seaborn.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec29\" class=\"Section2\"\u003e \u003ch2\u003ePatient characteristics\u003c/h2\u003e \u003cp\u003eA total of 50,920 adult ICU admissions requiring respiratory support were initially identified from the MIMIC-IV v2.2 database. After applying the predefined inclusion and exclusion criteria, 41,871 patients were retained in the 6-hour landmark risk set. As landmark-specific eligibility criteria were sequentially applied, cohort sizes decreased to 39,912, 36,472, and 31,367 patients at the 12, 18, and 24 hour landmarks, respectively.\u003c/p\u003e \u003cp\u003eThe Table\u0026nbsp;1 summarizes the demographic characteristics, physiological measurements, comorbidity burden, and treatment-related features of patients included at each landmark time point. All characteristics are reported descriptively.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBaseline characteristics of the study cohort at each landmark time\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVariable\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLandmark 6 h (n\u0026thinsp;=\u0026thinsp;41,871)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLandmark 12 h (n\u0026thinsp;=\u0026thinsp;39,912)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLandmark 18 h (n\u0026thinsp;=\u0026thinsp;36,472)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eLandmark 24 h (n\u0026thinsp;=\u0026thinsp;31,367)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAge\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e65 (53\u0026ndash;77)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e65 (53\u0026ndash;77)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e65 (53\u0026ndash;77)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e66 (53\u0026ndash;77)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eHeart rate, max\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e90 (79.5\u0026ndash;104.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e89 (78\u0026ndash;102)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e89 (78\u0026ndash;102)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e90 (79\u0026ndash;103)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSpO₂, min\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e95.5 (93\u0026ndash;98)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e95 (93\u0026ndash;97)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e95 (93\u0026ndash;97)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e95 (93\u0026ndash;97)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTemperature, max\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e36.9 (36.6\u0026ndash;37.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e36.9 (36.7\u0026ndash;37.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36.9 (36.7\u0026ndash;37.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e36.9 (36.7\u0026ndash;37.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eRespiratory rate, max\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e22 (19\u0026ndash;26)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e22 (19\u0026ndash;26)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e22 (19\u0026ndash;26)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e23 (20\u0026ndash;27)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMAP, non-invasive\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e83.7 (77.2\u0026ndash;91.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e81.9 (76.0\u0026ndash;88.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e82.2 (76.6\u0026ndash;88.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e82.1 (76.1\u0026ndash;88.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSex\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFemale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e19,030 (45.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e18,157 (45.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16,543 (45.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14,088 (44.9)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e22,841 (54.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e21,755 (54.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e19,929 (54.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17,279 (55.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eRace/ethnicity\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAsian\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,266 (3.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,210 (3.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1,111 (3.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e947 (3.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBlack\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4,159 (9.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3,974 (10.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3,555 (9.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3,011 (9.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHispanic/Latino\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,646 (3.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,585 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1,435 (3.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1,242 (4.0)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNative American\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e75 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e73 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e69 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e61 (0.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOther/Unknown\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e6,527 (15.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6,138 (15.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5,640 (15.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4,968 (15.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePacific Islander\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e66 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e62 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e54 (0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e44 (0.1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWhite\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e28,132 (67.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26,870 (67.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e24,608 (67.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21,094 (67.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eVasopressor use\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e36,064 (86.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e34,014 (85.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e31,427 (86.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e27,241 (86.8)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5,807 (13.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5,898 (14.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5,045 (13.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4,126 (13.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eContinuous renal replacement therapy\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e41,835 (99.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e39,813 (99.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36,329 (99.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e31,184 (99.4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e36 (0.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e99 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e143 (0.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e183 (0.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInvasive ventilation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e33,220 (79.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31,935 (80.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e30,370 (83.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e26,475 (84.4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8,651 (20.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7,977 (20.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6,102 (16.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4,892 (15.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eNon-invasive ventilation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e41,583 (99.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e39,661 (99.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36,259 (99.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e31,182 (99.4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e288 (0.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e251 (0.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e213 (0.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e185 (0.6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eHigh-flow oxygen\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e41,705 (99.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e39,713 (99.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e36,273 (99.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e31,156 (99.3)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e166 (0.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e199 (0.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e199 (0.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e211 (0.7)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGlasgow Coma Scale\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSevere (cat_1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,381 (3.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e787 (2.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e550 (1.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e518 (1.7)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModerate (cat_2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,907 (4.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,554 (3.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1,366 (3.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1,219 (3.9)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMild (cat_3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e38,583 (92.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e37,571 (94.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e34,556 (94.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e29,630 (94.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eElixhauser comorbidity\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLow (cat_1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16,159 (38.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15,279 (38.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e13,693 (37.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e11,352 (36.2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModerate (cat_2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16,208 (38.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15,516 (38.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e14,263 (39.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e12,404 (39.5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHigh (cat_3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e9,504 (22.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9,117 (22.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e8,516 (23.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7,611 (24.3)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAcross the landmark risk sets, the incidence of sepsis decreased progressively with increasing landmark time. Among patients included in the 6-hour landmark cohort, 621 sepsis events were observed among 41,871 patients, corresponding to an incidence of 1.48%.\u003c/p\u003e \u003cp\u003eAt subsequent landmarks, sepsis incidence declined to 0.55% (220 of 39,912) at 12 hours, 0.45% (164 of 36,472) at 18 hours, and 0.37% (117 of 31,367) at 24 hours (Table\u0026nbsp;2). This reduction reflects both attrition of higher-risk patients over time and the exclusion of patients who developed sepsis prior to each landmark.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eIncidence of Sepsis Across Landmark Windows\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eICU landmark (hours)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePatients at risk, n\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSepsis events, n (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEPV\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e41,871\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e621 (1.48)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e36.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e39,912\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e220 (0.55)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e12.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e36,472\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e164 (0.45)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e31,367\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e117 (0.37)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTraining and evaluation by Landmark\u003c/h3\u003e\n\u003cp\u003eThe model performance was evaluated separately at each landmark time using a standard two-stage training and evaluation framework. Cross-validated discrimination metrics were first estimated within the training data, followed by final performance assessment on an independent held-out test set.\u003c/p\u003e \u003cdiv id=\"Sec31\" class=\"Section2\"\u003e \u003ch2\u003eTwo stage classic approach\u003c/h2\u003e \u003cdiv id=\"Sec32\" class=\"Section3\"\u003e \u003ch2\u003eCross-validated model performance\u003c/h2\u003e \u003cp\u003eCross-validated discrimination varied across landmark times and modeling approaches (Table\u0026nbsp;3). Logistic regression demonstrated more balanced sensitivity and specificity across landmarks, whereas ensemble models consistently favoured high specificity at the expense of sensitivity.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCross-validated discrimination performance across ICU landmark times in the primary approach\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ehr\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003emodel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAUROC_\u003c/p\u003e \u003cp\u003emean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAUROC_\u003c/p\u003e \u003cp\u003esd\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUPRC_\u003c/p\u003e \u003cp\u003emean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSensitivity_\u003c/p\u003e \u003cp\u003emean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eSpecificity_\u003c/p\u003e \u003cp\u003emean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003en_patients\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003en_events\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLogistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e31403\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e465\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandomForest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e31403\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e465\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e31403\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e465\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLogistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e29920\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e174\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandomForest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e29920\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e174\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e29920\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e174\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLogistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e27386\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e121\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandomForest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e27386\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e121\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e27386\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e121\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLogistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e23604\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandomForest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e23604\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e23604\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec33\" class=\"Section4\"\u003e \u003ch2\u003eFinal test set performance\u003c/h2\u003e \u003cp\u003eThe final test set evaluation confirmed these trends (Table\u0026nbsp;4). At the 18-hour landmark, logistic regression achieved the highest overall discrimination (Fig.\u0026nbsp;3), with improved sensitivity and comparable specificity relative to other landmarks.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cb\u003eFinal test set performance of selected sepsis prediction models across ICU landmark times\u003c/b\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ehr\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003emodel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAUROC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAUPRC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003en_patients\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003en_events\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e10468\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e156\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLogistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e9992\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e18\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eLogistic\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.78\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.02\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.68\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.77\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.68\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e9086\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e43\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLogistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.718\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e7763\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec34\" class=\"Section3\"\u003e \u003ch2\u003eBalanced ensemble models\u003c/h2\u003e \u003cp\u003eSensitivity analyses using balanced ensemble models supported the robustness of the primary findings. Across landmarks, the 18-hour window consistently demonstrated the most favourable predictive performance, with the random forest ensemble achieving the best balance between discrimination (AUROC 0.77) and clinically relevant sensitivity while preserving acceptable specificity compared to others (Table\u0026nbsp;5). Although the optimal algorithm varied across landmarks in the sensitivity analysis, the relative ranking of landmark windows and the identification of 18 hours as the optimal prediction time remained unchanged, confirming the stability of the primary conclusions under alternative class-imbalance handling strategies.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel performance using the balanced ensemble approach (sensitivity analysis)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"10\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ehr\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAUROC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAUPRC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.02\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e18\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.77\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.02\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.65\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.74\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.01\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e0.65\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003e0.02\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e\n\u003ch3\u003eCalibration Performance\u003c/h3\u003e\n\u003cp\u003eCalibration was assessed for the best-performing model at each landmark using calibration plots and Brier scores. At the 18-hour landmarking the primary approach, the logistic regression model showed reasonable agreement between predicted and observed sepsis risk (Brier score\u0026thinsp;=\u0026thinsp;0.213), with mild overprediction at higher risk levels and good alignment in the low-risk range where most patients clustered. Other landmarks showed similar patterns, with greater instability at later times due to fewer events.\u003c/p\u003e \u003cp\u003eSensitivity analysis of balanced ensemble models indicated improved discrimination at some landmarks but more variable calibration, reflecting resampling and aggregation across sub-models (as shown in Fig. \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e of Additional file 3). Overall, findings support using the 18-hour model for risk stratification while emphasizing caution in interpreting absolute risk estimates in highly imbalanced settings.\u003c/p\u003e\n\u003ch3\u003eModel Interpretability\u003c/h3\u003e\n\u003cp\u003eTo enhance clinical interpretability of the selected models, feature importance analyses were conducted for the best-performing models identified at the 18-hour landmark in both the primary and sensitivity analyses.\u003c/p\u003e \u003cp\u003eFor the primary analysis, interpretability of the logistic regression model was assessed using a forest plot of adjusted odds ratios with 95% confidence intervals (Fig.\u0026nbsp;4.). This analysis highlighted physiologically plausible predictors of near-term sepsis, including markers of respiratory compromise, hemodynamic instability, and neurological status, with consistent directions of effect across covariates. The magnitude and direction of associations were clinically coherent, supporting the transparency and interpretability of the linear model.\u003c/p\u003e \u003cp\u003eIn the sensitivity analysis, permutation-based feature importance was derived for the random forest model selected at the 18-hour landmark (Fig.\u0026nbsp;5.). This approach identified a similar set of high-impact features, particularly vital sign extremes and indicators of organ support, suggesting concordance between linear and non-linear modeling approaches despite algorithmic difference.\u003c/p\u003e \u003cp\u003eTogether, these findings demonstrate that the predictive signal driving model performance was clinically meaningful and robust across modeling strategies. Interpretability analyses therefore reinforced the validity of the 18-hour landmark as a clinically actionable time point and supported the reliability of the identified predictors for early sepsis risk stratification.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cdiv id=\"Sec38\" class=\"Section2\"\u003e \u003ch2\u003ePrincipal findings\u003c/h2\u003e \u003cp\u003eIn this secondary data analysis of respiratory-supported ICU patients, dynamic machine-learning models were developed and evaluated to predict sepsis within a 6-hour horizon using a landmarking framework. Several important findings emerged. First, the incidence of sepsis declined across successive landmark windows, reflecting both clinical evolution and exclusion of earlier events. Second, model performance varied by landmark time, with the 18-hour landmark consistently demonstrating the most favourable balance between discrimination and clinically relevant operating characteristics. Third, although the best performing algorithms were different for each strategy (logistic regression in the primary analysis and random forest in the sensitivity analysis), the fact that the two studies converged upon the same landmark time point suggests that the predictive performance of the models was based on the temporal risk structure of sepsis development and not the specific classification method used. Finally, the calibration was performed and the key features were clinically interpretable and reflected respiratory, hemodynamic and neurological dysfunction that is consistent with the well-established pathophysiology of sepsis [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec39\" class=\"Section2\"\u003e \u003ch2\u003eComparison with existing literature\u003c/h2\u003e \u003cp\u003eEarly identification of sepsis in critically ill patients remains challenging due to heterogeneous clinical presentation and evolving physiology [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Research has shown that machine learning can provide effective methods for predicting sepsis in the ICU population, however, most of these studies utilize either static baseline data or continuous real-time monitoring without explicit temporal risk stratification [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Fleuren et al. in their systematic review and meta-analysis reported an average AUROC range from 0.68 to 0.99 for all ICU-based sepsis models [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] but they also found variability in calibration among the reviewed models as well as inconsistency in the reporting of the F1 measure, often leading to inflated impressions of clinical utility. In contrast to previous studies reported by Fleuren et al. showing near-perfect discrimination, our model achieved a more moderate AUROC of 0.78 when evaluated under real-world class imbalance. This finding suggests that discrimination alone may be insufficient to determine the clinical utility of a predictive model, particularly in settings where sepsis incidence is low and false alarms carry significant consequences.\u003c/p\u003e \u003cp\u003eBeyond overall discrimination in our analysis we found respiratory compromise, hemodynamic instability, temperature and neurological dysfunction as key contributors to predicting sepsis. These findings are consistent with prior interpretable machine learning studies by Nemati et al. [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] and Tang et al. [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], which highlighted oxygen saturation, respiratory rate, and temperature among the most influential predictors. The alignment of our model\u0026rsquo;s important features with established physiological markers of sepsis supports the biological plausibility and clinical relevance of our approach.\u003c/p\u003e \u003cp\u003eThe inclusion of Elixhauser comorbidity [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] as a strong contributor further supports the relevance of chronic disease burden in predisposing critically ill patients to infection-related organ dysfunction. This convergence across studies strengthens the face validity and clinical credibility of our model.\u003c/p\u003e \u003cp\u003eSimilar landmarking strategies have shown promise in other critical care contexts, including early infection detection and mortality risk modeling, supporting the relevance of temporally updated prediction frameworks [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Notably, our findings align with prior work showing improved predictive stability when models account for evolving patient trajectories rather than single-time-point measurements [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e \u003cdiv id=\"Sec40\" class=\"Section3\"\u003e \u003ch2\u003eClinical implications\u003c/h2\u003e \u003cp\u003eThe identification of the 18-hour landmark as the optimal prediction window has potential clinical relevance. At this time point, models achieved a favourable trade-off between discrimination, sensitivity, and specificity, suggesting that risk stratification may be most informative once initial stabilization and early therapeutic interventions have occurred [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Furthermore, the consistency of this landmark across primary and sensitivity analyses underscores the robustness of this temporal signal. Interpretability analyses highlighted clinically plausible predictors, including markers of respiratory compromise, presence of comorbidities, hemodynamic instability, and altered neurological status, reinforcing the face validity of the models and supporting their potential role in clinical decision support [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cb\u003eRobustness and sensitivity analyses\u003c/b\u003e \u003c/p\u003e \u003cp\u003eIn addition to the issues caused by large amounts of imbalanced data for predicting sepsis, a balanced ensemble modeling method was used as a sensitivity test, given the large degree of class imbalance [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] as a sensitivity analysis. Although different models were selected at each landmark, the results from all three landmarks showed generally similar clinical performance characteristics and discrimination relative to the primary analysis. Additionally, both methods found the 18-hour landmark to be the best time to predict sepsis, this further emphasizes the consistency of the temporal trend seen in the previous findings. These findings suggest that the predictive signal was not an artifact of sampling strategy or algorithm choice, but rather reflects underlying clinical risk dynamics.\u003c/p\u003e \u003cp\u003eOur research is strong from both clinical and methodological perspectives. First, we used a very large, fully characterized critical care data set, which had high resolution, clinically relevant data [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. The use of a dynamic landmarking model was able to capture the changing nature of sepsis risk in real-time, as opposed to the static predictions provided by many other models, and it aligned with recommendations regarding the application of time-sensitive machine learning in the field of critical care [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. To minimize the risk of information leakage and model over-fitting, we employed two methods of internal validation (patient level splitting and group based cross-validation), which are common among predictive modeling applications [\u003cspan additionalcitationids=\"CR33\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Additionally, we applied two types of interpretability analysis (forest plot and permutation feature importance) to increase the clarity of how clinical factors contribute to an individual's sepsis risk and to build clinician confidence through increasing their understanding of how these models\u0026rsquo; function [\u003cspan additionalcitationids=\"CR33 CR34\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSeveral limitations should be acknowledged. The retrospective design limits causal inference and may introduce selection bias. Consistent with systematic reviews in this field, many machine learning sepsis prediction models including ours are developed on single databases, which restricts generalizability and highlights the need for independent validation across diverse settings [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Although we employed standard techniques to mitigate class imbalance, the low incidence of sepsis constrained precision and positive predictive value, reflecting a common challenge in imbalanced clinical prediction tasks [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Additionally, certain potentially informative biomarkers (e.g., lactate, inflammatory markers) were unavailable or inconsistently measured in our dataset, which may have limited predictive performance compared with models incorporating broader laboratory data [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. Finally, integration into real-time clinical workflows and prospective evaluation remain necessary before practical deployment.\u003c/p\u003e \u003cp\u003eFuture work should prioritize external validation in multicenter and multinational cohorts to assess robustness and transportability, as underscored in recent meta-analyses [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Efforts to improve probability calibration and threshold optimization for clinically actionable decision thresholds are essential before implementation in decision-support systems [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Exploring adaptive updating strategies that incorporate additional longitudinal features and real-time data streams may further enhance predictive performance and clinical utility.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eA dynamic landmark-based machine learning approach to predict early sepsis in respiratory supported ICU patients was developed from routinely collected EHR data. The ability to update risk at clinically relevant times resulted in better capture of a patient's trajectory over time, and demonstrated that the 18 hour landmark provided the best trade-off between discrimination and clinical applicability. Model interpretability demonstrated reliance on physiological and logical indicators of respiratory, hemodynamic, and neurological dysfunction to support clinical relevance.\u003c/p\u003e \u003cp\u003eAlthough good discrimination was achieved in our model, low precision due to a large degree of class imbalance in the real world indicates that additional recalibration of probabilities, optimization of thresholds, and additional validation of the models externally is necessary before implementation as a clinical tool. In general, dynamic landmark-based prediction appears to be a promising and clinically appropriate strategy for sepsis risk stratification in critically ill patients.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial Intelligence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAUC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArea Under the Curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAUROC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArea Under the Receiver Operating Characteristic Curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAUPRC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArea Under the Precision-Recall Curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eBIDMC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eBeth Israel Deaconess Medical Center\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eConfidence Interval\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCV\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCross-Validation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCITI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCollaborative Institutional Training Initiative\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCRRT\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eContinuous Renal Replacement Therapy\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eEHR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eElectronic Health Record\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eEPV\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eEvents Per Variable\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eICU\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIntensive Care Unit\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLogistic Regression\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eML\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMachine Learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMIMIC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eIV-Medical Information Mart for Intensive Care IV\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eROC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eReceiver Operating Characteristic\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eqSOFA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eQuick Sequential Organ Failure Assessment\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSHAP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSHapley Additive exPlanations\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSOFA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSequential Organ Failure Assessment\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSpO₂\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ePeripheral Capillary Oxygen Saturation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eXGB\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eExtreme Gradient Boosting (XGBoost)\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors received no specific funding for this work.\u003c/p\u003e\u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003eThe MIMIC-IV database is a publicly available, de-identified critical care database. Access was granted after completion of the required data use training and credentialing process. The use of MIMIC-IV data is approved by the institutional review boards of the Massachusetts Institute of Technology and Beth Israel Deaconess Medical Center, and informed consent was waived due to the de-identified nature of the dataset.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication\u003c/strong\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eASA conceptualized the study, conducted the analysis, and drafted the manuscript. JHG supervised the project, contributed to the study design, and critically revised the manuscript. KSG, ST, YK, and RST assisted with data processing, statistical interpretation, and manuscript review. RD contributed to study design, methodology, and manuscript revision. All authors read and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003e The authors gratefully acknowledge the support provided by the SRM School of Public Health, Faculty of Medicine and Health Sciences, SRM Institute of Science and Technology (SRMIST), Kattankulathur. We also thank the PhysioNet team for maintaining open access to the MIMIC-IV database, which made this research possible.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe data that support the findings of this study are available from the MIMIC-IV database (PhysioNet) and require credentialed access and completion of a data use agreement.The full analysis pipeline, including cohort extraction, landmark dataset construction, feature engineering, model development, and evaluation scripts, is publicly available on GitHub at:[https://github.com/Sangenis11/sepsis-landmark-prediction](https:/github.com/Sangenis11/sepsis-landmark-prediction) .The repository contains all code necessary to reproduce the analyses, excluding patient-level data in accordance with the MIMIC data use agreement.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eRudd KE, Johnson SC, Agesa KM, et al. Global, regional, and national sepsis incidence and mortality, 1990\u0026ndash;2017: analysis for the Global Burden of Disease Study. Lancet. 2020;395(10219):200\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinger M, Deutschman CS, Seymour CW, et al. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA. 2016;315(8):801\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeymour CW, Gesten F, Prescott HC, et al. Time to treatment and mortality during mandated emergency care for sepsis. N Engl J Med. 2017;376(23):2235\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRaith EP, Udy AA, Bailey M, et al. Prognostic accuracy of SOFA score, SIRS criteria, and qSOFA score for in-hospital mortality among adults with suspected infection admitted to the intensive care unit. JAMA. 2017;317(3):290\u0026ndash;300.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLindner HA, Schamoni S, Kirschning T, et al. Ground truth labels challenge the validity of sepsis consensus definitions in critical illness. Crit Care. 2022;26:350.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFleuren LM, Klausch TLT, Zwager CL, et al. Machine learning for the prediction of sepsis: a systematic review and meta-analysis of diagnostic test accuracy. Intensive Care Med. 2020;46(3):383\u0026ndash;400.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoor M, Rieck B, Horn M, et al. Early prediction of sepsis in the ICU using machine learning: a systematic review. Front Med (Lausanne). 2021;8:607952.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Z, Luo L, Song D, et al. Early sepsis mortality prediction model based on interpretable machine learning: development and external validation. BMC Med Inf Decis Mak. 2024;24(1):46.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKomorowski M, Celi LA, Badawi O, et al. The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nat Med. 2018;24(11):1716\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarrett JK, Sweeting MJ, Wood AM. Dynamic risk prediction for cardiovascular disease: an illustration using the ARIC Study. In: Rao ASR, Pyne S, Rao CR, editors. Handbook of Statistics. Volume 36. Amsterdam: Elsevier; 2017. pp. 47\u0026ndash;65.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRizopoulos D, Molenberghs G, Lesaffre EMEH. Dynamic predictions with time-dependent covariates in survival analysis using joint modeling and landmarking. Biom J. 2017;59(6):1261\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoukheiber M, Moukheiber L, Moukheiber D, Hao S, Celi LA, Lee H. A temporal dataset for respiratory support in critically ill patients (version 1.1.0). PhysioNet. 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.13026/wewp-sj67\u003c/span\u003e\u003cspan address=\"10.13026/wewp-sj67\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnderson JR, Cain KC, Gelber RD. Analysis of survival by tumor response. J Clin Oncol. 1983;1:710\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDafni U. Landmark analysis at the 25-year landmark point. Circ Cardiovasc Qual Outcomes. 2011;4:363\u0026ndash;71.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeMers D, Wachs D. Physiology, mean arterial pressure. In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing; 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/books/NBK538226/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/books/NBK538226/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang Y, Zhu C, Tian C, et al. Early prediction of ventilator-associated pneumonia in critical care patients: a machine learning model. BMC Pulm Med. 2022;22:250.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. 2019;1:206\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShashikumar SP, Stanley MD, Sadiq I, et al. Early sepsis detection in critical care patients using multiscale physiological signals. J Electrocardiol. 2017;50(6):739\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHenry KE, Hager DN, Pronovost PJ, Saria S. A targeted real-time early warning score (TREWScore) for septic shock. Sci Transl Med. 2015;7(299):299ra122.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDesautels T, Calvert J, Hoffman J, et al. Prediction of sepsis in the ICU with minimal electronic health record data. JMIR Med Inf. 2016;4(3):e28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNemati S, Holder A, Razmi F, et al. An interpretable machine learning model for accurate prediction of sepsis in the ICU. Crit Care Med. 2018;46(4):547\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang J, Li J, Luo X, et al. Interpretable machine learning-based prediction of 28-day mortality in ICU sepsis patients. Front Public Health. 2024;12:1349928.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eElixhauser A, Steiner C, Harris DR, Coffey RM. Comorbidity measures for use with administrative data. Med Care. 1998;36(1):8\u0026ndash;27.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Houwelingen HC. Dynamic prediction by landmarking in event history analysis. Scand J Stat. 2007;34(1):70\u0026ndash;85.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVarkila MRJ, Lancia G, van Smeden M, et al. Early detection of ICU-acquired infections using high-frequency electronic health record data. BMC Med Inf Decis Mak. 2025;25:273. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12911-025-03031-6\u003c/span\u003e\u003cspan address=\"10.1186/s12911-025-03031-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Houwelingen HC, Putter H. Dynamic Prediction in Clinical Survival Analysis. Boca Raton: CRC; 2012.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSendak MP, D\u0026rsquo;Arcy J, Kashyap S, et al. A path for translation of machine learning products into healthcare delivery. NPJ Digit Med. 2020;3:1\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKawamoto K, Houlihan CA, Balas EA, Lobach DF. Improving clinical practice using clinical decision support systems. BMJ. 2005;330:765.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaito T, Rehmsmeier M. The precision\u0026ndash;recall plot is more informative than ROC plots for imbalanced datasets. PLoS ONE. 2015;10(3):e0118432.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe H, Garcia EA. Learning from imbalanced data. IEEE Trans Knowl Data Eng. 2009;21(9):1263\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJohnson AEW, Pollard TJ, Shen L, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10:1.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZubair M, Din I, Sarwar N, Elov B, Makhmudov S, Trabelsi Z. Revolutionizing sepsis diagnosis using machine learning and deep learning models: a systematic literature review. \u003cem\u003eBMC Infect Dis\u003c/em\u003e. 2025;25(1):1396. Published 2025 Oct 23. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12879-025-11423-2\u003c/span\u003e\u003cspan address=\"10.1186/s12879-025-11423-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang M, Zhong M, Cheng Y, Zhang T. Intelligent Prediction Platform for Sepsis Risk Based on Real-Time Dynamic Temporal Features: Design Study. JMIR Med Inf. 2025;13. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/74940\u003c/span\u003e\u003cspan address=\"10.2196/74940\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://medinform.jmir.org/2025/1/e74940\u003c/span\u003e\u003cspan address=\"https://medinform.jmir.org/2025/1/e74940\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. e74940, URL.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang SZ, Ding HY, Shen YM, et al. Harness machine learning for multiple prognoses prediction in sepsis patients: evidence from the MIMIC-IV database. BMC Med Inf Decis Mak. 2025;25:152. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12911-025-02976-y\u003c/span\u003e\u003cspan address=\"10.1186/s12911-025-02976-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu C, Li L, Huang W, et al. Interpretable Machine Learning for Early Prediction of Prognosis in Sepsis: A Discovery and Validation Study. Infect Dis Ther. 2022;11(3):1117\u0026ndash;32. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s40121-022-00628-6\u003c/span\u003e\u003cspan address=\"10.1007/s40121-022-00628-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYadgarov MY, Landoni G, Berikashvili LB, Polyakov PA, Kadantseva KK, Smirnova AV et al. Early detection of sepsis using machine learning algorithms: a systematic review and network meta-analysis. Front Med (Lausanne). 2024;11:1491358. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fmed.2024.1491358\u003c/span\u003e\u003cspan address=\"10.3389/fmed.2024.1491358\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.frontiersin.org/journals/medicine/articles/10.3389/fmed.2024.1491358\u003c/span\u003e\u003cspan address=\"https://www.frontiersin.org/journals/medicine/articles/10.3389/fmed.2024.1491358\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Z, Shu W, Li T, Zhang X, Chong W. Interpretable machine learning for predicting sepsis risk in emergency triage patients. Sci Rep. 2025;15(1):887. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-025-85121-z\u003c/span\u003e\u003cspan address=\"10.1038/s41598-025-85121-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Published 2025 Jan 6.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Sepsis, Machine learning, Landmark modelling, ICU, Dynamic prediction","lastPublishedDoi":"10.21203/rs.3.rs-8737800/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8737800/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eEarly recognition of sepsis in critically ill patients remains challenging due to dynamic physiological changes and nonspecific clinical presentation. Most prediction models rely on static or continuously updated data streams without explicitly accounting for evolving risk over clinically meaningful time intervals. The study aimed to develop and evaluate a landmark-based dynamic machine learning framework to predict sepsis within a 6-hour horizon among respiratory-supported intensive care unit (ICU) patients.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThis is a secondary analysis using data from the MIMIC-IV database. Adult intensive care unit patients receiving respiratory support were evaluated at four landmarks (6, 12, 18, and 24 hours). At each point, sepsis-free patients were used to predict sepsis onset within the next 6 hours. Models included logistic regression, random forest, and XGBoost. The patient\u0026ndash;level train\u0026ndash;test splitting and group cross-validation prevented information leakage. Performance was assessed using discrimination, classification metrics, and calibration. A balanced ensemble approach addressed class imbalance in sensitivity analysis, and interpretability was examined using permutation importance and regression effect estimates.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eA total of 41,871, 39,912, 36,472, and 31,367 patients were included at the 6, 12, 18, and 24-hour landmarks, respectively. Sepsis incidence declined from 1.48% to 0.37% across time points. Model performance varied, with the 18-hour landmark showing the best balance between discrimination and clinically meaningful operating characteristics. Logistic regression achieved the highest discrimination in the primary analysis (AUROC\u0026thinsp;=\u0026thinsp;0.78), while random forest performed best in sensitivity analyses (AUROC\u0026thinsp;=\u0026thinsp;0.77). Both consistently identified the 18-hour landmark as optimal, indicating that temporal risk structure outweighed algorithm choice. Calibration was checked overall but showed overestimation at higher predicted risks. Key predictors reflected respiratory, hemodynamic, neurological, and comorbidity factors.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eLandmark-based dynamic modelling provides a clinically interpretable and temporally informed strategy for early sepsis prediction in respiratory-supported intensive care unit patients. The consistent identification of the 18-hour window as the most informative prediction point suggests that intermediate ICU time frames may offer the best balance between timeliness and predictive stability. Further work should focus on recalibration, threshold optimization, and external validation before clinical implementation.\u003c/p\u003e","manuscriptTitle":"Dynamic Landmark-Based Prediction of Sepsis Using Interpretable and Balanced Machine Learning Models in Respiratory-Supported Critically ill Patients","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-25 06:34:35","doi":"10.21203/rs.3.rs-8737800/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3b11c430-4285-49dd-80c8-8bdfe760027c","owner":[],"postedDate":"March 25th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-25T06:34:35+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-25 06:34:35","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8737800","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8737800","identity":"rs-8737800","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.