Predicting infection risk in rheumatoid arthritis patients receiving biological or targeted synthetic disease-modifying anti-rheumatic drugs: an application of machine learning and healthcare big data

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Patients with rheumatoid arthritis (RA) initiating biologic or targeted synthetic disease-modifying antirheumatic drugs (b/ts DMARDs) face elevated risk of serious infections, necessitating tools for individualized risk stratification. Objectives Primary objective was to develop and validate a clinically interpretable machine learning (ML) model to predict 1-year risk of serious infection after b/ts DMARD initiation; secondary objectives were to estimate infection incidence and identify key predictors associated with risk. Methods We performed a retrospective cohort study using territory-wide EHR from Hong Kong’s Clinical Data Analysis and Reporting System (CDARS) for model development and internal validation, with external validation in the U.S. All of Us database. The outcome was first serious infection requiring hospitalization within 1 year. Candidate predictors included demographics, comorbidities, prior infections and medications, laboratory markers. Multiple ML algorithms were trained; model selection was based on AUROC, and interpretability was assessed using SHAP. Results A total of 3,159 patients from CDARS (8.8% with serious infections) and 1,845 from All of Us (2.8% with serious infections) were included. The model demonstrated the highest AUROC in internal validation (0.840, 95% CI: 0.793–0.888) and maintained robust performance in external validation (AUROC: 0.729, 95% CI: 0.665–0.793). Key predictors included prior infections, diabetes, b/ts DMARD type, and inflammatory markers. Rituximab was linked to the highest infection risk, while tofacitinib and upadacitinib had the lowest. Conclusion This study developed and validated an ML model using routine clinical data to predict serious infection risk in RA patients, supporting personalised treatment and proactive infection management.
Full text 183,877 characters · extracted from preprint-html · click to expand
Predicting infection risk in rheumatoid arthritis patients receiving biological or targeted synthetic disease-modifying anti-rheumatic drugs: an application of machine learning and healthcare big data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predicting infection risk in rheumatoid arthritis patients receiving biological or targeted synthetic disease-modifying anti-rheumatic drugs: an application of machine learning and healthcare big data Kuan Peng, Deliang Yang, Jiaqi Wang, Chin-Yao Shen, Michael Chun-Yuan Cheng, and 9 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9054871/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 13 You are reading this latest preprint version Abstract Background Patients with rheumatoid arthritis (RA) initiating biologic or targeted synthetic disease-modifying antirheumatic drugs (b/ts DMARDs) face elevated risk of serious infections, necessitating tools for individualized risk stratification. Objectives Primary objective was to develop and validate a clinically interpretable machine learning (ML) model to predict 1-year risk of serious infection after b/ts DMARD initiation; secondary objectives were to estimate infection incidence and identify key predictors associated with risk. Methods We performed a retrospective cohort study using territory-wide EHR from Hong Kong’s Clinical Data Analysis and Reporting System (CDARS) for model development and internal validation, with external validation in the U.S. All of Us database. The outcome was first serious infection requiring hospitalization within 1 year. Candidate predictors included demographics, comorbidities, prior infections and medications, laboratory markers. Multiple ML algorithms were trained; model selection was based on AUROC, and interpretability was assessed using SHAP. Results A total of 3,159 patients from CDARS (8.8% with serious infections) and 1,845 from All of Us (2.8% with serious infections) were included. The model demonstrated the highest AUROC in internal validation (0.840, 95% CI: 0.793–0.888) and maintained robust performance in external validation (AUROC: 0.729, 95% CI: 0.665–0.793). Key predictors included prior infections, diabetes, b/ts DMARD type, and inflammatory markers. Rituximab was linked to the highest infection risk, while tofacitinib and upadacitinib had the lowest. Conclusion This study developed and validated an ML model using routine clinical data to predict serious infection risk in RA patients, supporting personalised treatment and proactive infection management. Machine Learning Biological disease modifying anti-rheumatic drugs Rheumatoid arthritis Infection Figures Figure 1 Figure 2 Figure 3 Figure 4 Highlights This study developed and validated a machine learning model predicts 1-year serious infection risk after b/ts DMARD initiation. This model achieved AUROC 0.840 in internal and 0.729 in external validation datasets. Key predictors include prior infections, diabetes, b/ts DMARD type, and inflammatory markers. Rituximab associated with highest infection risk; tofacitinib and upadacitinib with lowest. Introduction Rheumatoid arthritis (RA) is one of the most prevalent chronic inflammatory joint diseases worldwide and is linked to excessive deaths and productivity loss 1 , 2 . Biologic or targeted synthetic disease-modifying antirheumatic drugs (b/ts DMARDs) have revolutionised the treatment paradigm of RA as the gold standard for managing patients whose response to conventional synthetic DMARDs is inadequate 3 . B/ts DMARDs work by targeting a wide range of inflammatory pathways, including but not limited to tumour necrosis factor inhibitors (TNFi) 4 , interleukin inhibitors (ILi) 5 , and Janus Kinase inhibitors (JAKi) 6 . B/ts DMARD therapy can effectively control inflammation, slow disease progression and joint deterioration, ultimately improving quality of life. However, as with all immunomodulatory therapies, these medicines are associated with an increased risk of infection 7 . It is estimated that infection is the third leading cause of premature death among patients with RA, contributing to 13.4% of all-cause mortality 8 . Consequently, the risk of infection has become one of the most important considerations for clinicians and patients when initiating b/ts DMARD therapy. One of the key identified risk factors for infection among patients with RA is the type of b/ts DMARD, however the current clinical treatment guideline has no recommendations for a particular b/ts DMARD therapy to minimise infection risk 3 . Prior studies have not identified any clinically significant difference among the treatment options, however this may be due to limited power or insufficient follow-up time to detect rare events like serious infections 7 , 9 , 10 . Other studies show that compared with TNFi, abatacept is associated with a lower infection risk [Hazard Ratio (HR) and 95% confidence interval (CI): 0.37 (0.18, 0.75)] 11 ; while tocilizumab is linked to a greater infection risk [HR and 95% CI: 1.39 (1.08, 1.79)] 12 . JAKi are targeted synthetic DMARDs found to be associated with an increased risk of herpes zoster infection [aHR and 95% CI: 2.37 (2.00, 2.80)] compared to TNFi 13 . The risk of infection can also vary significantly even among drugs with the same mode of action. A previous study reported that the risk of serious infection among patients treated with adalimumab, infliximab and etanercept were 2.61, 3.86, and 1.66 per 100 patient-years, respectively 14 . Other factors associated with serious infection include age, comorbidities (e.g., cancer), inflammation status [e.g., C-reactive protein (CRP), erythrocyte sedimentation rate (ESR)], medication use (e.g., opioid and glucocorticoids) 15 , 16 . In recent years, predictive models developed by machine learning (ML) based on large-scale, longitudinal electronic medical databases have played an important role in disease diagnosis, treatment choice, and prognosis estimation. Compared with conventional statistics, the unique advantages of ML include handling complex and large datasets, free from stringent assumptions, automated feature selection, improved predictive accuracy, and the ability to uncover hidden patterns and relationships in the data, leading to more accurate predictions and efficient decision-making 17 . In rheumatic diseases, ML has been applied to predict disease remission 18 , treatment response 19 , 20 , and assist in clinical diagnosis 21 ; however, to our knowledge, it has yet to be used to predict medication safety of b/ts DMARDs. In this study, we aimed to develop and validate an ML model using data derived from population-based electronic health records (EHR) databases to predict the risk of serious infection following the initiation of b/ts DMARDs. Method Data source We used the retrospective EHR database from the Hong Kong (HK) Clinical Data Analysis and Reporting System (CDARS) to develop the ML models. CDARS is a territory-wide EHR database developed and managed by the Hong Kong Hospital Authority (HA). The HA is a unique statutory body that manages all public hospitals and their ambulatory clinics, serving a population of 7.4 million through 43 hospitals and institutions, 49 specialist outpatient clinics, and 73 general outpatient clinics in HK. Established in 1993, the EHR database includes standardised and structured information on the population of HK, including demographics, hospital admission, prescriptions and diagnoses made from inpatient and outpatient settings, laboratory tests. All records have been de-identified and assigned unique reference keys to facilitate outcome tracing and data analysis while protecting patients’ privacy 22 . CDARS has been widely used for research and clinical management purposes. High-quality, published rheumatology studies thoroughly demonstrate coding accuracy and data quality 23 , 24 . We performed external validation using the All of Us database. All of Us Research Program EHR is a longitudinal cohort covering over 560,000 diverse participants of different races and ethnicities across the United States 25 . The EHR database contains conditions, drug exposure, lab measurements, and procedures that warrant the success of external validation. CDARS uses International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) for disease coding and British National Formulary (BNF) codes with generic drug names for prescriptions. All of Us employ both ICD-9-CM and ICD-10-CM for diseases, and Anatomical Therapeutic Chemical (ATC) codes with generic names for prescriptions. Supplementary Table 1 lists the specific diagnostic codes for infections and comorbidities. Target population Patients diagnosed with RA between 2010 to 2023 were identified from CDARS as a training cohort and extracted from All of Us between 2010 to 2022 as an external validation cohort. Inclusion and exclusion criteria consist of 1) age ≥ 18 years at disease onset, and 2) received at least one b/ts DMARDs treatment during the observation period 3) patients with any prescription of b/ts DMARDs before 2010 were excluded to ascertain new users. Biosimilar DMARDs were not differentiated as we assume they have comparable safety and efficacy profiles to their reference products 26 . Outcomes Target populations were followed up from the index date (prescription start date of the first recorded use of b/ts DMARDs) until the occurrence of the outcome, treatment discontinuation, switch of treatment to another b/ts DMARDs, death from any cause, observation end (one year after index date), or study end date (2023-12-31 for CDARS and 2022-12-31 for All of Us), whichever comes first. The primary outcome of interest was the first serious infection requiring hospitalisation that occurred within one year after the initiation of b/ts DMARDs. This infection had to take place either during the prescription period of b/ts DMARDs or within 30 days after the date of the last prescription (grace period). This is because b/ts DMARDs can have half-lives ranging from several days to weeks, leading to prolonged drug effects even after cessation. Failing to account for infections occurring shortly after treatment discontinuation might underestimate the true risk associated with b/ts DMARDs. We therefore applied a 30-day grace period that balances the need to capture relevant adverse events without extending the observation window excessively. The application of a grace period has been widely adopted in previous publications 27 , 28 . Infection outcomes included a wide spectrum of infection pathogens (bacterial/virus/fungi) and infection location (urinary tract infection, respiratory infection, etc). To ensure accurate and reliable serious infection outcomes, we further restricted study outcomes to infection-related hospitalisation requiring anti-microbial medications within ± 14 days of hospitalisation. Anti-microbial prescription records were identified using BNF code 5.1/5.2/5.3 in CDARS and ATC code J01/J02/J04/J05. Validation of infectious outcome using corresponding medications also confirms that the infection was significant enough and required medical treatment, distinguishing between minor and more serious infections. Predictive features Predictive features used for developing the model were demographic characteristics including age and sex; comorbidities measured prior to the index date (trace back to 2000); b/ts DMARDs products; clinical history of serious infection within past three years, routine laboratory tests (Platelet, Creatinine, C-Reactive Protein, and more—up to 27 tests) within 90 days before the index date; and daily dose of oral anti-inflammation and analgesic agents glucocorticoids, non-steroidal anti-inflammatory drug (NSAID), and opioid within 90 days before the index date. All computed daily dose were standardised by World Health Organization Defined Daily Dose ( Supplementary Table 2 ), described as the estimated average daily maintenance dose for primary indication in adults 29 . For laboratory tests, mean test results were applied if there were multiple test records. Analysis Data pre-processing and missing data imputation Variables with near-zero variance (where the fraction of unique values is less than 2% of the sample size) and highly correlated variables were removed. Laboratory tests with a missing data rate > 30% were excluded from the model construction. For the remaining variables with ≥ 70% data completeness, a series of imputation method including k-nearest neighbor imputation (KNN), Median imputation, Bagged Tree imputation, and MICE (Multiple Imputation by Chained Equations) were used to impute missing values. The performance of these methods was evaluated by visually comparing the density distributions of imputed values against the original observed data. The method yielding distributions most closely resembling the original will be selected. This approach is based on the rationale that observed and missing values should arise from the same underlying distribution under random missingness assumptions. A good imputation method should preserve this distributional structure rather than introduce bias or distortion. Our results ( Supplementary Fig. 1 ) showed that the KNN imputation best preserved the original distributional properties (shape, central tendency, and variability) of the data compared to other methods. Therefore, we retained KNN as our primary imputation method. All continuation variables were scaled and normalised. Supplementary Table 3 details the complete list of collected laboratory tests, and the corresponding completion rates. Model development and evaluation The dataset was split at a ratio of 8:2 into the training dataset and testing dataset. Random forest model, Least absolute shrinkage and selection operator (LASSO) regression 30 , Support vector machine 31 , XGBoost, and ensembling neural network using model averaging 32 algorithms as well as logistic regression were adapted to develop classifiers. Ten-fold cross-validation was used for internal validation. Down-sampling was employed to address the class imbalance between infected and non-infected patients 33 . Which randomly reduces the majority class (non-infected) to match the size of the minority class (infected), creating a balanced training dataset. As a result, the model outputs represent relative risk scores rather than calibrated probabilities, and a threshold of 0.5 was applied to classify patients into high-risk and low-risk groups. The area under the receiver operating characteristic curve (AUROC) was adapted as the main metric to assess the performance of the different ML algorithms. The model with the highest AUROC value was selected as the final model for prediction. Other metrics including accuracy, sensitivity, specificity, F1 score, positive predictive value, and negative predictive value were also reported. Model calibration We evaluated calibration using the Brier score and calibration curves, which plot observed outcome proportions against predicted probabilities across risk groups. To account for the artificial prevalence introduced by down-sampling during training, we applied a post-hoc Bayesian recalibration 34 to adjust the predicted probabilities to the probability reflecting the true prevalence. Performance metrics including the calibration intercept and slope were calculated to quantify goodness-of-fit. Decision Curve Analysis Decision Curve Analysis (DCA) 35 was performed to evaluate whether using the model to guide clinical decisions would benefit patients more than simple "treat all" strategy by computing metric \(NetBenefit=\raisebox{1ex}{$TP$}\!\left/\!\raisebox{-1ex}{$N$}\right.-\raisebox{1ex}{$FP$}\!\left/\!\raisebox{-1ex}{$N$}\right.(\frac{{P}_{t}}{1-{P}_{t}}\) ), where TP = true positive count, FP = false positive count, n = total number of subjects, and pₜ = threshold probability. The threshold probability is the minimum predicted risk at which intervention would be recommended. Patients whose predicted probability exceeds this threshold receive treatment. Subgroup analysis To evaluate the robustness and consistency of the model across different patient populations, we performed a subgroup analysis on the hold-out test set. We stratified the patients based on sex (Female vs. Male), age (≥ 60 years vs. <60 years), and history of previous infection. Explanation of ML results SHAP (Shapley Additive explanations) value is a technique used to elucidate the predictions generated by ML models. They facilitate the identification of how each feature contributes to the model's ultimate prediction. A positive SHAP value indicates that the feature contributes to the development of hospitalised infection. In contrast, a negative SHAP value means the feature is negatively associated with the development of hospitalised infection. We computed the SHAP value based on the training cohort. Features were ranked in descending order based on their importance (computed as mean absolute SHAP value). We visualised explanation diagrams at the dataset level, depicting the contribution of each feature to the prediction for each patient. To differentiate the infection risk across different b/ts DMARDs, the SHAP values of each b/ts DMARDs were further visualised in a violin plot in ascending order. Partial Dependence Plots (PDP) illustrate influence of biomarkers To enhance the interpretability of the random forest model, we generated PDPs 36 for the studied biomarkers, trimmed to the 5th–95th percentile range to exclude extreme values. PDPs visualize the marginal effect of a specific feature on the predicted probability of infection while averaging out the effects of all other variables. Additionally, we calculated quantitative metrics to characterize each biomarker's influence. The effect magnitude was defined as the range between maximum and minimum predicted probabilities across the PDP curve, representing the total impact of that biomarker on infection risk. To identify clinically meaningful thresholds, we computed the first derivative of the PDP curve and located the point of maximum slope. This "tipping point" indicates the biomarker value at which infection risk changes most rapidly, providing actionable thresholds for clinical decision-making. Sample size calculation To ensure sufficient sample sizes for the development of robust models, we calculated the minimum sample size required for predicting time-to-event outcomes clinically proposed by Riley et al 37 . External validation Following the same data processing methodologies as those used for the CDARS cohort, we developed a new model using the All of Us cohort as external validation. Glucocorticoids, NSAIDs, and opioids were excluded from the external validation by designating 0 to the three predictive features in the All of Us cohort during the model fitting process, as the prescription daily dose of these medications is not available in the All of Us database. Data analysis and data visualisation were conducted using R version 4.1.4 (R Foundation for Statistical Computing, Vienna, Austria) and cross-checked by three independent researchers (KP, JQW and DLY). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD) 38 checklists was reported to ensure quality assurance and reporting standards. Results Patient characteristics In total, we identified 3159 and 1845 patients newly prescribed with b/ts DMARDs in the training cohort CDARS and the validation cohort All of Us, respectively. Among them, 277 (8.8%) and 52 (2.8%) serious infections occurred during the study period. Predictive features used to train the model are summarised in Table 1 . Compared with the non-infection group, the infected group had poorer baseline health conditions including older age, higher inflammation level reflected by ESR and CRP values, increased comorbidity and more medication. Table 1 Baseline characteristics of the developing cohort (CDARS) and the validation cohort (All of Us) Description CDARS All of Us Not Infected Infected P value Not Infected Infected P value Number of patients 2882 277 1793 52 Demographic characteristics Male (%) 542 (18.8) 69 (24.9) 0.017 362 (20.2) 7 (13.5) 0.3078 Age (mean (SD)) 54.70 (13.55) 63.47 (14.29) < 0.001 54.16 (13.72) 50.83 (16.28) 0.086 Previous infection within 3 years (%) 278 (9.6) 79 (28.5) < 0.001 77 (4.3) 16 (30.8) < 0.001 Laboratory tests Alanine transaminase (U/L mean (SD)) 22.53 (19.63) 23.51 (21.14) 0.441 24.30 (15.73) 27.21 (20.96) 0.312 Albumin (g/L mean (SD)) 38.71 (4.59) 36.10 (5.86) < 0.001 39.43 (5.78) 35.61 (6.81) < 0.001 CRP (mg/dL mean (SD)) 2.10 (2.73) 3.41 (3.92) < 0.001 1.44 (2.31) 4.26 (4.62) < 0.001 Creatinine (umol/L mean (SD)) 66.01 (22.68) 78.88 (47.24) < 0.001 266.05 (672.83) 375.70 (705.85) 0.285 ESR (mm/hr mean (SD)) 54.60 (32.04) 62.43 (32.53) < 0.001 26.74 (22.95) 53.89 (30.78) < 0.001 Lymphocytes (% mean (SD)) 22.78 (8.49) 19.62 (8.75) < 0.001 24.80 (9.59) 21.12 (9.43) 0.018 Platelet (10^9/L mean (SD)) 300.38 (93.56) 279.51 (109.28) 0.001 266.75 (73.90) 292.17 (95.57) 0.034 Red blood cell (10^12/L mean (SD)) 4.18 (0.55) 3.91 (0.60) < 0.001 4.24 (0.49) 4.03 (0.61) 0.011 Bilirubin (umol/L mean (SD)) 7.93 (7.79) 8.65 (6.40) 0.155 7.89 (3.86) 8.25 (5.44) 0.58 Urea (mmol/L mean (SD)) 5.05 (1.73) 6.29 (4.16) < 0.001 5.77 (2.97) 5.71 (2.24) 0.907 White blood cell (10^9/L mean (SD)) 7.23 (2.50) 8.16 (5.66) < 0.001 7.37 (2.82) 8.73 (3.90) 0.003 Baseline comorbidities Coronary artery disease and heart failure (%) 94 (3.3) 34 (12.3) < 0.001 101 (5.6) 5 (9.6) 0.361 Hyperlipidaemia (%) 144 (5.0) 31 (11.2) < 0.001 22 (1.2) 2 (3.8) 0.307 Thyroid disorder (%) 95 (3.3) 16 (5.8) 0.049 3 (0.2) 0 (0.0) 1 Osteoporosis (%) 59 (2.0) 17 (6.1) < 0.001 310 (17.3) 12 (23.1) 0.369 Hypertension (%) 384 (13.3) 81 (29.2) < 0.001 787 (43.9) 30 (57.7) 0.067 Diabetes mellitus (%) 169 (5.9) 37 (13.4) < 0.001 368 (20.5) 16 (30.8) 0.105 Cerebrovascular disease (%) 65 (2.3) 12 (4.3) 0.053 94 (5.2) 6 (11.5) 0.096 Chronic obstructive pulmonary disease (%) 139 (4.8) 30 (10.8) < 0.001 101 (5.6) 5 (9.6) 0.361 Ulcers (%) 133 (4.6) 20 (7.2) 0.075 80 (4.5) 6 (11.5) 0.04 Cancer (%) 118 (4.1) 36 (13.0) < 0.001 10 (0.6) 0 (0.0) 1 Baseline prescription Prednisolone (mean (SD)) 0.27 (0.35) 0.36 (0.43) < 0.001 - - NSAID (mean (SD)) 0.70 (0.63) 0.59 (0.66) 0.005 - - Opioid (mean (SD)) 0.05 (0.17) 0.09 (0.22) < 0.001 - - B/ts DMARDs episodes < 0.001 0.058 Abatacept 114 (4.0) 17 (6.1) 116 (6.5) 2 (3.8) Adalimumab 503 (17.5) 33 (11.9) 597 (33.3) 20 (38.5) Baricitinib 110 (3.8) 31 (11.2) 3 (0.2) 0 (0.0) Certolizumab Pegol 100 (3.5) 8 (2.9) 49 (2.7) 1 (1.9) Etanercept 533 (18.5) 34 (12.3) 474 (26.4) 7 (13.5) Golimumab 354 (12.3) 14 (5.1) 23 (1.3) 1 (1.9) Infliximab 119 (4.1) 12 (4.3) 141 (7.9) 2 (3.8) Rituximab 164 (5.7) 45 (16.2) 172 (9.6) 11 (21.2) Sarilumab 41 (1.4) 1 (0.4) 4 (0.2) 1 (1.9) Tocilizumab 335 (11.6) 50 (18.1) 88 (4.9) 4 (7.7) Tofacitinib 450 (15.6) 29 (10.5) 107 (6.0) 3 (5.8) Upadacitinib 59 (2.0) 3 (1.1) 19 (1.1) 0 (0.0) B/ts DMARDs: Biologic or targeted synthetic disease-modifying antirheumatic drugs; CRP: C-reactive protein; ESR: Erythrocyte Sedimentation Rate; SD: Standard deviation; NSAID: non-steroidal anti-inflammatory drug. Model evaluation and external validation The classification accuracy on the testing dataset of six trained models is described in Fig. 1 -A and Supplementary Table 4 ; the random forest model shows the best predictive performance [AUROC: 0.840 (95%CI: 0.793–0.888)] in CDARS and was therefore selected as the final model for external validation in All of Us. The random forest model generated predictive accuracy [AUROC: 0.729 (95%CI: 0.665–0.793)] in the external validation cohort (All of Us, Fig. 1 -B ) . The developed model satisfied the minimum sample size (1756 individuals with 154 events) required for new model development given the resulting model performance (AUROC 0.840), number of predictive features (28) and incidence of serious infection (8.8%). Model calibration The random forest model demonstrated good overall performance with a Brier score of 0.073 ( Supplemental Fig. 2 ). While calibration analysis indicated some degree of risk underestimation, a common finding in machine learning models applied to imbalanced datasets (Intercept = 3.503, Slope = 2.496). Decision Curve Analysis Decision curve analysis demonstrated that the model provided net benefit compared to default strategies (treat-all or treat-none) across threshold probabilities ranging from 3.1% − 32.4%. At the threshold probability of 8.5% the model provided the greatest net benefit gain (0.0438) compared to treat-all strategy ( Supplemental Fig. 3 ). In clinical practice, this means that for every 100 patients evaluated, using this model to guide decision-making would result in approximately 4 additional patients correctly receiving needed treatment while avoiding unnecessary interventions. Subgroup analysis The model demonstrated consistent discrimination across all subgroups ( Supplemental Fig. 4 ), with AUROC values ≥ 0.73 in all categories. Performance was higher in females (AUROC = 0.859) than males (AUROC = 0.760), and similar between age groups (< 60 years: 0.819; ≥60 years: 0.778). Patients without prior infection history showed higher AUROC (0.840) compared to those with prior infection (0.732). Lower performance in males and patients with infection history may partly reflect smaller sample sizes (N = 131 and N = 72, respectively), resulting in wider confidence intervals. Explanations of ML results Figure 2 visualised the SHAP values that shows the importance in descending order and distribution of features across all data points. The colour represents the feature's value for a particular data point. Particularly, for binary variables such as comorbidity, a low value represents 0, and a high value represents 1; for multi-class factor variables such as b/ts DMARDs, each value represents a unique pharmaceutical product. The top three important features were previous infection within three years, followed by b/ts DMARDs use and diabetes mellitus. The SHAP values of b/ts DMARDs were plotted in violin plots, with the median value plotted in black points (Fig. 3 ). Among b/ts DMARDs, tofacitinib and upadacitinib have the lowest SHAP values indicating less predicted infection risk. In contrast, rituximab held the highest SHAP values linked to increased predicted infection risk. Similarly, categorised by therapeutic area, sarilumab has the lowest predicted risk of infection among non-TNFi agents, while tocilizumab and rituximab have the highest predicted risks of infection. Tofacitinib and upadacitinib show low predicted risks of infection among JAKi, while golimumab and etanercept have relatively lower predicted risks of infection among TNFi agents. Partial Dependence Plots (PDP) illustrate influence of biomarkers Figure 4 and Supplemental table 5 illustrate the marginal effect of each laboratory parameter on the predicted probability of infection when holding all other model features constant. CRP showed the largest effect range (0.0132) with a threshold at 7.1 mg/dL and positive slope (+ 0.0053). ESR threshold was 103.0 mm/hr with positive slope (+ 0.0005). Albumin had an effect range of 0.0071 with threshold at 30.3 g/L and negative slope (-0.0036). Lymphocytes showed effect range of 0.0089 with threshold at 23.6% and negative slope (-0.0012). White blood cell threshold was 9.63 ×10⁹/L with positive slope (+ 0.0032). Creatinine threshold was 95.3 µmol/L (+ 0.0008) and Urea was 3.02 mmol/L (+ 0.0040). Red blood cell had effect range 0.0094 with threshold at 3.45 ×10¹²/L and negative slope (-0.0305). Discussion For optimised medication safety management among patients with RA, the study attempted to predict serious infection risk among patients receiving b/ts DMARDs treatment. All predictive features were derived from routine clinical data, ensuring the predictor’s ease of implementation in practice. It performs well in the training cohort (AUROC 0.840), while maintaining acceptable performance in the external validation cohort (AUROC 0.729). Performance differences may be attributed to variations in racial and ethnic composition and coding systems between the two cohorts. Furthermore, differences in infection rates may reflect disparities in the granularity of infection classification and coding practices across the two databases. Nevertheless, the study findings still showed a strong generalisability and reproducibility of our model. A previous study, the Rheumatoid Arthritis Observation of Biologic Therapy (RABBIT) study, used multivariable generalised estimation equations to predict the absolute risk of serious infection among patients taking b/ts DMARDs 39 . That study was performed in 2011 and did not consider newer treatments such as JAKi and biomarkers. In contrast, we successfully identified implementable biomarkers that serve as risk prediction with ML; we also established straightforward cut-off thresholds to simplify infection risk stratification. Comparing the model performance using traditional logistic regression as reference (AUROC 0.830), the random forest model showed slightly better performance (AUROC difference: + 0.010), algorithms like XGBoost had a lower AUROC (AUROC difference: − 0.035). This finding is in line with prior research that has found that advanced ML algorithms do not necessarily lead to improved performance over traditional logistic regression 40 . Explanations of ML and PDPs results We elucidated model interpretability by ranking feature importance (Fig. 2 ). The risk factors that were the highest ranked variables contributing to the risk of infection included previous infection, diabetes mellitus, hypertension, and use of symptom control medications such as opioids, NSAIDs and steroids. These risk factors have all been documented in previous reviews 10 , 41 . We used PDP (Fig. 4 ) to quantify the continuous relationship between biomarker values and predicted infection probability, characterizing both the effect range and direction for each variable. We found that both elevated CRP and ESR contribute to the infection risk positively, consistent with their well-established role as classical biomarkers measuring inflammatory status, and mounting evidence supports their use in the diagnosis of infection at different sites 42 , 43 . Albumin exhibited a negative relationship with infection risk, which is expected given that hypoalbuminemia reflects both malnutrition and chronic inflammation; low albumin has been consistently associated with poor outcomes in infectious diseases 44 . The ratio of blood urea nitrogen to albumin is a multifaceted indicator of renal and inflammatory function; a high urea nitrogen and low albumin level has shown a strong correlation with serious infection 45 . Similarly, the lymphocytes (proportion to white blood cells) showed a negative slope with threshold at 23.6%, supporting the notion that lymphopenia compromises immunity against pathogens 46 . Haematological manifestations, including anaemia (low red blood cells) and thrombocytopenia (low platelet), are common among patients with RA and likely associated with the autoimmune disease itself and anti-rheumatic drugs 47 . Elevated creatinine levels could indicate kidney infection and failure; prior reviews have identified chronic kidney disease as a risk factor for acute community-acquired infections 48 . Collectively, these biomarkers capture multiple physiological dimensions span from inflammation, immune function, nutritional status, and organ function that contribute to infection susceptibility. Additionally, we computed the first derivative of each PDP curve to identify "tipping points," meaning threshold values at which infection risk changes most rapidly. While this does not constitute a validated diagnostic cut-off, it highlights biomarker ranges that may merit heightened clinical vigilance and could serve as candidate thresholds for future prospective evaluation. Infection risk associated with different b/ts DMARDs Our findings align with previous meta-analyses of 58 randomised controlled trials (RCTs), which detected that etanercept and golimumab present a reduced risk of serious infections in comparison to other TNFi agents 49 . This finding was further validated by a nationwide register-based cohort study in Sweden that recruited over 20,000 patients with RA. The study examined the risk of serious infection across 11 different b/ts DMARDs, revealing that etanercept and golimumab exhibit a superior safety profile compared to infliximab, adalimumab, and certolizumab 50 . Our findings indicate that the JAKi tofacitinib and upadacitinib presents a lower infection risk compared to other b/ts DMARDs, consistent with the Swedish cohort study 50 . However, existing literature reported that incidences of herpes zoster and tuberculosis are greater among those treated with tofacitinib compared to TNFi 51 , 52 . Such differences could arise from 1) our adaptation of a different definition covering a broader range of infection outcomes 2) ORAL-Surveillance study 51 which reported a higher risk of major cardiovascular events and cancers with tofacitinib compared with TNFi, JAKi are thus used cautiously in patients > 65, smokers, or those at high cardiovascular risk, leading to fewer high-risk patients receiving JAKi and potentially contributing to the observed lower infection risk. In terms of non-TNFi agents, our results are consistent with existing studies where rituximab targeting CD20 has the highest risk in comparison to sarilumab, an interleukin-6 inhibitor shows the lowest risk 50 , 53 . We found that another interleukin − 6 inhibitor (tocilizumab) showed a higher infection risk than sarilumab. This likely reflects potential indication bias: tocilizumab is often used in more refractory with greater comorbidity or steroid use, while the new sarilumab are selected for lower-risk patients, such prescribing pattern have been reported previously in Germany, with tocilizumab recipients being systematically sicker than TNFi recipients 54 . Strength and clinical implications Our study possesses several advantages. Firstly, it provides clinically actionable absolute probabilities (0-100%) of one-year serious infection risk upon clinical visit, enabling clinicians to stratify patients into high-risk and low-risk groups and tailor treatment strategies based on the calculated risk. For patients with increased vulnerability to serious infections who have yet to take b/ts DMARDs for disease activity control, clinicians can select b/ts DMARDs associated with the lowest infection risk, specifically tailored to the patient's profile and unique clinical characteristics. Additionally, clinicians can taper the doses of opioids, NSAIDs, and prednisolone to maintain infection risk within an acceptable range. This approach not only offers practical strategies for rheumatology treatment to balance the risks and benefits of b/ts DMARDs, but also actively involves patients in shared decision-making by informing them of their infection risk during clinical visits. For high-risk patients, clinicians can also recommend pre-treatment screenings for latent infections (e.g., hepatitis B, tuberculosis) and advise patients to receive vaccinations promptly to prevent potential infections. Secondly, the success of external validation in a United States population supports the broader clinical applicability of the model, rendering a more reliable tool for guiding treatment decisions in diverse settings. Thirdly, the robustness of serious infection identification was strengthened by using the ICD diagnostic code and anti-microbial prescription; this dual approach minimises misclassification by confirming clinical relevance and capturing true infection cases. Lastly, the model requires limited clinical information to function without sacrificing model performance, making it easy to build into the clinical decision system. Limitations This study has certain limitations. Infection risk is multifactorial and dynamic, shaped by environmental factors (living conditions, hygiene, healthcare access) that are difficult to quantify, and by time-dependent factors (recent travel, disease flares, medication changes) not captured by baseline features; while this limitation is common to most clinical prediction models. To minimize the influence of time-varying features and focus on the most clinically relevant period we restricted the observation window to 12 months; while this horizon may miss later events, it was deliberately chosen because it aligns with decision points for DMARD selection and early monitoring and captures the interval of highest infection risk after treatment initiation. Also, during external validation using the All of Us cohort, three medication-related features (glucocorticoids, NSAIDs, and opioids daily dose) were not available and were therefore set to zero during model fitting. This may have attenuated the model's performance in the external cohort, as these medications contributed to risk prediction in the original model. Despite this limitation, the model still achieved an AUROC of 0.729 in external validation, and subgroup analyses demonstrated consistent discrimination across patient subgroups (AUROC ≥ 0.73), suggesting reasonable robustness. Lastly, due to the nature of the EHR database and the aim of using predictors readily available at the point of care, important predictive features, such as disease activities and patient-reported outcomes were not included in the model development. Conclusion This study developed a ML model to identify patients with an elevated risk of serious infection using routine clinical data exclusively. It provides clinical insights on risk stratification of different b/ts DMARDs therapies, predictive cut-off threshold on biomarkers, and importance ranking of risk factors. The prediction tool has the potential to streamline treatment decisions using available clinical inputs and will help to identify high-risk patients for enhanced safety management, and patient education. Abbreviations RA Rheumatoid arthritis b/ts DMARDs Biologic or targeted synthetic disease-modifying antirheumatic drugs ML Machine learning EHR Electronic health records CDARS Clinical Data Analysis and Reporting System AUROC Area under the receiver operating characteristic curve SHAP Shapley Additive explanations CI Confidence interval HR Hazard ratio aHR Adjusted hazard ratio TNFi Tumour necrosis factor inhibitors TNFα Tumour necrosis factor alpha ILi Interleukin inhibitors IL-6 Interleukin-6 JAKi Janus Kinase inhibitors CRP C-reactive protein ESR Erythrocyte sedimentation rate HK Hong Kong HA Hospital Authority ICD-9-CM International Classification of Diseases, Ninth Revision, Clinical Modification ICD-10-CM International Classification of Diseases, Tenth Revision, Clinical Modification BNF British National Formulary ATC Anatomical Therapeutic Chemical NSAID Non-steroidal anti-inflammatory drug KNN K-nearest neighbor imputation MICE Multiple Imputation by Chained Equations LASSO Least absolute shrinkage and selection operator XGBoost Extreme Gradient Boosting DCA Decision Curve Analysis PDP Partial Dependence Plots TRIPOD Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis RCTs Randomised controlled trials RABBIT Rheumatoid Arthritis Observation of Biologic Therapy ARTIS Anti-Rheumatic Therapies in Sweden DREAM Dutch Rheumatoid Arthritis Monitoring WHO World Health Organization DDD Defined Daily Dose CD20 Cluster of differentiation 20 TP True positive FP False positive WBC White blood cell RBC Red blood cell COPD Chronic obstructive pulmonary disease SD Standard deviation HMRF Health and Medical Research Fund HKSAR Hong Kong Special Administrative Region RGC Research Grants Council ECS Early Career Scheme RIF Research Impact Fund ADAMS Advanced Data Analytics for Medical Science D²4H Laboratory of Data Discovery for Health EULAR European Alliance of Associations for Rheumatology Declarations Ethics approval and consent to participate Ethical approval for the research was granted by the Institutional Review Board of The University of Hong Kong / Hospital Authority Hong Kong West Cluster under reference numbers UW 21–338. The Institutional Review Board granted a waiver of informed patient consent. This was justified as the studies were exclusively observational, meaning no practical interventions were administered, and all patient-level data utilized were de-identified by the data custodian at the time of data extraction. Consent for publication Not applicable Competing Interests Xue Li reports a relationship with Janssen Pharmaceuticals Inc that includes: consulting or advisory. Xue Li reports a relationship with Pfizer that includes: consulting or advisory. Xue Li reports a relationship with Amgen Inc that includes: consulting or advisory. Xue Li reports a relationship with Merck Sharp & Dohme UK Ltd that includes: consulting or advisory. Xue Li reports a relationship with OPEN Health Communications LLP that includes: consulting or advisory. Xue Li reports a relationship with Office of Health Economics that includes: consulting or advisory. Xue Li reports a relationship with ADAMS Limited Hong Kong that includes: employment. XL received research grants from the Research Fund Secretariat of the Health Bureau, Health and Medical Research Fund (HMRF, HKSAR), Health and Medical Research Fund Fellowship Scheme (HMRF Fellowship, HKSAR), Research Grants Council Early Career Scheme (RGC/ECS, HKSAR), Research Grants Council Research Impact Fund (RGC/RIF, HKSAR), Commission grants from Hospital Authority of Hong Kong; educational and investigator initiate research fund from Janssen, Pfizer and Amgen; internal funding from the University of Hong Kong; consultancy fee from Pfizer, Merck Sharp & Dohme, Open Health, Office of Health Economics; she is also the former non-executive director of ADAMS Limited Hong Kong; all outside the submitted work. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Funding Enhanced Start-up Fund for new academic staff, LKS Faculty of Medicine, The University of Hong Kong; Internal Research Fund, Department of Medicine, School of Clinical Medicine, LKS Faculty of Medicine, The University of Hong Kong. Author Contribution KP, DLY, and JQW performed the data analysis. CYS, MCYC, and ECCL conducted the external validation. KP, SCWC, and XL conceptualized the study. KP and XL wrote the manuscript. SCWC, NLP, ICKW, CSL, JJG, QPZ, and IYKT contributed to the interpretation of results. XL supervised the study and obtained funding. All authors reviewed and approved the final manuscript. Acknowledgement Acknowledgements and affiliationsWe thank Lisa Lam for proofreading this paper. Data Availability The data that support the findings of this study are available from the Hospital Authority of Hong Kong, but restrictions apply to the availability of these data, which were used under license for the current study, and so are not publicly available. Data are however available from the authors upon reasonable request and with permission of the Hospital Authority of Hong Kong. References Joachim L, Jörn K, Bernhard M, et al. Mortality in rheumatoid arthritis: the impact of disease activity, treatment with glucocorticoids, TNFα inhibitors and rituximab. Ann Rheum Dis. 2015;74(2):415. Filipovic I, Walker D, Forster F, Curry AS. Quantifying the economic burden of productivity loss in rheumatoid arthritis. Rheumatology. 2011;50(6):1083–90. Josef SS, Robert BML, Sytske Anne B, et al. EULAR recommendations for the management of rheumatoid arthritis with synthetic and biological disease-modifying antirheumatic drugs: 2022 update. Ann Rheum Dis. 2023;82(1):3. Ma X, Xu S. TNF inhibitor therapy for rheumatoid arthritis. Biomed Rep. 2013;1(2):177–84. Ogata A, Kato Y, Higa S, Yoshizaki K. IL-6 inhibitor for the treatment of rheumatoid arthritis: A comprehensive review. Mod Rheumatol. 2019;29(2):258–67. Yamaoka K. Janus kinase inhibitors for rheumatoid arthritis. Curr Opin Chem Biol. 2016;32:29–33. Sepriano A, Kerschbaumer A, Bergstra SA, et al. Safety of synthetic and biological DMARDs: a systematic literature review informing the 2022 update of the EULAR recommendations for the management of rheumatoid arthritis. Ann Rheum Dis. 2023;82(1):107. van den Hoek J, Boshuizen HC, Roorda LD, et al. Mortality in patients with rheumatoid arthritis: a 15-year prospective cohort study. Rheumatol Int. 2017;37(4):487–93. Ozen G, Pedro S, England BR, Mehta B, Wolfe F, Michaud K. Risk of Serious Infection in Patients With Rheumatoid Arthritis Treated With Biologic Versus Nonbiologic Disease-Modifying Antirheumatic Drugs. ACR Open Rheumatol. 2019;1(7):424–32. Jani M, Barton A, Hyrich K. Prediction of infection risk in rheumatoid arthritis patients treated with biologics: are we any closer to risk stratification? Curr Opin Rheumatol. 2019;31(3):285–92. Ozen G, Pedro S, Schumacher R, Simon TA, Michaud K. Safety of abatacept compared with other biologic and conventional synthetic disease-modifying antirheumatic drugs in patients with rheumatoid arthritis: data from an observational study. Arthritis Res Ther. 2019;21(1):141. Giles JT, Sattar N, Gabriel S, et al. Cardiovascular Safety of Tocilizumab Versus Etanercept in Rheumatoid Arthritis: A Randomized Controlled Trial. Arthritis Rheumatol. 2020;72(1):31–40. Strand V, Ahadieh S, French J, et al. Systematic review and meta-analysis of serious infections with tofacitinib and biologic disease-modifying antirheumatic drug treatment in rheumatoid arthritis clinical trials. Arthritis Res Ther. 2015;17:362. van Dartel SA, Fransen J, Kievit W, et al. Difference in the risk of serious infections in patients with rheumatoid arthritis treated with adalimumab, infliximab and etanercept: results from the Dutch Rheumatoid Arthritis Monitoring (DREAM) registry. Ann Rheum Dis. 2013;72(6):895–900. Listing J, Gerhold K, Zink A. The risk of infections associated with rheumatoid arthritis, with its comorbidity and treatment. Rheumatology. 2013;52(1):53–61. George MD, Baker JF, Winthrop K, et al. Risk for Serious Infection With Low-Dose Glucocorticoids in Patients With Rheumatoid Arthritis: A Cohort Study. Ann Intern Med. 2020;173(11):870–8. Kim KJ, Tagkopoulos I. Application of machine learning in rheumatic disease research. Korean J Intern Med. 2019;34(4):708–22. Koo BS, Eun S, Shin K, et al. Machine learning model for identifying important clinical features for predicting remission in patients with rheumatoid arthritis treated with biologics. Arthritis Res Ther. 2021;23(1):178. Bouget V, Duquesne J, Hassler S et al. Machine learning predicts response to TNF inhibitors in rheumatoid arthritis: results on the ESPOIR and ABIRISK cohorts. RMD Open 2022;8(2). Duquesne J, Bouget V, Cournède PH, et al. Machine learning identifies a profile of inadequate responder to methotrexate in rheumatoid arthritis. Rheumatology. 2023;62(7):2402–9. Lim AJW, Tyniana CT, Lim LJ, et al. Robust SNP-based prediction of rheumatoid arthritis through machine-learning-optimized polygenic risk score. J Transl Med. 2023;21(1):92. Introduction. Caring for our community’s health. https://www.ha.org.hk/visitor/ha_visitor_index.asp?Parent_ID=10004&Content_ID=10008&Ver=HTML . Accessed July 6, 2024. Li X, Tong X, Yeung WWY, et al. Two-dose COVID-19 vaccination and possible arthritis flare among patients with rheumatoid arthritis in Hong Kong. Ann Rheum Dis. 2022;81(4):564–8. Chai Y, Luo H, Wong GHY, et al. Risk of self-harm after the diagnosis of psychiatric disorders in Hong Kong, 2000-10: a nested case-control study. Lancet Psychiatry. 2020;7(2):135–47. Bick AG, Metcalf GA, Mayo KR, et al. Genomic data in the All of Us Research Program. Nature. 2024;627(8003):340–6. Conran CA, Moreland LW. A review of biosimilars for rheumatoid arthritis. Curr Opin Pharmacol. 2022;64:102234. Curtis JR, Winthrop K, O'Brien C, Ndlovu MN, de Longueville M, Haraoui B. Use of a baseline risk score to identify the risk of serious infectious events in patients with rheumatoid arthritis during certolizumab pegol treatment. Arthritis Res Ther. 2017;19(1):276. Galloway JB, Hyrich KL, Mercer LK, et al. Risk of septic arthritis in patients with rheumatoid arthritis and the effect of anti-TNF therapy: results from the British Society for Rheumatology Biologics Register. Ann Rheum Dis. 2011;70(10):1810–4. ATC/DDD Index. 2024. https://atcddd.fhi.no/atc_ddd_index/ . Accessed Oct 29th, 2024. Tibshirani R. Regression Shrinkage and Selection via the Lasso. J Royal Stat Soc Ser B (Methodological). 1996;58(1):267–88. Cortes C, Vapnik V. Support-vector networks. Mach Learn. 1995;20(3):273–97. Ripley BD. Pattern Recognition and Neural Networks. Cambridge: Cambridge University Press; 1996. LemaÃŽtre G, Nogueira F, Aridas CKJJ. Imbalanced-learn: A python toolbox to tackle the curse of imbalanced datasets in machine learning. 2017;18(17):1–5. Pozzolo AD, Caelen O, Johnson RA, Bontempi G. Dec. Calibrating Probability with Undersampling for Unbalanced Classification. Paper presented at: 2015 IEEE Symposium Series on Computational Intelligence; 7–10 2015, 2015. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis making: Int J Soc Med Decis Mak. 2006;26(6):565–74. Friedman JH. Greedy Function Approximation: A Gradient Boosting Machine. Annals Stat. 2001;29(5):1189–232. Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. 2020;368:m441. Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med. 2015;162(1):55–63. Zink A, Manger B, Kaufmann J, et al. Evaluation of the RABBIT Risk Score for serious infections. Ann Rheum Dis. 2014;73(9):1673. Christodoulou E, Ma J, Collins GS, Steyerberg EW, Verbakel JY, Van Calster B. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12–22. Waseem A-A, Laurie T, James V, et al. The association between diabetes mellitus and incident infections: a systematic review and meta-analysis of observational studies. BMJ Open Diabetes Res Care. 2017;5(1):e000336. Lapić I, Padoan A, Bozzato D, Plebani M. Erythrocyte Sedimentation Rate and C-Reactive Protein in Acute Inflammation: Meta-Analysis of Diagnostic Accuracy Studies. Am J Clin Pathol. 2020;153(1):14–29. Lindsay CP, Olcott CW, Del Gaizo DJ. ESR and CRP are useful between stages of 2-stage revision for periprosthetic joint infection. Arthroplasty Today. 2017;3(3):183–6. Wiedermann CJ. Hypoalbuminemia as Surrogate and Culprit of Infections. Int J Mol Sci 2021;22(9). Ugajin M, Yamaki K, Iwamura N, Yagi T, Asano T. Blood urea nitrogen to serum albumin ratio independently predicts mortality and severity of community-acquired pneumonia. Int J Gen Med. 2012;5:583–9. Zhou YQ, Feng DY, Li WJ, et al. Lower neutrophil-to-lymphocyte ratio predicts high risk of multidrug-resistant Pseudomonas aeruginosa infection in patients with hospital-acquired pneumonia. Ther Clin Risk Manag. 2018;14:1863–9. Bowman SJ. Hematological manifestations of rheumatoid arthritis. Scand J Rheumatol. 2002;31(5):251–9. McDonald HI, Thomas SL, Nitsch D. Chronic kidney disease as a risk factor for acute community-acquired infections in high-income countries: a systematic review. BMJ Open. 2014;4(4):e004100. Minozzi S, Bonovas S, Lytras T, et al. Risk of infections using anti-TNF agents in rheumatoid arthritis, psoriatic arthritis, and ankylosing spondylitis: a systematic review and meta-analysis. Expert Opin Drug Saf. 2016;15(sup1):11–34. Frisell T, Bower H, Morin M, et al. Safety of biological and targeted synthetic disease-modifying antirheumatic drugs for rheumatoid arthritis as used in clinical practice: results from the ARTIS programme. Ann Rheum Dis. 2023;82(5):601. Ytterberg Steven R, Bhatt Deepak L, Mikuls Ted R, et al. Cardiovascular and Cancer Risk with Tofacitinib in Rheumatoid Arthritis. N Engl J Med. 2022;386(4):316–26. Frisell T, Bower H, Morin M, et al. Safety of biological and targeted synthetic disease-modifying antirheumatic drugs for rheumatoid arthritis as used in clinical practice: results from the ARTIS programme. Ann Rheum Dis. 2023;82(5):601–10. Grøn KL, Arkema EV, Glintborg B, et al. Risk of serious infections in patients with rheumatoid arthritis treated in routine care with abatacept, rituximab and tocilizumab in Denmark and Sweden. Ann Rheum Dis. 2019;78(3):320–7. Backhaus M, Kaufmann J, Richter C, et al. Comparison of tocilizumab and tumour necrosis factor inhibitors in rheumatoid arthritis: a retrospective analysis of 1603 patients managed in routine clinical practice. Clin Rheumatol. 2015;34(4):673–81. Additional Declarations Competing interest reported. Xue Li reports a relationship with Janssen Pharmaceuticals Inc that includes: consulting or advisory. Xue Li reports a relationship with Pfizer that includes: consulting or advisory. Xue Li reports a relationship with Amgen Inc that includes: consulting or advisory. Xue Li reports a relationship with Merck Sharp & Dohme UK Ltd that includes: consulting or advisory. Xue Li reports a relationship with OPEN Health Communications LLP that includes: consulting or advisory. Xue Li reports a relationship with Office of Health Economics that includes: consulting or advisory. Xue Li reports a relationship with ADAMS Limited Hong Kong that includes: employment. XL received research grants from the Research Fund Secretariat of the Health Bureau, Health and Medical Research Fund (HMRF, HKSAR), Health and Medical Research Fund Fellowship Scheme (HMRF Fellowship, HKSAR), Research Grants Council Early Career Scheme (RGC/ECS, HKSAR), Research Grants Council Research Impact Fund (RGC/RIF, HKSAR), Commission grants from Hospital Authority of Hong Kong; educational and investigator initiate research fund from Janssen, Pfizer and Amgen; internal funding from the University of Hong Kong; consultancy fee from Pfizer, Merck Sharp & Dohme, Open Health, Office of Health Economics; she is also the former non-executive director of ADAMS Limited Hong Kong; all outside the submitted work. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Supplementary Files BMCMIDMSupplementarydocuments.docx BMCMIDMGraphicabstract.png Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 15 Apr, 2026 Reviews received at journal 09 Apr, 2026 Reviews received at journal 28 Mar, 2026 Reviewers agreed at journal 25 Mar, 2026 Reviewers agreed at journal 24 Mar, 2026 Reviewers agreed at journal 22 Mar, 2026 Reviewers agreed at journal 21 Mar, 2026 Reviewers agreed at journal 20 Mar, 2026 Reviewers invited by journal 20 Mar, 2026 Editor assigned by journal 20 Mar, 2026 Editor invited by journal 16 Mar, 2026 Submission checks completed at journal 13 Mar, 2026 First submitted to journal 13 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9054871","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":611632691,"identity":"90189284-a091-4d55-a285-16567c8fa781","order_by":0,"name":"Kuan Peng","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Kuan","middleName":"","lastName":"Peng","suffix":""},{"id":611632692,"identity":"22791cac-6c38-4cce-90ed-ba42d83bb3ee","order_by":1,"name":"Deliang Yang","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Deliang","middleName":"","lastName":"Yang","suffix":""},{"id":611632693,"identity":"a38c4376-596a-4c74-af3a-932f361d9512","order_by":2,"name":"Jiaqi Wang","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Jiaqi","middleName":"","lastName":"Wang","suffix":""},{"id":611632694,"identity":"ccb51129-86bc-4516-bdff-3d8bbb5c8d9b","order_by":3,"name":"Chin-Yao Shen","email":"","orcid":"","institution":"National Cheng Kung University","correspondingAuthor":false,"prefix":"","firstName":"Chin-Yao","middleName":"","lastName":"Shen","suffix":""},{"id":611632695,"identity":"77438d4f-76a8-49ea-8bb7-49b60e3bc7ba","order_by":4,"name":"Michael Chun-Yuan Cheng","email":"","orcid":"","institution":"National Cheng Kung University","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"Chun-Yuan","lastName":"Cheng","suffix":""},{"id":611632696,"identity":"fd27af3e-2a04-4e4c-8950-4635a5ac7935","order_by":5,"name":"Shirley C.W. Chan","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Shirley","middleName":"C.W.","lastName":"Chan","suffix":""},{"id":611632697,"identity":"2baa1b09-49d5-439a-b9f7-35dde487568b","order_by":6,"name":"Iris Y.K. Tang","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Iris","middleName":"Y.K.","lastName":"Tang","suffix":""},{"id":611632698,"identity":"8b1932a6-1490-4af5-9a90-825e5f7cc895","order_by":7,"name":"Qingpeng Zhang","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Qingpeng","middleName":"","lastName":"Zhang","suffix":""},{"id":611632699,"identity":"256de274-d942-4aee-8956-13206fc28546","order_by":8,"name":"Edward Chia‑Cheng Lai","email":"","orcid":"","institution":"National Cheng Kung University","correspondingAuthor":false,"prefix":"","firstName":"Edward","middleName":"Chia‑Cheng","lastName":"Lai","suffix":""},{"id":611632700,"identity":"2f5ee0c7-be83-4e43-b1b6-0f52732228cd","order_by":9,"name":"Nicole L. Pratt","email":"","orcid":"","institution":"University of Adelaide","correspondingAuthor":false,"prefix":"","firstName":"Nicole","middleName":"L.","lastName":"Pratt","suffix":""},{"id":611632701,"identity":"5b5c4cb6-5cc4-4c58-ad30-2d375f9dbbc7","order_by":10,"name":"Ian Chi Kei Wong","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Ian","middleName":"Chi Kei","lastName":"Wong","suffix":""},{"id":611632702,"identity":"14c5d0c7-33ba-4147-9faa-82e57104f52b","order_by":11,"name":"Chak-sing Lau","email":"","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":false,"prefix":"","firstName":"Chak-sing","middleName":"","lastName":"Lau","suffix":""},{"id":611632703,"identity":"f0f349e4-3459-4ea4-aef1-042de4ac74d6","order_by":12,"name":"Jeff Jianfei Guo","email":"","orcid":"","institution":"University of Cincinnati","correspondingAuthor":false,"prefix":"","firstName":"Jeff","middleName":"Jianfei","lastName":"Guo","suffix":""},{"id":611632704,"identity":"9eb3e43c-0e33-4cb5-afb5-5e3f061686b2","order_by":13,"name":"Xue Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAoklEQVRIiWNgGAWjYLCCBwY2CRIMDGyMDURrSTBII1kLw2EStMjPyDF8kFBwPk9yRgLbwxnEaDE4c8bYIMHgdrG0RAK74QaitLD3mEkAtSTOk0hgk3xAlMOaecx/JBicI0ELw/EeM2CIHUicDdJCnMPOHCsGOiw5cWbPwzZJorwvPyN544cPf+wSZxxPPibZQ5TDEICEiBwFo2AUjIJRQAAAAEBEMkk0+MhQAAAAAElFTkSuQmCC","orcid":"","institution":"The University of Hong Kong","correspondingAuthor":true,"prefix":"","firstName":"Xue","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2026-03-07 02:53:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9054871/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9054871/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105433017,"identity":"80fee6ff-79b9-4f65-8352-18ae5c6a6a5f","added_by":"auto","created_at":"2026-03-26 02:57:41","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":183667,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003e(A) Area under the receiver operating characteristic (AUROC) curve for the six machine learning models on training cohort (B) Area under the receiver operating characteristic (AUROC) curve for the random forest model in validation set for training cohort (CDARS) and external validation cohort (All of Us).\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLASSO: least absolute shrinkage and selection operator; XGBoost: Extreme Gradient Boosting.\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/11fdb7037975e5cf3c2bbfb6.jpg"},{"id":105565962,"identity":"f55a9099-05f5-4aff-976e-c5e3cb5c7d84","added_by":"auto","created_at":"2026-03-27 12:54:53","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":240403,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eShapley Additive explanations (SHAP) value of predictive features ranked by importance in ascending order\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFeatures were ranked descending order based on feature importance computed as mean absolute SHAP values.\u003c/p\u003e\n\u003cp\u003eB/ts DMARDs: Biologic or targeted synthetic disease-modifying antirheumatic drugs; CRP: C-reactive protein; ESR: Erythrocyte Sedimentation Rate; NSAID: non-steroidal anti-inflammatory drug; RBC: Red blood cell; WBC: White blood cell; COPD: Chronic obstructive pulmonary disease\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/5d4a58ac9d4a17397c2c2c97.jpg"},{"id":105433018,"identity":"56cb1d16-8124-41b5-980f-f67a7c411594","added_by":"auto","created_at":"2026-03-26 02:57:41","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":97800,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe Shapley Additive explanations value across different biologic or targeted synthetic disease-modifying antirheumatic drugs\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eB/tsDMARDs: Biological or targeted synthetic disease-modifying anti-rheumatic drugs\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/5a3943be72dbe1a6ce87672c.jpg"},{"id":105433022,"identity":"80aeddca-643a-4173-a898-e30b228b5bc4","added_by":"auto","created_at":"2026-03-26 02:57:42","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":207842,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePartial Dependence Plots of Key Biomarkers Predicting Infection Risk\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe panels illustrate the marginal effect of each laboratory parameter on the predicted probability of infection (y-axis), holding all other model features constant. Solid Line represents the average predicted risk across the continuous range of the biomarker. Red dotted line indicates the threshold of maximum change (point of steepest slope), suggesting a potential critical clinical value where infection risk alters most significantly.\u003c/p\u003e\n\u003cp\u003eCRP, C-reactive protein; ESR, erythrocyte sedimentation rate.\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/a796ab96de35cd097751a0d2.jpg"},{"id":105569412,"identity":"cdf21c19-4961-4b5f-827e-f551c608c633","added_by":"auto","created_at":"2026-03-27 13:12:24","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2350571,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/e4c643a0-3000-4900-9a0d-2b6ef8deeab9.pdf"},{"id":105565471,"identity":"fde39117-8ce7-4a30-8c9d-eafcd43bf722","added_by":"auto","created_at":"2026-03-27 12:53:21","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":784965,"visible":true,"origin":"","legend":"","description":"","filename":"BMCMIDMSupplementarydocuments.docx","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/ac12edcba2bd9dd00b3e2d77.docx"},{"id":105565370,"identity":"d726278d-57a5-434f-9034-a72c75da2c24","added_by":"auto","created_at":"2026-03-27 12:53:03","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":329191,"visible":true,"origin":"","legend":"","description":"","filename":"BMCMIDMGraphicabstract.png","url":"https://assets-eu.researchsquare.com/files/rs-9054871/v1/d21c9df7697bb1690c1f69c6.png"}],"financialInterests":"Competing interest reported. Xue Li reports a relationship with Janssen Pharmaceuticals Inc that includes: consulting or advisory. Xue Li reports a relationship with Pfizer that includes: consulting or advisory. Xue Li reports a relationship with Amgen Inc that includes: consulting or advisory. Xue Li reports a relationship with Merck Sharp \u0026 Dohme UK Ltd that includes: consulting or advisory. Xue Li reports a relationship with OPEN Health Communications LLP that includes: consulting or advisory. Xue Li reports a relationship with Office of Health Economics that includes: consulting or advisory. Xue Li reports a relationship with ADAMS Limited Hong Kong that includes: employment. XL received research grants from the Research Fund Secretariat of the Health Bureau, Health and Medical Research Fund (HMRF, HKSAR), Health and Medical Research Fund Fellowship Scheme (HMRF Fellowship, HKSAR), Research Grants Council Early Career Scheme (RGC/ECS, HKSAR), Research Grants Council Research Impact Fund (RGC/RIF, HKSAR), Commission grants from Hospital Authority of Hong Kong; educational and investigator initiate research fund from Janssen, Pfizer and Amgen; internal funding from the University of Hong Kong; consultancy fee from Pfizer, Merck Sharp \u0026 Dohme, Open Health, Office of Health Economics; she is also the former non-executive director of ADAMS Limited Hong Kong; all outside the submitted work. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","formattedTitle":"Predicting infection risk in rheumatoid arthritis patients receiving biological or targeted synthetic disease-modifying anti-rheumatic drugs: an application of machine learning and healthcare big data","fulltext":[{"header":"Highlights","content":"\u003cul\u003e\n \u003cli\u003eThis study developed and validated a machine learning model predicts 1-year serious infection risk after b/ts DMARD initiation. This model achieved AUROC 0.840 in internal and 0.729 in external validation datasets.\u003c/li\u003e\n \u003cli\u003eKey predictors include prior infections, diabetes, b/ts DMARD type, and inflammatory markers.\u003c/li\u003e\n \u003cli\u003eRituximab associated with highest infection risk; tofacitinib and upadacitinib with lowest.\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"Introduction","content":"\u003cp\u003eRheumatoid arthritis (RA) is one of the most prevalent chronic inflammatory joint diseases worldwide and is linked to excessive deaths and productivity loss\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. Biologic or targeted synthetic disease-modifying antirheumatic drugs (b/ts DMARDs) have revolutionised the treatment paradigm of RA as the gold standard for managing patients whose response to conventional synthetic DMARDs is inadequate\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. B/ts DMARDs work by targeting a wide range of inflammatory pathways, including but not limited to tumour necrosis factor inhibitors (TNFi)\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, interleukin inhibitors (ILi)\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e, and Janus Kinase inhibitors (JAKi)\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. B/ts DMARD therapy can effectively control inflammation, slow disease progression and joint deterioration, ultimately improving quality of life. However, as with all immunomodulatory therapies, these medicines are associated with an increased risk of infection\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. It is estimated that infection is the third leading cause of premature death among patients with RA, contributing to 13.4% of all-cause mortality\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Consequently, the risk of infection has become one of the most important considerations for clinicians and patients when initiating b/ts DMARD therapy.\u003c/p\u003e \u003cp\u003eOne of the key identified risk factors for infection among patients with RA is the type of b/ts DMARD, however the current clinical treatment guideline has no recommendations for a particular b/ts DMARD therapy to minimise infection risk\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Prior studies have not identified any clinically significant difference among the treatment options, however this may be due to limited power or insufficient follow-up time to detect rare events like serious infections\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. Other studies show that compared with TNFi, abatacept is associated with a lower infection risk [Hazard Ratio (HR) and 95% confidence interval (CI): 0.37 (0.18, 0.75)]\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e; while tocilizumab is linked to a greater infection risk [HR and 95% CI: 1.39 (1.08, 1.79)]\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. JAKi are targeted synthetic DMARDs found to be associated with an increased risk of herpes zoster infection [aHR and 95% CI: 2.37 (2.00, 2.80)] compared to TNFi\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. The risk of infection can also vary significantly even among drugs with the same mode of action. A previous study reported that the risk of serious infection among patients treated with adalimumab, infliximab and etanercept were 2.61, 3.86, and 1.66 per 100 patient-years, respectively\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. Other factors associated with serious infection include age, comorbidities (e.g., cancer), inflammation status [e.g., C-reactive protein (CRP), erythrocyte sedimentation rate (ESR)], medication use (e.g., opioid and glucocorticoids)\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn recent years, predictive models developed by machine learning (ML) based on large-scale, longitudinal electronic medical databases have played an important role in disease diagnosis, treatment choice, and prognosis estimation. Compared with conventional statistics, the unique advantages of ML include handling complex and large datasets, free from stringent assumptions, automated feature selection, improved predictive accuracy, and the ability to uncover hidden patterns and relationships in the data, leading to more accurate predictions and efficient decision-making\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. In rheumatic diseases, ML has been applied to predict disease remission\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e, treatment response\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e, and assist in clinical diagnosis\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e; however, to our knowledge, it has yet to be used to predict medication safety of b/ts DMARDs. In this study, we aimed to develop and validate an ML model using data derived from population-based electronic health records (EHR) databases to predict the risk of serious infection following the initiation of b/ts DMARDs.\u003c/p\u003e"},{"header":"Method","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData source\u003c/h2\u003e \u003cp\u003eWe used the retrospective EHR database from the Hong Kong (HK) Clinical Data Analysis and Reporting System (CDARS) to develop the ML models. CDARS is a territory-wide EHR database developed and managed by the Hong Kong Hospital Authority (HA). The HA is a unique statutory body that manages all public hospitals and their ambulatory clinics, serving a population of 7.4\u0026nbsp;million through 43 hospitals and institutions, 49 specialist outpatient clinics, and 73 general outpatient clinics in HK. Established in 1993, the EHR database includes standardised and structured information on the population of HK, including demographics, hospital admission, prescriptions and diagnoses made from inpatient and outpatient settings, laboratory tests. All records have been de-identified and assigned unique reference keys to facilitate outcome tracing and data analysis while protecting patients\u0026rsquo; privacy\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. CDARS has been widely used for research and clinical management purposes. High-quality, published rheumatology studies thoroughly demonstrate coding accuracy and data quality\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. We performed external validation using the All of Us database. All of Us Research Program EHR is a longitudinal cohort covering over 560,000 diverse participants of different races and ethnicities across the United States\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. The EHR database contains conditions, drug exposure, lab measurements, and procedures that warrant the success of external validation.\u003c/p\u003e \u003cp\u003eCDARS uses International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) for disease coding and British National Formulary (BNF) codes with generic drug names for prescriptions. All of Us employ both ICD-9-CM and ICD-10-CM for diseases, and Anatomical Therapeutic Chemical (ATC) codes with generic names for prescriptions. \u003cb\u003eSupplementary Table\u0026nbsp;1\u003c/b\u003e lists the specific diagnostic codes for infections and comorbidities.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eTarget population\u003c/h3\u003e\n\u003cp\u003ePatients diagnosed with RA between 2010 to 2023 were identified from CDARS as a training cohort and extracted from All of Us between 2010 to 2022 as an external validation cohort. Inclusion and exclusion criteria consist of 1) age\u0026thinsp;\u0026ge;\u0026thinsp;18 years at disease onset, and 2) received at least one b/ts DMARDs treatment during the observation period 3) patients with any prescription of b/ts DMARDs before 2010 were excluded to ascertain new users. Biosimilar DMARDs were not differentiated as we assume they have comparable safety and efficacy profiles to their reference products\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003ch3\u003eOutcomes\u003c/h3\u003e\n\u003cp\u003eTarget populations were followed up from the index date (prescription start date of the first recorded use of b/ts DMARDs) until the occurrence of the outcome, treatment discontinuation, switch of treatment to another b/ts DMARDs, death from any cause, observation end (one year after index date), or study end date (2023-12-31 for CDARS and 2022-12-31 for All of Us), whichever comes first. The primary outcome of interest was the first serious infection requiring hospitalisation that occurred within one year after the initiation of b/ts DMARDs. This infection had to take place either during the prescription period of b/ts DMARDs or within 30 days after the date of the last prescription (grace period). This is because b/ts DMARDs can have half-lives ranging from several days to weeks, leading to prolonged drug effects even after cessation. Failing to account for infections occurring shortly after treatment discontinuation might underestimate the true risk associated with b/ts DMARDs. We therefore applied a 30-day grace period that balances the need to capture relevant adverse events without extending the observation window excessively. The application of a grace period has been widely adopted in previous publications\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e,\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Infection outcomes included a wide spectrum of infection pathogens (bacterial/virus/fungi) and infection location (urinary tract infection, respiratory infection, etc).\u003c/p\u003e \u003cp\u003eTo ensure accurate and reliable serious infection outcomes, we further restricted study outcomes to infection-related hospitalisation requiring anti-microbial medications within \u0026plusmn;\u0026thinsp;14 days of hospitalisation. Anti-microbial prescription records were identified using BNF code 5.1/5.2/5.3 in CDARS and ATC code J01/J02/J04/J05. Validation of infectious outcome using corresponding medications also confirms that the infection was significant enough and required medical treatment, distinguishing between minor and more serious infections.\u003c/p\u003e\n\u003ch3\u003ePredictive features\u003c/h3\u003e\n\u003cp\u003ePredictive features used for developing the model were demographic characteristics including age and sex; comorbidities measured prior to the index date (trace back to 2000); b/ts DMARDs products; clinical history of serious infection within past three years, routine laboratory tests (Platelet, Creatinine, C-Reactive Protein, and more\u0026mdash;up to 27 tests) within 90 days before the index date; and daily dose of oral anti-inflammation and analgesic agents glucocorticoids, non-steroidal anti-inflammatory drug (NSAID), and opioid within 90 days before the index date. All computed daily dose were standardised by World Health Organization Defined Daily Dose (\u003cb\u003eSupplementary Table\u0026nbsp;2\u003c/b\u003e), described as the estimated average daily maintenance dose for primary indication in adults\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. For laboratory tests, mean test results were applied if there were multiple test records.\u003c/p\u003e\n\u003ch3\u003eAnalysis\u003c/h3\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eData pre-processing and missing data imputation\u003c/h2\u003e \u003cp\u003eVariables with near-zero variance (where the fraction of unique values is less than 2% of the sample size) and highly correlated variables were removed. Laboratory tests with a missing data rate\u0026thinsp;\u0026gt;\u0026thinsp;30% were excluded from the model construction. For the remaining variables with \u0026ge;\u0026thinsp;70% data completeness, a series of imputation method including k-nearest neighbor imputation (KNN), Median imputation, Bagged Tree imputation, and MICE (Multiple Imputation by Chained Equations) were used to impute missing values. The performance of these methods was evaluated by visually comparing the density distributions of imputed values against the original observed data. The method yielding distributions most closely resembling the original will be selected. This approach is based on the rationale that observed and missing values should arise from the same underlying distribution under random missingness assumptions. A good imputation method should preserve this distributional structure rather than introduce bias or distortion. Our results (\u003cb\u003eSupplementary Fig.\u0026nbsp;1\u003c/b\u003e) showed that the KNN imputation best preserved the original distributional properties (shape, central tendency, and variability) of the data compared to other methods. Therefore, we retained KNN as our primary imputation method. All continuation variables were scaled and normalised. \u003cb\u003eSupplementary Table\u0026nbsp;3 details\u003c/b\u003e the complete list of collected laboratory tests, and the corresponding completion rates.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eModel development and evaluation\u003c/h3\u003e\n\u003cp\u003eThe dataset was split at a ratio of 8:2 into the training dataset and testing dataset. Random forest model, Least absolute shrinkage and selection operator (LASSO) regression\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e, Support vector machine\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e, XGBoost, and ensembling neural network using model averaging\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e algorithms as well as logistic regression were adapted to develop classifiers. Ten-fold cross-validation was used for internal validation. Down-sampling was employed to address the class imbalance between infected and non-infected patients\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. Which randomly reduces the majority class (non-infected) to match the size of the minority class (infected), creating a balanced training dataset. As a result, the model outputs represent relative risk scores rather than calibrated probabilities, and a threshold of 0.5 was applied to classify patients into high-risk and low-risk groups. The area under the receiver operating characteristic curve (AUROC) was adapted as the main metric to assess the performance of the different ML algorithms. The model with the highest AUROC value was selected as the final model for prediction. Other metrics including accuracy, sensitivity, specificity, F1 score, positive predictive value, and negative predictive value were also reported.\u003c/p\u003e\n\u003ch3\u003eModel calibration\u003c/h3\u003e\n\u003cp\u003eWe evaluated calibration using the Brier score and calibration curves, which plot observed outcome proportions against predicted probabilities across risk groups. To account for the artificial prevalence introduced by down-sampling during training, we applied a post-hoc Bayesian recalibration\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e to adjust the predicted probabilities to the probability reflecting the true prevalence. Performance metrics including the calibration intercept and slope were calculated to quantify goodness-of-fit.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eDecision Curve Analysis\u003c/h2\u003e \u003cp\u003eDecision Curve Analysis (DCA)\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e was performed to evaluate whether using the model to guide clinical decisions would benefit patients more than simple \"treat all\" strategy by computing metric \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(NetBenefit=\\raisebox{1ex}{$TP$}\\!\\left/\\!\\raisebox{-1ex}{$N$}\\right.-\\raisebox{1ex}{$FP$}\\!\\left/\\!\\raisebox{-1ex}{$N$}\\right.(\\frac{{P}_{t}}{1-{P}_{t}}\\)\u003c/span\u003e\u003c/span\u003e), where TP\u0026thinsp;=\u0026thinsp;true positive count, FP\u0026thinsp;=\u0026thinsp;false positive count, n\u0026thinsp;=\u0026thinsp;total number of subjects, and pₜ = threshold probability. The threshold probability is the minimum predicted risk at which intervention would be recommended. Patients whose predicted probability exceeds this threshold receive treatment.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eSubgroup analysis\u003c/h2\u003e \u003cp\u003eTo evaluate the robustness and consistency of the model across different patient populations, we performed a subgroup analysis on the hold-out test set. We stratified the patients based on sex (Female vs. Male), age (\u0026ge;\u0026thinsp;60 years vs. \u0026lt;60 years), and history of previous infection.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eExplanation of ML results\u003c/h2\u003e \u003cp\u003eSHAP (Shapley Additive explanations) value is a technique used to elucidate the predictions generated by ML models. They facilitate the identification of how each feature contributes to the model's ultimate prediction. A positive SHAP value indicates that the feature contributes to the development of hospitalised infection. In contrast, a negative SHAP value means the feature is negatively associated with the development of hospitalised infection. We computed the SHAP value based on the training cohort. Features were ranked in descending order based on their importance (computed as mean absolute SHAP value). We visualised explanation diagrams at the dataset level, depicting the contribution of each feature to the prediction for each patient. To differentiate the infection risk across different b/ts DMARDs, the SHAP values of each b/ts DMARDs were further visualised in a violin plot in ascending order.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003ePartial Dependence Plots (PDP) illustrate influence of biomarkers\u003c/h2\u003e \u003cp\u003eTo enhance the interpretability of the random forest model, we generated PDPs\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e for the studied biomarkers, trimmed to the 5th\u0026ndash;95th percentile range to exclude extreme values. PDPs visualize the marginal effect of a specific feature on the predicted probability of infection while averaging out the effects of all other variables. Additionally, we calculated quantitative metrics to characterize each biomarker's influence. The effect magnitude was defined as the range between maximum and minimum predicted probabilities across the PDP curve, representing the total impact of that biomarker on infection risk. To identify clinically meaningful thresholds, we computed the first derivative of the PDP curve and located the point of maximum slope. This \"tipping point\" indicates the biomarker value at which infection risk changes most rapidly, providing actionable thresholds for clinical decision-making.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eSample size calculation\u003c/h2\u003e \u003cp\u003eTo ensure sufficient sample sizes for the development of robust models, we calculated the minimum sample size required for predicting time-to-event outcomes clinically proposed by Riley et al\u003csup\u003e37\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eExternal validation\u003c/h2\u003e \u003cp\u003eFollowing the same data processing methodologies as those used for the CDARS cohort, we developed a new model using the All of Us cohort as external validation. Glucocorticoids, NSAIDs, and opioids were excluded from the external validation by designating 0 to the three predictive features in the All of Us cohort during the model fitting process, as the prescription daily dose of these medications is not available in the All of Us database.\u003c/p\u003e \u003cp\u003eData analysis and data visualisation were conducted using R version 4.1.4 (R Foundation for Statistical Computing, Vienna, Austria) and cross-checked by three independent researchers (KP, JQW and DLY). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD)\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e checklists was reported to ensure quality assurance and reporting standards.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003ePatient characteristics\u003c/h2\u003e \u003cp\u003eIn total, we identified 3159 and 1845 patients newly prescribed with b/ts DMARDs in the training cohort CDARS and the validation cohort All of Us, respectively. Among them, 277 (8.8%) and 52 (2.8%) serious infections occurred during the study period. Predictive features used to train the model are summarised in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Compared with the non-infection group, the infected group had poorer baseline health conditions including older age, higher inflammation level reflected by ESR and CRP values, increased comorbidity and more medication.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBaseline characteristics of the developing cohort (CDARS) and the validation cohort (All of Us)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003eCDARS\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c7\" namest=\"c5\"\u003e \u003cp\u003eAll of Us\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNot Infected\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eInfected\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eP value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNot Infected\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInfected\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eP value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of patients\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2882\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e277\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1793\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eDemographic characteristics\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMale (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e542 (18.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e69 (24.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.017\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e362 (20.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e7 (13.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.3078\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge (mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e54.70 (13.55)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e63.47 (14.29)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e54.16 (13.72)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e50.83 (16.28)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.086\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePrevious infection within 3 years (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e278 (9.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e79 (28.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e77 (4.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e16 (30.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eLaboratory tests\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlanine transaminase (U/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e22.53 (19.63)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e23.51 (21.14)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.441\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24.30 (15.73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e27.21 (20.96)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.312\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlbumin (g/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e38.71 (4.59)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e36.10 (5.86)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e39.43 (5.78)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e35.61 (6.81)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCRP (mg/dL mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.10 (2.73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3.41 (3.92)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.44 (2.31)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.26 (4.62)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCreatinine (umol/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e66.01 (22.68)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e78.88 (47.24)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e266.05 (672.83)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e375.70 (705.85)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.285\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eESR (mm/hr mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e54.60 (32.04)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e62.43 (32.53)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e26.74 (22.95)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e53.89 (30.78)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLymphocytes (% mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e22.78 (8.49)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e19.62 (8.75)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24.80 (9.59)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e21.12 (9.43)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.018\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePlatelet (10^9/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e300.38 (93.56)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e279.51 (109.28)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e266.75 (73.90)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e292.17 (95.57)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.034\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRed blood cell (10^12/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e4.18 (0.55)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3.91 (0.60)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4.24 (0.49)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4.03 (0.61)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.011\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBilirubin (umol/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.93 (7.79)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8.65 (6.40)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.155\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.89 (3.86)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8.25 (5.44)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.58\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUrea (mmol/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5.05 (1.73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.29 (4.16)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.77 (2.97)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.71 (2.24)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.907\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWhite blood cell (10^9/L mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.23 (2.50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8.16 (5.66)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7.37 (2.82)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e8.73 (3.90)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.003\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eBaseline comorbidities\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCoronary artery disease and heart failure (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e94 (3.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e34 (12.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e101 (5.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5 (9.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.361\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHyperlipidaemia (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e144 (5.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31 (11.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e22 (1.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e2 (3.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.307\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eThyroid disorder (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e95 (3.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e16 (5.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.049\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOsteoporosis (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59 (2.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17 (6.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e310 (17.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e12 (23.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.369\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHypertension (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e384 (13.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e81 (29.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e787 (43.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e30 (57.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.067\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDiabetes mellitus (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e169 (5.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e37 (13.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e368 (20.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e16 (30.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.105\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCerebrovascular disease (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e65 (2.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12 (4.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.053\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e94 (5.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6 (11.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.096\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChronic obstructive pulmonary disease (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e139 (4.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e30 (10.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e101 (5.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5 (9.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.361\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUlcers (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e133 (4.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e20 (7.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.075\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e80 (4.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e6 (11.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.04\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCancer (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e118 (4.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e36 (13.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10 (0.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eBaseline prescription\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePrednisolone (mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.27 (0.35)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.36 (0.43)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNSAID (mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.70 (0.63)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.59 (0.66)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOpioid (mean (SD))\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.05 (0.17)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.09 (0.22)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eB/ts DMARDs episodes\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.058\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAbatacept\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e114 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17 (6.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e116 (6.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e2 (3.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAdalimumab\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e503 (17.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e33 (11.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e597 (33.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e20 (38.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBaricitinib\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e110 (3.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e31 (11.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCertolizumab Pegol\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e100 (3.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8 (2.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e49 (2.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEtanercept\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e533 (18.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e34 (12.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e474 (26.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e7 (13.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGolimumab\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e354 (12.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14 (5.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e23 (1.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInfliximab\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e119 (4.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12 (4.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e141 (7.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e2 (3.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRituximab\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e164 (5.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e45 (16.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e172 (9.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e11 (21.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSarilumab\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e41 (1.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1 (0.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4 (0.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTocilizumab\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e335 (11.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50 (18.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e88 (4.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e4 (7.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTofacitinib\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e450 (15.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e29 (10.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e107 (6.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3 (5.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUpadacitinib\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e59 (2.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3 (1.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19 (1.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0 (0.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"7\"\u003eB/ts DMARDs: Biologic or targeted synthetic disease-modifying antirheumatic drugs; CRP: C-reactive protein; ESR: Erythrocyte Sedimentation Rate; SD: Standard deviation; NSAID: non-steroidal anti-inflammatory drug.\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eModel evaluation and external validation\u003c/h2\u003e \u003cp\u003eThe classification accuracy on the testing dataset of six trained models is described in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e-A and \u003cb\u003eSupplementary Table\u0026nbsp;4\u003c/b\u003e; the random forest model shows the best predictive performance [AUROC: 0.840 (95%CI: 0.793\u0026ndash;0.888)] in CDARS and was therefore selected as the final model for external validation in All of Us. The random forest model generated predictive accuracy [AUROC: 0.729 (95%CI: 0.665\u0026ndash;0.793)] in the external validation cohort (All of Us, Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e-B\u003cb\u003e)\u003c/b\u003e. The developed model satisfied the minimum sample size (1756 individuals with 154 events) required for new model development given the resulting model performance (AUROC 0.840), number of predictive features (28) and incidence of serious infection (8.8%).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eModel calibration\u003c/h2\u003e \u003cp\u003eThe random forest model demonstrated good overall performance with a Brier score of 0.073 (\u003cb\u003eSupplemental Fig.\u0026nbsp;2\u003c/b\u003e). While calibration analysis indicated some degree of risk underestimation, a common finding in machine learning models applied to imbalanced datasets (Intercept\u0026thinsp;=\u0026thinsp;3.503, Slope\u0026thinsp;=\u0026thinsp;2.496).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eDecision Curve Analysis\u003c/h2\u003e \u003cp\u003eDecision curve analysis demonstrated that the model provided net benefit compared to default strategies (treat-all or treat-none) across threshold probabilities ranging from 3.1% \u0026minus;\u0026thinsp;32.4%. At the threshold probability of 8.5% the model provided the greatest net benefit gain (0.0438) compared to treat-all strategy (\u003cb\u003eSupplemental Fig.\u0026nbsp;3\u003c/b\u003e). In clinical practice, this means that for every 100 patients evaluated, using this model to guide decision-making would result in approximately 4 additional patients correctly receiving needed treatment while avoiding unnecessary interventions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eSubgroup analysis\u003c/h2\u003e \u003cp\u003eThe model demonstrated consistent discrimination across all subgroups (\u003cb\u003eSupplemental Fig.\u0026nbsp;4\u003c/b\u003e), with AUROC values\u0026thinsp;\u0026ge;\u0026thinsp;0.73 in all categories. Performance was higher in females (AUROC\u0026thinsp;=\u0026thinsp;0.859) than males (AUROC\u0026thinsp;=\u0026thinsp;0.760), and similar between age groups (\u0026lt;\u0026thinsp;60 years: 0.819; \u0026ge;60 years: 0.778). Patients without prior infection history showed higher AUROC (0.840) compared to those with prior infection (0.732). Lower performance in males and patients with infection history may partly reflect smaller sample sizes (N\u0026thinsp;=\u0026thinsp;131 and N\u0026thinsp;=\u0026thinsp;72, respectively), resulting in wider confidence intervals.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eExplanations of ML results\u003c/h2\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e visualised the SHAP values that shows the importance in descending order and distribution of features across all data points. The colour represents the feature's value for a particular data point. Particularly, for binary variables such as comorbidity, a low value represents 0, and a high value represents 1; for multi-class factor variables such as b/ts DMARDs, each value represents a unique pharmaceutical product. The top three important features were previous infection within three years, followed by b/ts DMARDs use and diabetes mellitus. The SHAP values of b/ts DMARDs were plotted in violin plots, with the median value plotted in black points (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Among b/ts DMARDs, tofacitinib and upadacitinib have the lowest SHAP values indicating less predicted infection risk. In contrast, rituximab held the highest SHAP values linked to increased predicted infection risk. Similarly, categorised by therapeutic area, sarilumab has the lowest predicted risk of infection among non-TNFi agents, while tocilizumab and rituximab have the highest predicted risks of infection. Tofacitinib and upadacitinib show low predicted risks of infection among JAKi, while golimumab and etanercept have relatively lower predicted risks of infection among TNFi agents.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003ePartial Dependence Plots (PDP) illustrate influence of biomarkers\u003c/h2\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cb\u003eSupplemental table 5\u003c/b\u003e illustrate the marginal effect of each laboratory parameter on the predicted probability of infection when holding all other model features constant. CRP showed the largest effect range (0.0132) with a threshold at 7.1 mg/dL and positive slope (+\u0026thinsp;0.0053). ESR threshold was 103.0 mm/hr with positive slope (+\u0026thinsp;0.0005). Albumin had an effect range of 0.0071 with threshold at 30.3 g/L and negative slope (-0.0036). Lymphocytes showed effect range of 0.0089 with threshold at 23.6% and negative slope (-0.0012). White blood cell threshold was 9.63 \u0026times;10⁹/L with positive slope (+\u0026thinsp;0.0032). Creatinine threshold was 95.3 \u0026micro;mol/L (+\u0026thinsp;0.0008) and Urea was 3.02 mmol/L (+\u0026thinsp;0.0040). Red blood cell had effect range 0.0094 with threshold at 3.45 \u0026times;10\u0026sup1;\u0026sup2;/L and negative slope (-0.0305).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eFor optimised medication safety management among patients with RA, the study attempted to predict serious infection risk among patients receiving b/ts DMARDs treatment. All predictive features were derived from routine clinical data, ensuring the predictor\u0026rsquo;s ease of implementation in practice. It performs well in the training cohort (AUROC 0.840), while maintaining acceptable performance in the external validation cohort (AUROC 0.729). Performance differences may be attributed to variations in racial and ethnic composition and coding systems between the two cohorts. Furthermore, differences in infection rates may reflect disparities in the granularity of infection classification and coding practices across the two databases. Nevertheless, the study findings still showed a strong generalisability and reproducibility of our model. A previous study, the Rheumatoid Arthritis Observation of Biologic Therapy (RABBIT) study, used multivariable generalised estimation equations to predict the absolute risk of serious infection among patients taking b/ts DMARDs\u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e. That study was performed in 2011 and did not consider newer treatments such as JAKi and biomarkers. In contrast, we successfully identified implementable biomarkers that serve as risk prediction with ML; we also established straightforward cut-off thresholds to simplify infection risk stratification. Comparing the model performance using traditional logistic regression as reference (AUROC 0.830), the random forest model showed slightly better performance (AUROC difference: + 0.010), algorithms like XGBoost had a lower AUROC (AUROC difference: \u0026minus;\u0026thinsp;0.035). This finding is in line with prior research that has found that advanced ML algorithms do not necessarily lead to improved performance over traditional logistic regression\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003eExplanations of ML and PDPs results\u003c/h2\u003e \u003cp\u003eWe elucidated model interpretability by ranking feature importance (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The risk factors that were the highest ranked variables contributing to the risk of infection included previous infection, diabetes mellitus, hypertension, and use of symptom control medications such as opioids, NSAIDs and steroids. These risk factors have all been documented in previous reviews\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eWe used PDP (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) to quantify the continuous relationship between biomarker values and predicted infection probability, characterizing both the effect range and direction for each variable. We found that both elevated CRP and ESR contribute to the infection risk positively, consistent with their well-established role as classical biomarkers measuring inflammatory status, and mounting evidence supports their use in the diagnosis of infection at different sites\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e,\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. Albumin exhibited a negative relationship with infection risk, which is expected given that hypoalbuminemia reflects both malnutrition and chronic inflammation; low albumin has been consistently associated with poor outcomes in infectious diseases\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e. The ratio of blood urea nitrogen to albumin is a multifaceted indicator of renal and inflammatory function; a high urea nitrogen and low albumin level has shown a strong correlation with serious infection\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. Similarly, the lymphocytes (proportion to white blood cells) showed a negative slope with threshold at 23.6%, supporting the notion that lymphopenia compromises immunity against pathogens\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e. Haematological manifestations, including anaemia (low red blood cells) and thrombocytopenia (low platelet), are common among patients with RA and likely associated with the autoimmune disease itself and anti-rheumatic drugs\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e. Elevated creatinine levels could indicate kidney infection and failure; prior reviews have identified chronic kidney disease as a risk factor for acute community-acquired infections\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e. Collectively, these biomarkers capture multiple physiological dimensions span from inflammation, immune function, nutritional status, and organ function that contribute to infection susceptibility.\u003c/p\u003e \u003cp\u003eAdditionally, we computed the first derivative of each PDP curve to identify \"tipping points,\" meaning threshold values at which infection risk changes most rapidly. While this does not constitute a validated diagnostic cut-off, it highlights biomarker ranges that may merit heightened clinical vigilance and could serve as candidate thresholds for future prospective evaluation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section2\"\u003e \u003ch2\u003eInfection risk associated with different b/ts DMARDs\u003c/h2\u003e \u003cp\u003eOur findings align with previous meta-analyses of 58 randomised controlled trials (RCTs), which detected that etanercept and golimumab present a reduced risk of serious infections in comparison to other TNFi agents\u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. This finding was further validated by a nationwide register-based cohort study in Sweden that recruited over 20,000 patients with RA. The study examined the risk of serious infection across 11 different b/ts DMARDs, revealing that etanercept and golimumab exhibit a superior safety profile compared to infliximab, adalimumab, and certolizumab\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. Our findings indicate that the JAKi tofacitinib and upadacitinib presents a lower infection risk compared to other b/ts DMARDs, consistent with the Swedish cohort study\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e. However, existing literature reported that incidences of herpes zoster and tuberculosis are greater among those treated with tofacitinib compared to TNFi\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e,\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e. Such differences could arise from 1) our adaptation of a different definition covering a broader range of infection outcomes 2) ORAL-Surveillance study\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e which reported a higher risk of major cardiovascular events and cancers with tofacitinib compared with TNFi, JAKi are thus used cautiously in patients\u0026thinsp;\u0026gt;\u0026thinsp;65, smokers, or those at high cardiovascular risk, leading to fewer high-risk patients receiving JAKi and potentially contributing to the observed lower infection risk. In terms of non-TNFi agents, our results are consistent with existing studies where rituximab targeting CD20 has the highest risk in comparison to sarilumab, an interleukin-6 inhibitor shows the lowest risk\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e,\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e. We found that another interleukin\u0026thinsp;\u0026minus;\u0026thinsp;6 inhibitor (tocilizumab) showed a higher infection risk than sarilumab. This likely reflects potential indication bias: tocilizumab is often used in more refractory with greater comorbidity or steroid use, while the new sarilumab are selected for lower-risk patients, such prescribing pattern have been reported previously in Germany, with tocilizumab recipients being systematically sicker than TNFi recipients\u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003eStrength and clinical implications\u003c/h2\u003e \u003cp\u003eOur study possesses several advantages. Firstly, it provides clinically actionable absolute probabilities (0-100%) of one-year serious infection risk upon clinical visit, enabling clinicians to stratify patients into high-risk and low-risk groups and tailor treatment strategies based on the calculated risk. For patients with increased vulnerability to serious infections who have yet to take b/ts DMARDs for disease activity control, clinicians can select b/ts DMARDs associated with the lowest infection risk, specifically tailored to the patient's profile and unique clinical characteristics. Additionally, clinicians can taper the doses of opioids, NSAIDs, and prednisolone to maintain infection risk within an acceptable range. This approach not only offers practical strategies for rheumatology treatment to balance the risks and benefits of b/ts DMARDs, but also actively involves patients in shared decision-making by informing them of their infection risk during clinical visits. For high-risk patients, clinicians can also recommend pre-treatment screenings for latent infections (e.g., hepatitis B, tuberculosis) and advise patients to receive vaccinations promptly to prevent potential infections. Secondly, the success of external validation in a United States population supports the broader clinical applicability of the model, rendering a more reliable tool for guiding treatment decisions in diverse settings. Thirdly, the robustness of serious infection identification was strengthened by using the ICD diagnostic code and anti-microbial prescription; this dual approach minimises misclassification by confirming clinical relevance and capturing true infection cases. Lastly, the model requires limited clinical information to function without sacrificing model performance, making it easy to build into the clinical decision system.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec29\" class=\"Section2\"\u003e \u003ch2\u003eLimitations\u003c/h2\u003e \u003cp\u003eThis study has certain limitations. Infection risk is multifactorial and dynamic, shaped by environmental factors (living conditions, hygiene, healthcare access) that are difficult to quantify, and by time-dependent factors (recent travel, disease flares, medication changes) not captured by baseline features; while this limitation is common to most clinical prediction models. To minimize the influence of time-varying features and focus on the most clinically relevant period we restricted the observation window to 12 months; while this horizon may miss later events, it was deliberately chosen because it aligns with decision points for DMARD selection and early monitoring and captures the interval of highest infection risk after treatment initiation. Also, during external validation using the All of Us cohort, three medication-related features (glucocorticoids, NSAIDs, and opioids daily dose) were not available and were therefore set to zero during model fitting. This may have attenuated the model's performance in the external cohort, as these medications contributed to risk prediction in the original model. Despite this limitation, the model still achieved an AUROC of 0.729 in external validation, and subgroup analyses demonstrated consistent discrimination across patient subgroups (AUROC\u0026thinsp;\u0026ge;\u0026thinsp;0.73), suggesting reasonable robustness. Lastly, due to the nature of the EHR database and the aim of using predictors readily available at the point of care, important predictive features, such as disease activities and patient-reported outcomes were not included in the model development.\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study developed a ML model to identify patients with an elevated risk of serious infection using routine clinical data exclusively. It provides clinical insights on risk stratification of different b/ts DMARDs therapies, predictive cut-off threshold on biomarkers, and importance ranking of risk factors. The prediction tool has the potential to streamline treatment decisions using available clinical inputs and will help to identify high-risk patients for enhanced safety management, and patient education.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRheumatoid arthritis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eb/ts DMARDs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eBiologic or targeted synthetic disease-modifying antirheumatic drugs\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eML\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMachine learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eEHR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eElectronic health records\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCDARS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eClinical Data Analysis and Reporting System\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAUROC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArea under the receiver operating characteristic curve\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSHAP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eShapley Additive explanations\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eConfidence interval\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHazard ratio\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eaHR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAdjusted hazard ratio\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTNFi\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTumour necrosis factor inhibitors\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTNFα\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTumour necrosis factor alpha\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eILi\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInterleukin inhibitors\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eIL-6\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInterleukin-6\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eJAKi\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eJanus Kinase inhibitors\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCRP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eC-reactive protein\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eESR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eErythrocyte sedimentation rate\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHK\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHong Kong\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHospital Authority\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eICD-9-CM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInternational Classification of Diseases, Ninth Revision, Clinical Modification\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eICD-10-CM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInternational Classification of Diseases, Tenth Revision, Clinical Modification\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eBNF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eBritish National Formulary\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eATC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAnatomical Therapeutic Chemical\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNSAID\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNon-steroidal anti-inflammatory drug\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eKNN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eK-nearest neighbor imputation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMICE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMultiple Imputation by Chained Equations\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLASSO\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLeast absolute shrinkage and selection operator\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eXGBoost\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eExtreme Gradient Boosting\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDCA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDecision Curve Analysis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePDP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ePartial Dependence Plots\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTRIPOD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTransparent reporting of a multivariable prediction model for individual prognosis or diagnosis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRCTs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRandomised controlled trials\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRABBIT\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRheumatoid Arthritis Observation of Biologic Therapy\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eARTIS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAnti-Rheumatic Therapies in Sweden\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDREAM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDutch Rheumatoid Arthritis Monitoring\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eWHO\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eWorld Health Organization\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDDD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDefined Daily Dose\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCD20\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCluster of differentiation 20\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTrue positive\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eFalse positive\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eWBC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eWhite blood cell\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRBC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRed blood cell\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCOPD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eChronic obstructive pulmonary disease\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eStandard deviation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHMRF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHealth and Medical Research Fund\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHKSAR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHong Kong Special Administrative Region\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRGC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eResearch Grants Council\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eECS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eEarly Career Scheme\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRIF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eResearch Impact Fund\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eADAMS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAdvanced Data Analytics for Medical Science\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eD\u0026sup2;4H\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLaboratory of Data Discovery for Health\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eEULAR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eEuropean Alliance of Associations for Rheumatology\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eEthics approval and consent to participate\u003c/h2\u003e \u003cp\u003e Ethical approval for the research was granted by the Institutional Review Board of The University of Hong Kong / Hospital Authority Hong Kong West Cluster under reference numbers UW 21\u0026ndash;338. The Institutional Review Board granted a waiver of informed patient consent. This was justified as the studies were exclusively observational, meaning no practical interventions were administered, and all patient-level data utilized were de-identified by the data custodian at the time of data extraction.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication\u003c/strong\u003e \u003cp\u003eNot applicable\u003c/p\u003e \u003c/p\u003e\u003cp\u003e\u003ch2\u003eCompeting Interests\u003c/h2\u003e\u003cp\u003eXue Li reports a relationship with Janssen Pharmaceuticals Inc that includes: consulting or advisory. Xue Li reports a relationship with Pfizer that includes: consulting or advisory. Xue Li reports a relationship with Amgen Inc that includes: consulting or advisory. Xue Li reports a relationship with Merck Sharp \u0026amp; Dohme UK Ltd that includes: consulting or advisory. Xue Li reports a relationship with OPEN Health Communications LLP that includes: consulting or advisory. Xue Li reports a relationship with Office of Health Economics that includes: consulting or advisory. Xue Li reports a relationship with ADAMS Limited Hong Kong that includes: employment. XL received research grants from the Research Fund Secretariat of the Health Bureau, Health and Medical Research Fund (HMRF, HKSAR), Health and Medical Research Fund Fellowship Scheme (HMRF Fellowship, HKSAR), Research Grants Council Early Career Scheme (RGC/ECS, HKSAR), Research Grants Council Research Impact Fund (RGC/RIF, HKSAR), Commission grants from Hospital Authority of Hong Kong; educational and investigator initiate research fund from Janssen, Pfizer and Amgen; internal funding from the University of Hong Kong; consultancy fee from Pfizer, Merck Sharp \u0026amp; Dohme, Open Health, Office of Health Economics; she is also the former non-executive director of ADAMS Limited Hong Kong; all outside the submitted work. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eEnhanced Start-up Fund for new academic staff, LKS Faculty of Medicine, The University of Hong Kong; Internal Research Fund, Department of Medicine, School of Clinical Medicine, LKS Faculty of Medicine, The University of Hong Kong.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eKP, DLY, and JQW performed the data analysis. CYS, MCYC, and ECCL conducted the external validation. KP, SCWC, and XL conceptualized the study. KP and XL wrote the manuscript. SCWC, NLP, ICKW, CSL, JJG, QPZ, and IYKT contributed to the interpretation of results. XL supervised the study and obtained funding. All authors reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eAcknowledgements and affiliationsWe thank Lisa Lam for proofreading this paper.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe data that support the findings of this study are available from the Hospital Authority of Hong Kong, but restrictions apply to the availability of these data, which were used under license for the current study, and so are not publicly available. Data are however available from the authors upon reasonable request and with permission of the Hospital Authority of Hong Kong.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eJoachim L, J\u0026ouml;rn K, Bernhard M, et al. Mortality in rheumatoid arthritis: the impact of disease activity, treatment with glucocorticoids, TNFα inhibitors and rituximab. Ann Rheum Dis. 2015;74(2):415.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFilipovic I, Walker D, Forster F, Curry AS. Quantifying the economic burden of productivity loss in rheumatoid arthritis. Rheumatology. 2011;50(6):1083\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJosef SS, Robert BML, Sytske Anne B, et al. EULAR recommendations for the management of rheumatoid arthritis with synthetic and biological disease-modifying antirheumatic drugs: 2022 update. Ann Rheum Dis. 2023;82(1):3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa X, Xu S. TNF inhibitor therapy for rheumatoid arthritis. Biomed Rep. 2013;1(2):177\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOgata A, Kato Y, Higa S, Yoshizaki K. IL-6 inhibitor for the treatment of rheumatoid arthritis: A comprehensive review. Mod Rheumatol. 2019;29(2):258\u0026ndash;67.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYamaoka K. Janus kinase inhibitors for rheumatoid arthritis. Curr Opin Chem Biol. 2016;32:29\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSepriano A, Kerschbaumer A, Bergstra SA, et al. Safety of synthetic and biological DMARDs: a systematic literature review informing the 2022 update of the EULAR recommendations for the management of rheumatoid arthritis. Ann Rheum Dis. 2023;82(1):107.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan den Hoek J, Boshuizen HC, Roorda LD, et al. Mortality in patients with rheumatoid arthritis: a 15-year prospective cohort study. Rheumatol Int. 2017;37(4):487\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOzen G, Pedro S, England BR, Mehta B, Wolfe F, Michaud K. Risk of Serious Infection in Patients With Rheumatoid Arthritis Treated With Biologic Versus Nonbiologic Disease-Modifying Antirheumatic Drugs. ACR Open Rheumatol. 2019;1(7):424\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJani M, Barton A, Hyrich K. Prediction of infection risk in rheumatoid arthritis patients treated with biologics: are we any closer to risk stratification? Curr Opin Rheumatol. 2019;31(3):285\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOzen G, Pedro S, Schumacher R, Simon TA, Michaud K. Safety of abatacept compared with other biologic and conventional synthetic disease-modifying antirheumatic drugs in patients with rheumatoid arthritis: data from an observational study. Arthritis Res Ther. 2019;21(1):141.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiles JT, Sattar N, Gabriel S, et al. Cardiovascular Safety of Tocilizumab Versus Etanercept in Rheumatoid Arthritis: A Randomized Controlled Trial. Arthritis Rheumatol. 2020;72(1):31\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStrand V, Ahadieh S, French J, et al. Systematic review and meta-analysis of serious infections with tofacitinib and biologic disease-modifying antirheumatic drug treatment in rheumatoid arthritis clinical trials. Arthritis Res Ther. 2015;17:362.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Dartel SA, Fransen J, Kievit W, et al. Difference in the risk of serious infections in patients with rheumatoid arthritis treated with adalimumab, infliximab and etanercept: results from the Dutch Rheumatoid Arthritis Monitoring (DREAM) registry. Ann Rheum Dis. 2013;72(6):895\u0026ndash;900.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eListing J, Gerhold K, Zink A. The risk of infections associated with rheumatoid arthritis, with its comorbidity and treatment. Rheumatology. 2013;52(1):53\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeorge MD, Baker JF, Winthrop K, et al. Risk for Serious Infection With Low-Dose Glucocorticoids in Patients With Rheumatoid Arthritis: A Cohort Study. Ann Intern Med. 2020;173(11):870\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim KJ, Tagkopoulos I. Application of machine learning in rheumatic disease research. Korean J Intern Med. 2019;34(4):708\u0026ndash;22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoo BS, Eun S, Shin K, et al. Machine learning model for identifying important clinical features for predicting remission in patients with rheumatoid arthritis treated with biologics. Arthritis Res Ther. 2021;23(1):178.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBouget V, Duquesne J, Hassler S et al. Machine learning predicts response to TNF inhibitors in rheumatoid arthritis: results on the ESPOIR and ABIRISK cohorts. RMD Open 2022;8(2).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuquesne J, Bouget V, Courn\u0026egrave;de PH, et al. Machine learning identifies a profile of inadequate responder to methotrexate in rheumatoid arthritis. Rheumatology. 2023;62(7):2402\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim AJW, Tyniana CT, Lim LJ, et al. Robust SNP-based prediction of rheumatoid arthritis through machine-learning-optimized polygenic risk score. J Transl Med. 2023;21(1):92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIntroduction. Caring for our community\u0026rsquo;s health. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ha.org.hk/visitor/ha_visitor_index.asp?Parent_ID=10004\u0026amp;Content_ID=10008\u0026amp;Ver=HTML\u003c/span\u003e\u003cspan address=\"https://www.ha.org.hk/visitor/ha_visitor_index.asp?Parent_ID=10004\u0026amp;Content_ID=10008\u0026amp;Ver=HTML\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed July 6, 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi X, Tong X, Yeung WWY, et al. Two-dose COVID-19 vaccination and possible arthritis flare among patients with rheumatoid arthritis in Hong Kong. Ann Rheum Dis. 2022;81(4):564\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChai Y, Luo H, Wong GHY, et al. Risk of self-harm after the diagnosis of psychiatric disorders in Hong Kong, 2000-10: a nested case-control study. Lancet Psychiatry. 2020;7(2):135\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBick AG, Metcalf GA, Mayo KR, et al. Genomic data in the All of Us Research Program. Nature. 2024;627(8003):340\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eConran CA, Moreland LW. A review of biosimilars for rheumatoid arthritis. Curr Opin Pharmacol. 2022;64:102234.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCurtis JR, Winthrop K, O'Brien C, Ndlovu MN, de Longueville M, Haraoui B. Use of a baseline risk score to identify the risk of serious infectious events in patients with rheumatoid arthritis during certolizumab pegol treatment. Arthritis Res Ther. 2017;19(1):276.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGalloway JB, Hyrich KL, Mercer LK, et al. Risk of septic arthritis in patients with rheumatoid arthritis and the effect of anti-TNF therapy: results from the British Society for Rheumatology Biologics Register. Ann Rheum Dis. 2011;70(10):1810\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eATC/DDD Index. 2024. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://atcddd.fhi.no/atc_ddd_index/\u003c/span\u003e\u003cspan address=\"https://atcddd.fhi.no/atc_ddd_index/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed Oct 29th, 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTibshirani R. Regression Shrinkage and Selection via the Lasso. J Royal Stat Soc Ser B (Methodological). 1996;58(1):267\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCortes C, Vapnik V. Support-vector networks. Mach Learn. 1995;20(3):273\u0026ndash;97.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRipley BD. Pattern Recognition and Neural Networks. Cambridge: Cambridge University Press; 1996.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLema\u0026Atilde;Žtre G, Nogueira F, Aridas CKJJ. Imbalanced-learn: A python toolbox to tackle the curse of imbalanced datasets in machine learning. 2017;18(17):1\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePozzolo AD, Caelen O, Johnson RA, Bontempi G. Dec. Calibrating Probability with Undersampling for Unbalanced Classification. Paper presented at: 2015 IEEE Symposium Series on Computational Intelligence; 7\u0026ndash;10 2015, 2015.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis making: Int J Soc Med Decis Mak. 2006;26(6):565\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFriedman JH. Greedy Function Approximation: A Gradient Boosting Machine. Annals Stat. 2001;29(5):1189\u0026ndash;232.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRiley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. 2020;368:m441.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollins GS, Reitsma JB, Altman DG, Moons KG. Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med. 2015;162(1):55\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZink A, Manger B, Kaufmann J, et al. Evaluation of the RABBIT Risk Score for serious infections. Ann Rheum Dis. 2014;73(9):1673.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChristodoulou E, Ma J, Collins GS, Steyerberg EW, Verbakel JY, Van Calster B. A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J Clin Epidemiol. 2019;110:12\u0026ndash;22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWaseem A-A, Laurie T, James V, et al. The association between diabetes mellitus and incident infections: a systematic review and meta-analysis of observational studies. BMJ Open Diabetes Res Care. 2017;5(1):e000336.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLapić I, Padoan A, Bozzato D, Plebani M. Erythrocyte Sedimentation Rate and C-Reactive Protein in Acute Inflammation: Meta-Analysis of Diagnostic Accuracy Studies. Am J Clin Pathol. 2020;153(1):14\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLindsay CP, Olcott CW, Del Gaizo DJ. ESR and CRP are useful between stages of 2-stage revision for periprosthetic joint infection. Arthroplasty Today. 2017;3(3):183\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWiedermann CJ. Hypoalbuminemia as Surrogate and Culprit of Infections. Int J Mol Sci 2021;22(9).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUgajin M, Yamaki K, Iwamura N, Yagi T, Asano T. Blood urea nitrogen to serum albumin ratio independently predicts mortality and severity of community-acquired pneumonia. Int J Gen Med. 2012;5:583\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou YQ, Feng DY, Li WJ, et al. Lower neutrophil-to-lymphocyte ratio predicts high risk of multidrug-resistant Pseudomonas aeruginosa infection in patients with hospital-acquired pneumonia. Ther Clin Risk Manag. 2018;14:1863\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBowman SJ. Hematological manifestations of rheumatoid arthritis. Scand J Rheumatol. 2002;31(5):251\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcDonald HI, Thomas SL, Nitsch D. Chronic kidney disease as a risk factor for acute community-acquired infections in high-income countries: a systematic review. BMJ Open. 2014;4(4):e004100.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMinozzi S, Bonovas S, Lytras T, et al. Risk of infections using anti-TNF agents in rheumatoid arthritis, psoriatic arthritis, and ankylosing spondylitis: a systematic review and meta-analysis. Expert Opin Drug Saf. 2016;15(sup1):11\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrisell T, Bower H, Morin M, et al. Safety of biological and targeted synthetic disease-modifying antirheumatic drugs for rheumatoid arthritis as used in clinical practice: results from the ARTIS programme. Ann Rheum Dis. 2023;82(5):601.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYtterberg Steven R, Bhatt Deepak L, Mikuls Ted R, et al. Cardiovascular and Cancer Risk with Tofacitinib in Rheumatoid Arthritis. N Engl J Med. 2022;386(4):316\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrisell T, Bower H, Morin M, et al. Safety of biological and targeted synthetic disease-modifying antirheumatic drugs for rheumatoid arthritis as used in clinical practice: results from the ARTIS programme. Ann Rheum Dis. 2023;82(5):601\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGr\u0026oslash;n KL, Arkema EV, Glintborg B, et al. Risk of serious infections in patients with rheumatoid arthritis treated in routine care with abatacept, rituximab and tocilizumab in Denmark and Sweden. Ann Rheum Dis. 2019;78(3):320\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBackhaus M, Kaufmann J, Richter C, et al. Comparison of tocilizumab and tumour necrosis factor inhibitors in rheumatoid arthritis: a retrospective analysis of 1603 patients managed in routine clinical practice. Clin Rheumatol. 2015;34(4):673\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Machine Learning, Biological disease modifying anti-rheumatic drugs, Rheumatoid arthritis, Infection","lastPublishedDoi":"10.21203/rs.3.rs-9054871/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9054871/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003ePatients with rheumatoid arthritis (RA) initiating biologic or targeted synthetic disease-modifying antirheumatic drugs (b/ts DMARDs) face elevated risk of serious infections, necessitating tools for individualized risk stratification.\u003c/p\u003e\u003ch2\u003eObjectives\u003c/h2\u003e \u003cp\u003ePrimary objective was to develop and validate a clinically interpretable machine learning (ML) model to predict 1-year risk of serious infection after b/ts DMARD initiation; secondary objectives were to estimate infection incidence and identify key predictors associated with risk.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe performed a retrospective cohort study using territory-wide EHR from Hong Kong\u0026rsquo;s Clinical Data Analysis and Reporting System (CDARS) for model development and internal validation, with external validation in the U.S. All of Us database. The outcome was first serious infection requiring hospitalization within 1 year. Candidate predictors included demographics, comorbidities, prior infections and medications, laboratory markers. Multiple ML algorithms were trained; model selection was based on AUROC, and interpretability was assessed using SHAP.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eA total of 3,159 patients from CDARS (8.8% with serious infections) and 1,845 from All of Us (2.8% with serious infections) were included. The model demonstrated the highest AUROC in internal validation (0.840, 95% CI: 0.793\u0026ndash;0.888) and maintained robust performance in external validation (AUROC: 0.729, 95% CI: 0.665\u0026ndash;0.793). Key predictors included prior infections, diabetes, b/ts DMARD type, and inflammatory markers. Rituximab was linked to the highest infection risk, while tofacitinib and upadacitinib had the lowest.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThis study developed and validated an ML model using routine clinical data to predict serious infection risk in RA patients, supporting personalised treatment and proactive infection management.\u003c/p\u003e","manuscriptTitle":"Predicting infection risk in rheumatoid arthritis patients receiving biological or targeted synthetic disease-modifying anti-rheumatic drugs: an application of machine learning and healthcare big data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-26 02:57:37","doi":"10.21203/rs.3.rs-9054871/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-04-15T08:55:43+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-09T22:18:54+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-29T03:23:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"149681716800116963933022026708273667444","date":"2026-03-25T13:27:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"256660090695209226438807116002826586491","date":"2026-03-24T22:48:36+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"35569788762631966908237753464916234443","date":"2026-03-22T13:56:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"239755202923033248243491155091073933500","date":"2026-03-21T05:49:49+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"136306491678284191264454696083187774671","date":"2026-03-20T10:14:25+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-20T10:03:01+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-20T09:53:01+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-16T08:08:18+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-13T22:21:09+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Informatics and Decision Making","date":"2026-03-13T12:49:36+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"69316112-a220-45d2-848b-e1c68537ac55","owner":[],"postedDate":"March 26th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-07T09:39:06+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-26 02:57:37","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9054871","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9054871","identity":"rs-9054871","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00