Impact of Lifestyle Patterns on Breast Cancer Prognosis: Evidence from a UK Biobank Survival Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Impact of Lifestyle Patterns on Breast Cancer Prognosis: Evidence from a UK Biobank Survival Study Yishan Yao, Xiaobing Zhai, Zhichao Liang, Chi Kin Lam, Hui Xie, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8588282/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Background: Understanding how lifestyle habits influence breast cancer survival is crucial for improving long-term outcomes and guiding individualized care. The purpose of this study is to comprehensively understand how dietary habits, exercise frequency, sleep quality and the frequency of tobacco and alcohol use affect the survival of breast cancer patients by analyzing the synergistic effects of multiple lifestyle habits. Methods: We analyzed data from 14,901 female breast cancer patients in the UK Biobank, a large-scale population-based cohort. Three feature selection methods were applied to identify key prognostic variables. Five survival risk models were compared to investigate the relationship between breast cancer survival and lifestyle habits. Kaplan-Meier survival analysis and SHapley Additive exPlanations (SHAP) interpretability analysis were performed on risk scores calculated from the best-performing model. Results: Through feature selection, we identified that age, number of medications taken, Body Mass Index (BMI), alcohol intake, diastolic blood pressure and smoking status were pinpointed as critical characteristics influencing the survival of breast cancer patients. Among all models, the eXtreme Gradient Boosting (XGBoost) model demonstrated the highest predictive performance (3 years: AUC = 0.72, 6 years: AUC = 0.727, 9 years: AUC = 0.749). Kaplan-Meier analysis showed that patients classified as high-risk based on the median risk score had significantly worse survival outcomes ( P < 0.0001). SHAP analysis further confirmed the dominant influence of age and BMI on mortality risk. Conclusions: This study highlights the prognostic value of lifestyle habits in breast cancer survival. By integrating routine health indicators into interpretable machine learning models, our findings provide a practical foundation for individualized risk assessment and lifestyle-based intervention strategies. Survival analysis Kaplan-Meier XGBoost TimeROC Lifestyle Patterns Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Breast cancer remains a significant global health concern, with its incidence continuing to increase among women worldwide. It is now the most commonly diagnosed malignant tumor in women, second only to skin cancer, accounting for about 30% of all female cancers [ 1 , 2 ] . According to the latest report released in February 2024 by the International Agency for Research on Cancer (IARC) of the World Health Organization, there were an estimated 19.97 million new cancer cases globally in 2022, including more than 2.3 million new cases of breast cancer. Moreover, approximately 670,000 women died from breast cancer, which puts this disease as the most frequent oncological cause of death in women worldwide [ 3 ] . While substantial efforts have been devoted to improving early detection and treatment, understanding the factors that influence long term survival remains critically important. With the growing number of breast cancer survivors and their increasing motivation to adopt healthier lifestyles post-diagnosis [ 4 ] , identifying modifiable behavioral factors that may improve survival outcomes is essential. Such insights could contribute to the development of evidence-based, personalized survivorship strategies aimed at enhancing both longevity and quality of life. Breast cancer survival risk is influenced by a broad spectrum of modifiable lifestyle habits, many of which are increasingly prevalent in modern societies. Adverse habits such as smoking, excessive alcohol consumption, poor stress management, altered reproductive patterns, obesity, and physical inactivity have all significant impacts on the survival risk of breast cancer. Among these, alcohol consumption has been identified as a major contributor to both mortality and disability [ 5 ] , while smoking [ 6 ] facilitates cancer progression by promoting metastasis and impairing overall health. Chronic psychological stress [ 7 ] may indirectly increase the mortality risk by disrupting the endocrine and immune system functions. Additionally, the widespread issue of overweight and obesity [ 8 ] can lead to hormonal imbalances and chronic inflammation, both of which may serve as potential triggers for breast cancer, and a sedentary lifestyle [ 9 ] further compromises the body's immune defense. In addition to these factors, dietary patterns have received increasing attention. Previous studies have demonstrated that high glycemic index or insulin load diets after diagnosis are associated with increased risks of breast cancer (specific mortality and all-cause mortality), emphasizing the long-term impact of nutritional choices on survivorship [ 10 ] . Furthermore, disparities in breast cancer incidence, screening access, and survival outcomes across racial and ethnic groups highlight the importance of developing individualized and culturally sensitive interventions [ 11 ] . Although these modifiable lifestyle factors have been studied individually, they frequently co-occur and may interact in complex ways, potentially influencing outcomes in a synergistic manner. Therefore, an integrated assessment of multiple behaviors is essential for a more comprehensive understanding of their collective influence on breast cancer survival. Although some studies have explored the relationship between a single lifestyle and the survival status of breast cancer patients, studies that systematically evaluate the combined effects of multiple behaviors remain limited [ 12 , 13 ] . Research focusing solely on isolated factors may fail to capture the intricate interplay and cumulative impact of real-world lifestyle patterns. Moreover, prior studies are often constrained by limited sample sizes or population heterogeneity, which may restrict the reliability and generalizability of their findings. To fill this gap, this study utilizes data from the UK Biobank, a large and well-characterized prospective cohort, to investigate how multiple lifestyle habits jointly influence the survival outcomes of breast cancer patients. Specifically, we assess the synergistic effects of dietary habits, physical activity, sleep quality,tobacco and alcohol consumption using a comprehensive machine learning framework. By identifying key behavioral predictors and quantifying their relative contributions to patient prognosis, this study aims to provide actionable evidence to inform personalized rehabilitation strategies and promote healthier lifestyles among breast cancer survivors. Methods Study Population The UK Biobank is a large-scale prospective, population-based cohort study that recruited approximately 500,000 participants (229,041 men and 273,293 women) aged 40–69 years from 22 study evaluation centers across the UK between 2006 and 2010 [ 14 ] . All participants were followed through a link in the NHS’s electronic health record. At baseline, participants completed a self-administered touch-screen questionnaire that collected information on sociodemographic characteristics, health and medical history, and lifestyle-related exposures [ 15 ] . A detailed description of the recruitment procedures and population characteristics has been described elsewhere [ 16 ] . In addition to the questionnaire, participants also received physical measurements and provided blood samples, including blood specimens. All participants provided written informed consent, and the study was approved by the Northwest Multicenter Research Ethics Committee. Individuals who had requested withdraw from the UK Biobank cohort study were excluded from our analysis. Data from biological samples from the United Kingdom ( http://www.ukbiobank.ac.UK/ ) are available to all researchers upon request. Ascertainment of analytic population We restricted the study population to self-reported female participants with a history of breast cancer ( n = 14,901). In this study, we set the following strict inclusion criteria for these women with breast cancer to ensure the accuracy of the study subjects and the validity of the study: (a) Breast cancer diagnosed as the primary malignancy. (b) Death not on the same-day as diagnosis (i.e., survival time > 0). (c) Eventual therapeutic drugs administered to the patient are not intended to treat breast cancer. Exposure Breast cancer cases were identified using inpatient admission records sourced from the Hospital Episode Statistics (England), the Scottish Morbidity Record (Scotland), and the Patient Episode Database (Wales). These records were supplemented with death registration records provided by the National Health Service (NHS) system in England and Wales, as well as the information and statistics departments in Scotland. The diagnosis of breast cancer was based on the International Classification of Diseases, Tenth Revision (ICD-10) coding system, with the ICD codes used to identify breast cancer having been selected and validated by the UK Biobank Outcomes Adjudication Group. Cancer diagnosis data were obtained through the NHS number linkage for participants in England and Wales, and through the NHS Central Register for participants in Scotland. Participants were followed from the date of their baseline assessment until the earliest of the following: a diagnosis of breast cancer, diagnosis of a non-breast malignancy (excluding non-melanoma skin cancer), loss to follow-up, or the end of the study period. Participants who were diagnosed with any cancer other than breast cancer and non-melanoma skin cancer during follow-up were censored from the study on the date of their diagnosis of those other malignancies. Specifically, the ICD-10 codes for breast cancer patients encompass the following detailed classifications: C50.1, C50.2, C50.3, C50.4, C50.5, C50.6, C50.8, C50.9, D05.7, D05.9, and D48.9. These codes encompass both invasive and in situ subtypes, ensuring comprehensive case ascertainment. Covariates Based on our clinical knowledge and prior literature, we conducted a comprehensive analysis of prognostic factors for breast cancer and selected features for evaluation. Covariates include sociodemographic, lifestyle, and genetic factors previously associated with breast cancer, which are considered here as potential confounders in determining the association between exposure and outcome. It included patient information (age, sex, and year of diagnosis) and survival information (survival period and survival status at the cut-off date of follow-up), including 19 characteristics such as personal age, smoking, alcohol consumption, BMI, frequency of insomnia, maternal genetic disease, sibling genetic disease, stroke, race, fruit intake, diastolic blood pressure, age at menopause, and age at menarche. Baseline age (in years) was calculated from the date of birth and the date the cancer was diagnosed. For categorical variables, smoking status ("never", "before", "now"), alcohol intake ( "low frequency", "median frequency", "high frequency"), sleeplessness ("never/rarely", "sometimes", "usually"), nervous ("No", "Yes"), siblings with or without breast cancer ( "no breast cancer", "breast cancer"), mothers with or without breast cancer ( "no breast cancer", "breast cancers"), Caucasian by race ("No", "Yes"), fruit intake ( "not eating", "eating 1 to 2", "≥3"), irritability ( "No", "Yes"), chest pain ( "No", "Yes"), exercise time ( "<30 min", "≥30 min") were obtained by touch screen questionnaire. The proportion of participants with missing data on these covariates was low (< 3% of the sample), and confounding factors with missing data were uniformly attributed to the category "No" of categorical variables. For menopausal status at the time of recruitment, women were defined as premenopausal based on whether they had had periods in the previous year, while those who reported that their periods had stopped for at least one year were classified as postmenopausal. To minimize confounding due to treatment effects, 756 breast cancer patients who reported taking breast cancer-specific medications (e.g., capecitabine, tamoxifen, goserelin, anastrozole, fluorouracil, letrozole, cyclophosphamide, toremifene and exemestane) were excluded. Survival time was calculated from the data of breast cancer diagnosis to either the date of death or the end of follow-up (August 1, 2022). For participants who died, survival time was defined as the interval between diagnosis and death; for those who remained alive, it was defined as the time between diagnosis and August 1, 2022. Survival status was coded as "0" for alive and "1" for deceased. Ascertainment of cancer and death The death registration records were obtained from the National Health Service (NHS) system in England and Wales, as well as from the information and statistics departments in Scotland. These records provide the exact dates of death for the women who died during the study period. The follow-up continued until each woman passed away, was lost to follow-up (i.e., lost contact or withdrew from the study), or until the end of the cancer incidence follow-up period on August 1, 2022, whichever came first. The accuracy of the death registry has been verified through contact with the death registration offices. According to reports, the completeness of case identification by the UK Cancer Registration Service is approximately 98–99%, based on a study that linked routine cancer registrations with information in hospital episode statistics databases [ 17 ] . In this study, all death registration records pertain to deaths caused by breast cancer. Feature Selection To identify the most informative prognostic variables, we applied three feature selection methods: Stepwise Cox regression (Stepcox) [ 18 ] , Least absolute shrinkage and selection operator (Lasso) [ 19 ] and eXtreme Gradient Boosting (XGBoost) [ 20 ] . The intersection of features selected by these three methods was used to determine the final set of key variables. The rationale and implementation of each method are briefly summarized below. XGBoost employs a greedy algorithm combined with regularization techniques to iteratively select the optimal feature and split point during each tree node split, aiming to simplify the model structure and enhance generalization performance. Lasso, on the other hand, applies pressure to the regression coefficients through an L1 regularization term, forcing some coefficients to zero, thereby achieving effective feature selection and dimensionality reduction. Stepcox selects the optimal variable combination by gradually introducing or excluding variables based on statistical criteria such as Akaike Information Criterion (AIC) to identify the best-fitting model. Prediction Models In this study, we systematically evaluate the prediction performance based on five models: Cox Regression [ 21 ] , DeepSurv [ 22 ] , Enet [ 23 ] , GlmBoost [ 24 ] and XGBoost, aiming to identify the optimal prediction model. The specific mechanisms by which each model utilizes the characteristic variables of breast cancer patients to predict survival outcomes is introduced below. The Cox proportional hazards regression model assesses the impact of different factors on survival risk by analyzing the relationship between patients' characteristic variables and survival time. During the prediction process, the Cox model calculates a risk score for each patient based on their characteristic variables and uses this risk score to predict the patient's survival time or survival probability. DeepSurv is a survival analysis method that leverages feedforward neural networks to capture complex nonlinear relationships between input features and survival risk. During the prediction process, DeepSurv takes patients' characteristic variables as input, passes them through multiple layers of neural networks to perform nonlinear transformations, and ultimately outputs the predicted survival probability or risk score. XGBoost is a powerful gradient boosting algorithm that iteratively trains multiple weak learners (typically decision trees) and combines them to achieve robust predictive capability. During the prediction process, XGBoost makes progressive predictions through a sequence of trained decision trees based on the input characteristic variables. Each decision tree makes decisions based on specific feature split points, and the final prediction value is obtained by weighted summation of the predictions from all decision trees. Elastic Network (ENet) is a deep learning-based model that is particularly suitable for handling high-dimensional and sparse data. During the prediction process, ENet receives the input characteristic variables and processes them through multiple layers of neural networks for feature extraction and transformation. Gradient Boosting with Component-wise Linear Models (GLMBoost) is a predictive model that combines Generalized Linear Models (GLMs) with gradient boosting techniques. It iteratively trains multiple GLMs as weak learners using a gradient boosting algorithm and combines them to achieve robust predictive capability. During the prediction process, GLMBoost makes progressive predictions through a sequence of trained GLMs based on the input characteristic variables. Each GLM calculates a prediction value based on a specific linear combination and link function, and the final prediction value is obtained by weighted summation of the predictions from all GLMs. Statistical Analysis Participants were followed from the date of the baseline assessment center visit to the date of breast cancer registration (ICD-10 C50), date of death, date lost to follow-up, or end of cancer incidence follow-up (August 1, 2022). Women diagnosed with any cancer other than breast cancer (except nonmelanoma skin cancer) during follow-up were censored at the date of diagnosis [ 25 ] . In this study, we first selected eligible female participants with complete information from the UK Biobank database, resulting in a cohort characterized by 20 comprehensive demographic, clinical, and lifestyle-related variables. Feature selection was then conducted using Stepcox, Lasso, and XGBoost, resulting in the identification of intersection features common to all three methods. By taking the intersection of three feature selection methods, we can reduce the biases that may arise from using a single method, allowing for a more comprehensive evaluation of the importance of features and enabling the accurate identification of key factors. Model training and evaluation are carried out on the features obtained by intersection. We performed a survival analysis using patient and tumor information to identify characteristics that significantly affect patient survival. These analyses were performed after excluding patients with incomplete information. Subsequently, we systematically evaluate the prediction performance based on the five models (Cox regression, DeepSurv, ENet, GLMBoost, and XGBoost). By using the training set to train the model and the test set to verify the accuracy of the model, we ensure the rigor of the evaluation process. In order to comprehensively evaluate the prediction effect of each model, we construct a Time-dependent receiver operating characteristic curve (TimeROC) [ 26 ] curve, and use area under the curve (AUC) as the core performance evaluation index, which has a range from 0.5 to 1.0, and is the key statistic to measure the performance of a binary classifier. AUC values, and 95% confidence intervals (CIs) for each model were calculated for accurate model performance comparisons across years of data. A risk score was calculated for each patient after incorporating expression values for each feature, where weights reflected the relative contribution of each variable to overall survival risk.. Patients were then divided into low risk group and high risk group based on the median risk score. Kaplan–Meier survival analysis was used to compare survival differences between these groups in both the training cohort and four external validation cohorts. Finally, model interpretability was assessed using SHAP value [ 27 ] , which quantify the contribution of each feature to the predicted risk score. This analysis enabled the identification of key predictors that influence breast cancer prognosis. The overall analysis workflow are presented in Fig. 1 . Results Population Characteristics Among 502,164 participants in the UK biobank, a total of 16,540 cases (3.29%) of breast cancer were diagnosed among 273,175 female participants. To ensure data integrity and maintain methodological rigor, we rigorously excluded cases with incomplete information (see Fig. 2 for the exclusion criteria and screening process). After this meticulous screening process, a final total of 14,901 patients were included in this study. To validate the effectiveness and generalization ability of the model, the study population was randomly divided into two subsets: a training set ( n = 10431; 70%) and a validation set ( n = 4470; 30%). Subsequent data analysis revealed that participants with higher risk scores were more likely to exhibit a distinct set of risk factors. In particular, individuals in the high-risk group were more likely to have the following characteristics: older age, higher BMI, current smoking habits, frequent alcohol intake, usually sleeplessness problems, non-white ethnic background, little or no consumption of fresh fruit, history of irritability, experience chest tightness or chest pain, late age at menopause or late age at menarche, high age at first birth, relatively low Polygenic risk score (PRS). Baseline characteristics of study participants are presented in Table 1 . Furthermore, we present the annual number of death among breast cancer patients in Fig. 3 providing the survival status and mortality trends across the follow-up period, These statistics were subsequently used as input for downstream survival modeling. Table 1 Baseline characteristics of the study population Characteristic , High risk (N = 7205) Low risk (N = 7696) Overall (N = 14901) Age, mean(SD) 62.89(8.80) 53.93(7.96) 58.26(9.50) BMI, mean(SD) 27.41(5.07) 27.17(4.83) 27.28(4.95) Smoking status, No. (%) Never 4240(58.8%) 4413(57.3%) 8653(58.0%) Previous 2317(32.2%) 2759(35.8%) 5076(34.6%) Current 648(9.0%) 524(6.9%) 1172(7.4%) (Continued) Table.1 (Continued) Characteristic a,b High risk (N = 8217) Low risk (N = 8315) Overall (N = 16532) Alcohol intake, No. (%) Low frequency 1726(23.9%) 1829(23.8%) 3555(23.9%) Median frequency 2661(36.9%) 2961(38.5%) 5622(37.7%) High frequency 2818(39.2%) 2906(37.7%) 5724(38.4%) Sleeplessness, No. (%) Never/rarely 1324(18.4%) 1130(14.7%) 2454(16.5%) Sometimes 3525(48.9%) 3751(48.7%) 7276(48.8%) Usually 2356(32.7%) 2815(36.6%) 5171(34.7%) Nervous, No. (%) No 5365(74.5%) 5760(74.2%) 11075(74.3%) Yes 1840(25.5%) 1986(25.8%) 3826(25.7%) Siblings illnesses, No. (%) No Breast Cancer 6807(94.5%) 7133(92.7%) 13940(93.6%) Breast Cancer 398(5.5%) 563(7.3%) 961(6.4%) Mother illnesses, No. (%) No Breast Cancer 6383(88.6%) 6729(87.4%) 13112(88.0%) Breast Cancer 822(11.4%) 967(12.6%) 1789(12.0%) Ethnic White, No. (%) No 329(4.6%) 241(3.1%) 570(3.8%) Yes 6876(95.4%) 7455(96.9%) 14331(96.2%) Fruit intake, No. (%) Don’t eat 543(7.5%) 426(5.5%) 969(6.5%) Eat 1–2 3859(53.6%) 3866(50.2%) 7725(51.8%) Eat ≥ 3 2803(38.9%) 3404(44.3%) 6207(41.7%) Diastolic blood pressure, mean(SD) 81.20(10.22) 81.38(9.97) 81.28(10.09%) Irritability, No. (%) No 5461(75.8%) 5917(76.9%) 11378(76.4%) Yes 1744(24.2%) 1779(23.1%) 3523(23.6%) Chest pain, No. (%) No 6129(85.1%) 6606(85.8%) 12735(85.5%) Yes 1076(14.9%) 1090(14.2%) 2166(14.5%) Exercises duration, No. (%) < 30 min 4591(63.7%) 4980(64.7%) 9571(64.2%) ≥ 30 min 2614(36.3%) 2716(35.3%) 5330(35.8%) Menopause age, mean(SD) 50.09(3.82) 49.55(4.21) 49.81(4.04%) Periods started age, mean(SD) 12.90(1.58) 12.89(1.58) 12.90(1.58) First birth age, mean(SD) 25.38(3.83) 25.33(3.67) 25.35(3.75) Treatment number, mean(SD) 2.63(2.71) 3.06(2.69) 2.85(2.71) Standard PRS, mean(SD) 0.40(0.98) 0.47(0.97) 0.44(0.97) a Characteristics of eligible female participants. BMI ¼ body mass index; PRS ¼ polygenic risk score; SD ¼ standard deviation. b Mean (SD) for continuous variables and n (%) for categorical variables. To improve the prediction performance of the model, we adopt a comprehensive feature selection strategy, which combines the advantages of Lasso regression, Stepcox regression and XGBoost. Specifically, by computing the intersection of the feature sets provided by each of the three methods individually, we obtained a rigorous model evaluation process. In the Lasso regression analysis, we construct a penalized Cox proportional hazards model and determine the optimal penalty parameter (Lambda. Min) using cross-validation (CV) technique, which is finally determined to be 0.00087. To show the results of Lasso regression, we provide two charts: one is the Lasso regression coefficient path chart (Fig. 4 A), which intuitively shows which features are selected by the model and the size of their coefficients by revealing the sparsity of the contribution of each feature to the model; and the cross-validation error plot (Fig. 4 B), which shows how the model’s performance varies across different penalty values and supports the selection of the optimal lambda. For the XGBoost method, we first calculate the importance score of each variable and select the top 10 most influential features based on their contribution to predictive accuracy. To visually demonstrate the importance of these features, we create a feature importance ranking diagram (Fig. 4 C), which intuitively displays the relative importance of each feature in the model through a clear distribution of weights, providing a solid foundation for subsequent model interpretation and optimization. Using the Stepcox method, we have identified features with significant impacts ( P < 0.05), enabling a more targeted understanding of prognostic variables. These features were further refined and incorporated into downstream modeling processes (Fig. 4 D). By integrating the results from three distinct feature selection methods (Lasso regression, stepwise Cox regression, and XGBoost) and applying advanced technical means such as cross-validation, stepwise regression and feature importance evaluation, we identified six overlapping features consistently selected across all methods. These key features include age, BMI, smoking status, number of medications, alcohol intake and diastolic blood pressure (Fig. 5 ). Through the visual display of a Venn diagram, we can observe that although three different feature screening strategies are used, these six features are consistently identified in all methods and constitute the core part of feature intersection. Based on this finding, we have reason to speculate that these six characteristics may play a particularly significant role in predicting or influencing the survival of breast cancer patients. This inference not only emphasizes the importance of these features in the prognosis of breast cancer, but also provides valuable clues and directions for further study of the pathogenesis of breast cancer and personalized treatment strategies. Model Evaluation To thorouphly evaluate the time-varying characteristics of the prediction accuracy of the model, we adopted the TimeROC method to ensure the reliability, stability and time sensitivity of the evaluation results. Specifically, we randomly divided the dataset into a training set and a test set, and constructed a TimeROC curve for both subsets across different time intervals. This curve can intuitively show the trade-off relationship between sensitivity (true positive rate) and specificity (false positive rate) of the model at different time points and different thresholds, thereby enabling a more nuanced understanding of the model’s predictive performance over time. Furthermore, we calculated the area AUC values for three key time points of 3 years, 6 years and 9 years respectively (Fig. 6 ). By comparing the AUC values across different time intervals, we aimed to better understand the performance differences of the model for different prediction periods, thus providing a scientific basis for the subsequent model application. In the comparative analysis of survival predictions for breast cancer patients at 3-year, 6-year, and 9-year marks, the XGBoost model consistently demonstrated superior and stable predictive performance. Figure 6 presents the specific AUC values for five models, with the length of the error values represented by the length of the lines. Specifically, both in the training and validation sets, the AUC values achieved by the XGBoost model were significantly higher than those of the Cox model, DeepSurv model, GlmBoost model, and Enet model. This result demonstrates the accuracy and reliability of the XGBoost model in breast cancer survival prediction, with its advantages being evident in predicting 3-year, 6-year, and even 9-year survival rates. As the prediction time span extended, the XGBoost model not only maintained its efficient and accurate predictive ability but also exhibited stable performance in cross-validation, with AUC values consistently remaining at a high level. Therefore, within the scope of this study, the XGBoost is recognized as the most efficient, accurate, and suitable tool for long-term survival prediction in breast cancer patients.It is expected to offer sustained and reliable support for clinical decision-making and individualized patient management. We further use the XGBoost model to predict the risk score for each sample. Based on the median risk score derived from the training set, patients were divided into high risk and low risk groups based on the median risk score. The Kaplan-Meier survival analysis showed that, in the test cohort, patients in the high risk group had significantly poorer survival outcomes compared to those in the low risk group ( P < 0.0001, Fig. 7 A). To better understand the decision-making process of the model, we ranked the importance of the features by calculating and analyzing the SHAP value. This analysis enabled us to quantify the contribution of each feature to the predicted risk score, thereby identifying the variables with the greatest influence on survival prediction. Since age and menopausal age show a large number of positive and negative impact points (Fig. 7 B), it shows that age and BMI have a significant and variable impact on the results predicted by the model under different observation values. For diastolic blood pressure and alcohol intake, although two color dots were also shown, the distribution of dots was more concentrated, indicating that the effect was relatively small. Then, we give the importance ranking of the 6 selected features (Fig. 7 C). It indicates that, in this study, age and BMI may be the most pivotal factors influencing the survival duration of breast cancer patients. Discussion In this study, we conducted survival analysis on data from breast cancer patients in the UK Biobank, which also provides extensive information on personal health factors such as lifestyle habits. We employed three feature selection methods to identify key factors and utilized five survival analysis algorithms to establish predictive models for 3-year, 6-year, and 9-year survival outcomes. Ultimately, the XGBoost model demonstrated the best performance. Furthermore, we adopted the Kaplan-Meier method to explore and visually present the impact of various factors on the survival duration of breast cancer patients. Additionally, we utilized SHAP analysis to ascertain which features had the most significant influence on the model's predictions. In addition, we conducted a comprehensive evaluation of five different machine learning models in terms of their performance in predicting survival outcomes using complex medical data. Among them, DeepSurv, as a deep learning algorithm, demonstrated its unique advantages, particularly in modeling survival data that exhibit both linear and nonlinear risk characteristics, showing superiority compared to the traditional Cox Proportional Hazards algorithm. This finding is consistent with the viewpoint expressed in literature [ 23 ] , which states that the DeepSurv model can be seen as an extension of the Faraggi-Simon network and experimentally verifies that DeepSurv outperforms the Cox algorithm when handling survival data containing both linear and nonlinear risk functions. However, it is worth noting that despite DeepSurv's breakthrough in performance, it did not surpass the XGBoost model in this study. This result aligns with the research conducted by Léo Grinsztajn et al. [ 28 ] , which reported that tree-based methods are easier to achieve good predictions on tabular data than deep learning. Therefore, deep learning algorithms have an advantage in capturing complex nonlinear relationships, while machine learning algorithms, especially ensemble learning algorithms, excel in processing large-scale data and improving prediction accuracy and stability. In practical applications, the choice of algorithm depends on the nature of the specific problem, the characteristics of the data, and the available computational resources. We noted that the Kaplan-Meier curve results from both the training and validation cohorts showed that the survival probability of the high-risk group was significantly lower than that of the low-risk group, with P -value < 0.0001. This finding not only verifies the validity of the features screened in our study, but also further emphasizes the importance of these features in breast cancer survival analysis. In this study, we utilized three different feature selection methods to consistently identify 6 key variablesthat potentially play critical roles in the prognosis of breast cancer patients. SHAP interpretability analysis further confirmed the prognostic importance of these features, as each was shown to have a substantial positive or negative impact on the model’s predictions. Although features such as fruit intake and psychological stress (e.g., nervousness) were not selected as key features, they demonstrated moderate importance during the initial screening process, indicating that these factors also exert a certain influence on the prognosis of breast cancer patients. Interestingly, our research findings indicate that the genetic situations of siblings and mothers do not have a significant impact on the survival prognosis of breast cancer patients. This may be attributed to the fact that the development of breast cancer involves multigene inheritance and complex interactions with environmental factors, making the contribution of single genetic factors relatively limited in overall prognosis assessment and thereby affecting their significance in prognostic analysis. Beyond traditional prognostic factors, our analysis revealed a noteworthy association between the use of non–breast cancer medications and increased mortality risk. Specifically, our research data clearly shows that patients who take more non breast cancer drugs face a higher risk of mortality than those who take less of these drugs. Prior evidence suggests that certain non–breast cancer drugs may contribute to tumor progression [ 29 , 30 ] , or may interact adversely with oncologic treatments [ 31 – 34 ] . These interactions can reduce the therapeutic effect, aggravate side effects, or even trigger new health problems, thus directly or indirectly affecting the survival of patients. In our in-depth analysis of how age influences survival predictions in breast cancer patients, SHAP charts presented a key finding: elderly individuals exhibit a significant negative trend in the impact on the prediction model, while young individuals show a positive trend. This result clearly indicates that the prognosis of elderly breast cancer patients is significantly worse than that of young patients in our data set. This finding is consistent with previous studies, which have reported substantial age-related differences in survival, often attributed to factors such as reduced physiological reserve, comorbidities, and increased susceptibility to treatment-related complications in older adults [ 35 ] . Further supporting this observation, additional research has identified age ≥ 60 as an independent predictor of poor prognosis in female breast cancer patients [ 36 ] . Together, these findings reinforce the notion that increasing age is a critical factor associated with elevated mortality risk following breast cancer diagnosis. Moreover, the adverse impact of high BMI on breast cancer prognosis was also demonstrated. This conclusion is consistent with previous studies, which also found a strong association between high levels of obesity and overall mortality, breast cancer-specific mortality, recurrence, and distant metastasis in breast cancer patients, which further supports the association between obesity and poor prognosis [ 37 ] . These findings further emphasize the crucial role of BMI in the assessment of breast cancer prognosis and provide new insights for future exploration of breast cancer prevention and treatment strategies. Our study also explored the impact of other characteristics on individual breast cancer survival risk. Among these, high smoking frequency and early menopausal age have been confirmed to decrease survival for breast cancer patients. Frequent smoking behavior not only elevates the risk of developing cancer but also adversely affects the prognosis of cancer patients [ 5 ] . Although the effect of alcohol intake on survival was relatively modest in our analysis, we observed that lower levels of alcohol consumption were unexpectedly associated with worse survival outcomes. This may be due to studies suggesting that moderate alcohol consumption may have a certain protective effect on cardiovascular health, thereby indirectly influencing the survival of breast cancer patients [ 38 ] , this conclusion still requires further verification. From a clinical perspective, our model provides an e accessible and cost-effective approach to predict survival outcomes in breast cancer patients. By utilizing only comprehensive personal health factors as features, our prediction model achieved accuracy rates of 72%, 73%, and 75% respectively in the 3-year, 6-year, and 9-year assessments. This result fully demonstrates that, even under resource-limited conditions, by reasonably selecting and utilizing key health indicators, we can still provide extremely valuable references for the prognosis assessment of breast cancer patients. However, this study also has certain limitations. Firstly, the relatively small sample size may limit the generalizability and reliability of the research results. Secondly, we did not consider all factors that may affect the survival of breast cancer patients, such as genetics, lifestyle, and socioeconomic status, which are important prognostic factors for breast cancer. Finally, in terms of data processing, although combining "No" and "Unknown" into a single category can enhance data consistency and analytical robustness, it also has obvious drawbacks, including information loss, potential bias introduction, and reduced data transparency. Future research will focus on validating the results of this study across multiple cohorts, while expanding the sample size and incorporating more factors, such as socioeconomic status and specific medical test data, into the analysis, as well as conducting prospective randomized clinical trials to verify our findings. Conclusion This study conducted an in-depth analysis of the synergistic effects among individual health-related factors and systematically demonstrated how lifestyle habits collectively influence the survival outcomes of breast cancer patients. The survival analysis model developed herein demonstrates promising performance in predicting the survival duration of breast cancer patients. It not only provides robust scientific evidence for the prognosis assessment of breast cancer patients but also provides solid guidance for personalized habit and lifestyle adjustment, possessing significant clinical guidance value. Abbreviations AIC (Akaike Information Criterion) AUC (area under the curve) BMI (Body Mass Index) CIs (confidence intervals) CV (cross-validation) ENet (Elastic Network) GlmBoost (Gradient Boosting with Component-wise Linear Models) GLMs (Generalized Linear Models) Lasso (Least absolute shrinkage and selection operator) NHS (National Health Service) PRS (Polygenic risk score) SHAP (SHapley Additive exPlanations) Stepcox (Stepwise Cox regression) TimeROC (Time-dependent receiver operating curve) XGBoost (eXtreme Gradient Boosting) Declarations Author Contributions Yishan Yao (Conceptualization; Formal analysis; Investigation; Methodology; Visualization; Writing-original draft; Writing-review & editing), Xiaobing Zhai (Formal analysis;Writing-review & editing), Zhichao Liang (Formal analysis; Writing-review & editing), Chi Kin Lam (Supervision; Writing-review & editing), Hui Xie (Supervision), Tong, H.H.Y. (Data Curation), Ritse Mann (Supervision), Tong Tong (Writing-review & editing), Yan Chen (Supervision), Muzhen He (Supervision), Kefeng Li (Data Curation), Tao Tan (Writing-review & editing; Supervision; Conceptualization) Funding The funding sources played no role in the study design, data collection, data analysis, and interpretation of results or the decisions made in preparation and submission of the article. UK Biobank has received core funding from the Wellcome Trust medical charity, Medical Research Council, Department of Health, Scottish Government and the Northwest Regional Development Agency, the Welsh Government, British Heart Foundation, Cancer Research UK and Diabetes UK, and National Institute for Health and Care Research. The details of UK Biobank core funding and additional funding are reported at https://www.ukbiobank.ac.uk/learn-more-about-uk-biobank/about-us/our-funding. Conflicts of Interest The authors declare no potential conflicts of interest. Availability of Data and Materials Statements This research has been conducted using data from the UK Biobank under Application Number 99946. The UK Biobank data are not publicly available due to participant privacy and data sharing restrictions. Researchers can apply for access to the data through the UK Biobank resource at https://www.ukbiobank.ac.uk. Ethical approval: The details of UK Biobank research ethics approval are elaborated at https://www.ukbiobank.ac.uk/learn-more-about-uk-biobank/about-us/ethics. References Giaquinto AN, Sung H, Miller KD, Kramer JL, Newman LA, Minihan A, et al. Breast Cancer Stat 2022 CA: cancer J Clin. 2022;72(6):524–41. Loibl S, Poortmans P, Morrow M, Denkert C, Curigliano G. Breast cancer. Lancet (London England). 2021;397(10286):1750–69. Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I et al. Global cancer statistics 2022:GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. Cancer J Clin. 2024. Ellsworth RE, et al. Impact of lifestyle factors on prognosis among breast cancer survivors in the USA. Expert Rev pharmacoeconomics outcomes Res vol. 2012;12(4):451–64. 10.1586/erp.12.37 . GBD 2016 Alcohol Collaborators. Lancet (London England). 2018;392(10152):1015–35. https://doi.org/10.1016/S0140-6736(18)31310-2 . Alcohol use and burden for 195 countries and territories, 1990–2016: a systematic analysis for the Global Burden of Disease Study 2016. Passarelli MN, Newcomb PA, Hampton JM, et al. Cigarette Smoking Before and After Breast Cancer Diagnosis: Mortality From Breast Cancer and Smoking-Related Diseases. J Clin Oncol. 2016;34(12):1315–22. 10.1200/JCO.2015.63.9328 . Printz C. Anxiety and depression may contribute to an increased risk of death in some cancers. Cancer. 2017;123(13):2389–2389. Dehesh T, Fadaghi S, Seyedi M, et al. The relation between obesity and breast cancer risk in women by considering menstruation status and geographical variations: a systematic review and meta-analysis. BMC Womens Health. 2023;23:392. https://doi.org/10.1186/s12905-023-02543-5 . Soldato D, Michiels S, Havas J, Di Meglio A, Pagliuca M, Franzoi MA, Pistilli B, Iyengar NM, Cottu P, Lerebours F, Coutant C, Bertaut A, Tredan O, Vanlemmens L, Jouannaud C, Hrab I, Everhard S, Martin AL, André F, Vaz-Luis I, Jones LW. Dose/Exposure Relationship of Exercise and Distant Recurrence in Primary Breast Cancer. J Clin Oncol. 2024;42(25):3022–32. Epub 2024 Jun 5. PMID: 38838281; PMCID: PMC11361355. Maryam S, Farvid RM, Tamimi EM, Poole, Wendy Y, Chen BA, Rosner WC, Willett MD, Holmes A. Heather Eliassen; Postdiagnostic Dietary Glycemic Index, Glycemic Load, Dietary Insulin Index, and Insulin Load and Breast Cancer Survival. Cancer Epidemiol Biomarkers Prev 1 Febr. 2021;30(2):335–43. DeSantis CE, Ma J, Goding Sauer A, Newman LA, Jemal A. Breast cancer statistics, 2017, racial disparity in mortality by state. CA Cancer J Clin. 2017;67:439–48. Kreklau A et al. An Observational Study on Breast Cancer Survival and Lifestyle Related Risk Factors. In vivo (Athens, Greece) vol. 35,2 (2021): 1007–1015. 10.21873/invivo.12344 Parada H Jr, et al. Lifestyle Patterns and Survival Following Breast Cancer in the Carolina Breast Cancer Study. Epidemiol (Cambridge Mass) vol. 2019;30(1):83–92. 10.1097/EDE.0000000000000933 . Sudlow C, Gallacher J, Allen N, Beral V, Burton P, Danesh J, et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015;12(3):e1001779. https:// Foster HME, Celis-Morales CA, Nicholl BI, Petermann-Rocha F, Pell JP, Gill JMR, et al. The effect of socioeconomic deprivation on the association between an extended measurement of unhealthy lifestyle factors and health outcomes: a prospective analysis of the UK Biobank cohort. Lancet Public Health. 2018;3(12):e576–85. Moller H, Richards S, Hanchett N, Riaz SP, Luchtenborg M, Holmberg L, et al. Completeness ofcase ascertainment and survival time error in English cancer registries: impact on 1-year survival estimates. Br J Cancer. 2011;105:170–6. Hastie TJ, Pregibon D. In: Chambers S, Hastie TJ, editors. Generalized linear models. Chapter 6 of Statistical Models. Wadsworth & Brooks/Cole; 1992. Tibshirani R. Regression Shrinkage and Selection via the Lasso. J royal Stat Soc Ser b-methodological. 1996;58:267–88. Chen T, Guestrin C, Data Mining. XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and (2016): n. pag. Cox DR. Regression models and life-tables[J]. J Roy Stat Soc: Ser B (Methodol). 1972;34(2):187–202. Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol. 2018;18(1):24. 10.1186/s12874-018-0482-1 . PMID: 29482517; PMCID: PMC5828433. Paszke A, Chaurasia A, Kim S et al. ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation[J]. 2016. 10.48550/arXiv.1606.02147 Peter Buehlmann and Bin Yu. Boosting with the L2 loss: regression and classification. J Am Stat Assoc. 2003;98:324–39. Guo W, Key TJ, Reeves GK. Adiposity and breast cancer risk in postmenopausal women: results from the UK Biobank prospective cohort. Int J Cancer. 2018;143:1037–46. Kamarudin AN, Cox T, Kolamunnage-Dona R. Time-dependent ROC curve analysis in medical research: current methods and applications. BMC Med Res Methodol. 2017;17:53. https://doi.org/10.1186/s12874-017-0332-6 . Lundberg S. (2017). A unified approach to interpreting model predictions. arXiv preprint arXiv:170507874. Grinsztajn L, Oyallon E. Why do tree-based modelsstill outperform deep learning on tabular data? 2022, arXiv:2207.08815. Collins JA, Blake JM, Crosignani PG. Breast cancer risk with postmenopausal hormonal treatment. Hum Reprod Update. 2005;11:545–60. Støer NC, et al. Drug use and cancer risk: a drug-wide association study (DWAS) in Norway. CancerEpidemiol Biomark Prev. 2021;30:682–9. Baker AF, Dorr RT. Drug interactions with the taxanes: clinical implications. Cancer Treat Rev. 2001;27:221–33. Hamy A-S, et al. Celecoxib with neoadjuvant chemotherapy for breast cancer might worsen outcomes differentially by cox-2 expression and er status: exploratory analysis of the REMAGUS02 trial. JCO. 2019;37:624–35. Tyler T. Drug interactions in metastatic breast cancer. J Oncol Pharm Pr. 2011;17:236–45. Bibi R, et al. Prevalence of potential drug-drug interactions in breast cancer patients and determination of their risk factors. J Oncol Pharm Pr. 2021;27:1616–22. Zhou Y, Wen Y, Xiang Z, Ma J, Lin Y, Huang Y, Chen C. Cancer Survival Trends in Southeastern China, 2011–2021: A Population-Based Study. Clin Epidemiol. 2024;16:45–56. PMID: 38318284; PMCID: PMC10840559. Chen HL, Zhou MQ, Tian W, Meng KX, He HF. Effect of Age on Breast Cancer Patient Prognoses: A Population-Based Study Using the SEER 18 Database. PLoS ONE. 2016;11(10):e0165409. 10.1371/journal.pone.0165409 . PMID: 27798652; PMCID: PMC5087840. Pang Y, Wei Y, Kartsonaki C. Associations ofadiposity and weight change with recurrence and survival in breast cancer patients: a systematic review and meta-analysis. Breast Cancer. 2022;29:575–88. https://doi.org/10.1007/s12282-022-01355-z . Newcomb PA, Kampman E, Trentham-Dietz A, et al. Alcohol consumption before and after breast cancer diagnosis: associations with survival from breast cancer, cardiovascular disease, and other causes. J Clin Oncol. 2013;31(16):1939–46. 10.1200/JCO.2012.46.5765 . Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major Revision 17 Feb, 2026 Reviewers agreed at journal 28 Jan, 2026 Reviewers invited by journal 15 Jan, 2026 Editor assigned by journal 13 Jan, 2026 First submitted to journal 12 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8588282","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":575287848,"identity":"4f602ccc-73d3-4e83-8475-f896a5650d54","order_by":0,"name":"Yishan Yao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA40lEQVRIie3PMYvCMBTA8UjguTzNmuJRv0Kk0M+TUHBSObilg6CgtIOIq4IfwtGxEsgUd8d2d7DbTXeCq3LpbQ75ze/Pe48Qz3tD0F6cylv6OwPKqlKmU3fSRZMMtpZK1kYqSmvcSchHca+TURlsEIJqSRschkUUIcBE6I5J1RwIy1fS8ctcVXvEL6G7w4s6fhBuzwfXFi2unLcOmsQXZYEIPnYkXGUchXgknyqjTZIEApRS7RYYk2YJGjrYFkXEKCRcWoPOX/r5pi7rnyIEpk/1dzoNWb7+O3mC/xv3PM/zXroDJUVJkHB4R5EAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-5403-0887","institution":"Macao Polytechnic University","correspondingAuthor":true,"prefix":"","firstName":"Yishan","middleName":"","lastName":"Yao","suffix":""},{"id":575287849,"identity":"8c96a934-78f3-4fed-93cf-af3b670ec9d7","order_by":1,"name":"Xiaobing Zhai","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Xiaobing","middleName":"","lastName":"Zhai","suffix":""},{"id":575287850,"identity":"65444c67-66cb-410b-b25f-add806ce7c82","order_by":2,"name":"Zhichao Liang","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Zhichao","middleName":"","lastName":"Liang","suffix":""},{"id":575287851,"identity":"e2f7295a-2777-4c8d-b630-333ad2801e24","order_by":3,"name":"Chi Kin Lam","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Chi","middleName":"Kin","lastName":"Lam","suffix":""},{"id":575287852,"identity":"6b90774c-eea0-4f4a-9ab7-b7214e748679","order_by":4,"name":"Hui Xie","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Hui","middleName":"","lastName":"Xie","suffix":""},{"id":575287853,"identity":"f3a107a7-f384-4387-ae80-60acc4133407","order_by":5,"name":"Henry H Y Tong","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Henry","middleName":"H Y","lastName":"Tong","suffix":""},{"id":575287854,"identity":"64a547cb-8889-4fce-bcf9-6e6c21d150ee","order_by":6,"name":"Tong Tong","email":"","orcid":"","institution":"Fuzhou University","correspondingAuthor":false,"prefix":"","firstName":"Tong","middleName":"","lastName":"Tong","suffix":""},{"id":575287855,"identity":"04537e23-6c54-4698-90a6-8bfb0a572a46","order_by":7,"name":"Yan Chen","email":"","orcid":"","institution":"University of Nottingham","correspondingAuthor":false,"prefix":"","firstName":"Yan","middleName":"","lastName":"Chen","suffix":""},{"id":575287856,"identity":"c417c742-6773-4149-b507-906c8b94b7f0","order_by":8,"name":"Ritse M. Mann","email":"","orcid":"","institution":"Radboud University: Radboud Universiteit","correspondingAuthor":false,"prefix":"","firstName":"Ritse","middleName":"M.","lastName":"Mann","suffix":""},{"id":575287857,"identity":"b501e94e-f6a7-4e66-9703-849e28e2f2f2","order_by":9,"name":"Muzhen He","email":"","orcid":"","institution":"Fujian Provincial Hospital","correspondingAuthor":false,"prefix":"","firstName":"Muzhen","middleName":"","lastName":"He","suffix":""},{"id":575287858,"identity":"c78ff2c6-03b5-4e03-9ee7-cb00f2831bb4","order_by":10,"name":"Kefeng Li","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Kefeng","middleName":"","lastName":"Li","suffix":""},{"id":575287859,"identity":"510b0efb-51af-49d7-9f85-9ab41f4c33b1","order_by":11,"name":"Tao Tan","email":"","orcid":"","institution":"Macao Polytechnic University","correspondingAuthor":false,"prefix":"","firstName":"Tao","middleName":"","lastName":"Tan","suffix":""}],"badges":[],"createdAt":"2026-01-13 06:57:49","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8588282/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8588282/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100689549,"identity":"0d64113f-84bc-4a45-b2f1-a28dcf46e0dd","added_by":"auto","created_at":"2026-01-20 13:42:52","extension":"xml","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":13754,"visible":true,"origin":"","legend":"","description":"","filename":"brcaBRCAD2600044.xml","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/82c51e94bc786406a774db4a.xml"},{"id":100689374,"identity":"0928bb0c-8292-47a5-bbcd-b4dd67222dc5","added_by":"auto","created_at":"2026-01-20 13:41:44","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":927,"visible":true,"origin":"","legend":"","description":"","filename":"BRCAD260004425300.go.xml","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/ae51da40a05f58f2ef356323.xml"},{"id":100689222,"identity":"eeee4f51-54b3-4509-90c7-c66df13236f1","added_by":"auto","created_at":"2026-01-20 13:40:07","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":785,"visible":true,"origin":"","legend":"","description":"","filename":"BRCAD2600044Import.xml","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/1c4006aff9e1fa3668be4d69.xml"},{"id":100689411,"identity":"56157574-8e62-460f-aa11-2e95c669860b","added_by":"auto","created_at":"2026-01-20 13:41:50","extension":"xml","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":128825,"visible":true,"origin":"","legend":"","description":"","filename":"BRCAD26000440enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/6aa6e6b494ca933bb73c2186.xml"},{"id":100689748,"identity":"d12a9365-32b9-4808-93b9-8b4e59a34af6","added_by":"auto","created_at":"2026-01-20 13:45:56","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":154786,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/1f08cab4b9258ac123030321.png"},{"id":100689557,"identity":"827e3572-da66-4f02-864d-f814db19062a","added_by":"auto","created_at":"2026-01-20 13:43:09","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":24009,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/724b8212585d080f4feb046d.png"},{"id":100689665,"identity":"3998ca11-93a2-480d-9ec5-117f7cf3d466","added_by":"auto","created_at":"2026-01-20 13:45:19","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45321,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/ed7c3ad4c2d1190508bf0d67.png"},{"id":100689554,"identity":"28b91b69-d1c3-4c28-b1ad-8e7cc4f2d7c1","added_by":"auto","created_at":"2026-01-20 13:43:06","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":124206,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/131a1e771643701cb68f9ace.png"},{"id":100689427,"identity":"f19c255f-446c-4897-8786-6122493d819e","added_by":"auto","created_at":"2026-01-20 13:42:12","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":52861,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/54c381c63ab7e1502ced58ea.png"},{"id":100689746,"identity":"8fddac03-0753-42a5-822f-95a809cca7e8","added_by":"auto","created_at":"2026-01-20 13:45:52","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36033,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/264743e3f01c0577dd3db10c.png"},{"id":100689448,"identity":"21de1155-0839-41b6-9018-7307620e60ae","added_by":"auto","created_at":"2026-01-20 13:42:18","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":111369,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/c862b73ed3cccb62b724a66b.png"},{"id":100689603,"identity":"3b3e3bb9-6afc-4578-940b-8a3e023bd5fb","added_by":"auto","created_at":"2026-01-20 13:43:38","extension":"xml","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":123746,"visible":true,"origin":"","legend":"","description":"","filename":"BRCAD26000440structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/5ca4bf5f78ecc8d5bf976895.xml"},{"id":100689987,"identity":"ba7dc366-4e98-4490-bab5-b6d6b462a0be","added_by":"auto","created_at":"2026-01-20 13:48:37","extension":"html","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":137310,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/686d31bd03793b6e484f5f27.html"},{"id":100689792,"identity":"4bbb11ce-1717-440e-bf17-3f810caaa916","added_by":"auto","created_at":"2026-01-20 13:47:08","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":635065,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow overview with data preprocessing, feature selection, model building, and model evaluation. \u003cstrong\u003eA\u003c/strong\u003e Data preprocessing. All patients are screened according to the inclusion criteria. \u003cstrong\u003eB\u003c/strong\u003e Conduct feature selection based on individual health factors. Three feature selection methods are applied to screen 19 features, and the features obtained from their intersection are taken as key features. \u003cstrong\u003eC\u003c/strong\u003e Model Building and Evaluation. Five survival models were constructed with key features, and the model evaluation was carried out by using TimeROC and KM curves, and the interpretability analysis was carried out by using SHAP diagrams.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/b1a6bc2ab052ab6c42699d2f.jpeg"},{"id":100689521,"identity":"7d3fcd62-c286-4ce3-8403-3098a90cdb86","added_by":"auto","created_at":"2026-01-20 13:42:40","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":81161,"visible":true,"origin":"","legend":"\u003cp\u003eRegarding the process of data preprocessing and dataset partitioning.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/b40f5da71c7f260dcf5c0a3f.png"},{"id":100796458,"identity":"a946c363-e38b-443d-9fbf-4282abc22cfe","added_by":"auto","created_at":"2026-01-21 13:43:18","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":168204,"visible":true,"origin":"","legend":"\u003cp\u003eThe aggregated statistical data on annual deaths.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/73e2a0a3725695769b0e4d85.jpeg"},{"id":100689711,"identity":"64087679-cb2e-4a06-ad9d-dc9f345bc92f","added_by":"auto","created_at":"2026-01-20 13:45:26","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":593360,"visible":true,"origin":"","legend":"\u003cp\u003eFeature selection. \u003cstrong\u003eA \u003c/strong\u003ePathway plot of the Lasso regression coefficients \u003cstrong\u003eB\u003c/strong\u003e Impact plot of penalty parameter of the Lasso regression \u003cstrong\u003eC\u003c/strong\u003e Characteristics importance ranking plot of XGBoost \u003cstrong\u003eD\u003c/strong\u003e Characteristics \u003cem\u003eP\u003c/em\u003e-value table of Stepcox\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/78f2350281d100ce0cd72b01.jpeg"},{"id":100689604,"identity":"dcb0070b-4a4a-498c-bf8c-956d9570a8d0","added_by":"auto","created_at":"2026-01-20 13:43:38","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":239636,"visible":true,"origin":"","legend":"\u003cp\u003eNetwork Venn diagram for three feature selection methods\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/cd2db4a04b241c55036ec470.png"},{"id":100689561,"identity":"15480051-877b-4e7a-9360-814d093fe529","added_by":"auto","created_at":"2026-01-20 13:43:11","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":132550,"visible":true,"origin":"","legend":"\u003cp\u003eThe ROC curves for the test cohort at three, six, and nine years.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/6f19117d9a8f97c5d433576a.png"},{"id":100689426,"identity":"11bdb0ea-64f5-434e-a4cd-91f4bf0d23ed","added_by":"auto","created_at":"2026-01-20 13:42:11","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":465967,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA\u003c/strong\u003e The Kaplan-Meier curves for the high risk and low risk groups of the test cohort\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB \u003c/strong\u003eFeature ranking-distinguishing feature value of SHAP interpretability analysis\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eC \u003c/strong\u003eFeature ranking-not distinguishing feature value of SHAP interpretability analysis\u003c/p\u003e","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/e5a169968e725a8eb922ab2b.jpeg"},{"id":100797968,"identity":"a3a5fb8b-ad77-4a01-b1b7-66b5eb9f3a35","added_by":"auto","created_at":"2026-01-21 13:51:58","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3213778,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8588282/v1/22c6991a-149e-42ef-a2bc-b4c746702e19.pdf"}],"financialInterests":"","formattedTitle":"Impact of Lifestyle Patterns on Breast Cancer Prognosis: Evidence from a UK Biobank Survival Study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eBreast cancer remains a significant global health concern, with its incidence continuing to increase among women worldwide. It is now the most commonly diagnosed malignant tumor in women, second only to skin cancer, accounting for about 30% of all female cancers\u003csup\u003e[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]\u003c/sup\u003e. According to the latest report released in February 2024 by the International Agency for Research on Cancer (IARC) of the World Health Organization, there were an estimated 19.97\u0026nbsp;million new cancer cases globally in 2022, including more than 2.3\u0026nbsp;million new cases of breast cancer. Moreover, approximately 670,000 women died from breast cancer, which puts this disease as the most frequent oncological cause of death in women worldwide\u003csup\u003e[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]\u003c/sup\u003e. While substantial efforts have been devoted to improving early detection and treatment, understanding the factors that influence long term survival remains critically important. With the growing number of breast cancer survivors and their increasing motivation to adopt healthier lifestyles post-diagnosis\u003csup\u003e[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]\u003c/sup\u003e, identifying modifiable behavioral factors that may improve survival outcomes is essential. Such insights could contribute to the development of evidence-based, personalized survivorship strategies aimed at enhancing both longevity and quality of life.\u003c/p\u003e \u003cp\u003eBreast cancer survival risk is influenced by a broad spectrum of modifiable lifestyle habits, many of which are increasingly prevalent in modern societies. Adverse habits such as smoking, excessive alcohol consumption, poor stress management, altered reproductive patterns, obesity, and physical inactivity have all significant impacts on the survival risk of breast cancer. Among these, alcohol consumption has been identified as a major contributor to both mortality and disability\u003csup\u003e[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]\u003c/sup\u003e, while smoking\u003csup\u003e[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]\u003c/sup\u003e facilitates cancer progression by promoting metastasis and impairing overall health. Chronic psychological stress\u003csup\u003e[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/sup\u003e may indirectly increase the mortality risk by disrupting the endocrine and immune system functions. Additionally, the widespread issue of overweight and obesity\u003csup\u003e[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/sup\u003e can lead to hormonal imbalances and chronic inflammation, both of which may serve as potential triggers for breast cancer, and a sedentary lifestyle\u003csup\u003e[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]\u003c/sup\u003e further compromises the body's immune defense. In addition to these factors, dietary patterns have received increasing attention. Previous studies have demonstrated that high glycemic index or insulin load diets after diagnosis are associated with increased risks of breast cancer (specific mortality and all-cause mortality), emphasizing the long-term impact of nutritional choices on survivorship\u003csup\u003e[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/sup\u003e. Furthermore, disparities in breast cancer incidence, screening access, and survival outcomes across racial and ethnic groups highlight the importance of developing individualized and culturally sensitive interventions\u003csup\u003e[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]\u003c/sup\u003e. Although these modifiable lifestyle factors have been studied individually, they frequently co-occur and may interact in complex ways, potentially influencing outcomes in a synergistic manner. Therefore, an integrated assessment of multiple behaviors is essential for a more comprehensive understanding of their collective influence on breast cancer survival.\u003c/p\u003e \u003cp\u003eAlthough some studies have explored the relationship between a single lifestyle and the survival status of breast cancer patients, studies that systematically evaluate the combined effects of multiple behaviors remain limited\u003csup\u003e[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/sup\u003e. Research focusing solely on isolated factors may fail to capture the intricate interplay and cumulative impact of real-world lifestyle patterns. Moreover, prior studies are often constrained by limited sample sizes or population heterogeneity, which may restrict the reliability and generalizability of their findings. To fill this gap, this study utilizes data from the UK Biobank, a large and well-characterized prospective cohort, to investigate how multiple lifestyle habits jointly influence the survival outcomes of breast cancer patients. Specifically, we assess the synergistic effects of dietary habits, physical activity, sleep quality,tobacco and alcohol consumption using a comprehensive machine learning framework. By identifying key behavioral predictors and quantifying their relative contributions to patient prognosis, this study aims to provide actionable evidence to inform personalized rehabilitation strategies and promote healthier lifestyles among breast cancer survivors.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Population\u003c/h2\u003e \u003cp\u003eThe UK Biobank is a large-scale prospective, population-based cohort study that recruited approximately 500,000 participants (229,041 men and 273,293 women) aged 40\u0026ndash;69 years from 22 study evaluation centers across the UK between 2006 and 2010\u003csup\u003e[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]\u003c/sup\u003e. All participants were followed through a link in the NHS\u0026rsquo;s electronic health record. At baseline, participants completed a self-administered touch-screen questionnaire that collected information on sociodemographic characteristics, health and medical history, and lifestyle-related exposures\u003csup\u003e[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u003c/sup\u003e. A detailed description of the recruitment procedures and population characteristics has been described elsewhere\u003csup\u003e[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]\u003c/sup\u003e. In addition to the questionnaire, participants also received physical measurements and provided blood samples, including blood specimens. All participants provided written informed consent, and the study was approved by the Northwest Multicenter Research Ethics Committee. Individuals who had requested withdraw from the UK Biobank cohort study were excluded from our analysis. Data from biological samples from the United Kingdom (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.ukbiobank.ac.UK/\u003c/span\u003e\u003cspan address=\"http://www.ukbiobank.ac.UK/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) are available to all researchers upon request.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eAscertainment of analytic population\u003c/h3\u003e\n\u003cp\u003eWe restricted the study population to self-reported female participants with a history of breast cancer (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;14,901). In this study, we set the following strict inclusion criteria for these women with breast cancer to ensure the accuracy of the study subjects and the validity of the study: (a) Breast cancer diagnosed as the primary malignancy. (b) Death not on the same-day as diagnosis (i.e., survival time\u0026thinsp;\u0026gt;\u0026thinsp;0). (c) Eventual therapeutic drugs administered to the patient are not intended to treat breast cancer.\u003c/p\u003e\n\u003ch3\u003eExposure\u003c/h3\u003e\n\u003cp\u003eBreast cancer cases were identified using inpatient admission records sourced from the Hospital Episode Statistics (England), the Scottish Morbidity Record (Scotland), and the Patient Episode Database (Wales). These records were supplemented with death registration records provided by the National Health Service (NHS) system in England and Wales, as well as the information and statistics departments in Scotland. The diagnosis of breast cancer was based on the International Classification of Diseases, Tenth Revision (ICD-10) coding system, with the ICD codes used to identify breast cancer having been selected and validated by the UK Biobank Outcomes Adjudication Group. Cancer diagnosis data were obtained through the NHS number linkage for participants in England and Wales, and through the NHS Central Register for participants in Scotland. Participants were followed from the date of their baseline assessment until the earliest of the following: a diagnosis of breast cancer, diagnosis of a non-breast malignancy (excluding non-melanoma skin cancer), loss to follow-up, or the end of the study period. Participants who were diagnosed with any cancer other than breast cancer and non-melanoma skin cancer during follow-up were censored from the study on the date of their diagnosis of those other malignancies. Specifically, the ICD-10 codes for breast cancer patients encompass the following detailed classifications: C50.1, C50.2, C50.3, C50.4, C50.5, C50.6, C50.8, C50.9, D05.7, D05.9, and D48.9. These codes encompass both invasive and in situ subtypes, ensuring comprehensive case ascertainment.\u003c/p\u003e \n\u003ch3\u003eCovariates\u003c/h3\u003e\n\u003cp\u003eBased on our clinical knowledge and prior literature, we conducted a comprehensive analysis of prognostic factors for breast cancer and selected features for evaluation. Covariates include sociodemographic, lifestyle, and genetic factors previously associated with breast cancer, which are considered here as potential confounders in determining the association between exposure and outcome. It included patient information (age, sex, and year of diagnosis) and survival information (survival period and survival status at the cut-off date of follow-up), including 19 characteristics such as personal age, smoking, alcohol consumption, BMI, frequency of insomnia, maternal genetic disease, sibling genetic disease, stroke, race, fruit intake, diastolic blood pressure, age at menopause, and age at menarche. Baseline age (in years) was calculated from the date of birth and the date the cancer was diagnosed. For categorical variables, smoking status (\"never\", \"before\", \"now\"), alcohol intake ( \"low frequency\", \"median frequency\", \"high frequency\"), sleeplessness (\"never/rarely\", \"sometimes\", \"usually\"), nervous (\"No\", \"Yes\"), siblings with or without breast cancer ( \"no breast cancer\", \"breast cancer\"), mothers with or without breast cancer ( \"no breast cancer\", \"breast cancers\"), Caucasian by race (\"No\", \"Yes\"), fruit intake ( \"not eating\", \"eating 1 to 2\", \"\u0026ge;3\"), irritability ( \"No\", \"Yes\"), chest pain ( \"No\", \"Yes\"), exercise time ( \"\u0026lt;30 min\", \"\u0026ge;30 min\") were obtained by touch screen questionnaire. The proportion of participants with missing data on these covariates was low (\u0026lt;\u0026thinsp;3% of the sample), and confounding factors with missing data were uniformly attributed to the category \"No\" of categorical variables. For menopausal status at the time of recruitment, women were defined as premenopausal based on whether they had had periods in the previous year, while those who reported that their periods had stopped for at least one year were classified as postmenopausal. To minimize confounding due to treatment effects, 756 breast cancer patients who reported taking breast cancer-specific medications (e.g., capecitabine, tamoxifen, goserelin, anastrozole, fluorouracil, letrozole, cyclophosphamide, toremifene and exemestane) were excluded. Survival time was calculated from the data of breast cancer diagnosis to either the date of death or the end of follow-up (August 1, 2022). For participants who died, survival time was defined as the interval between diagnosis and death; for those who remained alive, it was defined as the time between diagnosis and August 1, 2022. Survival status was coded as \"0\" for alive and \"1\" for deceased.\u003c/p\u003e\n\u003ch3\u003eAscertainment of cancer and death\u003c/h3\u003e\n\u003cp\u003eThe death registration records were obtained from the National Health Service (NHS) system in England and Wales, as well as from the information and statistics departments in Scotland. These records provide the exact dates of death for the women who died during the study period. The follow-up continued until each woman passed away, was lost to follow-up (i.e., lost contact or withdrew from the study), or until the end of the cancer incidence follow-up period on August 1, 2022, whichever came first. The accuracy of the death registry has been verified through contact with the death registration offices. According to reports, the completeness of case identification by the UK Cancer Registration Service is approximately 98\u0026ndash;99%, based on a study that linked routine cancer registrations with information in hospital episode statistics databases\u003csup\u003e[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]\u003c/sup\u003e. In this study, all death registration records pertain to deaths caused by breast cancer.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eFeature Selection\u003c/h2\u003e \u003cp\u003eTo identify the most informative prognostic variables, we applied three feature selection methods: Stepwise Cox regression (Stepcox)\u003csup\u003e[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]\u003c/sup\u003e, Least absolute shrinkage and selection operator (Lasso)\u003csup\u003e[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]\u003c/sup\u003e and eXtreme Gradient Boosting (XGBoost)\u003csup\u003e[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]\u003c/sup\u003e. The intersection of features selected by these three methods was used to determine the final set of key variables. The rationale and implementation of each method are briefly summarized below.\u003c/p\u003e \u003cp\u003eXGBoost employs a greedy algorithm combined with regularization techniques to iteratively select the optimal feature and split point during each tree node split, aiming to simplify the model structure and enhance generalization performance. Lasso, on the other hand, applies pressure to the regression coefficients through an L1 regularization term, forcing some coefficients to zero, thereby achieving effective feature selection and dimensionality reduction. Stepcox selects the optimal variable combination by gradually introducing or excluding variables based on statistical criteria such as Akaike Information Criterion (AIC) to identify the best-fitting model.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePrediction Models\u003c/h3\u003e\n\u003cp\u003eIn this study, we systematically evaluate the prediction performance based on five models: Cox Regression\u003csup\u003e[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]\u003c/sup\u003e, DeepSurv\u003csup\u003e[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]\u003c/sup\u003e, Enet\u003csup\u003e[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/sup\u003e, GlmBoost\u003csup\u003e[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]\u003c/sup\u003e and XGBoost, aiming to identify the optimal prediction model. The specific mechanisms by which each model utilizes the characteristic variables of breast cancer patients to predict survival outcomes is introduced below.\u003c/p\u003e \u003cp\u003eThe Cox proportional hazards regression model assesses the impact of different factors on survival risk by analyzing the relationship between patients' characteristic variables and survival time. During the prediction process, the Cox model calculates a risk score for each patient based on their characteristic variables and uses this risk score to predict the patient's survival time or survival probability.\u003c/p\u003e \u003cp\u003eDeepSurv is a survival analysis method that leverages feedforward neural networks to capture complex nonlinear relationships between input features and survival risk. During the prediction process, DeepSurv takes patients' characteristic variables as input, passes them through multiple layers of neural networks to perform nonlinear transformations, and ultimately outputs the predicted survival probability or risk score.\u003c/p\u003e \u003cp\u003eXGBoost is a powerful gradient boosting algorithm that iteratively trains multiple weak learners (typically decision trees) and combines them to achieve robust predictive capability. During the prediction process, XGBoost makes progressive predictions through a sequence of trained decision trees based on the input characteristic variables. Each decision tree makes decisions based on specific feature split points, and the final prediction value is obtained by weighted summation of the predictions from all decision trees.\u003c/p\u003e \u003cp\u003eElastic Network (ENet) is a deep learning-based model that is particularly suitable for handling high-dimensional and sparse data. During the prediction process, ENet receives the input characteristic variables and processes them through multiple layers of neural networks for feature extraction and transformation.\u003c/p\u003e \u003cp\u003eGradient Boosting with Component-wise Linear Models (GLMBoost) is a predictive model that combines Generalized Linear Models (GLMs) with gradient boosting techniques. It iteratively trains multiple GLMs as weak learners using a gradient boosting algorithm and combines them to achieve robust predictive capability. During the prediction process, GLMBoost makes progressive predictions through a sequence of trained GLMs based on the input characteristic variables. Each GLM calculates a prediction value based on a specific linear combination and link function, and the final prediction value is obtained by weighted summation of the predictions from all GLMs.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eParticipants were followed from the date of the baseline assessment center visit to the date of breast cancer registration (ICD-10 C50), date of death, date lost to follow-up, or end of cancer incidence follow-up (August 1, 2022). Women diagnosed with any cancer other than breast cancer (except nonmelanoma skin cancer) during follow-up were censored at the date of diagnosis\u003csup\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn this study, we first selected eligible female participants with complete information from the UK Biobank database, resulting in a cohort characterized by 20 comprehensive demographic, clinical, and lifestyle-related variables. Feature selection was then conducted using Stepcox, Lasso, and XGBoost, resulting in the identification of intersection features common to all three methods. By taking the intersection of three feature selection methods, we can reduce the biases that may arise from using a single method, allowing for a more comprehensive evaluation of the importance of features and enabling the accurate identification of key factors. Model training and evaluation are carried out on the features obtained by intersection. We performed a survival analysis using patient and tumor information to identify characteristics that significantly affect patient survival. These analyses were performed after excluding patients with incomplete information.\u003c/p\u003e \u003cp\u003eSubsequently, we systematically evaluate the prediction performance based on the five models (Cox regression, DeepSurv, ENet, GLMBoost, and XGBoost). By using the training set to train the model and the test set to verify the accuracy of the model, we ensure the rigor of the evaluation process. In order to comprehensively evaluate the prediction effect of each model, we construct a Time-dependent receiver operating characteristic curve (TimeROC)\u003csup\u003e[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]\u003c/sup\u003e curve, and use area under the curve (AUC) as the core performance evaluation index, which has a range from 0.5 to 1.0, and is the key statistic to measure the performance of a binary classifier. AUC values, and 95% confidence intervals (CIs) for each model were calculated for accurate model performance comparisons across years of data.\u003c/p\u003e \u003cp\u003eA risk score was calculated for each patient after incorporating expression values for each feature, where weights reflected the relative contribution of each variable to overall survival risk.. Patients were then divided into low risk group and high risk group based on the median risk score. Kaplan\u0026ndash;Meier survival analysis was used to compare survival differences between these groups in both the training cohort and four external validation cohorts.\u003c/p\u003e \u003cp\u003eFinally, model interpretability was assessed using SHAP value\u003csup\u003e[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]\u003c/sup\u003e, which quantify the contribution of each feature to the predicted risk score. This analysis enabled the identification of key predictors that influence breast cancer prognosis. The overall analysis workflow are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003ePopulation Characteristics\u003c/h2\u003e \u003cp\u003eAmong 502,164 participants in the UK biobank, a total of 16,540 cases (3.29%) of breast cancer were diagnosed among 273,175 female participants. To ensure data integrity and maintain methodological rigor, we rigorously excluded cases with incomplete information (see Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e for the exclusion criteria and screening process). After this meticulous screening process, a final total of 14,901 patients were included in this study. To validate the effectiveness and generalization ability of the model, the study population was randomly divided into two subsets: a training set (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;10431; 70%) and a validation set (\u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4470; 30%).\u003c/p\u003e \u003cp\u003eSubsequent data analysis revealed that participants with higher risk scores were more likely to exhibit a distinct set of risk factors. In particular, individuals in the high-risk group were more likely to have the following characteristics: older age, higher BMI, current smoking habits, frequent alcohol intake, usually sleeplessness problems, non-white ethnic background, little or no consumption of fresh fruit, history of irritability, experience chest tightness or chest pain, late age at menopause or late age at menarche, high age at first birth, relatively low Polygenic risk score (PRS). Baseline characteristics of study participants are presented in Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Furthermore, we present the annual number of death among breast cancer patients in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e providing the survival status and mortality trends across the follow-up period, These statistics were subsequently used as input for downstream survival modeling.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBaseline characteristics of the study population\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCharacteristic\u003ca class=\"FNLink\" href=\"#Fn1\" id=\"#FNLinkFn1\"\u003e\u003c/a\u003e\u003csup\u003e,\u003c/sup\u003e\u003ca class=\"FNLink\" href=\"#Fn2\" id=\"#FNLinkFn2\"\u003e\u003c/a\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh risk\u003c/p\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;7205)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLow risk\u003c/p\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;7696)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOverall\u003c/p\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;14901)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAge, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e62.89(8.80)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e53.93(7.96)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e58.26(9.50)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eBMI, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e27.41(5.07)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e27.17(4.83)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e27.28(4.95)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSmoking status, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNever\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e4240(58.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4413(57.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e8653(58.0%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePrevious\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2317(32.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2759(35.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5076(34.6%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCurrent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e648(9.0%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e524(6.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1172(7.4%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e(Continued)\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eTable.1\u003c/b\u003e (Continued)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCharacteristic\u003csup\u003ea,b\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHigh risk\u003c/p\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;8217)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLow risk\u003c/p\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;8315)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOverall\u003c/p\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;16532)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlcohol intake, No. (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLow frequency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1726(23.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1829(23.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3555(23.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMedian frequency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2661(36.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2961(38.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5622(37.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHigh frequency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2818(39.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2906(37.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5724(38.4%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSleeplessness, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNever/rarely\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1324(18.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1130(14.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2454(16.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSometimes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3525(48.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3751(48.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7276(48.8%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUsually\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2356(32.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2815(36.6%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5171(34.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eNervous, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5365(74.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5760(74.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e11075(74.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1840(25.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1986(25.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3826(25.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSiblings illnesses, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo Breast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6807(94.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7133(92.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e13940(93.6%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBreast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e398(5.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e563(7.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e961(6.4%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMother illnesses, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo Breast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6383(88.6%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6729(87.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e13112(88.0%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBreast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e822(11.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e967(12.6%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1789(12.0%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eEthnic White, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e329(4.6%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e241(3.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e570(3.8%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6876(95.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7455(96.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e14331(96.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eFruit intake, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDon\u0026rsquo;t eat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e543(7.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e426(5.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e969(6.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEat 1\u0026ndash;2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3859(53.6%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3866(50.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7725(51.8%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEat\u0026thinsp;\u0026ge;\u0026thinsp;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2803(38.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3404(44.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6207(41.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eDiastolic blood pressure, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e81.20(10.22)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e81.38(9.97)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e81.28(10.09%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eIrritability, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5461(75.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5917(76.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e11378(76.4%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1744(24.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1779(23.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e3523(23.6%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eChest pain, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6129(85.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6606(85.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e12735(85.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1076(14.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1090(14.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2166(14.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eExercises duration, No. (%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;30 min\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e4591(63.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4980(64.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9571(64.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;30 min\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2614(36.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2716(35.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5330(35.8%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMenopause age, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e50.09(3.82)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e49.55(4.21)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e49.81(4.04%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePeriods started age, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12.90(1.58)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e12.89(1.58)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e12.90(1.58)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eFirst birth age, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e25.38(3.83)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e25.33(3.67)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e25.35(3.75)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTreatment number, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2.63(2.71)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.06(2.69)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2.85(2.71)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eStandard PRS, mean(SD)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.40(0.98)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.47(0.97)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.44(0.97)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e\u003cp\u003e\u003csup\u003ea\u003c/sup\u003e Characteristics of eligible female participants. BMI \u0026frac14; body mass index; PRS \u0026frac14; polygenic risk score; SD \u0026frac14; standard deviation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003csup\u003eb\u003c/sup\u003e Mean (SD) for continuous variables and n (%) for categorical variables.\u003c/p\u003e \u003cp\u003eTo improve the prediction performance of the model, we adopt a comprehensive feature selection strategy, which combines the advantages of Lasso regression, Stepcox regression and XGBoost. Specifically, by computing the intersection of the feature sets provided by each of the three methods individually, we obtained a rigorous model evaluation process.\u003c/p\u003e \u003cp\u003eIn the Lasso regression analysis, we construct a penalized Cox proportional hazards model and determine the optimal penalty parameter (Lambda. Min) using cross-validation (CV) technique, which is finally determined to be 0.00087. To show the results of Lasso regression, we provide two charts: one is the Lasso regression coefficient path chart (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA), which intuitively shows which features are selected by the model and the size of their coefficients by revealing the sparsity of the contribution of each feature to the model; and the cross-validation error plot (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB), which shows how the model\u0026rsquo;s performance varies across different penalty values and supports the selection of the optimal lambda. For the XGBoost method, we first calculate the importance score of each variable and select the top 10 most influential features based on their contribution to predictive accuracy. To visually demonstrate the importance of these features, we create a feature importance ranking diagram (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC), which intuitively displays the relative importance of each feature in the model through a clear distribution of weights, providing a solid foundation for subsequent model interpretation and optimization. Using the Stepcox method, we have identified features with significant impacts (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05), enabling a more targeted understanding of prognostic variables. These features were further refined and incorporated into downstream modeling processes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eBy integrating the results from three distinct feature selection methods (Lasso regression, stepwise Cox regression, and XGBoost) and applying advanced technical means such as cross-validation, stepwise regression and feature importance evaluation, we identified six overlapping features consistently selected across all methods. These key features include age, BMI, smoking status, number of medications, alcohol intake and diastolic blood pressure (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eThrough the visual display of a Venn diagram, we can observe that although three different feature screening strategies are used, these six features are consistently identified in all methods and constitute the core part of feature intersection. Based on this finding, we have reason to speculate that these six characteristics may play a particularly significant role in predicting or influencing the survival of breast cancer patients. This inference not only emphasizes the importance of these features in the prognosis of breast cancer, but also provides valuable clues and directions for further study of the pathogenesis of breast cancer and personalized treatment strategies.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eModel Evaluation\u003c/h2\u003e \u003cp\u003eTo thorouphly evaluate the time-varying characteristics of the prediction accuracy of the model, we adopted the TimeROC method to ensure the reliability, stability and time sensitivity of the evaluation results. Specifically, we randomly divided the dataset into a training set and a test set, and constructed a TimeROC curve for both subsets across different time intervals. This curve can intuitively show the trade-off relationship between sensitivity (true positive rate) and specificity (false positive rate) of the model at different time points and different thresholds, thereby enabling a more nuanced understanding of the model\u0026rsquo;s predictive performance over time. Furthermore, we calculated the area AUC values for three key time points of 3 years, 6 years and 9 years respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). By comparing the AUC values across different time intervals, we aimed to better understand the performance differences of the model for different prediction periods, thus providing a scientific basis for the subsequent model application.\u003c/p\u003e \u003cp\u003eIn the comparative analysis of survival predictions for breast cancer patients at 3-year, 6-year, and 9-year marks, the XGBoost model consistently demonstrated superior and stable predictive performance. Figure\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e presents the specific AUC values for five models, with the length of the error values represented by the length of the lines. Specifically, both in the training and validation sets, the AUC values achieved by the XGBoost model were significantly higher than those of the Cox model, DeepSurv model, GlmBoost model, and Enet model. This result demonstrates the accuracy and reliability of the XGBoost model in breast cancer survival prediction, with its advantages being evident in predicting 3-year, 6-year, and even 9-year survival rates. As the prediction time span extended, the XGBoost model not only maintained its efficient and accurate predictive ability but also exhibited stable performance in cross-validation, with AUC values consistently remaining at a high level. Therefore, within the scope of this study, the XGBoost is recognized as the most efficient, accurate, and suitable tool for long-term survival prediction in breast cancer patients.It is expected to offer sustained and reliable support for clinical decision-making and individualized patient management.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe further use the XGBoost model to predict the risk score for each sample. Based on the median risk score derived from the training set, patients were divided into high risk and low risk groups based on the median risk score. The Kaplan-Meier survival analysis showed that, in the test cohort, patients in the high risk group had significantly poorer survival outcomes compared to those in the low risk group (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001, Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003eTo better understand the decision-making process of the model, we ranked the importance of the features by calculating and analyzing the SHAP value. This analysis enabled us to quantify the contribution of each feature to the predicted risk score, thereby identifying the variables with the greatest influence on survival prediction. Since age and menopausal age show a large number of positive and negative impact points (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB), it shows that age and BMI have a significant and variable impact on the results predicted by the model under different observation values. For diastolic blood pressure and alcohol intake, although two color dots were also shown, the distribution of dots was more concentrated, indicating that the effect was relatively small. Then, we give the importance ranking of the 6 selected features (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC). It indicates that, in this study, age and BMI may be the most pivotal factors influencing the survival duration of breast cancer patients.\u003c/p\u003e "},{"header":"Discussion","content":"\u003cp\u003eIn this study, we conducted survival analysis on data from breast cancer patients in the UK Biobank, which also provides extensive information on personal health factors such as lifestyle habits. We employed three feature selection methods to identify key factors and utilized five survival analysis algorithms to establish predictive models for 3-year, 6-year, and 9-year survival outcomes. Ultimately, the XGBoost model demonstrated the best performance. Furthermore, we adopted the Kaplan-Meier method to explore and visually present the impact of various factors on the survival duration of breast cancer patients. Additionally, we utilized SHAP analysis to ascertain which features had the most significant influence on the model's predictions.\u003c/p\u003e \u003cp\u003eIn addition, we conducted a comprehensive evaluation of five different machine learning models in terms of their performance in predicting survival outcomes using complex medical data. Among them, DeepSurv, as a deep learning algorithm, demonstrated its unique advantages, particularly in modeling survival data that exhibit both linear and nonlinear risk characteristics, showing superiority compared to the traditional Cox Proportional Hazards algorithm. This finding is consistent with the viewpoint expressed in literature\u003csup\u003e[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/sup\u003e, which states that the DeepSurv model can be seen as an extension of the Faraggi-Simon network and experimentally verifies that DeepSurv outperforms the Cox algorithm when handling survival data containing both linear and nonlinear risk functions. However, it is worth noting that despite DeepSurv's breakthrough in performance, it did not surpass the XGBoost model in this study. This result aligns with the research conducted by L\u0026eacute;o Grinsztajn et al.\u003csup\u003e[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]\u003c/sup\u003e, which reported that tree-based methods are easier to achieve good predictions on tabular data than deep learning. Therefore, deep learning algorithms have an advantage in capturing complex nonlinear relationships, while machine learning algorithms, especially ensemble learning algorithms, excel in processing large-scale data and improving prediction accuracy and stability. In practical applications, the choice of algorithm depends on the nature of the specific problem, the characteristics of the data, and the available computational resources.\u003c/p\u003e \u003cp\u003eWe noted that the Kaplan-Meier curve results from both the training and validation cohorts showed that the survival probability of the high-risk group was significantly lower than that of the low-risk group, with \u003cem\u003eP\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.0001. This finding not only verifies the validity of the features screened in our study, but also further emphasizes the importance of these features in breast cancer survival analysis.\u003c/p\u003e \u003cp\u003eIn this study, we utilized three different feature selection methods to consistently identify 6 key variablesthat potentially play critical roles in the prognosis of breast cancer patients. SHAP interpretability analysis further confirmed the prognostic importance of these features, as each was shown to have a substantial positive or negative impact on the model\u0026rsquo;s predictions. Although features such as fruit intake and psychological stress (e.g., nervousness) were not selected as key features, they demonstrated moderate importance during the initial screening process, indicating that these factors also exert a certain influence on the prognosis of breast cancer patients.\u003c/p\u003e \u003cp\u003eInterestingly, our research findings indicate that the genetic situations of siblings and mothers do not have a significant impact on the survival prognosis of breast cancer patients. This may be attributed to the fact that the development of breast cancer involves multigene inheritance and complex interactions with environmental factors, making the contribution of single genetic factors relatively limited in overall prognosis assessment and thereby affecting their significance in prognostic analysis.\u003c/p\u003e \u003cp\u003eBeyond traditional prognostic factors, our analysis revealed a noteworthy association between the use of non\u0026ndash;breast cancer medications and increased mortality risk. Specifically, our research data clearly shows that patients who take more non breast cancer drugs face a higher risk of mortality than those who take less of these drugs. Prior evidence suggests that certain non\u0026ndash;breast cancer drugs may contribute to tumor progression\u003csup\u003e[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]\u003c/sup\u003e, or may interact adversely with oncologic treatments\u003csup\u003e[\u003cspan additionalcitationids=\"CR32 CR33\" citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/sup\u003e. These interactions can reduce the therapeutic effect, aggravate side effects, or even trigger new health problems, thus directly or indirectly affecting the survival of patients.\u003c/p\u003e \u003cp\u003eIn our in-depth analysis of how age influences survival predictions in breast cancer patients, SHAP charts presented a key finding: elderly individuals exhibit a significant negative trend in the impact on the prediction model, while young individuals show a positive trend. This result clearly indicates that the prognosis of elderly breast cancer patients is significantly worse than that of young patients in our data set. This finding is consistent with previous studies, which have reported substantial age-related differences in survival, often attributed to factors such as reduced physiological reserve, comorbidities, and increased susceptibility to treatment-related complications in older adults\u003csup\u003e[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/sup\u003e. Further supporting this observation, additional research has identified age\u0026thinsp;\u0026ge;\u0026thinsp;60 as an independent predictor of poor prognosis in female breast cancer patients\u003csup\u003e[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]\u003c/sup\u003e. Together, these findings reinforce the notion that increasing age is a critical factor associated with elevated mortality risk following breast cancer diagnosis.\u003c/p\u003e \u003cp\u003eMoreover, the adverse impact of high BMI on breast cancer prognosis was also demonstrated. This conclusion is consistent with previous studies, which also found a strong association between high levels of obesity and overall mortality, breast cancer-specific mortality, recurrence, and distant metastasis in breast cancer patients, which further supports the association between obesity and poor prognosis\u003csup\u003e[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]\u003c/sup\u003e. These findings further emphasize the crucial role of BMI in the assessment of breast cancer prognosis and provide new insights for future exploration of breast cancer prevention and treatment strategies.\u003c/p\u003e \u003cp\u003eOur study also explored the impact of other characteristics on individual breast cancer survival risk. Among these, high smoking frequency and early menopausal age have been confirmed to decrease survival for breast cancer patients. Frequent smoking behavior not only elevates the risk of developing cancer but also adversely affects the prognosis of cancer patients\u003csup\u003e[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]\u003c/sup\u003e. Although the effect of alcohol intake on survival was relatively modest in our analysis, we observed that lower levels of alcohol consumption were unexpectedly associated with worse survival outcomes. This may be due to studies suggesting that moderate alcohol consumption may have a certain protective effect on cardiovascular health, thereby indirectly influencing the survival of breast cancer patients\u003csup\u003e[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u003c/sup\u003e, this conclusion still requires further verification.\u003c/p\u003e \u003cp\u003eFrom a clinical perspective, our model provides an e accessible and cost-effective approach to predict survival outcomes in breast cancer patients. By utilizing only comprehensive personal health factors as features, our prediction model achieved accuracy rates of 72%, 73%, and 75% respectively in the 3-year, 6-year, and 9-year assessments. This result fully demonstrates that, even under resource-limited conditions, by reasonably selecting and utilizing key health indicators, we can still provide extremely valuable references for the prognosis assessment of breast cancer patients.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eHowever, this study also has certain limitations. Firstly, the relatively small sample size may limit the generalizability and reliability of the research results. Secondly, we did not consider all factors that may affect the survival of breast cancer patients, such as genetics, lifestyle, and socioeconomic status, which are important prognostic factors for breast cancer. Finally, in terms of data processing, although combining \"No\" and \"Unknown\" into a single category can enhance data consistency and analytical robustness, it also has obvious drawbacks, including information loss, potential bias introduction, and reduced data transparency. Future research will focus on validating the results of this study across multiple cohorts, while expanding the sample size and incorporating more factors, such as socioeconomic status and specific medical test data, into the analysis, as well as conducting prospective randomized clinical trials to verify our findings.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study conducted an in-depth analysis of the synergistic effects among individual health-related factors and systematically demonstrated how lifestyle habits collectively influence the survival outcomes of breast cancer patients. The survival analysis model developed herein demonstrates promising performance in predicting the survival duration of breast cancer patients. It not only provides robust scientific evidence for the prognosis assessment of breast cancer patients but also provides solid guidance for personalized habit and lifestyle adjustment, possessing significant clinical guidance value.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003e\u003cem\u003eAIC\u003c/em\u003e (Akaike Information Criterion)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAUC\u003c/em\u003e (area under the curve)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eBMI\u003c/em\u003e (Body Mass Index)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eCIs\u003c/em\u003e (confidence intervals)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eCV\u003c/em\u003e (cross-validation)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eENet\u003c/em\u003e (Elastic Network)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eGlmBoost\u003c/em\u003e (Gradient Boosting with Component-wise Linear Models)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eGLMs\u003c/em\u003e (Generalized Linear Models)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eLasso\u003c/em\u003e (Least absolute shrinkage and selection operator)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eNHS\u003c/em\u003e (National Health Service)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003ePRS\u003c/em\u003e (Polygenic risk score)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSHAP\u003c/em\u003e (SHapley Additive exPlanations)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eStepcox\u003c/em\u003e (Stepwise Cox regression)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eTimeROC\u003c/em\u003e (Time-dependent receiver operating curve)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eXGBoost\u0026nbsp;\u003c/em\u003e(eXtreme Gradient Boosting)\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eYishan Yao (Conceptualization; Formal analysis; Investigation; Methodology; Visualization; Writing-original draft; Writing-review \u0026amp; editing), Xiaobing Zhai (Formal analysis;Writing-review \u0026amp; editing), Zhichao Liang (Formal analysis; Writing-review \u0026amp; editing), Chi Kin Lam (Supervision; Writing-review \u0026amp; editing), Hui Xie (Supervision), Tong, H.H.Y. (Data Curation), Ritse Mann (Supervision), Tong Tong (Writing-review \u0026amp; editing), Yan Chen (Supervision), Muzhen He (Supervision), Kefeng Li (Data Curation), Tao Tan (Writing-review \u0026amp; editing; Supervision; Conceptualization)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe funding sources played no role in the study design, data collection, data analysis, and interpretation of results or the decisions made in preparation and submission of the article.\u003c/p\u003e\n\u003cp\u003eUK Biobank has received core funding from the Wellcome Trust medical charity, Medical Research Council, Department of Health, Scottish Government and the Northwest Regional Development Agency, the Welsh Government, British Heart Foundation, Cancer Research UK and Diabetes UK, and National Institute for Health and Care Research.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe details of UK Biobank core funding and additional funding are reported at https://www.ukbiobank.ac.uk/learn-more-about-uk-biobank/about-us/our-funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no potential conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of Data and Materials Statements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research has been conducted using data from the UK Biobank under Application Number 99946. The UK Biobank data are not publicly available due to participant privacy and data sharing restrictions. Researchers can apply for access to the data through the UK Biobank resource at https://www.ukbiobank.ac.uk.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical approval:\u003c/strong\u003e The details of UK Biobank research ethics approval are elaborated at https://www.ukbiobank.ac.uk/learn-more-about-uk-biobank/about-us/ethics.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGiaquinto AN, Sung H, Miller KD, Kramer JL, Newman LA, Minihan A, et al. Breast Cancer Stat 2022 CA: cancer J Clin. 2022;72(6):524\u0026ndash;41.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoibl S, Poortmans P, Morrow M, Denkert C, Curigliano G. Breast cancer. Lancet (London England). 2021;397(10286):1750\u0026ndash;69.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I et al. Global cancer statistics 2022:GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. Cancer J Clin. 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEllsworth RE, et al. Impact of lifestyle factors on prognosis among breast cancer survivors in the USA. Expert Rev pharmacoeconomics outcomes Res vol. 2012;12(4):451\u0026ndash;64. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1586/erp.12.37\u003c/span\u003e\u003cspan address=\"10.1586/erp.12.37\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGBD 2016 Alcohol Collaborators. Lancet (London England). 2018;392(10152):1015\u0026ndash;35. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/S0140-6736(18)31310-2\u003c/span\u003e\u003cspan address=\"10.1016/S0140-6736(18)31310-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Alcohol use and burden for 195 countries and territories, 1990\u0026ndash;2016: a systematic analysis for the Global Burden of Disease Study 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePassarelli MN, Newcomb PA, Hampton JM, et al. Cigarette Smoking Before and After Breast Cancer Diagnosis: Mortality From Breast Cancer and Smoking-Related Diseases. J Clin Oncol. 2016;34(12):1315\u0026ndash;22. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1200/JCO.2015.63.9328\u003c/span\u003e\u003cspan address=\"10.1200/JCO.2015.63.9328\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrintz C. Anxiety and depression may contribute to an increased risk of death in some cancers. Cancer. 2017;123(13):2389\u0026ndash;2389.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDehesh T, Fadaghi S, Seyedi M, et al. The relation between obesity and breast cancer risk in women by considering menstruation status and geographical variations: a systematic review and meta-analysis. BMC Womens Health. 2023;23:392. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12905-023-02543-5\u003c/span\u003e\u003cspan address=\"10.1186/s12905-023-02543-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSoldato D, Michiels S, Havas J, Di Meglio A, Pagliuca M, Franzoi MA, Pistilli B, Iyengar NM, Cottu P, Lerebours F, Coutant C, Bertaut A, Tredan O, Vanlemmens L, Jouannaud C, Hrab I, Everhard S, Martin AL, Andr\u0026eacute; F, Vaz-Luis I, Jones LW. Dose/Exposure Relationship of Exercise and Distant Recurrence in Primary Breast Cancer. J Clin Oncol. 2024;42(25):3022\u0026ndash;32. Epub 2024 Jun 5. PMID: 38838281; PMCID: PMC11361355.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaryam S, Farvid RM, Tamimi EM, Poole, Wendy Y, Chen BA, Rosner WC, Willett MD, Holmes A. Heather Eliassen; Postdiagnostic Dietary Glycemic Index, Glycemic Load, Dietary Insulin Index, and Insulin Load and Breast Cancer Survival. Cancer Epidemiol Biomarkers Prev 1 Febr. 2021;30(2):335\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeSantis CE, Ma J, Goding Sauer A, Newman LA, Jemal A. Breast cancer statistics, 2017, racial disparity in mortality by state. CA Cancer J Clin. 2017;67:439\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKreklau A et al. An Observational Study on Breast Cancer Survival and Lifestyle Related Risk Factors. In vivo (Athens, Greece) vol. 35,2 (2021): 1007\u0026ndash;1015. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.21873/invivo.12344\u003c/span\u003e\u003cspan address=\"10.21873/invivo.12344\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eParada H Jr, et al. Lifestyle Patterns and Survival Following Breast Cancer in the Carolina Breast Cancer Study. Epidemiol (Cambridge Mass) vol. 2019;30(1):83\u0026ndash;92. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/EDE.0000000000000933\u003c/span\u003e\u003cspan address=\"10.1097/EDE.0000000000000933\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSudlow C, Gallacher J, Allen N, Beral V, Burton P, Danesh J, et al. UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Med. 2015;12(3):e1001779.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ehttps://\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003c/span\u003e\u003cspan address=\"http://www.ukbiobank.ac.uk/enable-your-research/about-our-data/questionnaire-data\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFoster HME, Celis-Morales CA, Nicholl BI, Petermann-Rocha F, Pell JP, Gill JMR, et al. The effect of socioeconomic deprivation on the association between an extended measurement of unhealthy lifestyle factors and health outcomes: a prospective analysis of the UK Biobank cohort. Lancet Public Health. 2018;3(12):e576\u0026ndash;85.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoller H, Richards S, Hanchett N, Riaz SP, Luchtenborg M, Holmberg L, et al. Completeness ofcase ascertainment and survival time error in English cancer registries: impact on 1-year survival estimates. Br J Cancer. 2011;105:170\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHastie TJ, Pregibon D. In: Chambers S, Hastie TJ, editors. Generalized linear models. Chapter 6 of Statistical Models. Wadsworth \u0026amp; Brooks/Cole; 1992.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTibshirani R. Regression Shrinkage and Selection via the Lasso. J royal Stat Soc Ser b-methodological. 1996;58:267\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen T, Guestrin C, Data Mining. XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and (2016): n. pag.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCox DR. Regression models and life-tables[J]. J Roy Stat Soc: Ser B (Methodol). 1972;34(2):187\u0026ndash;202.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol. 2018;18(1):24. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12874-018-0482-1\u003c/span\u003e\u003cspan address=\"10.1186/s12874-018-0482-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 29482517; PMCID: PMC5828433.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaszke A, Chaurasia A, Kim S et al. ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation[J]. 2016.\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.48550/arXiv.1606.02147\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.1606.02147\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeter Buehlmann and Bin Yu. Boosting with the L2 loss: regression and classification. J Am Stat Assoc. 2003;98:324\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo W, Key TJ, Reeves GK. Adiposity and breast cancer risk in postmenopausal women: results from the UK Biobank prospective cohort. Int J Cancer. 2018;143:1037\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKamarudin AN, Cox T, Kolamunnage-Dona R. Time-dependent ROC curve analysis in medical research: current methods and applications. BMC Med Res Methodol. 2017;17:53. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12874-017-0332-6\u003c/span\u003e\u003cspan address=\"10.1186/s12874-017-0332-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundberg S. (2017). A unified approach to interpreting model predictions. arXiv preprint arXiv:170507874.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrinsztajn L, Oyallon E. Why do tree-based modelsstill outperform deep learning on tabular data? 2022, arXiv:2207.08815.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollins JA, Blake JM, Crosignani PG. Breast cancer risk with postmenopausal hormonal treatment. Hum Reprod Update. 2005;11:545\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSt\u0026oslash;er NC, et al. Drug use and cancer risk: a drug-wide association study (DWAS) in Norway. CancerEpidemiol Biomark Prev. 2021;30:682\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaker AF, Dorr RT. Drug interactions with the taxanes: clinical implications. Cancer Treat Rev. 2001;27:221\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHamy A-S, et al. Celecoxib with neoadjuvant chemotherapy for breast cancer might worsen outcomes differentially by cox-2 expression and er status: exploratory analysis of the REMAGUS02 trial. JCO. 2019;37:624\u0026ndash;35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTyler T. Drug interactions in metastatic breast cancer. J Oncol Pharm Pr. 2011;17:236\u0026ndash;45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBibi R, et al. Prevalence of potential drug-drug interactions in breast cancer patients and determination of their risk factors. J Oncol Pharm Pr. 2021;27:1616\u0026ndash;22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Y, Wen Y, Xiang Z, Ma J, Lin Y, Huang Y, Chen C. Cancer Survival Trends in Southeastern China, 2011\u0026ndash;2021: A Population-Based Study. Clin Epidemiol. 2024;16:45\u0026ndash;56. PMID: 38318284; PMCID: PMC10840559.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen HL, Zhou MQ, Tian W, Meng KX, He HF. Effect of Age on Breast Cancer Patient Prognoses: A Population-Based Study Using the SEER 18 Database. PLoS ONE. 2016;11(10):e0165409. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0165409\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0165409\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PMID: 27798652; PMCID: PMC5087840.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePang Y, Wei Y, Kartsonaki C. Associations ofadiposity and weight change with recurrence and survival in breast cancer patients: a systematic review and meta-analysis. Breast Cancer. 2022;29:575\u0026ndash;88. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s12282-022-01355-z\u003c/span\u003e\u003cspan address=\"10.1007/s12282-022-01355-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNewcomb PA, Kampman E, Trentham-Dietz A, et al. Alcohol consumption before and after breast cancer diagnosis: associations with survival from breast cancer, cardiovascular disease, and other causes. J Clin Oncol. 2013;31(16):1939\u0026ndash;46. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1200/JCO.2012.46.5765\u003c/span\u003e\u003cspan address=\"10.1200/JCO.2012.46.5765\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"breast-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"brca","sideBox":"Learn more about [Breast Cancer](http://link.springer.com/journal/12282)","snPcode":"12282","submissionUrl":"https://www.editorialmanager.com/brca/default2.aspx","title":"Breast Cancer","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Survival analysis, Kaplan-Meier, XGBoost, TimeROC, Lifestyle Patterns","lastPublishedDoi":"10.21203/rs.3.rs-8588282/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8588282/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cb\u003eBackground:\u003c/b\u003e\u003c/p\u003e \u003cp\u003eUnderstanding how lifestyle habits influence breast cancer survival is crucial for improving long-term outcomes and guiding individualized care. The purpose of this study is to comprehensively understand how dietary habits, exercise frequency, sleep quality and the frequency of tobacco and alcohol use affect the survival of breast cancer patients by analyzing the synergistic effects of multiple lifestyle habits.\u003c/p\u003e\u003cp\u003e\u003cb\u003eMethods:\u003c/b\u003e\u003c/p\u003e \u003cp\u003eWe analyzed data from 14,901 female breast cancer patients in the UK Biobank, a large-scale population-based cohort. Three feature selection methods were applied to identify key prognostic variables. Five survival risk models were compared to investigate the relationship between breast cancer survival and lifestyle habits. Kaplan-Meier survival analysis and SHapley Additive exPlanations (SHAP) interpretability analysis were performed on risk scores calculated from the best-performing model.\u003c/p\u003e\u003cp\u003e\u003cb\u003eResults:\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThrough feature selection, we identified that age, number of medications taken, Body Mass Index (BMI), alcohol intake, diastolic blood pressure and smoking status were pinpointed as critical characteristics influencing the survival of breast cancer patients. Among all models, the eXtreme Gradient Boosting (XGBoost) model demonstrated the highest predictive performance (3 years: AUC\u0026thinsp;=\u0026thinsp;0.72, 6 years: AUC\u0026thinsp;=\u0026thinsp;0.727, 9 years: AUC\u0026thinsp;=\u0026thinsp;0.749). Kaplan-Meier analysis showed that patients classified as high-risk based on the median risk score had significantly worse survival outcomes (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). SHAP analysis further confirmed the dominant influence of age and BMI on mortality risk.\u003c/p\u003e\u003cp\u003e\u003cb\u003eConclusions:\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThis study highlights the prognostic value of lifestyle habits in breast cancer survival. By integrating routine health indicators into interpretable machine learning models, our findings provide a practical foundation for individualized risk assessment and lifestyle-based intervention strategies.\u003c/p\u003e","manuscriptTitle":"Impact of Lifestyle Patterns on Breast Cancer Prognosis: Evidence from a UK Biobank Survival Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-20 11:28:38","doi":"10.21203/rs.3.rs-8588282/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major Revision","date":"2026-02-18T00:45:34+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2026-01-29T00:54:13+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-01-15T23:46:10+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-13T15:40:07+00:00","index":"","fulltext":""},{"type":"submitted","content":"Breast Cancer","date":"2026-01-13T01:57:12+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"breast-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"brca","sideBox":"Learn more about [Breast Cancer](http://link.springer.com/journal/12282)","snPcode":"12282","submissionUrl":"https://www.editorialmanager.com/brca/default2.aspx","title":"Breast Cancer","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"0b8b3d1e-76f1-4db0-b864-864f1a7f9334","owner":[],"postedDate":"January 20th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-17T07:45:04+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-20 11:28:38","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8588282","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8588282","identity":"rs-8588282","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.