Machine learning survival models trained on clinical data to identify high risk patients with hormone responsive HER2 negative breast cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Machine learning survival models trained on clinical data to identify high risk patients with hormone responsive HER2 negative breast cancer Annarita Fanizzi, Domenico Pomarico, Alessandro Rizzo, Samantha Bove, and 12 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2238591/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 26 May, 2023 Read the published version in Scientific Reports → Version 1 posted 11 You are reading this latest preprint version Abstract For endocrine-positive Her2 negative breast cancer patients at an early stage, the benefit of adding chemotherapy to adjuvant endocrine therapy is controversial. Several genomic tests are available on the market but are very expensive. Therefore, there is the urgent need to explore novel reliable and less expensive prognostic tools in this setting. In this paper, we shown a machine learning survival model to estimate Invasive Disease-Free Events trained on clinical and histological data commonly collected in clinical practice. We collected clinical and cytohistological outcomes of 145 patients referred to Istituto Tumori “Giovanni Paolo II”. Three machine learning survival models are compared with the Cox proportional hazards regression according to time-dependent performance metrics evaluated in cross-validation. The c-index at 10 years obtained by random survival forest, gradient boosting, and component-wise gradient boosting is stabled with or without feature selection at approximately 0.68 in average respect to 0.57 obtained to Cox model. Moreover, machine learning survival models have accurately discriminated low- and high-risk patients, and so a large group which can be spared additional chemotherapy to hormone therapy. The preliminary results obtained by including only clinical determinants are encouraging. The integrated use of data already collected in clinical practice for routine diagnostic investigations, if properly analyzed, can reduce time and costs of the genomic tests. Biological sciences/Cancer Health sciences/Oncology Health sciences/Risk factors Machine learning Survival Breast Cancer her2 negative patients hormone responsive patients decision support system Figures Figure 1 Figure 2 Figure 3 Figure 4 1. Introduction For breast cancer (BC) patients with endocrine-positive and HER2 negative early-stage disease, the benefit of adding chemotherapy to adjuvant endocrine therapy is controversial. These patients are at risk for being undertreated or overtreated with endocrine therapy and chemotherapy, and tests are required to save an important number of patients from the potentially harmful side effects of chemotherapy; in particular, several studies have showed that a non-negligible proportion of BC patients, especially those with a hormone receptor-positive and lymph node-negative disease, could only be effectively treated with hormone therapy alone [ 1 , 2 ]. The use of adjuvant chemotherapy for estrogen receptor (ER) – positive, HER2-negative BC patients has been investigated by an impressive number of studies aimed at measuring its efficiency in a predictive manner [ 3 ]. Such studies range from genomic tests [ 4 , 5 ] to sophisticated artificial intelligence models [ 6 ], with the purpose of describing the benefit gained by each patient undergoing a specific therapy. Recent years have witnessed the availability of several molecular tests which have received long-standing recommendations in clinical guidelines [ 7 ]. In particular, the use of gene signatures has provided a standardized reproducible and quantitative tool able to define the risk of distant recurrent for ER-positive, HER2-negative early BC. Nevertheless, the adoption in the clinical practice of these decisional support tools requires a careful analysis of their cost-effectiveness, because genomics tests have an important cost and not all centers are provided with laboratories performing this type of analyses. This issue is currently driving the studies aimed the achievement of the same information by means of less expensive procedure. In general, new interdisciplinary approaches are emerging in survival analysis, which aim to analyze data commonly collected in the clinical practice and drive the therapeutic choices. Indeed, in clinical practice, medical oncologists are increasingly using prediction tools available online, such as PREDICT, Adjuvant!, and CancerMath to guide systemic adjuvant treatment [ 8 ]. The online tools provide personalized 10-year overall survival estimates for the adjuvant treatment setting by basing their predictions on patient data (e.g. age) and tumour characteristics (e.g. size, nodal status, ER-status and grade), but they perform well at the population level, but exhibit a high degree of discordance in the intermediate and poor prognosis groups [ 9 , 10 ]. Furthermore, some models have been proposed for the estimation of disease-free survival with classic approaches [ 11 , 12 ], but works aimed at predicting high-risk patients who might actually benefit from additional chemotherapy to hormone therapy is missing. A wide variety of techniques is currently available, ranging from classical non-parametric Kaplan-Meier descriptive curves to extensions of the semi-parametric inferential Cox model. A limit of classical algorithms is the difficulty to model high dimensionality. Recently, machine learning techniques applied to survival tasks allow to overcome this issue [ 13 – 17 ]. Indeed, the classical Cox regression is a parametric model based on a probabilistic estimation whose prediction performances depend on parameters associated with each feature. Therefore, the difficulty in identifying an accurate probabilistic model increases in parallel with the increase in the included number of features considered. On the contrary, machine learning survival model, such as random forest and gradient boosting survival models, are non-parametric methods whose performances depend on the size of the training set. In practice, the latter models do not impose any hypothesis on the probabilistic distribution, thus allowing to properly model nonlinearities and interaction effects in a data driven approach [ 18 ]. If, on the one hand, these limitations are pursued to achieve the explainability of black-box machine learning survival models [ 19 – 21 ], on the other their overcoming guarantees higher performances [ 18 ]. In this work, we propose a model for estimating disease-free survival with respect to invasive events for patients with endocrine-positive and HER2 negative BC, which are potentially candidates for genomic testing, which are potentially candidates for genomic testing, if only hormone therapy is carry out. Our preliminary study is configured to identify low- and high-risk patients and assess the chance of achieving comparable performances with genomic test but exploiting much cheaper and already available clinical data. Three machine learning survival models are compared with the Cox proportional hazards regression according to time-dependent classification performance metrics [ 22 ]. Once the best machine learning survival model was chosen, we evaluated the correlation risk score obtained from our model with that of a genetic test performed on a sample of independent patients. 2. Results 2.1. Enrolled Patients and Features Our dataset is composed by clinical and cytohistological outcomes of 145 patients our extracted from our database of approximately 900 patients registered for a first BC diagnosis in the period 1997–2019 and referred to Istituto Tumori “Giovanni Paolo II” in Bari (Italy). The inclusion criteria for collecting such database were: absence of primary chemotherapy for BC, ab initio non-metastatic patient. Then, according to the genomic test eligibility criteria defined by the decree of the Ministry of Health of May 2021 [ 23 ], i.e. early stage tumor, patients not at high or low risk of recurrence with hormone responsive and HER2 negative BC, 145 patients wase extracted. In line with the aim of our work, we considered the only patients who did not undergo that chemotherapy. In fact, for patients who have undergone chemotherapy, the absence of a second event could be due to a patient-specific positive prognostic profile and not necessarily to an effect of the therapeutic treatment. In other words, there may be patients who would not have relapsed even if they had not undergone additional chemotherapy. In this work, we specifically focus our attention on breast cancer-related invasive disease events (IDEs), which include local recurrence, the appearance of distant visceral and soft tissue metastases, contralateral invasive breast cancer or a second primary tumor [ 24 ]. Collected features were the age at diagnosis, tumor size (diameter: T1a, T1b, T1c, T2, T3, T4), histological subtype (ductal, lobular, other), type of surgery (quadrantectomy/mastectomy), estrogen receptor expression (ER, %), progesterone receptor expression (PgR, %), cellular marker for proliferation (Ki67, %), histological grade (grading, Elston–Ellis scale: G1, G2, G3), human epidermal growth factor receptor-2 score (HER2/neu: 0 + , 1 + , 2 + ), the number of metastatic and eradicated lymph nodes, lymph nodes dissection (no/yes), sentinel lymph node (no, negative, positive), lymph nodes stage (N: 0, 1, 2, 3), in situ component (absent, G1, G2, G3, present but not typed), lymphovascular invasion (absent, focal, extensive, present but not typed), multiplicity (no/yes) and previous tumors (no/yes). The data set is described in Table 1 . The set of predictive features is composed by \(N=18\) prognostic factors, typically considered by clinicians during the first tumor diagnosis and related surgery. The missing data recovery has been implemented by means of the Python package musingly (v. 0.2.0). A separate dataset, composed by 27 patients endowed with EndoPredict® (EP) scores and undergoing surgery during 2021 in our institute, is exploited for further evaluations of our survival estimation. EP is a gene expression test for patients with ER-positive and HER2-negative early-stage BC, both node-negative and node-positive (N0, N1, micrometastasis). It is a second-generation test that combines a molecular score of 12 genes with tumor size and lymph node status [ 5 ]. This genomic test has entered the clinical practice of the Istituto Tumori 'Giovanni Paolo II' in Bari since 2021 and it is used by the Breast Care team when the clinical case is highly doubtful. Table 1 Observed patients’ statistics according to considered features. Features Counts (%) Features Counts (%) Overall 145 (100) Lymph Nodes Stage Type of Surgery N0 114 (78.6) quadrantectomy 108 (74.5) N1 26 (17.9) mastectomy 37 (25.5) N2 2 (1.4) In Situ Component N3 1 (0.7) absent 101 (69.7) NA 2 (1.4) G1 7 (4.8) Lymph Node Dissection G2 8 (5.5) no 34 (23.5) G3 4 (2.8) yes 104 (71.7) present, not typed 25 (17.2) NA 7 (4.8) HER2/neu Sentinel Lymph Node 0 + 70 (48.3) no 97 (66.9) 1 + 56 (38.6) negative 34 (23.5) 2 + 15 (10.3) positive 8 (5.5) NA 4 (2.8) NA 6 (4.1) Multiplicity Grading no 118 (81.4) G1 20 (13.8) yes 27 (18.6) G2 104 (71.7) Diameter G3 17 (11.7) T1b 26 (17.9) NA 4 (2.8) T1c 68 (46.9) Lymphovascular Invasion T2 42 (29.0) absent 98 (67.6) T3 1 (0.7) focal 33 (22.8) T4 3 (2.1) extensive 4 (2.8) NA 5 (3.5) present, not typed 10 (6.9) Histologic Type Previous Tumors ductal 107 (73.8) no 140 (96.6) lobular 20 (13.8) yes 5 (3.5) other 18 (12.4) Median [ \({{q}}_{0},{{q}}_{1},{{q}}_{3},{{q}}_{4}\) ] Median [ \({{q}}_{0},{{q}}_{1},{{q}}_{3},{{q}}_{4}\) ] ER 80 [9, 70, 90, 100] Age 59 [32, 49, 67, 86] PgR 60 [0, 20, 90, 98] Metastatic Lymph Nodes 0 [0, 0, 0, 10] Ki67 12 [1, 5, 20, 70] Eradicated Lymph Nodes 16 [0, 2, 23, 40] Table 2 Summary of the sample labels variation in the time period comprised between 20 and 160 months. Invasive disease events 7 12 17 20 27 29 30 30 33 35 38 39 41 43 44 Control cases 136 130 125 121 112 109 107 103 91 82 75 64 47 31 20 Months 20 31 40 51 60 70 80 90 100 110 120 130 140 150 160 2.2. Time Dependent Classification Random survival forest feature importance is resumed in Fig. 1 . Their calculation is nested in the 20 rounds of 5-fold cross-validations to avoid any influence imposed by a single evaluation with a fixed training set. To take into account the included statistical variation, each feature weight is described by its average and standard deviation. We select those features characterized by a weight greater than 0.01, thus yielding: Ki67, PgR, age, ER, eradicated lymph nodes and diameter. The ability of machine learning survival algorithms to model data high dimensionality with respect to the classical CPH regression [ 18 ] is shown by comparing the time behavior of the considered metrics (see Appendix B) when all features are included (N = 18) or just the six selected ones. In Fig. 2 , the upper panels are related with the first case, while the lower panels to the latter one. At 5 and 10 years, time period usually considered for follow-up in clinical practice, the performances of the CPH model are on average lower than those of machine learning models by at least 10 percentage points, when all the features are considered. However, when just a subset of features is considered, consisting in the most important ones, the average difference in performance is halved, signaling a better condition for the CPH model, while the performance of machine learning survival models tends to remain unchanged with respect to the number of features involved. In order to define a parsimonious model, i.e. to use just the right number of predictors needed to explain the model well, the following analyses will be carried out on the results obtained by considering the selected subset of features. The sensitivity and specificity balanced performance (see Fig. 6 in Appendix C) corresponding to the 5 years time frame are equal to 0.62–0.65 for the three machine learning survival models, while the much lower one of CPH shows approximately 0.55 for the same balanced metrics pair. If we consider 10 years after the first BC, the CPH model still shows the lowest performance, while the machine learning survival models RSF and GB are characterized by a balanced aforementioned metric pair equal to 0.63–0.65, while CGB emerges as the best performing classifier with 0.67 for the balanced sensitivity and specificity pair. Once the 10 years time frame is kept fixed, we establish a threshold for each model according to the median of the ones selected by the Youden index optimization (see Fig. 6 in Appendix 6). The average score of each patient over the 20 rounds is then adopted to assign each case to a high or low risk category. These strata are further characterized by means of the Kaplan-Meier curves shown in Fig. 3 . The p-values confirm that CGB implements the best discrimination between high or low risk patients, while CPH is much less efficient than the machine learning survival models. 2.3. Correlation with EndoPredict® Scores The similarity measure of our risk estimation with the one predicted by EP is evaluated over an independent test consisting of 27 patients. The last step of our preliminary study consists in the calculation of the Pearson correlation coefficients between the hazards (see Appendix A) estimated by our best performing CGB survival curves at 10 years after the first BC diagnosis and the risk scores provided to our institution by the EP software exploiting genetic data. A scores statistics for the separate dataset is obtained by testing the sample set of 27 patients on 20 rounds 5-fold cross-validation for the model trained on 145 patients, such that a variation in the training is obtained to gain a wider statistic. The hazard value corresponding to 10 years after the first breast cancer diagnosis is then deduced, whose overall correlation statistics is shown in Fig. 4 . The violin plot in the left panel is characterized by a sufficient stability around the median value, imposing the slope of the trend line in the bubble plot of the right panel. The performance of EP declared by the authors in terms of c-index is equal to 0.753 for the prediction of distant recurrences within 10 years [ 5 ]. As emerged from our results, CGB shows the highest correlations, equal in average to 0.42, with respect to the remaining survival models, instead resulting in average uncorrelated. We underline that such correlations are time independent by definition for CGB, GB and CPH, because the corresponding hazard functions assume time independent parameters, such that features and time are independent variables. 3. Discussion The high recurrence rate characterizing BC patients has prompted to the adoption of post-operative treatments, including adjuvant chemotherapy. At the same time, the risk of overtreatment in this patient population has supported the development of tools able to perform a proper risk-benefit assessment and to guide the “decision-making” process [ 1 , 2 , 7 ]. The role of molecular data has become increasingly important in guiding therapeutic decisions in this setting. Nevertheless, there is the urgent need to explore novel reliable and less expensive prognostic tools. Particular attention deserver hormone-responsive, HER2 negative BC patients for which the prescription of an adjunctive chemotherapy hormone therapy is often highly doubtful. Recently, the genomic tests play a key rule into assess the benefit provided by the addition of chemotherapy, but are very expensive and their cost-effectiveness needs to be neglected. In this paper, we proposed a machine learning approach to estimate disease free survival. To date, a plethora of predictive models have been developed to estimate disease-free survival with respect to recurrence breast cancer by solving a classification task [ 12 , 24 , 25 ] or focusing on survival [ 12 , 26 ]. However, it is known that, despite not in common cases, anticancer drugs can cause second tumors, correlated with chemotherapy [ 27 ]. Therefore, recently, in the adjuvant clinical trial setting for breast cancer, experts proposed to adopt only one term, that is Invasive Disease-Free Survival, to refer to composite events, such as local and distant recurrence, contralateral invasive breast cancers, second primary tumors and death [ 28 ]. Recent works on survival model for invasive events prediction and its variants have been freshly proposed [ 29 , 30 ] and were based on the exploitation of patients’ characteristics related to demographics, diagnosis, pathology and therapy. Among these, machine learning algorithms represent a novel, promising tool. The usefulness of survival analysis inspired by machine learning algorithms is currently assessed by interdisciplinary studies, because of its improved ability to take into account high-dimensional data with respect to classical methods [ 18 ]. Such new approaches include a much higher complexity which require otherwise an accurate feature selection [ 4 ]. Our preliminary study aims to evaluate the potential of machine learning models trained on commonly clinical features for predicting of IDEs for patients with endocrine-positive and HER2 negative BC. This tool, which has been built starting from information frequently collected in clinical practice, could replace the genomic profiling tests, notoriously more expensive in terms of application times and costs, when their application is not available [ 4 , 5 ]. In this preliminary work, we have compared three machine learning survival models with the classical approach, i.e. Cox proportional hazards regression, to predict IDES endocrine-positive and HER2 negative BC and, thus, identify low- and high-risk patients. The c-index obtained within the same time frame by CGB, GB and RSF is stable with or without feature selection at approximately 0.67–0.68 in average. Considering that EP test declares a c-index equal to 0.753 [ 5 ], the obtained preliminary results including only clinical determinants are encouraging. We subsequently verified on an independent subset of patients who performed EP tests, whether the decision suggested by the model we trained was in agreement with the result of the EP test. Even if the latter takes into account genetic information not included in our survival models, we observe a sufficient similarity of the risk estimation for an invasive disease event corresponding to 10 years after the first BC, as measured by Pearson correlation. Indeed, the proposed model showed a significant agreement with the result of the EP test meaning that clinical features, if properly elaborated, could express part of the information expressed by the genomic test. However, the main advantage is the much cheaper and already available data exploited in our scheme, providing a sufficient information as measured by the correlation similarity. If we consider the time period comprised between 5 and 10 years, a stable behavior of the mean performances emerges for the machine learning survival models. Moreover, they shown significantly higher risk estimation performance than the classical Cox model. In addition, the machine learning survival models have shown a significant ability to predict the IDEs in both early and late periods (5 and 10 years, respectively), to accurately discriminate patients at low or high risk, and to detect a large risk patient group with positive outcome after 10 year with only 5 years of endocrine therapy To the best of our knowledge, studies aimed at developing a predictive model of IDFS for patients with endocrine-positive and HER2 negative BC are lacking. Therefore, we believe that a comparison between our results with those obtained with more generic state-of-the-art models that estimate overall survival or ides but trained on heterogenous populations, can be mistaking. What we consider interesting instead is the comparability of the forecast results of the survival disease carried out with genomic tests, as previously discussed. Another software called Pam50 (Prosigna®) adopting 50 genes and engineered for distant recurrences declares a c-index related with the 3 years time frame equal to 0.72 [ 31 ], a value which is comparable with our performances. Although the general performance does not yet allow a clinical application of the model, the experimental results encourage future developments aimed at introducing features of a different nature. Therefore, our hypothesis is that it is possible to define a machine learning model trained on data commonly collected in clinical practice that could accurately surrogate the genomic test, finally, reducing the cost of healthcare without compromising patient care, and significantly impacting clinical governance. Further developments will be focused on the inclusion of radiomic features [ 32 , 33 ] as well as some radiologic indices extracted from clinical reports. Indeed, the inclusion of imaging data is investigated in radiogenomics to reduce costs of genetic tests, thus proving an information resource that has to be comprised in a high-dimensional setting. The usage of structured electronic health records is limited to relatively small datasets, while recent investigations are exploiting natural language processing to take free-text clinic notes as input [ 34 , 35 ]. Moreover, for the models adopted in survival analysis, much general hypotheses have been formulated to take into account the time dependence of regression coefficients, thus giving up the assumed independence of features in factorized hazard functions [ 36 – 39 ], which is partly observed just for RSF among the implemented machine learning survival models. 4. Materials And Methods 4.1. Survival Analysis for Risk Estimation In this work, we propose the application of machine learning survival models to evaluate the risk of an invasive disease event for each patient. Such models are formulated in Appendix A, implemented by using the Python package scikit-survival (v. 0.17.1) and listed as follows: Random survival forest (RSF); Gradient boosting (GB); Component-wise gradient boosting (CGB). A comparison of these machine learning methodologies with the well-known Cox proportional hazards (CPH) are performed. The selection of most important features is executed by means of the Python package eli5 (v. 0.11), which provides a way to compute feature importance by measuring how concordance index (c-index) decreases when a feature is not available. In the survival framework, the remotion of the relationship of a certain feature with the survival time is executed by random shuffling of its values: the weight of each feature is quantified by the drop on average of the c-index [ 40 ]. The performance metrics are estimated in a time dependent approach [ 22 ] and they are obtained by adopting both the whole set of features and just those characterized by a sufficient weight in the aforementioned importance evaluation. To understand the variation in time of some metrics exploited to assess the inferential power of a survival model, we have to rephrase it as a classifier yielding a time varying score equal to the complement to one of the survival probability $${F}_{i}\left(t\right)=1-{S}_{i}\left(t\right)$$ 1, with \(i=1,\dots ,M\) labelling each patient in the sample, with \(M\) equal to the total number of patients. The observed event time is $${Z}_{i}=\text{m}\text{i}\text{n}\{{T}_{i},{C}_{i}\}$$ 2, where \({T}_{i}\) denotes the time of invasive disease onset and \({C}_{i}\) the censoring time. In this way a time dependent disease status \({D}_{i}\left(t\right)\) takes value 0 until \({tC}_{i}\) , as shown in Table 2 starting from 20 months to avoid too unbalanced data sets. The precise formulation of the adopted metrics consisting in the time dependent version of the area under the receiver operating characteristic (ROC) curve (AUC) and c-index is presented in Appendix B. These quantities are able to describe time by time how the survival models capture the patient’s status behavior with respect to the used features. The optimization of the available parameters was implemented on ReCaS datacenter [ 41 ]. Later on, the analysis considered just the fixed time frame corresponding to 5 and 10 years after the first BC diagnosis. The research of an optimal threshold balancing sensitivity and specificity is based on the Youden index [ 4 ], whose maximization drives the solution achievement. To measure the efficiency of our scheme in estimating patients risks, the comparison with EP scores associated with a separate set of 27 patients is implemented. Such patients are endowed with the feature subset selected in the described procedure. Declarations Author Contributions: Conceptualization, A.F., D.P. and R.M.; methodology, A.F. and D.P.; software, D.P.; validation, A.F., D.P., V.L. and R.M.; formal analysis, A.F. and D.P.; investigation, A.F., D.P., V.L. and R.M.; resources, V.D., D.L.F., L.R., A.Z. and V.L.; data curation, D.P., S.B. and N.P.; writing—original draft preparation, A.F., D.P., A.R., S.B., M.C.C. and R.M.; writing—review and editing, V.D., D.L.F., A.L., F.G., M.I.P., L.R., P.T. and A.Z.; visualization, A.F. and D.P.; supervision, A.L., F.G., M.I.P., L.R., A.Z., V.L. and R.M.; project administration, V.L. and R.M.; funding acquisition, V.L. and R.M. All authors have read and agreed to the published version of the manuscript. Funding: This work was supported by funding from the Italian Ministry of Health “Ricerca Finalizzata 2018”. Institutional Review Board Statement: Institutional Review Board Statement: The study received approval from the Scientific Board of Istituto Tumori “Giovanni Paolo II”—Bari, Italy and was carried out in accordance with the Declaration of Helsinki’s standards. The authors affiliated to the Istituto Tumori “Giovanni Paolo II” RCCS, Bari are responsible for the views expressed in this article, which do not necessarily represent the ones of the Institute. Informed Consent Statement: ‘Informed consent’ for publication was waived by the Scientific Board of Istituto Tumori ‘Giovanni Paolo II’, Bari, for data related to the cohort of patients, as this study is retrospective and involves minimal risk. Data Availability Statement: The raw data supporting the conclusions of this article will be made available by the corresponding author, without undue reservation. Conflicts of Interest: The authors declare no conflict of interest. References Paik, S.; Shak, S.; Tang, G.; et al. A Multigene Assay to Predict Recurrence of Tamoxifen-Treated, Node-Negative Breast Cancer. N. Engl. J. Med. 2004, 351:2817–26. Sparano J.A.; Gray R.J.; Makower, D.F.; et al. Adjuvant Chemotherapy Guided by a 21-Gene Expression Assay in Breast Cancer. N. Engl. J. Med. 2018, 379(2): 111–121. Buus, R.; Sestak, I.; Kronenwett, R.; et al. Molecular Drivers of Oncotype DX, Prosigna, EndoPredict, and the Breast Cancer Index: A TransATAC Study. J. Clin. Oncol. 2021, 39(2): 126–135. Filipits, M.; Rudas, M.; Jakesz, R.; et al. A New Molecular Predictor of Distant Recurrence in ER-Positive, HER2-Negative Breast Cancer Adds Independent Information to Conventional Clinical Risk Factors. Clin. Cancer. Res. 2011, 17(18): 6012–20. EndoPredict Clinical Dossier. Available online: https://myriad.com/managed-care/endopredict-clinical-dossier/ (accessed on 23 March 2022). Banna, G.L.; Gomes, F.; Maltese, G.; et al. An electronic tool for frailty and fitness assessment in the immunotherapy era. Arg. Geriat. Oncol. 2021; 6: 7–14. Cardoso, F.; Kyriakides, S.; Ohno, S.; et al. Early breast cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up. Ann. Oncol. 2019, 30(8): 1194–1220. Erratum in: Ann. Oncol. 2019 , 30(10): 1674. Erratum in: Ann. Oncol. 2021 , 32(2): 284. Engelhardt, E. G.; van den Broek, A. J.; Linn, S. C.; Wishart, G. C.; Emiel, J. T.; van de Velde, A. O.; Schmidt, M. K.; et al. Accuracy of the online prognostication tools PREDICT and Adjuvant! for early-stage breast cancer patients younger than 50 years. Euro. J. Canc. 2017, 78, 37–44. Laas, E.; Mallon, P.; Delomenie, M.; Gardeux, V.; Pierga, J. Y.; Cottu, P.; Reyal, F.; et al. Are we able to predict survival in ER-positive HER2-negative breast cancer? A comparison of web-based models. Brit. J. Canc. 2015, 112(5), 912–917. Fanizzi, A.; Pomarico, D.; Paradiso, A.V.; Lorusso, V.; Massafra, R.; et al. Predicting of Sentinel Lymph Node Status in Breast Cancer Patients with Clinically Negative Nodes: A Validation Study. Cancers 2021 Lambertini, M.; Pinto, A.C.; Ameye, L.; et al. The prognostic performance of Adjuvant! Online and Nottingham Prognostic Index in young breast cancer patients. Brit. J. Canc. 2016, 115: 1471–1478. Wu, X.; Ye Y.; Barcenas, C.H.; et al. Personalized Prognostic Prediction Models for Breast Cancer Recurrence and Survival Incorporating Multidimensional Data. JNCI J. Nat. Canc. Inst. 2017, 109(7). Wang, P.; Li, Y.; Reddy, C.K. Machine Learning for Survival Analysis: A Survey. ACM Comp. Surv. 2019, 51. Ishwaran, H.; Kogalur, U.B.; Blackstone E.H; et al. Random Survival Forests. Ann. Appl. Stat. 2008, 2(3):841–860. Hothorn, T.; Buhlmann, P.; Dudoit, S.; et al. Survival ensembles. Biostatistics 2006, 7(3): 355–373. Li, H.; Luan, Y. Boosting proportional hazards models using smoothing splines, with applications to high-dimensional microarray data. Bioinformatics 2005, 21(10): 2403–2409. He, K.; Li, Y.; Zhu, J.; et al. Component-wise gradient boosting and false discovery control in survival analysis with high-dimensional covariates. Bioinformatics 2016, 32(1): 50–57. Moncada-Torres, A.; van Maaren, M.C.; Hendriks, M.P.; et al. Explainable machine learning can outperform Cox regression predictions and provide insights in breast cancer survival. Sci. Rep. 2021, 11: 6968. Kovalev, M.S.; Utkin, L.V.; Kasimov, E.M. SurvLIME: A method for explaining machine learning survival models. Know.-Bas. Sys. 2020, 203: 106164. Utkin, L.V.; Satyukov, E.D.; Konstantinov, A.V. SurvNAM: The machine learning survival model explanation. Neu. Net. 2022, 147: 81–102. Kuruc, F.; Binder, H.; Hess, M. Stratified neural networks in a time-to-event setting. Brief. Bio. 2022, 23(1): 1–11. Kamarudin, A.N.; Cox, T.; Kolamunnage-Dona, R. Time-dependent ROC curve analysis in medical research: current methods and applications. BMC Med. Res. Met. 2017, 17: 53. DECRETO 18 maggio 2021 - Gazzetta Ufficiale. Available online: https://www.gazzettaufficiale.it/eli/id/2021/07/07/21A04069/sg (accessed on 23 March 2022). Massafra, R.; Latorre, A.; Fanizzi, A.; et al. A Clinical Decision Support System for Predicting Invasive Breast Cancer Recurrence: Preliminary Results. Front. Oncol. 2021 11: 1–13. Tseng, Y.J.; et al. Predicting breast cancer metastasis by using serum biomarkers and clinicopathological data with machine learning technologies. Int. J. Med. Inform. 2019, 128: 79–86. Li, J.; et al. Predicting breast cancer 5-year survival using machine learning: A systematic review. PLoS One 2021, 16: 1–24. Zou, L.; Pei, L.; Hu, Y.; Ying, L.; Bei, P. The incidence and risk factors of related lymphedema for breast cancer survivors post operation: a 2 year follow up prospective cohort study. Breast Cancer 2018, 25: 309–314. Hudis, C.A.; et al. Proposal for standardized definitions for efficacy end points in adjuvant breast cancer trials: The STEEP system. J. Clin. Oncol. 2007, 25, 2127–2132. Demoor-Goldschmidt, C.; De Vathaire, F. Review of risk factors of secondary cancers among cancer survivors. Br. J. Radiol. 2019, 92: 1–8. Fu, B.; et al. Predicting Invasive Disease-Free Survival for Early Stage Breast Cancer Patients Using Follow-Up Clinical Data. 2019, 66: 2053–2064. Gnant, M.; Filipits, M.; Greil, R.; et al. Predicting distant recurrence in receptor-positive breast cancer patients with limited clinicopathological risk: using the PAM50 Risk of Recurrence score in 1478 postmenopausal patients of the ABCSG-8 trial treated with adjuvant endocrine therapy alone. Ann. Onc. 2014, 25: 339–345. Bai, H.X.; Lee, A.M.; Yang, L.; et al. Imaging genomics in cancer research: limitations and promises. Br. J. Radiol. 2016, 89: 20151030. Grimm, L.J.; Mazurowski M.A. Breast Cancer Radiogenomics: Current Status and Future Directions. Acad. Radiol. 2020, 27(1): 39–46. Wang, H.; Li, Y.; Khan, A.S.; Luo, Y. Prediction of breast cancer distant recurrence using natural language processing and knowledge-guided convolutional neural network. Art. Int. Med. 2020, 110:101977. Sanyal, J.; Tariq, A.; Kurian, A.W.; Rubin, D.; Banerjee, I. Weakly supervised temporal model for prediction of breast cancer distant recurrence. Sci. Rep. 2021, 11:9461. Murphy, S.A.; Sen, P.K. Time-dependent coefficients in a Cox-type regression model. Stoch. Proc. Appl. 1991, 39:153–180. Murphy, S.A. Testing for a Time Dependent Coefficient in Cox's Regression Model. Scand. J. Stat. 1993, 20:35–50. Zhang, Z.; Reinikainen, J.; Adeleke, K.A.; et al. Time-varying covariates and coefficients in Cox regression models. Ann. Transl. Med. 2018, 6(7):121. Thomas, L.; Reyes, E.M. Tutorial: Survival Estimation for Cox Regression Models with Time-Varying Coefficients Using SAS and R. J. Stat. Soft. 2014, 61:1–23. Harrell, F.E.; Califf, R.M.; Pryor, D.B.; Lee, K.L.; Rosati, R.A. Evaluating the yield of medical tests, J. Am. Med. Ass. 1982, 247: 2543–2546. ReCaS Bari. Available online: https://www.recas-bari.it/index.php/en/ (accessed on 24 March 2022). Cox, D.R. Regression Models and Life-Tables. J. Roy. Stat. Soc. 1972, 34: 187–220. Additional Declarations No competing interests reported. Supplementary Files SUPPLEMENTARYMATERIALS.docx Cite Share Download PDF Status: Published Journal Publication published 26 May, 2023 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Major revision 06 Apr, 2023 Reviews received at journal 05 Apr, 2023 Reviews received at journal 07 Mar, 2023 Reviewers agreed at journal 24 Feb, 2023 Reviews received at journal 15 Feb, 2023 Reviewers agreed at journal 09 Feb, 2023 Reviewers invited by journal 07 Feb, 2023 Editor assigned by journal 20 Nov, 2022 Editor invited by journal 08 Nov, 2022 Submission checks completed at journal 08 Nov, 2022 First submitted to journal 04 Nov, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2238591","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":150552353,"identity":"7b753690-0a03-4c13-bd48-cc37106f1d2e","order_by":0,"name":"Annarita Fanizzi","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Annarita","middleName":"","lastName":"Fanizzi","suffix":""},{"id":150552355,"identity":"9207798b-6675-4f24-b873-4296954a4e60","order_by":1,"name":"Domenico Pomarico","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Domenico","middleName":"","lastName":"Pomarico","suffix":""},{"id":150552358,"identity":"65bff33b-a914-4b28-b278-599a640d1764","order_by":2,"name":"Alessandro Rizzo","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Oncologia Per la Presa in Carico Globale del Paziente Oncologico “Don Tonino Bello”, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Alessandro","middleName":"","lastName":"Rizzo","suffix":""},{"id":150552362,"identity":"f4fb4520-f7f9-4450-9298-edca322a0610","order_by":3,"name":"Samantha Bove","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYFACHhBhAyIYDzAwSEBFDCDieLSkgZmoWnDrAcschmmBizDgtEa+gffgY56K83Lm7IcPHPjwy0Ken4H34OeKAgYZexxaDA7wJRvznLltbNmTlnBwZp+E4cwGvmTJM3gcBpQyk5zZdjtxww0eg8O8PRIJBgd4DCQb8GiRbwBp+XcORYvxT3xaGA7wmEl8bDgA0cLzA6zFDK8tBof5kg0+HEs2NjgD8ksD0C/NfGmWDQYSPDwHcDisvffgg4QaOzmD44cPPvjwp06en7338M2GPzb27A04rGFG5jC2wUUkcPkEHfwhVuEoGAWjYBSMJAAAAFVSQ0JqR+0AAAAASUVORK5CYII=","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Samantha","middleName":"","lastName":"Bove","suffix":""},{"id":150552364,"identity":"8fe3d010-e15e-47cf-98a7-b1e2142c3868","order_by":4,"name":"Maria Colomba Comes","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Maria","middleName":"Colomba","lastName":"Comes","suffix":""},{"id":150552366,"identity":"689b55c3-6f0a-4f72-8519-4596529a6089","order_by":5,"name":"Vittorio Didonna","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Vittorio","middleName":"","lastName":"Didonna","suffix":""},{"id":150552367,"identity":"d5b9ea89-f514-4310-be40-a77ec6233366","order_by":6,"name":"Francesco Giotta","email":"","orcid":"","institution":"Unità Operativa Complessa di Oncologia Medica, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Francesco","middleName":"","lastName":"Giotta","suffix":""},{"id":150552368,"identity":"4099377d-718a-4454-8520-75b78b228c48","order_by":7,"name":"Daniele La Forgia","email":"","orcid":"","institution":"I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Daniele","middleName":"La","lastName":"Forgia","suffix":""},{"id":150552369,"identity":"90f708b5-88b5-4d07-8e2c-0ae0cb892e4c","order_by":8,"name":"Agnese Latorre","email":"","orcid":"","institution":"Unità Operativa Complessa di Oncologia Medica, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Agnese","middleName":"","lastName":"Latorre","suffix":""},{"id":150552370,"identity":"934a6abd-f0be-4bc4-990b-9b2f5fab2ac2","order_by":9,"name":"Maria Irene Pastena","email":"","orcid":"","institution":"Unità Operativa Complessa di Anatomia Patologica, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Maria","middleName":"Irene","lastName":"Pastena","suffix":""},{"id":150552371,"identity":"3773e16b-30b3-46c0-bee6-022de1daf2d2","order_by":10,"name":"Nicole Petruzzellis","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nicole","middleName":"","lastName":"Petruzzellis","suffix":""},{"id":150552372,"identity":"f8522b10-205f-4862-9576-4648d2a74968","order_by":11,"name":"Lucia Rinaldi","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Oncologia Per la Presa in Carico Globale del Paziente Oncologico “Don Tonino Bello”, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lucia","middleName":"","lastName":"Rinaldi","suffix":""},{"id":150552373,"identity":"f19d7091-165d-4b34-a59f-c1222f5934ef","order_by":12,"name":"Pasquale Tamborra","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Pasquale","middleName":"","lastName":"Tamborra","suffix":""},{"id":150552374,"identity":"b4a90eae-2117-4967-ba8c-9d62a91fe5b3","order_by":13,"name":"Alfredo Zito","email":"","orcid":"","institution":"Unità Operativa Complessa di Anatomia Patologica, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Alfredo","middleName":"","lastName":"Zito","suffix":""},{"id":150552375,"identity":"d1233d80-69d7-4e4a-92f3-d7bc876c6142","order_by":14,"name":"Vito Lorusso","email":"","orcid":"","institution":"Unità Operativa Complessa di Oncologia Medica, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Vito","middleName":"","lastName":"Lorusso","suffix":""},{"id":150552376,"identity":"d4d78efb-7d2f-4c63-b200-23ce56ebec0a","order_by":15,"name":"Raffaella Massafra","email":"","orcid":"","institution":"Struttura Semplice Dipartimentale di Fisica Sanitaria, I.R.C.C.S. Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Raffaella","middleName":"","lastName":"Massafra","suffix":""}],"badges":[],"createdAt":"2022-11-04 14:59:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2238591/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2238591/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-023-35344-9","type":"published","date":"2023-05-26T20:56:05+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":28925307,"identity":"428c9135-a3ff-4855-82fd-0969e17fe3f8","added_by":"auto","created_at":"2022-11-10 22:30:40","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":55264,"visible":true,"origin":"","legend":"\u003cp\u003eRandom survival forest features importance averaged over 20 rounds of 5-fold cross-validation. Error bars correspond to standard deviation and the red dashed line represents the imposed threshold for the mean feature weight.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-2238591/v1/e973e990b9fa934d83539bc8.png"},{"id":28925308,"identity":"60781a60-e85f-41df-8442-e4e814d8e419","added_by":"auto","created_at":"2022-11-10 22:30:40","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":396060,"visible":true,"origin":"","legend":"\u003cp\u003eDescription of the time behavior of the performance metrics, area under the ROC curve (AUC) and concordance index (c-index). Top panels are referred to the exploitation of the whole features set, while the bottom ones exploit just selected features. Solid lines join the mean points evaluated every 10 months, while shaded regions correspond to standard deviations.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-2238591/v1/aef3431c8aa210dacd05b6a4.png"},{"id":28925100,"identity":"b0fca0dd-8763-45a0-bb07-a5470694c6b1","added_by":"auto","created_at":"2022-11-10 22:22:40","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":101324,"visible":true,"origin":"","legend":"","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-2238591/v1/3bb54c6ba52d180bfb778201.png"},{"id":28925099,"identity":"3cfa548b-559f-4bb4-8620-ad16290ff0dd","added_by":"auto","created_at":"2022-11-10 22:22:40","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":60900,"visible":true,"origin":"","legend":"\u003cp\u003eRepresentation of Pearson correlation statistics. In the left panel the distribution of correlations between the hazards yielded by the best performing CGB survival curves at 10 years and EP scores; In the right panel a bubble plot of scores and hazards in a rescaled version with respect to their maximum values, where each bubble ray measures the standard deviation over the 20 rounds, while the dashed red line slope is equal to the median value in the left panel.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-2238591/v1/c033eb4a5cd9adf4c53fbde5.png"},{"id":44730088,"identity":"f9f8e147-f124-4dc7-b49b-a719d1486344","added_by":"auto","created_at":"2023-10-16 21:26:10","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1008179,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2238591/v1/8a930f8d-5421-4fd8-ba1d-95b5ebb631e1.pdf"},{"id":28925102,"identity":"d933be16-c84f-4765-88f7-2b0e6f5a0c8c","added_by":"auto","created_at":"2022-11-10 22:22:40","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":364598,"visible":true,"origin":"","legend":"","description":"","filename":"SUPPLEMENTARYMATERIALS.docx","url":"https://assets-eu.researchsquare.com/files/rs-2238591/v1/0b82fd2fb2bb91f017ba6e13.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine learning survival models trained on clinical data to identify high risk patients with hormone responsive HER2 negative breast cancer","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eFor breast cancer (BC) patients with endocrine-positive and HER2 negative early-stage disease, the benefit of adding chemotherapy to adjuvant endocrine therapy is controversial. These patients are at risk for being undertreated or overtreated with endocrine therapy and chemotherapy, and tests are required to save an important number of patients from the potentially harmful side effects of chemotherapy; in particular, several studies have showed that a non-negligible proportion of BC patients, especially those with a hormone receptor-positive and lymph node-negative disease, could only be effectively treated with hormone therapy alone [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. The use of adjuvant chemotherapy for estrogen receptor (ER) \u0026ndash; positive, HER2-negative BC patients has been investigated by an impressive number of studies aimed at measuring its efficiency in a predictive manner [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Such studies range from genomic tests [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] to sophisticated artificial intelligence models [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], with the purpose of describing the benefit gained by each patient undergoing a specific therapy. Recent years have witnessed the availability of several molecular tests which have received long-standing recommendations in clinical guidelines [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. In particular, the use of gene signatures has provided a standardized reproducible and quantitative tool able to define the risk of distant recurrent for ER-positive, HER2-negative early BC. Nevertheless, the adoption in the clinical practice of these decisional support tools requires a careful analysis of their cost-effectiveness, because genomics tests have an important cost and not all centers are provided with laboratories performing this type of analyses. This issue is currently driving the studies aimed the achievement of the same information by means of less expensive procedure.\u003c/p\u003e \u003cp\u003eIn general, new interdisciplinary approaches are emerging in survival analysis, which aim to analyze data commonly collected in the clinical practice and drive the therapeutic choices. Indeed, in clinical practice, medical oncologists are increasingly using prediction tools available online, such as PREDICT, Adjuvant!, and CancerMath to guide systemic adjuvant treatment [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The online tools provide personalized 10-year overall survival estimates for the adjuvant treatment setting by basing their predictions on patient data (e.g. age) and tumour characteristics (e.g. size, nodal status, ER-status and grade), but they perform well at the population level, but exhibit a high degree of discordance in the intermediate and poor prognosis groups [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Furthermore, some models have been proposed for the estimation of disease-free survival with classic approaches [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], but works aimed at predicting high-risk patients who might actually benefit from additional chemotherapy to hormone therapy is missing.\u003c/p\u003e \u003cp\u003eA wide variety of techniques is currently available, ranging from classical non-parametric Kaplan-Meier descriptive curves to extensions of the semi-parametric inferential Cox model. A limit of classical algorithms is the difficulty to model high dimensionality. Recently, machine learning techniques applied to survival tasks allow to overcome this issue [\u003cspan additionalcitationids=\"CR14 CR15 CR16\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Indeed, the classical Cox regression is a parametric model based on a probabilistic estimation whose prediction performances depend on parameters associated with each feature. Therefore, the difficulty in identifying an accurate probabilistic model increases in parallel with the increase in the included number of features considered. On the contrary, machine learning survival model, such as random forest and gradient boosting survival models, are non-parametric methods whose performances depend on the size of the training set. In practice, the latter models do not impose any hypothesis on the probabilistic distribution, thus allowing to properly model nonlinearities and interaction effects in a data driven approach [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. If, on the one hand, these limitations are pursued to achieve the explainability of black-box machine learning survival models [\u003cspan additionalcitationids=\"CR20\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], on the other their overcoming guarantees higher performances [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn this work, we propose a model for estimating disease-free survival with respect to invasive events for patients with endocrine-positive and HER2 negative BC, which are potentially candidates for genomic testing, which are potentially candidates for genomic testing, if only hormone therapy is carry out. Our preliminary study is configured to identify low- and high-risk patients and assess the chance of achieving comparable performances with genomic test but exploiting much cheaper and already available clinical data. Three machine learning survival models are compared with the Cox proportional hazards regression according to time-dependent classification performance metrics [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Once the best machine learning survival model was chosen, we evaluated the correlation risk score obtained from our model with that of a genetic test performed on a sample of independent patients.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e"},{"header":"2. Results","content":"\u003cdiv class=\"Section2\" id=\"Sec3\"\u003e\n \u003ch2\u003e2.1. Enrolled Patients and Features\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eOur dataset is composed by clinical and cytohistological outcomes of 145 patients our extracted from our database of approximately 900 patients registered for a first BC diagnosis in the period 1997\u0026ndash;2019 and referred to Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo; in Bari (Italy). The inclusion criteria for collecting such database were: absence of primary chemotherapy for BC, ab initio non-metastatic patient. Then, according to the genomic test eligibility criteria defined by the decree of the Ministry of Health of May 2021 [\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e], i.e. early stage tumor, patients not at high or low risk of recurrence with hormone responsive and HER2 negative BC, 145 patients wase extracted. In line with the aim of our work, we considered the only patients who did not undergo that chemotherapy. In fact, for patients who have undergone chemotherapy, the absence of a second event could be due to a patient-specific positive prognostic profile and not necessarily to an effect of the therapeutic treatment. In other words, there may be patients who would not have relapsed even if they had not undergone additional chemotherapy.\u003c/p\u003e\n \u003cp\u003eIn this work, we specifically focus our attention on breast cancer-related invasive disease events (IDEs), which include local recurrence, the appearance of distant visceral and soft tissue metastases, contralateral invasive breast cancer or a second primary tumor [\u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e\n \u003cp\u003eCollected features were the age at diagnosis, tumor size (diameter: T1a, T1b, T1c, T2, T3, T4), histological subtype (ductal, lobular, other), type of surgery (quadrantectomy/mastectomy), estrogen receptor expression (ER, %), progesterone receptor expression (PgR, %), cellular marker for proliferation (Ki67, %), histological grade (grading, Elston\u0026ndash;Ellis scale: G1, G2, G3), human epidermal growth factor receptor-2 score (HER2/neu: 0\u003csup\u003e+\u003c/sup\u003e, 1\u003csup\u003e+\u003c/sup\u003e, 2\u003csup\u003e+\u003c/sup\u003e), the number of metastatic and eradicated lymph nodes, lymph nodes dissection (no/yes), sentinel lymph node (no, negative, positive), lymph nodes stage (N: 0, 1, 2, 3), in situ component (absent, G1, G2, G3, present but not typed), lymphovascular invasion (absent, focal, extensive, present but not typed), multiplicity (no/yes) and previous tumors (no/yes). The data set is described in Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e. The set of predictive features is composed by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N=18\\)\u003c/span\u003e\u003c/span\u003e prognostic factors, typically considered by clinicians during the first tumor diagnosis and related surgery. The missing data recovery has been implemented by means of the Python package musingly (v. 0.2.0).\u003c/p\u003e\n \u003cp\u003eA separate dataset, composed by 27 patients endowed with EndoPredict\u0026reg; (EP) scores and undergoing surgery during 2021 in our institute, is exploited for further evaluations of our survival estimation. EP is a gene expression test for patients with ER-positive and HER2-negative early-stage BC, both node-negative and node-positive (N0, N1, micrometastasis). It is a second-generation test that combines a molecular score of 12 genes with tumor size and lymph node status [\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e]. This genomic test has entered the clinical practice of the Istituto Tumori \u0026apos;Giovanni Paolo II\u0026apos; in Bari since 2021 and it is used by the Breast Care team when the clinical case is highly doubtful.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable border=\"1\" id=\"Tab1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eObserved patients\u0026rsquo; statistics according to considered features.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eFeatures\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCounts (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eFeatures\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCounts (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eOverall\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e145 (100)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eLymph Nodes Stage\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eType of Surgery\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eN0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e114 (78.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003equadrantectomy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e108 (74.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eN1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e26 (17.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003emastectomy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e37 (25.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eN2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2 (1.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eIn Situ Component\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eN3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eabsent\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e101 (69.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2 (1.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eG1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7 (4.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eLymph Node Dissection\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eG2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8 (5.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eno\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e34 (23.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eG3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4 (2.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eyes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e104 (71.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epresent, not typed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e25 (17.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7 (4.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eHER2/neu\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eSentinel Lymph Node\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0\u003csup\u003e+\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e70 (48.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eno\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e97 (66.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1\u003csup\u003e+\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e56 (38.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003enegative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e34 (23.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2\u003csup\u003e+\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e15 (10.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epositive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8 (5.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4 (2.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6 (4.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMultiplicity\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eGrading\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eno\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e118 (81.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eG1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e20 (13.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eyes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e27 (18.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eG2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e104 (71.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eDiameter\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eG3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e17 (11.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eT1b\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e26 (17.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4 (2.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eT1c\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e68 (46.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eLymphovascular Invasion\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eT2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e42 (29.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eabsent\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e98 (67.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eT3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1 (0.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003efocal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e33 (22.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eT4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e3 (2.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eextensive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4 (2.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNA\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5 (3.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epresent, not typed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e10 (6.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eHistologic Type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrevious Tumors\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eductal\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e107 (73.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eno\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e140 (96.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003elobular\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e20 (13.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eyes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5 (3.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eother\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e18 (12.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMedian [\u003c/strong\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({{q}}_{0},{{q}}_{1},{{q}}_{3},{{q}}_{4}\\)\u003c/span\u003e\u003c/span\u003e\u003cstrong\u003e]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMedian [\u003c/strong\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({{q}}_{0},{{q}}_{1},{{q}}_{3},{{q}}_{4}\\)\u003c/span\u003e\u003c/span\u003e\u003cstrong\u003e]\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eER\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e80 [9, 70, 90, 100]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eAge\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e59 [32, 49, 67, 86]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003ePgR\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e60 [0, 20, 90, 98]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMetastatic Lymph Nodes\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0 [0, 0, 0, 10]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eKi67\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12 [1, 5, 20, 70]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eEradicated Lymph Nodes\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e16 [0, 2, 23, 40]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\u003cbr\u003e\u003c/div\u003e\u0026nbsp;\u003ctable border=\"1\" id=\"Tab2\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eSummary of the sample labels variation in the time period comprised between 20 and 160 months.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"16\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eInvasive disease events\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e17\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e27\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e29\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e30\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e30\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e33\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e35\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e39\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e41\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e43\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e44\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eControl cases\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e136\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e130\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e125\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e121\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e112\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e109\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e107\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e103\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e75\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e64\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e47\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e31\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMonths\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e31\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e60\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e90\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e100\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e110\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e120\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e130\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e140\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e150\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e160\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv class=\"Section2\" id=\"Sec4\"\u003e\n \u003ch2\u003e2.2. Time Dependent Classification\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eRandom survival forest feature importance is resumed in Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e. Their calculation is nested in the 20 rounds of 5-fold cross-validations to avoid any influence imposed by a single evaluation with a fixed training set. To take into account the included statistical variation, each feature weight is described by its average and standard deviation. We select those features characterized by a weight greater than 0.01, thus yielding: Ki67, PgR, age, ER, eradicated lymph nodes and diameter.\u003c/p\u003e\n \u003cp\u003eThe ability of machine learning survival algorithms to model data high dimensionality with respect to the classical CPH regression [\u003cspan class=\"CitationRef\"\u003e18\u003c/span\u003e] is shown by comparing the time behavior of the considered metrics (see Appendix B) when all features are included (N\u0026thinsp;=\u0026thinsp;18) or just the six selected ones. In Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e, the upper panels are related with the first case, while the lower panels to the latter one. At 5 and 10 years, time period usually considered for follow-up in clinical practice, the performances of the CPH model are on average lower than those of machine learning models by at least 10 percentage points, when all the features are considered. However, when just a subset of features is considered, consisting in the most important ones, the average difference in performance is halved, signaling a better condition for the CPH model, while the performance of machine learning survival models tends to remain unchanged with respect to the number of features involved. In order to define a parsimonious model, i.e. to use just the right number of predictors needed to explain the model well, the following analyses will be carried out on the results obtained by considering the selected subset of features.\u003c/p\u003e\n \u003cp\u003eThe sensitivity and specificity balanced performance (see Fig.\u0026nbsp;6 in Appendix C) corresponding to the 5 years time frame are equal to 0.62\u0026ndash;0.65 for the three machine learning survival models, while the much lower one of CPH shows approximately 0.55 for the same balanced metrics pair. If we consider 10 years after the first BC, the CPH model still shows the lowest performance, while the machine learning survival models RSF and GB are characterized by a balanced aforementioned metric pair equal to 0.63\u0026ndash;0.65, while CGB emerges as the best performing classifier with 0.67 for the balanced sensitivity and specificity pair.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eOnce the 10 years time frame is kept fixed, we establish a threshold for each model according to the median of the ones selected by the Youden index optimization (see Fig. 6 in Appendix 6). The average score of each patient over the 20 rounds is then adopted to assign each case to a high or low risk category. These strata are further characterized by means of the Kaplan-Meier curves shown in Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. The p-values confirm that CGB implements the best discrimination between high or low risk patients, while CPH is much less efficient than the machine learning survival models.\u003c/p\u003e\n \u003c/div\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv class=\"Section2\" id=\"Sec5\"\u003e\n \u003ch2\u003e2.3. Correlation with EndoPredict\u0026reg; Scores\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eThe similarity measure of our risk estimation with the one predicted by EP is evaluated over an independent test consisting of 27 patients. The last step of our preliminary study consists in the calculation of the Pearson correlation coefficients between the hazards (see Appendix A) estimated by our best performing CGB survival curves at 10 years after the first BC diagnosis and the risk scores provided to our institution by the EP software exploiting genetic data.\u003c/p\u003e\n \u003cp\u003eA scores statistics for the separate dataset is obtained by testing the sample set of 27 patients on 20 rounds 5-fold cross-validation for the model trained on 145 patients, such that a variation in the training is obtained to gain a wider statistic. The hazard value corresponding to 10 years after the first breast cancer diagnosis is then deduced, whose overall correlation statistics is shown in Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e. The violin plot in the left panel is characterized by a sufficient stability around the median value, imposing the slope of the trend line in the bubble plot of the right panel.\u003c/p\u003e\n \u003cp\u003eThe performance of EP declared by the authors in terms of c-index is equal to 0.753 for the prediction of distant recurrences within 10 years [\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e]. As emerged from our results, CGB shows the highest correlations, equal in average to 0.42, with respect to the remaining survival models, instead resulting in average uncorrelated. We underline that such correlations are time independent by definition for CGB, GB and CPH, because the corresponding hazard functions assume time independent parameters, such that features and time are independent variables.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"3. Discussion","content":"\u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe high recurrence rate characterizing BC patients has prompted to the adoption of post-operative treatments, including adjuvant chemotherapy. At the same time, the risk of overtreatment in this patient population has supported the development of tools able to perform a proper risk-benefit assessment and to guide the \u0026ldquo;decision-making\u0026rdquo; process [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. The role of molecular data has become increasingly important in guiding therapeutic decisions in this setting. Nevertheless, there is the urgent need to explore novel reliable and less expensive prognostic tools. Particular attention deserver hormone-responsive, HER2 negative BC patients for which the prescription of an adjunctive chemotherapy hormone therapy is often highly doubtful. Recently, the genomic tests play a key rule into assess the benefit provided by the addition of chemotherapy, but are very expensive and their cost-effectiveness needs to be neglected. In this paper, we proposed a machine learning approach to estimate disease free survival.\u003c/p\u003e \u003cp\u003eTo date, a plethora of predictive models have been developed to estimate disease-free survival with respect to recurrence breast cancer by solving a classification task [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] or focusing on survival [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. However, it is known that, despite not in common cases, anticancer drugs can cause second tumors, correlated with chemotherapy [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Therefore, recently, in the adjuvant clinical trial setting for breast cancer, experts proposed to adopt only one term, that is Invasive Disease-Free Survival, to refer to composite events, such as local and distant recurrence, contralateral invasive breast cancers, second primary tumors and death [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Recent works on survival model for invasive events prediction and its variants have been freshly proposed [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] and were based on the exploitation of patients\u0026rsquo; characteristics related to demographics, diagnosis, pathology and therapy. Among these, machine learning algorithms represent a novel, promising tool.\u003c/p\u003e \u003cp\u003eThe usefulness of survival analysis inspired by machine learning algorithms is currently assessed by interdisciplinary studies, because of its improved ability to take into account high-dimensional data with respect to classical methods [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Such new approaches include a much higher complexity which require otherwise an accurate feature selection [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOur preliminary study aims to evaluate the potential of machine learning models trained on commonly clinical features for predicting of IDEs for patients with endocrine-positive and HER2 negative BC. This tool, which has been built starting from information frequently collected in clinical practice, could replace the genomic profiling tests, notoriously more expensive in terms of application times and costs, when their application is not available [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In this preliminary work, we have compared three machine learning survival models with the classical approach, i.e. Cox proportional hazards regression, to predict IDES endocrine-positive and HER2 negative BC and, thus, identify low- and high-risk patients.\u003c/p\u003e \u003cp\u003eThe c-index obtained within the same time frame by CGB, GB and RSF is stable with or without feature selection at approximately 0.67\u0026ndash;0.68 in average. Considering that EP test declares a c-index equal to 0.753 [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], the obtained preliminary results including only clinical determinants are encouraging. We subsequently verified on an independent subset of patients who performed EP tests, whether the decision suggested by the model we trained was in agreement with the result of the EP test. Even if the latter takes into account genetic information not included in our survival models, we observe a sufficient similarity of the risk estimation for an invasive disease event corresponding to 10 years after the first BC, as measured by Pearson correlation. Indeed, the proposed model showed a significant agreement with the result of the EP test meaning that clinical features, if properly elaborated, could express part of the information expressed by the genomic test. However, the main advantage is the much cheaper and already available data exploited in our scheme, providing a sufficient information as measured by the correlation similarity.\u003c/p\u003e \u003cp\u003eIf we consider the time period comprised between 5 and 10 years, a stable behavior of the mean performances emerges for the machine learning survival models. Moreover, they shown significantly higher risk estimation performance than the classical Cox model.\u003c/p\u003e \u003cp\u003eIn addition, the machine learning survival models have shown a significant ability to predict the IDEs in both early and late periods (5 and 10 years, respectively), to accurately discriminate patients at low or high risk, and to detect a large risk patient group with positive outcome after 10 year with only 5 years of endocrine therapy\u003c/p\u003e \u003cp\u003eTo the best of our knowledge, studies aimed at developing a predictive model of IDFS for patients with endocrine-positive and HER2 negative BC are lacking. Therefore, we believe that a comparison between our results with those obtained with more generic state-of-the-art models that estimate overall survival or ides but trained on heterogenous populations, can be mistaking. What we consider interesting instead is the comparability of the forecast results of the survival disease carried out with genomic tests, as previously discussed. Another software called Pam50 (Prosigna\u0026reg;) adopting 50 genes and engineered for distant recurrences declares a c-index related with the 3 years time frame equal to 0.72 [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e], a value which is comparable with our performances.\u003c/p\u003e \u003cp\u003eAlthough the general performance does not yet allow a clinical application of the model, the experimental results encourage future developments aimed at introducing features of a different nature. Therefore, our hypothesis is that it is possible to define a machine learning model trained on data commonly collected in clinical practice that could accurately surrogate the genomic test, finally, reducing the cost of healthcare without compromising patient care, and significantly impacting clinical governance.\u003c/p\u003e \u003cp\u003eFurther developments will be focused on the inclusion of radiomic features [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e] as well as some radiologic indices extracted from clinical reports. Indeed, the inclusion of imaging data is investigated in radiogenomics to reduce costs of genetic tests, thus proving an information resource that has to be comprised in a high-dimensional setting.\u003c/p\u003e \u003cp\u003eThe usage of structured electronic health records is limited to relatively small datasets, while recent investigations are exploiting natural language processing to take free-text clinic notes as input [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. Moreover, for the models adopted in survival analysis, much general hypotheses have been formulated to take into account the time dependence of regression coefficients, thus giving up the assumed independence of features in factorized hazard functions [\u003cspan additionalcitationids=\"CR37 CR38\" citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], which is partly observed just for RSF among the implemented machine learning survival models.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e"},{"header":"4. Materials And Methods","content":"\u003cdiv class=\"Section2\" id=\"Sec8\"\u003e\n \u003cp\u003e4.1. Survival Analysis for Risk Estimation\u003c/p\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eIn this work, we propose the application of machine learning survival models to evaluate the risk of an invasive disease event for each patient. Such models are formulated in Appendix A, implemented by using the Python package scikit-survival (v. 0.17.1) and listed as follows:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003eRandom survival forest (RSF);\u003c/li\u003e\n \u003cli\u003eGradient boosting (GB);\u003c/li\u003e\n \u003cli\u003eComponent-wise gradient boosting (CGB).\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/div\u003e\n \u003cp\u003eA comparison of these machine learning methodologies with the well-known Cox proportional hazards (CPH) are performed.\u003cbr\u003e\u003cbr\u003eThe selection of most important features is executed by means of the Python package eli5 (v. 0.11), which provides a way to compute feature importance by measuring how concordance index (c-index) decreases when a feature is not available. In the survival framework, the remotion of the relationship of a certain feature with the survival time is executed by random shuffling of its values: the weight of each feature is quantified by the drop on average of the c-index [\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e]. The performance metrics are estimated in a time dependent approach [\u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e] and they are obtained by adopting both the whole set of features and just those characterized by a sufficient weight in the aforementioned importance evaluation. To understand the variation in time of some metrics exploited to assess the inferential power of a survival model, we have to rephrase it as a classifier yielding a time varying score equal to the complement to one of the survival probability\u003c/p\u003e\n \u003cdiv class=\"Equation\" id=\"Equ1\"\u003e\n \u003cp class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e$${F}_{i}\\left(t\\right)=1-{S}_{i}\\left(t\\right)$$\u003c/p\u003e\n \u003cp class=\"EquationNumber\"\u003e1,\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003ewith \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i=1,\\dots ,M\\)\u003c/span\u003e\u003c/span\u003elabelling each patient in the sample, with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(M\\)\u003c/span\u003e\u003c/span\u003e equal to the total number of patients. The observed event time is\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"Equation\" id=\"Equ2\"\u003e\n \u003cp class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e$${Z}_{i}=\\text{m}\\text{i}\\text{n}\\{{T}_{i},{C}_{i}\\}$$\u003c/p\u003e\u003cp class=\"EquationNumber\"\u003e2,\u003c/p\u003e\u003c/div\u003e\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({T}_{i}\\)\u003c/span\u003e\u003c/span\u003e denotes the time of invasive disease onset and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({C}_{i}\\)\u003c/span\u003e\u003c/span\u003e the censoring time. In this way a time dependent disease status \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({D}_{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e takes value 0 until \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({t\u0026lt;T}_{i}\\)\u003c/span\u003e\u003c/span\u003e, while it shifts to 1 afterwards [\u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e]. The censored patients, not showing any event, are removed when \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({t\u0026gt;C}_{i}\\)\u003c/span\u003e\u003c/span\u003e, as shown in Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e starting from 20 months to avoid too unbalanced data sets.\u003c/p\u003e\u003cp\u003eThe precise formulation of the adopted metrics consisting in the time dependent version of the area under the receiver operating characteristic (ROC) curve (AUC) and c-index is presented in Appendix B. These quantities are able to describe time by time how the survival models capture the patient\u0026rsquo;s status behavior with respect to the used features. The optimization of the available parameters was implemented on ReCaS datacenter [\u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e]. Later on, the analysis considered just the fixed time frame corresponding to 5 and 10 years after the first BC diagnosis. The research of an optimal threshold balancing sensitivity and specificity is based on the Youden index [\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e], whose maximization drives the solution achievement. To measure the efficiency of our scheme in estimating patients risks, the comparison with EP scores associated with a separate set of 27 patients is implemented. Such patients are endowed with the feature subset selected in the described procedure.\u003c/p\u003e\u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e Conceptualization, A.F., D.P. and R.M.; methodology, A.F. and D.P.; software, D.P.; validation, A.F., D.P., V.L. and R.M.; formal analysis, A.F. and D.P.; investigation, A.F., D.P., V.L. and R.M.; resources, V.D., D.L.F., L.R., A.Z. and V.L.; data curation, D.P., S.B. and N.P.; writing\u0026mdash;original draft preparation, A.F., D.P., A.R., S.B., M.C.C. and R.M.; writing\u0026mdash;review and editing, V.D., D.L.F., A.L., F.G., M.I.P., L.R., P.T. and A.Z.; visualization, A.F. and D.P.; supervision, A.L., F.G., M.I.P., L.R., A.Z., V.L. and R.M.; project administration, V.L. and R.M.; funding acquisition, V.L. and R.M. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u0026nbsp;\u003c/strong\u003eThis work was supported by funding from the Italian Ministry of Health \u0026ldquo;Ricerca Finalizzata 2018\u0026rdquo;.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInstitutional Review Board Statement:\u003c/strong\u003e Institutional Review Board Statement: The study received approval from the Scientific Board of Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo;\u0026mdash;Bari, Italy and was carried out in accordance with the Declaration of Helsinki\u0026rsquo;s standards. The authors affiliated to the Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo; RCCS, Bari are responsible for the views expressed in this article, which do not necessarily represent the ones of the Institute.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInformed Consent Statement:\u003c/strong\u003e \u0026lsquo;Informed consent\u0026rsquo; for publication was waived by the Scientific Board of Istituto Tumori \u0026lsquo;Giovanni Paolo II\u0026rsquo;, Bari, for data related to the cohort of patients, as this study is retrospective and involves minimal risk.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement:\u0026nbsp;\u003c/strong\u003eThe raw data supporting the conclusions of this article will be made available by the corresponding author, without undue reservation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest:\u003c/strong\u003e The authors declare no conflict of interest.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003e\u003cspan\u003ePaik, S.; Shak, S.; Tang, G.; et al. A Multigene Assay to Predict Recurrence of Tamoxifen-Treated, Node-Negative Breast Cancer. N. Engl. J. Med. 2004, 351:2817\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSparano J.A.; Gray R.J.; Makower, D.F.; et al. Adjuvant Chemotherapy Guided by a 21-Gene Expression Assay in Breast Cancer. N. Engl. J. Med. 2018, 379(2): 111\u0026ndash;121.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBuus, R.; Sestak, I.; Kronenwett, R.; et al. Molecular Drivers of Oncotype DX, Prosigna, EndoPredict, and the Breast Cancer Index: A TransATAC Study. J. Clin. Oncol. 2021, 39(2): 126\u0026ndash;135.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eFilipits, M.; Rudas, M.; Jakesz, R.; et al. A New Molecular Predictor of Distant Recurrence in ER-Positive, HER2-Negative Breast Cancer Adds Independent Information to Conventional Clinical Risk Factors. Clin. Cancer. Res. 2011, 17(18): 6012\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eEndoPredict Clinical Dossier. Available online: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://myriad.com/managed-care/endopredict-clinical-dossier/\u003c/span\u003e\u003c/span\u003e (accessed on 23 March 2022).\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBanna, G.L.; Gomes, F.; Maltese, G.; et al. An electronic tool for frailty and fitness assessment in the immunotherapy era. Arg. Geriat. Oncol. 2021; 6: 7\u0026ndash;14.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eCardoso, F.; Kyriakides, S.; Ohno, S.; et al. Early breast cancer: ESMO Clinical Practice Guidelines for diagnosis, treatment and follow-up. \u003cem\u003eAnn. Oncol.\u003c/em\u003e 2019, 30(8): 1194\u0026ndash;1220. Erratum in: \u003cem\u003eAnn. Oncol.\u003c/em\u003e \u003cstrong\u003e2019\u003c/strong\u003e, 30(10): 1674. Erratum in: \u003cem\u003eAnn. Oncol.\u003c/em\u003e \u003cstrong\u003e2021\u003c/strong\u003e, 32(2): 284.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eEngelhardt, E. G.; van den Broek, A. J.; Linn, S. C.; Wishart, G. C.; Emiel, J. T.; van de Velde, A. O.; Schmidt, M. K.; et al. Accuracy of the online prognostication tools PREDICT and Adjuvant! for early-stage breast cancer patients younger than 50 years. Euro. J. Canc. 2017, 78, 37\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eLaas, E.; Mallon, P.; Delomenie, M.; Gardeux, V.; Pierga, J. Y.; Cottu, P.; Reyal, F.; et al. Are we able to predict survival in ER-positive HER2-negative breast cancer? A comparison of web-based models. Brit. J. Canc. 2015, 112(5), 912\u0026ndash;917.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eFanizzi, A.; Pomarico, D.; Paradiso, A.V.; Lorusso, V.; Massafra, R.; et al. Predicting of Sentinel Lymph Node Status in Breast Cancer Patients with Clinically Negative Nodes: A Validation Study. Cancers 2021\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eLambertini, M.; Pinto, A.C.; Ameye, L.; et al. The prognostic performance of Adjuvant! Online and Nottingham Prognostic Index in young breast cancer patients. Brit. J. Canc. 2016, 115: 1471\u0026ndash;1478.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eWu, X.; Ye Y.; Barcenas, C.H.; et al. Personalized Prognostic Prediction Models for Breast Cancer Recurrence and Survival Incorporating Multidimensional Data. JNCI J. Nat. Canc. Inst. 2017, 109(7).\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eWang, P.; Li, Y.; Reddy, C.K. Machine Learning for Survival Analysis: A Survey. ACM Comp. Surv. 2019, 51.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eIshwaran, H.; Kogalur, U.B.; Blackstone E.H; et al. Random Survival Forests. Ann. Appl. Stat. 2008, 2(3):841\u0026ndash;860.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHothorn, T.; Buhlmann, P.; Dudoit, S.; et al. Survival ensembles. Biostatistics 2006, 7(3): 355\u0026ndash;373.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eLi, H.; Luan, Y. Boosting proportional hazards models using smoothing splines, with applications to high-dimensional microarray data. Bioinformatics 2005, 21(10): 2403\u0026ndash;2409.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHe, K.; Li, Y.; Zhu, J.; et al. Component-wise gradient boosting and false discovery control in survival analysis with high-dimensional covariates. Bioinformatics 2016, 32(1): 50\u0026ndash;57.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMoncada-Torres, A.; van Maaren, M.C.; Hendriks, M.P.; et al. Explainable machine learning can outperform Cox regression predictions and provide insights in breast cancer survival. Sci. Rep. 2021, 11: 6968.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eKovalev, M.S.; Utkin, L.V.; Kasimov, E.M. SurvLIME: A method for explaining machine learning survival models. Know.-Bas. Sys. 2020, 203: 106164.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eUtkin, L.V.; Satyukov, E.D.; Konstantinov, A.V. SurvNAM: The machine learning survival model explanation. Neu. Net. 2022, 147: 81\u0026ndash;102.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eKuruc, F.; Binder, H.; Hess, M. Stratified neural networks in a time-to-event setting. Brief. Bio. 2022, 23(1): 1\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eKamarudin, A.N.; Cox, T.; Kolamunnage-Dona, R. Time-dependent ROC curve analysis in medical research: current methods and applications. BMC Med. Res. Met. 2017, 17: 53.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eDECRETO 18 maggio 2021 - Gazzetta Ufficiale. Available online: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gazzettaufficiale.it/eli/id/2021/07/07/21A04069/sg\u003c/span\u003e\u003c/span\u003e (accessed on 23 March 2022).\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMassafra, R.; Latorre, A.; Fanizzi, A.; et al. A Clinical Decision Support System for Predicting Invasive Breast Cancer Recurrence: Preliminary Results. Front. Oncol. 2021 11: 1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eTseng, Y.J.; et al. Predicting breast cancer metastasis by using serum biomarkers and clinicopathological data with machine learning technologies. Int. J. Med. Inform. 2019, 128: 79\u0026ndash;86.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eLi, J.; et al. Predicting breast cancer 5-year survival using machine learning: A systematic review. PLoS One 2021, 16: 1\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eZou, L.; Pei, L.; Hu, Y.; Ying, L.; Bei, P. The incidence and risk factors of related lymphedema for breast cancer survivors post operation: a 2 year follow up prospective cohort study. Breast Cancer 2018, 25: 309\u0026ndash;314.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHudis, C.A.; et al. Proposal for standardized definitions for efficacy end points in adjuvant breast cancer trials: The STEEP system. J. Clin. Oncol. 2007, 25, 2127\u0026ndash;2132.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eDemoor-Goldschmidt, C.; De Vathaire, F. Review of risk factors of secondary cancers among cancer survivors. Br. J. Radiol. 2019, 92: 1\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eFu, B.; et al. Predicting Invasive Disease-Free Survival for Early Stage Breast Cancer Patients Using Follow-Up Clinical Data. 2019, 66: 2053\u0026ndash;2064.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eGnant, M.; Filipits, M.; Greil, R.; et al. Predicting distant recurrence in receptor-positive breast cancer patients with limited clinicopathological risk: using the PAM50 Risk of Recurrence score in 1478 postmenopausal patients of the ABCSG-8 trial treated with adjuvant endocrine therapy alone. Ann. Onc. 2014, 25: 339\u0026ndash;345.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eBai, H.X.; Lee, A.M.; Yang, L.; et al. Imaging genomics in cancer research: limitations and promises. Br. J. Radiol. 2016, 89: 20151030.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eGrimm, L.J.; Mazurowski M.A. Breast Cancer Radiogenomics: Current Status and Future Directions. Acad. Radiol. 2020, 27(1): 39\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eWang, H.; Li, Y.; Khan, A.S.; Luo, Y. Prediction of breast cancer distant recurrence using natural language processing and knowledge-guided convolutional neural network. Art. Int. Med. 2020, 110:101977.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eSanyal, J.; Tariq, A.; Kurian, A.W.; Rubin, D.; Banerjee, I. Weakly supervised temporal model for prediction of breast cancer distant recurrence. Sci. Rep. 2021, 11:9461.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMurphy, S.A.; Sen, P.K. Time-dependent coefficients in a Cox-type regression model. \u003cem\u003eStoch. Proc. Appl.\u003c/em\u003e 1991, 39:153\u0026ndash;180.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eMurphy, S.A. Testing for a Time Dependent Coefficient in Cox\u0026apos;s Regression Model. Scand. J. Stat. 1993, 20:35\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eZhang, Z.; Reinikainen, J.; Adeleke, K.A.; et al. Time-varying covariates and coefficients in Cox regression models. Ann. Transl. Med. 2018, 6(7):121.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eThomas, L.; Reyes, E.M. Tutorial: Survival Estimation for Cox Regression Models with Time-Varying Coefficients Using SAS and R. J. Stat. Soft. 2014, 61:1\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eHarrell, F.E.; Califf, R.M.; Pryor, D.B.; Lee, K.L.; Rosati, R.A. Evaluating the yield of medical tests, J. Am. Med. Ass. 1982, 247: 2543\u0026ndash;2546.\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eReCaS Bari. Available online: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.recas-bari.it/index.php/en/\u003c/span\u003e\u003c/span\u003e (accessed on 24 March 2022).\u003c/span\u003e\u003c/li\u003e\n \u003cli\u003e\u003cspan\u003eCox, D.R. Regression Models and Life-Tables. J. Roy. Stat. Soc. 1972, 34: 187\u0026ndash;220.\u003c/span\u003e\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Machine learning Survival, Breast Cancer, her2 negative patients, hormone responsive patients, decision support system, ","lastPublishedDoi":"10.21203/rs.3.rs-2238591/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2238591/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eFor endocrine-positive Her2 negative breast cancer patients at an early stage, the benefit of adding chemotherapy to adjuvant endocrine therapy is controversial. Several genomic tests are available on the market but are very expensive. Therefore, there is the urgent need to explore novel reliable and less expensive prognostic tools in this setting. In this paper, we shown a machine learning survival model to estimate Invasive Disease-Free Events trained on clinical and histological data commonly collected in clinical practice. We collected clinical and cytohistological outcomes of 145 patients referred to Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo;. Three machine learning survival models are compared with the Cox proportional hazards regression according to time-dependent performance metrics evaluated in cross-validation. The c-index at 10 years obtained by random survival forest, gradient boosting, and component-wise gradient boosting is stabled with or without feature selection at approximately 0.68 in average respect to 0.57 obtained to Cox model. Moreover, machine learning survival models have accurately discriminated low- and high-risk patients, and so a large group which can be spared additional chemotherapy to hormone therapy. The preliminary results obtained by including only clinical determinants are encouraging. The integrated use of data already collected in clinical practice for routine diagnostic investigations, if properly analyzed, can reduce time and costs of the genomic tests.\u003c/p\u003e","manuscriptTitle":"Machine learning survival models trained on clinical data to identify high risk patients with hormone responsive HER2 negative breast cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-11-10 22:22:35","doi":"10.21203/rs.3.rs-2238591/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-04-06T05:49:29+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-04-05T07:46:28+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-03-07T14:53:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"766a4cf5-e585-4c30-8523-22a074c4e118","date":"2023-02-24T16:42:45+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-02-15T17:19:51+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"d16cda43-36d7-4ed4-8d73-c85216755ce3","date":"2023-02-09T17:17:05+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-02-07T19:55:40+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-11-20T14:50:19+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2022-11-08T20:51:45+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-11-08T20:37:02+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2022-11-04T14:45:58+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0023fa72-33a4-482d-8bea-3bc5b8666282","owner":[],"postedDate":"November 10th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":16823145,"name":"Biological sciences/Cancer"},{"id":16823146,"name":"Health sciences/Oncology"},{"id":16823147,"name":"Health sciences/Risk factors"}],"tags":[],"updatedAt":"2023-10-16T21:10:13+00:00","versionOfRecord":{"articleIdentity":"rs-2238591","link":"https://doi.org/10.1038/s41598-023-35344-9","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2023-05-26 20:56:05","publishedOnDateReadable":"May 26th, 2023"},"versionCreatedAt":"2022-11-10 22:22:35","video":"","vorDoi":"10.1038/s41598-023-35344-9","vorDoiUrl":"https://doi.org/10.1038/s41598-023-35344-9","workflowStages":[]},"version":"v1","identity":"rs-2238591","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2238591","identity":"rs-2238591","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.