Machine Learning for Optimal Individual Survival Prediction in Resectable Upper Gastrointestinal Cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning for Optimal Individual Survival Prediction in Resectable Upper Gastrointestinal Cancer Jin-On Jung, Nerma Crnovrsanin, Naita Maren Wirsik, Henrik Nienhüser, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1318132/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract Purpose Surgical oncologists are frequently confronted with the question of expected long-term prognosis. The aim of this study was to apply machine learning algorithms to optimize survival prediction after oncological resection of gastroesophageal cancers. Methods Eligible patients underwent oncological resection of gastric or distal esophageal cancer between 2001 and 2020 at Heidelberg University Hospital, Department of General Surgery. Machine learning methods such as multi-task logistic regression and survival forests were compared with usual algorithms to establish an individual estimation. Results The study included 117 variables with a total of 1,360 patients. The overall missingness was 1.3%. Out of eight machine learning algorithms, the random survival forest (RSF) performed best with a concordance-index of 0.736 and an integrated Brier score of 0.166. The RSF demonstrated a mean area under the curve (AUC) of 0.814 over a time period of 10 years after diagnosis. The most important long-term outcome predictor was lymph node ratio with a mean AUC of 0.730. A numeric risk score was calculated by the RSF for each patient and three risk groups were defined accordingly. Median survival time was 18.8 months in the high-risk group, 44.6 months in the medium-risk group and above 10 years in the low-risk group. Conclusion The results of this study suggest that RSF is most appropriate to accurately answer the question of long-term prognosis. Furthermore, we could establish a compact risk score model with 20 input parameters and thus provide a clinical tool to improve prediction of oncological outcome after upper gastrointestinal surgery. gastric cancer esophageal cancer machine learning survival analysis oncological outcome Figures Figure 1 Figure 2 Figure 3 Introduction The demand for adequate prediction of oncological outcome after upper gastrointestinal surgery will increase due to a rising global incidence of esophageal and gastric carcinomas (Malhotra et al. 2017 ). In this regard, patient profiles are heterogenous due to varying in-hospital courses and additional information about clinical and histopathological characteristics which make the individual prognosis difficult to predict. While most patient characteristics are hardly improvable and are rather predetermined by the nature of the disease, there might be critical periods during the treatment of a malignancy such as the phase of oncological resection when many fundamental conditions for long-term outcome are set. This way, the in-hospital stay for surgery is literally an incisive event in the clinical course of one individual patient. However, surgical morbidity can significantly influence the overall and long-term survival in gastric cancer (Kulig et al. 2021 ). Eventually, clinicians are confronted with a patient curious to know what the information gathered so far from surgery actually means for his or her individual prognosis. While the patient is then frequently referred to the oncologist who will further supervise the prospective course, there may be a more concrete estimation to offer at the time of discharge. Surgeons with an oncologic focus need the tools to provide these answers to adequately care for their patients. The motivation to accurately predict the individual prognosis of one patient also derives from the growing interest to assess individualized patient profiles in times of personalized medicine (Hu and Steingrimsson 2018 ). Due to the complexity of each patient case, differences in individual prognosis could eventually implicate different (adjuvant) therapeutic regimens. In this context, the vast amount of clinical data should be viewed as another component of “omics”-based technology and as another step towards personalized treatment. Machine learning techniques can help to determine significant associations based on complex and vast data (Kourou et al. 2015 ). Retrospective machine learning studies have already delivered suggestive data that artificial intelligence may be promising in case of multiple potential factors with unclear relations to each other. Since the proposal of the proportional hazards model by David Cox in 1972 (Cox 1972 ) there were many advances to improve survival prediction, especially since computational possibilities have increased. Among those, machine learning algorithms have recently shown promising results (Zhu et al. 2020 ). According to the current literature there are many works regarding the automated prediction of time-dependent, censored data with machine learning methods. Besides survival being the most typical and classic scenario, there have also been approaches to predict heart failure (Panahiazar et al. 2015 ), kidney transplant durability (Sekercioglu et al. 2021 ) and many other settings. Spooner et al. were able to demonstrate that machine learning algorithms for survival analysis were applicable to the clinical task of predicting dementia (Spooner et al. 2020 ). Based on two separate datasets, it was possible to reach a concordance index (c-index) of 0.82 and 0.93, respectively. However, the prediction of oncological outcome after diagnosis of a malignant disease is still the most common target variable in machine learning studies. To this date, there are only few medical publications in the current literature dealing with machine learning models for survival analysis of oncologically resected upper gastrointestinal cancers. Akcay et al. have analyzed gastric cancer patients after chemoradiation with a comparable methodology (Akcay et al. 2020 ) and tested a random forest algorithm to predict tumor relapse in terms of distant metastases or peritoneal recurrence. Compared to other algorithms such as logistic regression, multilayer perceptron and extreme gradient boosting the random forest scored highest with an area under the curve (AUC) of 0.97 for peritoneal recurrence. However, the random forest algorithm did not show convincing results for overall survival with an AUC of only 0.59 and the authors did not apply a random survival forest methodology. Jiang et al. have utilized a LASSO cox analysis to predict oncological outcome of gastric cancer patients analyzing the radiomic signature of PET computer tomography (Jiang et al. 2018 ). It is notable that the authors evaluated imaging data and reached a powerful prediction for disease-free survival and overall survival with a c-index of 0.786. From a methodological point of view, Pölsterl et al. have found that feature extraction in the setting of survival prediction is not properly applicable in case of small sample sizes such as approximately 500 patients (Pölsterl et al. 2016 ). On the other hand, in case of larger sample sizes with more than 2500 individuals feature extraction performs similarly to feature selection methods. Due to this reason the effect of feature extraction is supposedly maximized in-between the mentioned case numbers. To our knowledge, the topic of survival prediction in case of oncologically resected upper gastrointestinal cancer is not sufficiently covered by machine learning methods. However, the opportunity of pre- and postoperatively supervising an individual patient course offers an ideal setting to predict long-term outcome based on the collected data. The aim of this study was therefore to apply various machine learning algorithms to optimize survival prediction after oncological resection of gastroesophageal cancer compared to previously established methods. Material And Methods Data collection and follow-up Eligible subjects of this study were patients with distal esophageal, gastroesophageal junction or gastric cancer who underwent oncological resection between September 2001 and December 2020 at Heidelberg University Hospital, Department of General Surgery. All patients provided written consent for data collection and analysis. The data was initially collected prospectively in a clinical database. The trial protocol was approved by the ethics committee at the University of Heidelberg (committee’s approval: S-635/2013) and was performed in accordance with the Declaration of Helsinki, Good Clinical Practices as well as local ethics and legal requirements. To acquire long-term survival data, all patients were systematically followed-up via continuous surveys. The relevant survival data for this patient collective was gathered until October 2020. Inclusion and exclusion criteria Out of the main collective, only those patients were further taken into account with a preoperatively diagnosed and histologically proven adenocarcinoma of the distal esophagus, the esophagogastric junction or stomach. Patients were excluded who had squamous cell carcinoma at any location. Also, only radical oncological resections were evaluated as opposed to exploratory laparotomies with palliative treatment. Furthermore, the type of oncological resection was limited to (sub-)total gastrectomy, gastrectomy with transhiatal extension, proximal gastrectomy, abdominothoracic esophagectomy and discontinuity resection of the esophagus. Variables with a missingness greater than 50% were generally excluded from further analysis. Also, patients were excluded with a case-specific missingness of more than 10%. Data structure Eventually, the dataset had a total of 117 features and 1,360 patients. Out of the whole dataset, 92 variables were categorical (with 77 Boolean variables) and 25 variables were numeric. Table 1 gives a short overview of the parameters including all dependent variables (or end points, respectively). There were obvious survival predictors and other variables which had to be excluded due to collinearity (see Supplementary Figure 1). Regarding survival data, 53.5% of the records were censored with an equal distribution over the whole timespan. Supplementary Figure 2 demonstrates the censored data added on top of the patients that verifiably passed away and plotted against the follow-up period. Missingness and imputation After the application of the above-mentioned criteria, there was a remaining missingness in 2,033 datapoints and thus an overall missingness of 1.3% without any duplicate values (see Supplementary Figure 3). For imputation, missingness at random was assumed and multiple imputations with n = 1000 iterations via IterativeImputer from the Sci-kit learn package (Pedregosa et al. 2011) were performed on Python 3.9 (Van Rossum G and Drake FL 2009). Statistical analysis All statistical analyses were performed on Python 3.9 with packages such as scikit-learn 0.24.2 by Pedregosa et al. (Pedregosa et al. 2011). To compare various machine learning algorithms the package PySurvival by Fotso et al. (Stephane Fotso 2019) was implemented. The performances of the different algorithms were scored according to the c-index (Uno et al. 2011). The inverse probability of censoring weights (IPCW c-index) was considered as an alternative to the standard c-index which is independent of the distribution of censored cases in the test data. This was relevant due to the comparably high prevalence of censored time points in the present dataset (see Supplementary Figure 2). To establish another performance score, the integrated Brier-score (further abbreviated as IBS) was utilized (Steyerberg et al. 2010). The IBS was eventually plotted to demonstrate prediction error with a cut-off limit of 0.25 considered as critical. Furthermore, the actual survival function was plotted against the predicted function by the model. Finally, time-dependent evaluation of the area under the curve (AUC) was performed for selected machine learning ensembles provided by the scikit-survival package version 0.15.1 (Pölsterl 2020). Feature selection We did not perform feature selection before application of algorithms since previous works by Spooner et al. have shown that feature selection in this scenario does not significantly improve model performance (Spooner et al. 2020). The authors have applied seven different feature selection algorithms before running 5-fold cross validation with 5 repeats. The results did not show any significant improvement of test statistics and most machine learning algorithms have a feature selection method internalized already. However, we identifiedal the 20 most relevant predictors after fitting the corresponding machine learning algorithm via permutation-based importance evaluation provided by ELI5 (Arya et al. 2020). Machine learning algorithms The classic Cox proportional hazards model introduced by David Cox in 1972 is probably the most popular survival prediction model with an easy to interpret statistic (Cox 1972). We also tested a non-linear Cox proportional hazards model (also called DeepSurv) which is a multi-layer perceptron (Katzman et al. 2018). However, the authors’ work is also known for the focus on recommender functions and treatment decision making. The Linear Multi-Task Logistic Regression (L-MTLR) was introduced by Yu et al. in 2011 and operates on the basis of multiple logistic regressions (Yu et al. 2011). It may be considered as another general alternative to the Cox regression model. To increase flexibility, the Neural Multi-Task Logistic Regression (N-MTLR) was introduced in 2019 (Stephane Fotso 2019). To fully represent all available survival prediction models, we also included one parametric model, the Gompertz model, as the relatively strongest test statistic compared to the Weibull and Exponential model which are not demonstrated in this work. Last but not least, several random survival forest models were analyzed including the classic Random Survival Forest (RSF) by Ishwaran et al. (Ishwaran et al. 2008). The Conditional Survival Forest model was developed by Wright et al. to improve splitting of the RSF by applying maximally selected rank statistics for the split point selection (Wright et al. 2017). The Extra Trees incorporated in PySurvival are an extension of the Extremely Randomized Trees (Geurts et al. 2006) which is a supervised learning method with almost totally randomized decision trees. The model is eventually independent of the output values of the learning sample and impresses with its high computational efficiency. Hyperparameter optimization We performed hyperparameter optimization through defining the closest neighbor parameter values for each algorithm. Thereafter, the optimization was executed via Halving Grid Search included in the latest version of scikit-learn (Pedregosa et al. 2011). Halving Grid Search operates via successive halving of the candidate hyperparameters and their effect on the final test statistic as measured by the c-index. Although Halving Grid Search was not yet applied in many scientific works, it has already been shown that it yields equivalent results for hyperparameter tuning compared to usual brute-force Grid Searching, however, with a much higher computation efficiency (Sraitih et al. 2021). Results The study included 1,360 patients with oncological resection between 2001 and 2020 and 117 variables. Table 2 summarizes all patient characteristics which were mainly evaluated as potentially relevant predictors for the machine learning models except for obvious outcome parameters. First of all, we tested the introduced machine learning algorithms which are demonstrated with the according test statistic in Figure 1 . The standard Cox proportional hazards (CPH) model is demonstrated in Figure 1a and reached a c-index of 0.645 and an integrated Brier score (or IBS) of 0.221. The prediction error curve is depicted in the last column of Figure 1 with the values for the Brier score as an integral. The non-linear Cox proportional hazards model (see Figure 1b ) was able to improve the statistic with a c-index of 0.681 and an IBS of 0.194. Note that while both Cox proportional hazards models initially have a comparably good score, the integral reaches the critical limit of 0.25 as the timespan reaches 10 years after cancer diagnosis. Compared to the first two calculations, the c-index could not be improved by linear multi-task logistic regression (see Figure 1c , c-index = 0.673, IBS = 0.229). However, the neural multi-task logistic regression reached a clearly better prediction of the actual survival function with a root mean squared error (RMSE) of the actual predicted survival curve of 9.188 (see Figure 1d , c-index = 0.672, IBS = 0.254). Note that both regression methods also exceed the IBS cut-off value of 0.25 which is generally considered as an acceptable limit. The Gompertz model as an example for a parametric model was not able to significantly improve the test statistic and showed relatively weak results (see Figure 1e , c-index = 0.677, IBS = 0.194). However, the CPH models could be outperformed by all three survival forest methods with the RSF being the strongest prediction model (see Figure 1h , c-index = 0.736, IBS = 0.166). The Extra Survival Trees (see Figure 1g , c-index = 0.736, IBS = 0.167) and the Conditional Survival Trees (see Figure 1f , c-index = 0.726, IBS = 0.166) showed similarly strong results which could still not outperform the RSF even after hyperparameter optimization. The accurate prediction by RSF is again underlined by the direct comparison between actual and predicted survival function demonstrated in the second column. Here, the root-mean-square error (RMSE) of the RSF was calculated as 6.224. To establish a more practical and compact approach, we selected the most important features identified by the RSF. Table 3 shows the permutation-based importance of the 20 most important parameters identified by the RSF model. While the individual weights of the predictors are not remarkably high (except for lymph node ratio), the RSF model can still build the prediction based on all variables of the dataset. The weights of the predictors are to be interpreted in a fashion that for instance the omittance of lymph node ratio in the model would evoke a change of the resulting c-index by 0.118 with the specified 95% confidence range. The time-dependent area under the curve (AUC) is separately demonstrated for the six most important predictors. It is remarkable that the AUC factors such as duration of intensive care, postoperative complications and intraoperative blood loss lose their predictive value rapidly as time passes. In contrast, the significance of lymph node ratio increases postoperatively and stays stable on an AUC level above 0.65 more than 5 years after cancer diagnosis (see Figure 2a ). While predictors have a time-dependent importance, the prediction models also show a time dependent accuracy. Of note, the RSF algorithm outperforms the CPH model on the time-dependent scale with an AUC of 0.821 (as opposed to 0.720 for the CPH model, see Figure 2b ). The remaining machine learning models also showed better time-dependent performances than CPH but were not able to outperform the RSF algorithm. While all models were able to predict survival with a very high score above 0.9 in the very first months, the predictions generally tended to decrease in accuracy while time progressed. However, the long-term survival prediction was most successful according to the RSF model. Figure 3a shows a risk scoring model based on the test statistic calculated by the RSF algorithm. A numeric risk score is assigned to each patient ranging from 4.7 to 7.1. Three different colors were utilized to depict the low-, medium- and high-risk group. The differentiation of three groups was performed manually based on the distribution of risk scores. Figure 3b shows the survival curves of all individuals classified by the scoring system as low-, medium- or high-risk within the same predefined test group. The three survival curves differ significantly according to the log-rank test (p < 0.0001) with the low-risk group having a 5-year survival rate of 73.18%, the medium-group showing 45.39% and the high-risk group finally 14.87%. Median survival time was 18.754 months in the high-risk group, 44.557 months in the medium-risk group and incalculable in the low-risk group since the survival rate did not fall below 50% in the observed long-term time interval of 10 years. Also, the survival curves can be separated into more groups to enable a further stratification of risk groups (not shown). Finally, the 20 most relevant predictors from the permutation-based importance scoring (see Table 3) were selected to establish a more compact RSF model. Supplementary Figure 4 shows the distribution of the risk scores resulting from the compact RSF model. A risk score between 0 and 5.3 was considered as low-risk, a score between 5.3 and 6.1 as medium-risk and a score above 6.1 as high-risk for decease after resection. The according survival curves of the three different groups from the test cohort are demonstrated in Supplementary Figure 5 . The low-risk group had an incalculable median survival longer than 10 years, the medium-risk group had a median survival of 85.639 months and the high-risk group 20.721 months. Discussion To improve the prediction of overall survival after gastroesophageal cancer resections we tested eight machine learning algorithms in this retrospective survival analysis to determine the best prediction model for oncological outcome. The Cox proportional hazards (CPH) model is by far the most frequently applied model to evaluate significant survival predictors with coefficient metrics that are easy to interpret. However, the CPH model relies on a manual selection of the most important features with potential omittance of important features. Therefore, the CPH model may not be able to handle complex and vast data with multidimensional features. Machine learning methods have shown since their application in econometrics and biosciences that it is especially useful for prediction problems (Jordan and Mitchell 2015 ). Using the c-index as a measure of performance, the random survival forest (RSF) model achieved the most accurate prediction with a c-index of 0.736. This value is generally considered in the setting of time-dependent censored data such as long-term survival as a good to strong prediction model. Moreover, we could show that the RSF algorithm was able to outperform the CPH model as well as all other machine learning models. In our setting, the RSF is robust which is proofed by 10-fold cross-validation. We therefore believe that the RSF model is the most suitable machine learning algorithm to predict oncological outcome after curative resection of gastroesophageal adenocarcinoma. The strength as well as the limitation of this study is the availability of detailed and high-resolution data for each patient. On the one hand, it was possible to establish an excellent prediction model under these circumstances. However, it might not be feasible in every setting to extract all presented data for each patient who is treated in a surgical hospital. To establish a more practical approach, we selected the 20 most important features identified by the RSF. By reducing the necessary input parameters from 117 to 20 and thus avoiding any tedious or unnecessary data management, it may be possible to make oncological predictions more practicable for the surgical oncologist. The resulting test statistic of this compact RSF model performs in a comparable strength and could be applied in the future. In both cases, extended or compact RSF, the test statistic outperformed the Cox proportional hazards model. It is also imaginable to establish an online application similar to the surgical risk calculator by the American College of Surgeons (Bilimoria et al. 2013 ) enabling clinicians to enter the 20 variables anonymously since no identifying information is needed to handle the algorithm. The application can then generate an individual assessment based on the information that is present at the time of discharge. Given the anonymous information, it would also be possible to show specific survival curves according to the trained model and the risk group that the patient belongs to. This is statistically legitimate since the plotted survival curve represents the survival function for all patients with the same risk score range over the time span of 10 years. The application of an individual and numeric risk score summarizes all available clinical and histopathological data that is predictive for long-term oncological outcome. Future clinical trials could rely on this risk score to find differences for overall survival which is the actual outcome parameter of primary interest. Thus, it is imaginable that subgroups could be treated with or without adjuvant therapy not according to single features such as TNM-status but based on the affiliation to the risk groups presented in this study. Likewise, high-risk patients with worst prognosis could be analyzed separately to eventually achieve improvements in surveillance and treatment. In this context, it is urgent to identify and characterize high-risk patients with unfavorable histopathological and postoperative parameters who still show a comparably good survival. To this date, it is not properly understood if the mere fact that the tumor was resected may nevertheless have an impact on overall survival for this specific patient collective and it is debatable if immunological reactions play a role. It is also necessary to implement a generally accepted risk score such as the presented machine learning model regardless of resectional status to offer a prognostic tool also for palliative situations. This study is outstanding due to a systematic follow-up, high data quality and a thorough, state-of-the-art imputation with application of documented Python packages. All reported machine learning models are retrievable and reproducible via open source and have been cited accordingly. However, the clinical relevance of the reported RSF model as well as the other machine learning models needs to be clarified in prospective analyses with inevitably necessary external validation. Furthermore, the value differences between the several survival forest models are limited to some hundredths of the c-index and the discussion of these differences may be of limited clinical relevance. Nevertheless, we could show in a supervised learning setting that the algorithms were able to automatically select and extract predictors that are verifiably relevant for long-term prognosis. Declarations Funding The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Competing Interests The authors have no relevant financial or non-financial interests to disclose. Author Contributions All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Jin-On Jung, Nerma Crnovrsanin, Naita Maren Wirsik, Henrik Nienhüser, Leila Peters and Thomas Schmidt. The first draft of the manuscript was written by Jin-On Jung and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript. Ethics Approval All procedures followed were in accordance with the ethical standards of the responsible committee on human experimentation (institutional and national) and with the Helsinki Declaration of 1964 and later versions. The trial protocol was approved by the ethics committee at the University of Heidelberg (committee’s approval: S-635/2013). Consenct to participate Informed consent was obtained from all individual participants included in the study. References Akcay M, Etiz D, Celik O, 2020. Prediction of Survival and Recurrence Patterns by Machine Learning in Gastric Cancer Cases Undergoing Radiation Therapy and Chemotherapy. Adv. Radiat. Oncol. 5, 1179–1187. https://doi.org/10.1016/j.adro.2020.07.007 Arya V, Bellamy RKE, Chen P-Y, Dhurandhar A, Hind M, Hoffman SC, Houde S, Liao QV, Luss R, Mourad S, Pedemonte P, Raghavendra R, Richards JT, Sattigeri P, Shanmugam K, Singh M, Varshney KR, Wei D, Zhang Y, 2020. AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models. J. Mach. Learn. Res. 21, 1–6. Bilimoria KY, Liu Y, Paruch JL, Zhou L, Kmiecik TE, Ko CY, Cohen ME, 2013. Development and evaluation of the universal ACS NSQIP surgical risk calculator: A decision aid and informed consent tool for patients and surgeons. J. Am. Coll. Surg. 217, 833-842.e3. https://doi.org/10.1016/j.jamcollsurg.2013.07.385 Cox DR, 1972. Regression Models and Life-Tables. J. R. Stat. Soc. Ser. B 34, 187–202. https://doi.org/10.1111/j.2517-6161.1972.tb00899.x Geurts P, Ernst D, Wehenkel L, 2006. Extremely randomized trees. Mach. Learn. 63, 3–42. https://doi.org/10.1007/s10994-006-6226-1 Hu C, Steingrimsson JA, 2018. Personalized Risk Prediction in Clinical Oncology Research: Applications and Practical Issues Using Survival Trees and Random Forests. J. Biopharm. Stat. 28, 333–349. https://doi.org/10.1080/10543406.2017.1377730 Ishwaran H, Kogalur UB, Blackstone EH, Lauer MS, 2008. Random survival forests. Ann. Appl. Stat. 2, 841–860. https://doi.org/10.1214/08-AOAS169 Jiang Y, Yuan Q, Lv W, Xi S, Huang W, Sun Z, Chen H, Zhao L, Liu W, Hu Y, Lu L, Ma J, Li T, Yu J, Wang Q, Li G, 2018. Radiomic signature of 18F fluorodeoxyglucose PET/CT for prediction of gastric cancer survival and chemotherapeutic benefits. Theranostics 8, 5915–5928. https://doi.org/10.7150/thno.28018 Jordan MI, Mitchell TM, 2015. Machine learning: Trends, perspectives, and prospects. Science (80-. ). https://doi.org/10.1126/science.aaa8415 Katzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y, 2018. DeepSurv: Personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med. Res. Methodol. 18. https://doi.org/10.1186/s12874-018-0482-1 Kourou K, Exarchos TP, Exarchos KP, Karamouzis M V., Fotiadis DI, 2015. Machine learning applications in cancer prognosis and prediction. Comput. Struct. Biotechnol. J. 13, 8–17. https://doi.org/10.1016/J.CSBJ.2014.11.005 Kulig P, Nowakowski P, Sierzȩga M, Pach R, Majewska O, Markiewicz A, Kołodziejczyk P, Kulig J, Richter P, 2021. Analysis of prognostic factors affecting short-term and long-term outcomes of gastric cancer resection. Anticancer Res. https://doi.org/10.21873/anticanres.15140 Malhotra GK, Yanala U, Ravipati A, Follet M, Vijayakumar M, Are C, 2017. Global trends in esophageal cancer. J. Surg. Oncol. 115, 564–579. https://doi.org/10.1002/jso.24592 Panahiazar M, Taslimitehrani V, Pereira N, Pathak J, 2015. Using EHRs and Machine Learning for Heart Failure Survival Analysis, in: Studies in Health Technology and Informatics. IOS Press, pp. 40–44. https://doi.org/10.3233/978-1-61499-564-7-40 Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay É, Pedregosa, F. and Varoquaux G, Gramfort, A. and Michel V, Thirion, B. and Grisel O, Blondel, M. and Prettenhofer P, Weiss, R. and Dubourg V, Vanderplas, J. and Passos A, Cournapeau, D. and Brucher M, Perrot M, Duchesnay E, 2011. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 12, 2825–2830. Pölsterl S, 2020. Scikit-survival: A library for time-to-event analysis built on top of scikit-learn. J. Mach. Learn. Res. 21, 1–6. Pölsterl S, Conjeti S, Navab N, Katouzian A, 2016. Survival analysis for high-dimensional, heterogeneous medical data: Exploring feature extraction as an alternative to feature selection. Artif. Intell. Med. 72, 1–11. https://doi.org/10.1016/j.artmed.2016.07.004 Sekercioglu N, Fu R, Kim SJ, Mitsakakis N, 2021. Machine learning for predicting long-term kidney allograft survival: a scoping review. Ir. J. Med. Sci. https://doi.org/10.1007/s11845-020-02332-1 Spooner A, Chen E, Sowmya A, Sachdev P, Kochan NA, Trollor J, Brodaty H, 2020. A comparison of machine learning methods for survival analysis of high-dimensional clinical data for dementia prediction. Sci. Rep. 10, 20410. https://doi.org/10.1038/s41598-020-77220-w Sraitih M, Jabrane Y, El Hassani AH, 2021. An automated system for ECG arrhythmia detection using machine learning techniques. J. Clin. Med. 10, 5450. https://doi.org/10.3390/jcm10225450 Stephane Fotso, 2019. PySurvival: Open source package for Survival Analysis modeling. https://www.pysurvival.io. Steyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW, 2010. Assessing the performance of prediction models: A framework for traditional and novel measures. Epidemiology. https://doi.org/10.1097/EDE.0b013e3181c30fb2 Uno H, Cai T, Pencina MJ, D’Agostino RB, Wei LJ, 2011. On the C-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Stat. Med. 30, 1105–1117. https://doi.org/10.1002/sim.4154 Van Rossum G, Drake FL, 2009. Python 3 Reference Manual. Scotts Val. CA Creat. Wright MN, Dankowski T, Ziegler A, 2017. Unbiased split variable selection for random survival forests using maximally selected rank statistics. Stat. Med. 36, 1272–1284. https://doi.org/10.1002/sim.7212 Yu CN, Greiner R, Lin HC, Baracos V, 2011. Learning patient-specific cancer survival distributions as a sequence of dependent regressors, in: Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011, NIPS 2011. Zhu W, Xie L, Han J, Guo X, 2020. The application of deep learning in cancer prognosis prediction. Cancers (Basel). https://doi.org/10.3390/cancers12030603 Tables Table 1 Overview of independent and outcome variables. Biometric variables height, weight, body mass index (BMI), age, sex. Preoperative variables past medical history (cardiovascular, pulmonary, metabolic and renal preconditions), tumor diagnosis, preceding malignant disease, cTNM classification, histology (Laurén type, signet cell component), neoadjuvant therapy (components, radiotherapy, completeness). (Intra-)operative variables time between diagnosis and resection, type of operation, extent of resection, anatomical reconstruction, duration of surgery, intraoperative complication, blood loss and transfusion. Postoperative variables days on ICU and ward, postoperative complications (according to Clavien-Dindo and additional 38 binarily classified types), pTNM classification, lymph node ratio (positive lymph nodes divided by resected), grading, R-status, histology (Laurén type, signet cell component, tumor regression), post-discharge problems. Outcome variables vital status, overall survival, no evidence of disease, time until tumor relapse, 30-day mortality and in-hospital mortality. Table 2 Overview of selected patient characteristics and imputation methods. ASA = American Society of Anesthesiology, ICU = intensive care unit. n = 1,360 Mean / Median / Frequency, (95% confidence interval) Missingness Biometric variables Weight 77.7 kg (54.0 - 104.0) 0.0% Height 1.72 m (1.58 - 1.87) 1.0% Age 62.35 years (41.0 - 80.0) 0.0% Sex male (71.4%), female (28.6%) 0.0% Preoperative variables ASA score 1 (1.7%), 2 (47.9%), 3 (46.9%), 4 (2.3%) 1.2% Past medical history see Supplementary Table 1 max. 0.2% cT status 1 (7.1%), 2 (20.1%), 3 (57.2%), 4 (10.9%) 4.7% cN status 0 (35.7%), 1 (61.5%), + (0.4%) 2.4% cM status 0 (87.7%), 1 (11.9%) 0.4% Neoadjuvant therapy yes (55.0%), no (45.0%) 0.0% - Chemotherapeutics see Supplementary Table 1 max. 2.1% - Radiation yes (3.5%), no (96.2%) 0.4% (Intra-)operative variables Time diagnosis to resection 82.84 days (11.0 - 168.1) 0.1% One vs. two cavity surgery one (71.1%), two (28.9%) 0.0% Operation type Subtotal gastrectomy (21.5%) Total gastrectomy (26.4%) Transhiatal extended gastrectomy (21.8%) Ivor-Lewis esophagectomy (27.8%) Other types (2.6%) 0.0% (Locally) extended resection yes (28.5%), no (71.2%) 0.3% Intraoperative complications yes (7.8%), no (92.1%) 0.1% Duration of surgery 275.29 minutes (150.0 - 479.3) 0.9% Intraoperative blood loss 622.93 milliliters (100.0 - 1,500.0) 19.2% Postoperative variables Duration of stay on ICU 6.89 days (0.0 - 29.0) 4.7% Duration of stay on ward 20.26 days (9.0 - 49.6) 0.7% Postoperative complications see Supplementary Table 1 max. 0.2% pT status 0/1 (22.3%), 2 (14.2%), 3 (47.3%), 4 (16.2%) 0.0% pN status 0 (42.1%), 1 (16.3%), 2 (15.0%), 3 (26.5%) 0.1% Lymph node ratio 18.25% (0.0 - 74.1) 0.1% pM status 0 (89.7%), 1 (10.3%) 0.1% R status 0 (81.5%), any 1/2/X (18.5%) 0.0% Outcome variables Status dead vs. alive/censored alive (51.5%), dead (46.1%) 2.4% Documented survival time 39.69 months (2.89 - 109.88) 1.0% Table 3 Most important 20 predictors based on permutation-based importance scoring for random survival forest. ICU = intensive care unit, ASA = American Society of Anesthesiology. Predictive Feature Weight Lower 95% CI Upper 95% CI Lymph node ratio (positive / total) 0.1183 0.1016 0.1350 Intraoperative blood loss 0.0060 -0.0009 0.0129 (y)pT4 status 0.0057 0.0002 0.0112 Age (at diagnosis) 0.0049 0.0035 0.0063 (y)pT3 status 0.0042 -0.0004 0.0088 Intraoperative peritoneal carcinosis 0.0040 0.0033 0.0047 Postoperative sepsis 0.0038 -0.0003 0.0079 cM+ status 0.0036 -0.0009 0.0081 Duration of in-hospital stay 0.0033 0.0028 0.0038 Duration of stay on ICU 0.0033 0.0004 0.0062 Major postoperative complications 0.0032 0.0020 0.0044 Any R+ status 0.0026 -0.0024 0.0076 Local R1 status 0.0021 0.0004 0.0038 (y)pM1 status 0.0020 0.0007 0.0033 (y)pT1 status 0.0014 0.0011 0.0017 Intraoperative blood transfusion 0.0013 0.0007 0.0019 Severe pre-existing diseases 0.0006 0.0002 0.0010 Siewert type I junction cancer 0.0006 -0.0004 0.0016 Cardiac postoperative adverse events 0.0005 0.0001 0.0009 ASA grade 3 0.0005 0.0000 0.0010 Supplementary Files suppfigures.docx supptables.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 24 Feb, 2022 Reviewers invited by journal 21 Feb, 2022 Editor invited by journal 03 Feb, 2022 Editor assigned by journal 03 Feb, 2022 First submitted to journal 01 Feb, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1318132","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":85541597,"identity":"49c623af-ea1f-4fdb-8684-884c78bac1fe","order_by":0,"name":"Jin-On Jung","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1ElEQVRIiWNgGAWjYBACxmYwxQbEzAeAhIQMKVrYEkBaeEixkMcATBJUx9zOe/BxAQNf4nb+NZ9f3aix4GFgP3x0A36H8SUbz2BgS9w54+0265xjQIfxpKXdwK+Fx0yaB6hlw42z24xz2IBaJHjMiNVy5plxzj+StJzvYX6c20acFmNjHgM24w032MyYc/skeNgI+cWw/4zhY56KY7Ibzh9+/DnnW50cP/vhY/i1NIBIg2PASExgkwCx2fApBwF5CFXDwMB/gPkDIdWjYBSMglEwMgEARztBHlcnhGEAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-6121-0114","institution":"University Hospital Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jin-On","middleName":"","lastName":"Jung","suffix":""},{"id":85541598,"identity":"a153059c-72f6-4787-8c8c-2aec1a85443e","order_by":1,"name":"Nerma Crnovrsanin","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nerma","middleName":"","lastName":"Crnovrsanin","suffix":""},{"id":85541599,"identity":"b9e5f4ce-96ba-4513-b9c0-cb7d95eac2f8","order_by":2,"name":"Naita Maren Wirsik","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Naita","middleName":"Maren","lastName":"Wirsik","suffix":""},{"id":85541600,"identity":"3cd04edf-a659-410f-aea8-e1b276bf0196","order_by":3,"name":"Henrik Nienhüser","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Henrik","middleName":"","lastName":"Nienhüser","suffix":""},{"id":85541601,"identity":"6020405b-16f3-47e2-b717-f6a24d5d5ab1","order_by":4,"name":"Leila Peters","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Leila","middleName":"","lastName":"Peters","suffix":""},{"id":85541602,"identity":"0cf3ec8e-a814-49f5-b60b-f0d0d4d6bba0","order_by":5,"name":"Felix Popp","email":"","orcid":"","institution":"University Hospital Cologne: Uniklinik Koln","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Felix","middleName":"","lastName":"Popp","suffix":""},{"id":85541603,"identity":"b12160a4-ca1b-4fc3-9a13-55bea9be398b","order_by":6,"name":"André Schulze","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"André","middleName":"","lastName":"Schulze","suffix":""},{"id":85541604,"identity":"574c261a-b0f8-45dc-84d8-d019d376c92d","order_by":7,"name":"Martin Wagner","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Martin","middleName":"","lastName":"Wagner","suffix":""},{"id":85541605,"identity":"3e19b603-5f2d-4722-ae06-5c73f2af7891","order_by":8,"name":"Beat Peter Müller-Stich","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Beat","middleName":"Peter","lastName":"Müller-Stich","suffix":""},{"id":85541606,"identity":"224c7f82-65a3-4ee1-802a-4720e4ab6c17","order_by":9,"name":"Markus Wolfgang Büchler","email":"","orcid":"","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Markus","middleName":"Wolfgang","lastName":"Büchler","suffix":""},{"id":85541607,"identity":"c60e92c5-404a-40ee-b192-b5aa28a95549","order_by":10,"name":"Thomas Schmidt","email":"","orcid":"https://orcid.org/0000-0002-7166-3675","institution":"UniversitätsKlinikum Heidelberg: UniversitatsKlinikum Heidelberg","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Thomas","middleName":"","lastName":"Schmidt","suffix":""}],"badges":[],"createdAt":"2022-02-01 15:01:40","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1318132/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1318132/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":18543202,"identity":"62830101-3430-41e9-bc95-55340e27a1c1","added_by":"auto","created_at":"2022-02-23 20:42:07","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":924277,"visible":true,"origin":"","legend":"\u003cp\u003ePlease See image above for figure legend.\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-1318132/v1/fd051b16f721c670c9ddf32b.jpg"},{"id":18542978,"identity":"2fcddbf3-56f4-42fc-949b-5f97f43f30a5","added_by":"auto","created_at":"2022-02-23 20:39:07","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":390884,"visible":true,"origin":"","legend":"\u003cp\u003ePlease See image above for figure legend.\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-1318132/v1/f658746ca75e0736d441ea1c.jpg"},{"id":18543429,"identity":"e5681390-0d7e-4abd-971a-095f311d4a3e","added_by":"auto","created_at":"2022-02-23 20:45:07","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":318906,"visible":true,"origin":"","legend":"\u003cp\u003ePlease See image above for figure legend.\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-1318132/v1/716f141b95e7f7e11a701418.jpg"},{"id":18543430,"identity":"cacb9c70-663b-4188-ab40-ffd5485feb33","added_by":"auto","created_at":"2022-02-23 20:45:10","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":746594,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1318132/v1/64cddd50-95a8-4ad2-b642-c764e638ea5b.pdf"},{"id":18542982,"identity":"f975ed14-266e-4148-aaa2-37c09a1f3565","added_by":"auto","created_at":"2022-02-23 20:39:07","extension":"docx","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":1712660,"visible":true,"origin":"","legend":"","description":"","filename":"suppfigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-1318132/v1/6c902ec83877b762ef7042dc.docx"},{"id":18543203,"identity":"d73dc9a2-a087-47cb-9b36-91cee2a51486","added_by":"auto","created_at":"2022-02-23 20:42:07","extension":"docx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":23600,"visible":true,"origin":"","legend":"","description":"","filename":"supptables.docx","url":"https://assets-eu.researchsquare.com/files/rs-1318132/v1/3c30cf4553126a31742cb9e0.docx"}],"financialInterests":"","formattedTitle":"Machine Learning for Optimal Individual Survival Prediction in Resectable Upper Gastrointestinal Cancer","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe demand for adequate prediction of oncological outcome after upper gastrointestinal surgery will increase due to a rising global incidence of esophageal and gastric carcinomas (Malhotra et al. \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). In this regard, patient profiles are heterogenous due to varying in-hospital courses and additional information about clinical and histopathological characteristics which make the individual prognosis difficult to predict. While most patient characteristics are hardly improvable and are rather predetermined by the nature of the disease, there might be critical periods during the treatment of a malignancy such as the phase of oncological resection when many fundamental conditions for long-term outcome are set. This way, the in-hospital stay for surgery is literally an incisive event in the clinical course of one individual patient. However, surgical morbidity can significantly influence the overall and long-term survival in gastric cancer (Kulig et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Eventually, clinicians are confronted with a patient curious to know what the information gathered so far from surgery actually means for his or her individual prognosis. While the patient is then frequently referred to the oncologist who will further supervise the prospective course, there may be a more concrete estimation to offer at the time of discharge. Surgeons with an oncologic focus need the tools to provide these answers to adequately care for their patients.\u003c/p\u003e \u003cp\u003eThe motivation to accurately predict the individual prognosis of one patient also derives from the growing interest to assess individualized patient profiles in times of personalized medicine (Hu and Steingrimsson \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Due to the complexity of each patient case, differences in individual prognosis could eventually implicate different (adjuvant) therapeutic regimens. In this context, the vast amount of clinical data should be viewed as another component of \u0026ldquo;omics\u0026rdquo;-based technology and as another step towards personalized treatment.\u003c/p\u003e \u003cp\u003eMachine learning techniques can help to determine significant associations based on complex and vast data (Kourou et al. \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). Retrospective machine learning studies have already delivered suggestive data that artificial intelligence may be promising in case of multiple potential factors with unclear relations to each other. Since the proposal of the proportional hazards model by David Cox in 1972 (Cox \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e1972\u003c/span\u003e) there were many advances to improve survival prediction, especially since computational possibilities have increased. Among those, machine learning algorithms have recently shown promising results (Zhu et al. \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). According to the current literature there are many works regarding the automated prediction of time-dependent, censored data with machine learning methods. Besides survival being the most typical and classic scenario, there have also been approaches to predict heart failure (Panahiazar et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2015\u003c/span\u003e), kidney transplant durability (Sekercioglu et al. \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) and many other settings. Spooner et al. were able to demonstrate that machine learning algorithms for survival analysis were applicable to the clinical task of predicting dementia (Spooner et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Based on two separate datasets, it was possible to reach a concordance index (c-index) of 0.82 and 0.93, respectively.\u003c/p\u003e \u003cp\u003eHowever, the prediction of oncological outcome after diagnosis of a malignant disease is still the most common target variable in machine learning studies. To this date, there are only few medical publications in the current literature dealing with machine learning models for survival analysis of oncologically resected upper gastrointestinal cancers. Akcay et al. have analyzed gastric cancer patients after chemoradiation with a comparable methodology (Akcay et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) and tested a random forest algorithm to predict tumor relapse in terms of distant metastases or peritoneal recurrence. Compared to other algorithms such as logistic regression, multilayer perceptron and extreme gradient boosting the random forest scored highest with an area under the curve (AUC) of 0.97 for peritoneal recurrence. However, the random forest algorithm did not show convincing results for overall survival with an AUC of only 0.59 and the authors did not apply a random survival forest methodology. Jiang et al. have utilized a LASSO cox analysis to predict oncological outcome of gastric cancer patients analyzing the radiomic signature of PET computer tomography (Jiang et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). It is notable that the authors evaluated imaging data and reached a powerful prediction for disease-free survival and overall survival with a c-index of 0.786. From a methodological point of view, P\u0026ouml;lsterl et al. have found that feature extraction in the setting of survival prediction is not properly applicable in case of small sample sizes such as approximately 500 patients (P\u0026ouml;lsterl et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). On the other hand, in case of larger sample sizes with more than 2500 individuals feature extraction performs similarly to feature selection methods. Due to this reason the effect of feature extraction is supposedly maximized in-between the mentioned case numbers.\u003c/p\u003e \u003cp\u003eTo our knowledge, the topic of survival prediction in case of oncologically resected upper gastrointestinal cancer is not sufficiently covered by machine learning methods. However, the opportunity of pre- and postoperatively supervising an individual patient course offers an ideal setting to predict long-term outcome based on the collected data. The aim of this study was therefore to apply various machine learning algorithms to optimize survival prediction after oncological resection of gastroesophageal cancer compared to previously established methods.\u003c/p\u003e"},{"header":"Material And Methods","content":"\u003ch2\u003e\u003cem\u003e\u003cu\u003eData collection and\u0026nbsp;\u003c/u\u003e\u003c/em\u003e\u003cem\u003e\u003cu\u003efollow-up\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eEligible subjects of this study were patients with distal esophageal, gastroesophageal junction or gastric cancer who underwent oncological resection between September 2001 and December 2020 at Heidelberg University Hospital, Department of General Surgery. All patients provided written consent for data collection and analysis. The data was initially collected prospectively in a clinical database. The trial protocol was approved by the ethics committee at the University of Heidelberg (committee\u0026rsquo;s approval:\u0026nbsp;S-635/2013) and was performed in accordance with the Declaration of Helsinki, Good Clinical Practices as well as local ethics and legal requirements. To acquire long-term survival data, all patients were systematically followed-up via continuous surveys. The relevant survival data for this patient collective was gathered until October 2020.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eInclusion and exclusion criteria\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eOut of the main collective, only those patients were further taken into account with a preoperatively diagnosed and histologically proven adenocarcinoma of the distal esophagus, the esophagogastric junction or stomach. Patients were excluded who had squamous cell carcinoma at any location. Also, only radical oncological resections were evaluated as opposed to exploratory laparotomies with palliative treatment. Furthermore, the type of oncological resection was limited to (sub-)total gastrectomy, gastrectomy with transhiatal extension, proximal gastrectomy, abdominothoracic esophagectomy and discontinuity resection of the esophagus. Variables with a missingness greater than 50% were generally excluded from further analysis. Also, patients were excluded with a case-specific missingness of more than 10%.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eData structure\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eEventually, the dataset had a total of 117 features and 1,360 patients. Out of the whole dataset, 92 variables were categorical (with 77 Boolean variables) and 25 variables were numeric. Table 1 gives a short overview of the parameters including all dependent variables (or end points, respectively). There were obvious survival predictors and other variables which had to be excluded due to collinearity (see Supplementary Figure 1).\u0026nbsp;Regarding survival data, 53.5% of the records were censored with an equal distribution over the whole timespan. Supplementary Figure 2\u0026nbsp;demonstrates the censored data added on top of the patients that verifiably passed away and plotted against the follow-up period.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eMissingness and imputation\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eAfter the application of the above-mentioned criteria, there was a remaining missingness in 2,033 datapoints and thus an overall missingness of 1.3% without any duplicate values (see Supplementary Figure 3). For imputation, missingness at random was assumed and multiple imputations with n = 1000 iterations via IterativeImputer from the Sci-kit learn package (Pedregosa et al. 2011) were performed on Python 3.9 (Van Rossum G and Drake FL 2009).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eStatistical analysis\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eAll statistical analyses were performed on Python 3.9 with packages such as scikit-learn 0.24.2 by Pedregosa et al. (Pedregosa et al. 2011). To compare various machine learning algorithms the package PySurvival by Fotso et al. (Stephane Fotso 2019) was implemented. The performances of the different algorithms were scored according to the c-index (Uno et al. 2011). The inverse probability of censoring weights (IPCW c-index) was considered as an alternative to the standard c-index which is independent of the distribution of censored cases in the test data. This was relevant due to the comparably high prevalence of censored time points in the present dataset (see Supplementary Figure 2). To establish another performance score, the integrated Brier-score (further abbreviated as IBS) was utilized (Steyerberg et al. 2010). The IBS was eventually plotted to demonstrate prediction error with a cut-off limit of 0.25 considered as critical. Furthermore, the actual survival function was plotted against the predicted function by the model. \u0026nbsp;Finally, time-dependent evaluation of the area under the curve (AUC) was performed for selected machine learning ensembles provided by the scikit-survival package version 0.15.1 (P\u0026ouml;lsterl 2020).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eFeature selection\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eWe did not perform feature selection before application of algorithms since previous works by Spooner et al. have shown that feature selection in this scenario does not significantly improve model performance (Spooner et al. 2020). The authors have applied seven different feature selection algorithms before running 5-fold cross validation with 5 repeats. The results did not show any significant improvement of test statistics and most machine learning algorithms have a feature selection method internalized already. However, we identifiedal the 20 most relevant predictors after fitting the corresponding machine learning algorithm via permutation-based importance evaluation provided by ELI5 (Arya et al. 2020).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eMachine learning algorithms\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eThe classic Cox proportional hazards model introduced by David Cox in 1972 is probably the most popular survival prediction model with an easy to interpret statistic (Cox 1972). We also tested a non-linear Cox proportional hazards model (also called DeepSurv) which is a multi-layer perceptron (Katzman et al. 2018). However, the authors\u0026rsquo; work is also known for the focus on recommender functions and treatment decision making. The Linear Multi-Task Logistic Regression (L-MTLR) was introduced by Yu et al. in 2011 and operates on the basis of multiple logistic regressions (Yu et al. 2011). It may be considered as another general alternative to the Cox regression model. To increase flexibility, the Neural Multi-Task Logistic Regression (N-MTLR) was introduced in 2019 (Stephane Fotso 2019). To fully represent all available survival prediction models, we also included one parametric model, the Gompertz model, as the relatively strongest test statistic compared to the Weibull and Exponential model which are not demonstrated in this work. Last but not least, several random survival forest models were analyzed including the classic Random Survival Forest (RSF) by Ishwaran et al. (Ishwaran et al. 2008). The Conditional Survival Forest model was developed by Wright et al. to improve splitting of the RSF by applying maximally selected rank statistics for the split point selection (Wright et al. 2017). The Extra Trees incorporated in PySurvival are an extension of the Extremely Randomized Trees (Geurts et al. 2006) which is a supervised learning method with almost totally randomized decision trees. The model is eventually independent of the output values of the learning sample and impresses with its high computational efficiency.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003e\u003cu\u003eHyperparameter optimization\u003c/u\u003e\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eWe performed hyperparameter optimization through defining the closest neighbor parameter values for each algorithm. Thereafter, the optimization was executed via Halving Grid Search included in the latest version of scikit-learn (Pedregosa et al. 2011). Halving Grid Search operates via successive halving of the candidate hyperparameters and their effect on the final test statistic as measured by the c-index. Although Halving Grid Search was not yet applied in many scientific works, it has already been shown that it yields equivalent results for hyperparameter tuning compared to usual brute-force Grid Searching, however, with a much higher computation efficiency (Sraitih et al. 2021).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eThe study included 1,360 patients with oncological resection between 2001 and 2020 and 117 variables. \u003cstrong\u003eTable 2\u003c/strong\u003e summarizes all patient characteristics which were mainly evaluated as potentially relevant predictors for the machine learning models except for obvious outcome parameters.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFirst of all, we tested the introduced machine learning algorithms which are demonstrated with the according test statistic in \u003cstrong\u003eFigure 1\u003c/strong\u003e. The standard Cox proportional hazards (CPH) model is demonstrated in \u003cstrong\u003eFigure 1a\u003c/strong\u003e and reached a c-index of 0.645 and an integrated Brier score (or IBS) of 0.221. The prediction error curve is depicted in the last column of \u003cstrong\u003eFigure 1\u003c/strong\u003e with the values for the Brier score as an integral. The non-linear Cox proportional hazards model (see \u003cstrong\u003eFigure 1b\u003c/strong\u003e) was able to improve the statistic with a c-index of 0.681 and an IBS of 0.194. Note that while both Cox proportional hazards models initially have a comparably good score, the integral reaches the critical limit of 0.25 as the timespan reaches 10 years after cancer diagnosis.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCompared to the first two calculations, the c-index could not be improved by linear multi-task logistic regression (see \u003cstrong\u003eFigure 1c\u003c/strong\u003e, c-index = 0.673, IBS = 0.229). However, the neural multi-task logistic regression reached a clearly better prediction of the actual survival function with a root mean squared error (RMSE) of the actual predicted survival curve of 9.188 (see \u003cstrong\u003eFigure 1d\u003c/strong\u003e, c-index = 0.672, IBS = 0.254). Note that both regression methods also exceed the IBS cut-off value of 0.25 which is generally considered as an acceptable limit. The Gompertz model as an example for a parametric model was not able to significantly improve the test statistic and showed relatively weak results (see \u003cstrong\u003eFigure 1e\u003c/strong\u003e, c-index = 0.677, IBS = 0.194).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHowever, the CPH models could be outperformed by all three survival forest methods with the RSF being the strongest prediction model (see \u003cstrong\u003eFigure 1h\u003c/strong\u003e, c-index = 0.736, IBS = 0.166). The Extra Survival Trees (see \u003cstrong\u003eFigure 1g\u003c/strong\u003e, c-index = 0.736, IBS = 0.167) and the Conditional Survival Trees (see \u003cstrong\u003eFigure 1f\u003c/strong\u003e, c-index = 0.726, IBS = 0.166) showed similarly strong results which could still not outperform the RSF even after hyperparameter optimization. The accurate prediction by RSF is again underlined by the direct comparison between actual and predicted survival function demonstrated in the second column. Here, the root-mean-square error (RMSE) of the RSF was calculated as 6.224.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo establish a more practical and compact approach, we selected the most important features identified by the RSF. \u003cstrong\u003eTable 3\u003c/strong\u003e shows the permutation-based importance of the 20 most important parameters identified by the RSF model. While the individual weights of the predictors are not remarkably high (except for lymph node ratio), the RSF model can still build the prediction based on all variables of the dataset. The weights of the predictors are to be interpreted in a fashion that for instance the omittance of lymph node ratio in the model would evoke a change of the resulting c-index by 0.118 with the specified 95% confidence range.\u003c/p\u003e\n\u003cp\u003eThe time-dependent area under the curve (AUC) is separately demonstrated for the six most important predictors. It is remarkable that the AUC factors such as duration of intensive care, postoperative complications and intraoperative blood loss lose their predictive value rapidly as time passes. In contrast, the significance of lymph node ratio increases postoperatively and stays stable on an AUC level above 0.65 more than 5 years after cancer diagnosis (see \u003cstrong\u003eFigure 2a\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWhile predictors have a time-dependent importance, the prediction models also show a time dependent accuracy. Of note, the RSF algorithm outperforms the CPH model on the time-dependent scale with an AUC of 0.821 (as opposed to 0.720 for the CPH model, see\u003cstrong\u003e\u0026nbsp;Figure 2b\u003c/strong\u003e). The remaining machine learning models also showed better time-dependent performances than CPH but were not able to outperform the RSF algorithm. While all models were able to predict survival with a very high score above 0.9 in the very first months, the predictions generally tended to decrease in accuracy while time progressed. However, the long-term survival prediction was most successful according to the RSF model.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 3a\u0026nbsp;\u003c/strong\u003eshows a risk scoring model based on the test statistic calculated by the RSF algorithm. A numeric risk score is assigned to each patient ranging from 4.7 to 7.1. Three different colors were utilized to depict the low-, medium- and high-risk group. The differentiation of three groups was performed manually based on the distribution of risk scores. \u003cstrong\u003eFigure 3b\u003c/strong\u003e shows the survival curves of all individuals classified by the scoring system as low-, medium- or high-risk within the same predefined test group. The three survival curves differ significantly according to the log-rank test (p \u0026lt; 0.0001) with the low-risk group having a 5-year survival rate of 73.18%, the medium-group showing 45.39% and the high-risk group finally 14.87%. Median survival time was 18.754 months in the high-risk group, 44.557 months in the medium-risk group and incalculable in the low-risk group since the survival rate did not fall below 50% in the observed long-term time interval of 10 years. Also, the survival curves can be separated into more groups to enable a further stratification of risk groups (not shown).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFinally, the 20 most relevant predictors from the permutation-based importance scoring (see Table 3) were selected to establish a more compact RSF model. \u003cstrong\u003eSupplementary Figure 4\u003c/strong\u003e shows the distribution of the risk scores resulting from the compact RSF model. A risk score between 0 and 5.3 was considered as low-risk, a score between 5.3 and 6.1 as medium-risk and a score above 6.1 as high-risk for decease after resection. The according survival curves of the three different groups from the test cohort are demonstrated in\u0026nbsp;\u003cstrong\u003eSupplementary Figure 5\u003c/strong\u003e. The low-risk group had an incalculable median survival longer than 10 years, the medium-risk group had a median survival of 85.639 months and the high-risk group 20.721 months.\u003cbr\u003e\u0026nbsp;\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eTo improve the prediction of overall survival after gastroesophageal cancer resections we tested eight machine learning algorithms in this retrospective survival analysis to determine the best prediction model for oncological outcome. The Cox proportional hazards (CPH) model is by far the most frequently applied model to evaluate significant survival predictors with coefficient metrics that are easy to interpret. However, the CPH model relies on a manual selection of the most important features with potential omittance of important features. Therefore, the CPH model may not be able to handle complex and vast data with multidimensional features. Machine learning methods have shown since their application in econometrics and biosciences that it is especially useful for prediction problems (Jordan and Mitchell \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). Using the c-index as a measure of performance, the random survival forest (RSF) model achieved the most accurate prediction with a c-index of 0.736. This value is generally considered in the setting of time-dependent censored data such as long-term survival as a good to strong prediction model. Moreover, we could show that the RSF algorithm was able to outperform the CPH model as well as all other machine learning models. In our setting, the RSF is robust which is proofed by 10-fold cross-validation. We therefore believe that the RSF model is the most suitable machine learning algorithm to predict oncological outcome after curative resection of gastroesophageal adenocarcinoma.\u003c/p\u003e \u003cp\u003eThe strength as well as the limitation of this study is the availability of detailed and high-resolution data for each patient. On the one hand, it was possible to establish an excellent prediction model under these circumstances. However, it might not be feasible in every setting to extract all presented data for each patient who is treated in a surgical hospital. To establish a more practical approach, we selected the 20 most important features identified by the RSF. By reducing the necessary input parameters from 117 to 20 and thus avoiding any tedious or unnecessary data management, it may be possible to make oncological predictions more practicable for the surgical oncologist. The resulting test statistic of this compact RSF model performs in a comparable strength and could be applied in the future. In both cases, extended or compact RSF, the test statistic outperformed the Cox proportional hazards model. It is also imaginable to establish an online application similar to the surgical risk calculator by the American College of Surgeons (Bilimoria et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2013\u003c/span\u003e) enabling clinicians to enter the 20 variables anonymously since no identifying information is needed to handle the algorithm. The application can then generate an individual assessment based on the information that is present at the time of discharge. Given the anonymous information, it would also be possible to show specific survival curves according to the trained model and the risk group that the patient belongs to. This is statistically legitimate since the plotted survival curve represents the survival function for all patients with the same risk score range over the time span of 10 years.\u003c/p\u003e \u003cp\u003eThe application of an individual and numeric risk score summarizes all available clinical and histopathological data that is predictive for long-term oncological outcome. Future clinical trials could rely on this risk score to find differences for overall survival which is the actual outcome parameter of primary interest. Thus, it is imaginable that subgroups could be treated with or without adjuvant therapy not according to single features such as TNM-status but based on the affiliation to the risk groups presented in this study. Likewise, high-risk patients with worst prognosis could be analyzed separately to eventually achieve improvements in surveillance and treatment. In this context, it is urgent to identify and characterize high-risk patients with unfavorable histopathological and postoperative parameters who still show a comparably good survival. To this date, it is not properly understood if the mere fact that the tumor was resected may nevertheless have an impact on overall survival for this specific patient collective and it is debatable if immunological reactions play a role. It is also necessary to implement a generally accepted risk score such as the presented machine learning model regardless of resectional status to offer a prognostic tool also for palliative situations.\u003c/p\u003e \u003cp\u003eThis study is outstanding due to a systematic follow-up, high data quality and a thorough, state-of-the-art imputation with application of documented Python packages. All reported machine learning models are retrievable and reproducible via open source and have been cited accordingly. However, the clinical relevance of the reported RSF model as well as the other machine learning models needs to be clarified in prospective analyses with inevitably necessary external validation. Furthermore, the value differences between the several survival forest models are limited to some hundredths of the c-index and the discussion of these differences may be of limited clinical relevance. Nevertheless, we could show in a supervised learning setting that the algorithms were able to automatically select and extract predictors that are verifiably relevant for long-term prognosis.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003e\u003cem\u003eFunding\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eThe authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003eCompeting Interests\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eThe authors have no relevant financial or non-financial interests to disclose.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003eAuthor Contributions\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eAll authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Jin-On Jung, Nerma Crnovrsanin, Naita Maren Wirsik, Henrik Nienh\u0026uuml;ser, Leila Peters and Thomas Schmidt. The first draft of the manuscript was written by Jin-On Jung and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003eEthics Approval\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eAll procedures followed were in accordance with the ethical standards of the responsible committee on human experimentation (institutional and national) and with the Helsinki Declaration of 1964 and later versions. The trial protocol was approved by the ethics committee at the University of Heidelberg (committee\u0026rsquo;s approval:\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eS-635/2013).\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003e\u003cem\u003eConsenct to participate\u003c/em\u003e\u003c/h2\u003e\n\u003cp\u003eInformed consent was obtained from all individual participants included in the study.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eAkcay M, Etiz D, Celik O, 2020. Prediction of Survival and Recurrence Patterns by Machine Learning in Gastric Cancer Cases Undergoing Radiation Therapy and Chemotherapy. Adv. Radiat. Oncol. 5, 1179\u0026ndash;1187. https://doi.org/10.1016/j.adro.2020.07.007\u003c/li\u003e\n \u003cli\u003eArya V, Bellamy RKE, Chen P-Y, Dhurandhar A, Hind M, Hoffman SC, Houde S, Liao QV, Luss R, Mourad S, Pedemonte P, Raghavendra R, Richards JT, Sattigeri P, Shanmugam K, Singh M, Varshney KR, Wei D, Zhang Y, 2020. AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models. J. Mach. Learn. Res. 21, 1\u0026ndash;6.\u003c/li\u003e\n \u003cli\u003eBilimoria KY, Liu Y, Paruch JL, Zhou L, Kmiecik TE, Ko CY, Cohen ME, 2013. Development and evaluation of the universal ACS NSQIP surgical risk calculator: A decision aid and informed consent tool for patients and surgeons. J. Am. Coll. Surg. 217, 833-842.e3. https://doi.org/10.1016/j.jamcollsurg.2013.07.385\u003c/li\u003e\n \u003cli\u003eCox DR, 1972. Regression Models and Life-Tables. J. R. Stat. Soc. Ser. B 34, 187\u0026ndash;202. https://doi.org/10.1111/j.2517-6161.1972.tb00899.x\u003c/li\u003e\n \u003cli\u003eGeurts P, Ernst D, Wehenkel L, 2006. Extremely randomized trees. Mach. Learn. 63, 3\u0026ndash;42. https://doi.org/10.1007/s10994-006-6226-1\u003c/li\u003e\n \u003cli\u003eHu C, Steingrimsson JA, 2018. Personalized Risk Prediction in Clinical Oncology Research: Applications and Practical Issues Using Survival Trees and Random Forests. J. Biopharm. Stat. 28, 333\u0026ndash;349. https://doi.org/10.1080/10543406.2017.1377730\u003c/li\u003e\n \u003cli\u003eIshwaran H, Kogalur UB, Blackstone EH, Lauer MS, 2008. Random survival forests. Ann. Appl. Stat. 2, 841\u0026ndash;860. https://doi.org/10.1214/08-AOAS169\u003c/li\u003e\n \u003cli\u003eJiang Y, Yuan Q, Lv W, Xi S, Huang W, Sun Z, Chen H, Zhao L, Liu W, Hu Y, Lu L, Ma J, Li T, Yu J, Wang Q, Li G, 2018. Radiomic signature of 18F fluorodeoxyglucose PET/CT for prediction of gastric cancer survival and chemotherapeutic benefits. Theranostics 8, 5915\u0026ndash;5928. https://doi.org/10.7150/thno.28018\u003c/li\u003e\n \u003cli\u003eJordan MI, Mitchell TM, 2015. Machine learning: Trends, perspectives, and prospects. Science (80-. ). https://doi.org/10.1126/science.aaa8415\u003c/li\u003e\n \u003cli\u003eKatzman JL, Shaham U, Cloninger A, Bates J, Jiang T, Kluger Y, 2018. DeepSurv: Personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med. Res. Methodol. 18. https://doi.org/10.1186/s12874-018-0482-1\u003c/li\u003e\n \u003cli\u003eKourou K, Exarchos TP, Exarchos KP, Karamouzis M V., Fotiadis DI, 2015. Machine learning applications in cancer prognosis and prediction. Comput. Struct. Biotechnol. J. 13, 8\u0026ndash;17. https://doi.org/10.1016/J.CSBJ.2014.11.005\u003c/li\u003e\n \u003cli\u003eKulig P, Nowakowski P, Sierzȩga M, Pach R, Majewska O, Markiewicz A, Kołodziejczyk P, Kulig J, Richter P, 2021. Analysis of prognostic factors affecting short-term and long-term outcomes of gastric cancer resection. Anticancer Res. https://doi.org/10.21873/anticanres.15140\u003c/li\u003e\n \u003cli\u003eMalhotra GK, Yanala U, Ravipati A, Follet M, Vijayakumar M, Are C, 2017. Global trends in esophageal cancer. J. Surg. Oncol. 115, 564\u0026ndash;579. https://doi.org/10.1002/jso.24592\u003c/li\u003e\n \u003cli\u003ePanahiazar M, Taslimitehrani V, Pereira N, Pathak J, 2015. Using EHRs and Machine Learning for Heart Failure Survival Analysis, in: Studies in Health Technology and Informatics. IOS Press, pp. 40\u0026ndash;44. https://doi.org/10.3233/978-1-61499-564-7-40\u003c/li\u003e\n \u003cli\u003ePedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay \u0026Eacute;, Pedregosa, F. and Varoquaux G, Gramfort, A. and Michel V, Thirion, B. and Grisel O, Blondel, M. and Prettenhofer P, Weiss, R. and Dubourg V, Vanderplas, J. and Passos A, Cournapeau, D. and Brucher M, Perrot M, Duchesnay E, 2011. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 12, 2825\u0026ndash;2830.\u003c/li\u003e\n \u003cli\u003eP\u0026ouml;lsterl S, 2020. Scikit-survival: A library for time-to-event analysis built on top of scikit-learn. J. Mach. Learn. Res. 21, 1\u0026ndash;6.\u003c/li\u003e\n \u003cli\u003eP\u0026ouml;lsterl S, Conjeti S, Navab N, Katouzian A, 2016. Survival analysis for high-dimensional, heterogeneous medical data: Exploring feature extraction as an alternative to feature selection. Artif. Intell. Med. 72, 1\u0026ndash;11. https://doi.org/10.1016/j.artmed.2016.07.004\u003c/li\u003e\n \u003cli\u003eSekercioglu N, Fu R, Kim SJ, Mitsakakis N, 2021. Machine learning for predicting long-term kidney allograft survival: a scoping review. Ir. J. Med. Sci. https://doi.org/10.1007/s11845-020-02332-1\u003c/li\u003e\n \u003cli\u003eSpooner A, Chen E, Sowmya A, Sachdev P, Kochan NA, Trollor J, Brodaty H, 2020. A comparison of machine learning methods for survival analysis of high-dimensional clinical data for dementia prediction. Sci. Rep. 10, 20410. https://doi.org/10.1038/s41598-020-77220-w\u003c/li\u003e\n \u003cli\u003eSraitih M, Jabrane Y, El Hassani AH, 2021. An automated system for ECG arrhythmia detection using machine learning techniques. J. Clin. Med. 10, 5450. https://doi.org/10.3390/jcm10225450\u003c/li\u003e\n \u003cli\u003eStephane Fotso, 2019. PySurvival: Open source package for Survival Analysis modeling. https://www.pysurvival.io.\u003c/li\u003e\n \u003cli\u003eSteyerberg EW, Vickers AJ, Cook NR, Gerds T, Gonen M, Obuchowski N, Pencina MJ, Kattan MW, 2010. Assessing the performance of prediction models: A framework for traditional and novel measures. Epidemiology. https://doi.org/10.1097/EDE.0b013e3181c30fb2\u003c/li\u003e\n \u003cli\u003eUno H, Cai T, Pencina MJ, D\u0026rsquo;Agostino RB, Wei LJ, 2011. On the C-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Stat. Med. 30, 1105\u0026ndash;1117. https://doi.org/10.1002/sim.4154\u003c/li\u003e\n \u003cli\u003eVan Rossum G, Drake FL, 2009. Python 3 Reference Manual. Scotts Val. CA Creat.\u003c/li\u003e\n \u003cli\u003eWright MN, Dankowski T, Ziegler A, 2017. Unbiased split variable selection for random survival forests using maximally selected rank statistics. Stat. Med. 36, 1272\u0026ndash;1284. https://doi.org/10.1002/sim.7212\u003c/li\u003e\n \u003cli\u003eYu CN, Greiner R, Lin HC, Baracos V, 2011. Learning patient-specific cancer survival distributions as a sequence of dependent regressors, in: Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011, NIPS 2011.\u003c/li\u003e\n \u003cli\u003eZhu W, Xie L, Han J, Guo X, 2020. The application of deep learning in cancer prognosis prediction. Cancers (Basel). https://doi.org/10.3390/cancers12030603\u003c/li\u003e\n \u003cli\u003e\u0026nbsp;\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003ch3 style=\"text-align: center;\"\u003eTable 1\u0026nbsp;\u003c/h3\u003e\n\u003cp style=\"text-align: center;\"\u003eOverview of independent and outcome variables.\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellpadding=\"0\" cellspacing=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003eBiometric variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003eheight, weight, body mass index (BMI), age, sex.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003ePreoperative variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003epast medical history (cardiovascular, pulmonary, metabolic and renal preconditions), tumor diagnosis, preceding malignant disease, cTNM classification, histology (Laur\u0026eacute;n type, signet cell component), neoadjuvant therapy (components, radiotherapy, completeness).\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003e(Intra-)operative variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003etime between diagnosis and resection, type of operation, extent of resection, anatomical reconstruction, duration of surgery, intraoperative complication, blood loss and transfusion.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003ePostoperative variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003edays on ICU and ward, postoperative complications (according to Clavien-Dindo and additional 38 binarily classified types), pTNM classification, lymph node ratio (positive lymph nodes divided by resected), grading, R-status, histology (Laur\u0026eacute;n type, signet cell component, tumor regression), post-discharge problems.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003eOutcome variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003evital status, overall survival, no evidence of disease, time until tumor relapse, 30-day mortality and in-hospital mortality.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"text-align: center;\"\u003e\u003cstrong\u003eTable 2\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"text-align: center;\"\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eOverview of selected patient characteristics and imputation methods. ASA = American Society of Anesthesiology, ICU = intensive care unit.\u003cbr\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellpadding=\"0\" cellspacing=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"29.6849087893864%\"\u003e\n \u003cp\u003e\u003cem\u003en = 1,360\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"56.38474295190713%\"\u003e\n \u003cp\u003e\u003cstrong\u003eMean / Median / Frequency, (95% confidence interval)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"13.930348258706468%\"\u003e\n \u003cp\u003e\u003cstrong\u003eMissingness\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003eBiometric variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eWeight\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e77.7 kg (54.0 - 104.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eHeight\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e1.72 m (1.58 - 1.87)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e1.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eAge\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e62.35 years (41.0 - 80.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eSex\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003emale (71.4%), female (28.6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003ePreoperative variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eASA score\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e1 (1.7%), 2 (47.9%), 3 (46.9%), 4 (2.3%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e1.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003ePast medical history\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e\u003cem\u003esee Supplementary Table 1\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003emax. 0.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003ecT status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e1 (7.1%), 2 (20.1%), 3 (57.2%), 4 (10.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e4.7%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003ecN status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e0 (35.7%), 1 (61.5%), + (0.4%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e2.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003ecM status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e0 (87.7%), 1 (11.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eNeoadjuvant therapy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003eyes (55.0%), no (45.0%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003e- Chemotherapeutics\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e\u003cem\u003esee Supplementary Table 1\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003emax. 2.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003e- Radiation\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003eyes (3.5%), no (96.2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003e(Intra-)operative variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eTime diagnosis to resection\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e82.84 days (11.0 - 168.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eOne vs. two cavity surgery\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003eone (71.1%), two (28.9%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eOperation type\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003eSubtotal gastrectomy (21.5%)\u003c/p\u003e\n \u003cp\u003eTotal gastrectomy (26.4%)\u003c/p\u003e\n \u003cp\u003eTranshiatal extended gastrectomy (21.8%)\u003c/p\u003e\n \u003cp\u003eIvor-Lewis esophagectomy (27.8%)\u003c/p\u003e\n \u003cp\u003eOther types (2.6%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003e(Locally) extended resection\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003eyes (28.5%), no (71.2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.3%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eIntraoperative complications\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003eyes (7.8%), no (92.1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eDuration of surgery\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e275.29 minutes (150.0 - 479.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.9%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eIntraoperative blood loss\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e622.93 milliliters (100.0 - 1,500.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e19.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003ePostoperative variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eDuration of stay on ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e6.89 days (0.0 - 29.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e4.7%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eDuration of stay on ward\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e20.26 days (9.0 - 49.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.7%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003ePostoperative complications\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e\u003cem\u003esee Supplementary Table 1\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003emax. 0.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003epT status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e0/1 (22.3%), 2 (14.2%), 3 (47.3%), 4 (16.2%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003epN status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e0 (42.1%), 1 (16.3%), 2 (15.0%), 3 (26.5%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eLymph node ratio\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e18.25% (0.0 - 74.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003epM status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e0 (89.7%), 1 (10.3%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eR status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e0 (81.5%), any 1/2/X (18.5%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e0.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cstrong\u003eOutcome variables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eStatus dead vs. alive/censored\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003ealive (51.5%), dead (46.1%)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e2.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"29.6849087893864%\"\u003e\n \u003cp\u003eDocumented survival time\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"56.38474295190713%\"\u003e\n \u003cp\u003e39.69 months (2.89 - 109.88)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"13.930348258706468%\"\u003e\n \u003cp\u003e1.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"3\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"text-align: center;\"\u003e\u003cstrong\u003eTable 3\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"text-align: center;\"\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eMost important 20 predictors based on permutation-based importance scoring for random survival forest. ICU = intensive care unit, ASA = American Society of Anesthesiology.\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellpadding=\"0\" cellspacing=\"0\" width=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003e\u003cstrong\u003ePredictive Feature\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e\u003cstrong\u003eWeight\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e\u003cstrong\u003eLower 95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e\u003cstrong\u003eUpper 95% CI\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eLymph node ratio (positive / total)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.1183\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.1016\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.1350\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eIntraoperative blood loss\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0060\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e-0.0009\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0129\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003e(y)pT4 status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0057\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0002\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0112\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eAge (at diagnosis)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0049\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0035\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0063\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003e(y)pT3 status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0042\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e-0.0004\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0088\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eIntraoperative peritoneal carcinosis\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0040\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0033\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0047\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003ePostoperative sepsis\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0038\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e-0.0003\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0079\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003ecM+ status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0036\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e-0.0009\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0081\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eDuration of in-hospital stay\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0033\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0028\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0038\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eDuration of stay on ICU\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0033\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0004\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0062\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eMajor postoperative complications\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0032\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0044\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eAny R+ status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0026\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e-0.0024\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0076\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eLocal R1 status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0021\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0004\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0038\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003e(y)pM1 status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0007\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0033\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003e(y)pT1 status\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0014\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0011\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0017\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eIntraoperative blood transfusion\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0013\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0007\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0019\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eSevere pre-existing diseases\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0006\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0002\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0010\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eSiewert type I junction cancer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0006\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e-0.0004\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0016\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eCardiac postoperative adverse events\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0005\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0009\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" width=\"39.168110918544194%\"\u003e\n \u003cp\u003eASA grade 3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0005\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0000\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" width=\"20.27729636048527%\"\u003e\n \u003cp\u003e0.0010\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"4\" valign=\"top\" width=\"100%\"\u003e\n \u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"journal-of-cancer-research-and-clinical-oncology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jocr","sideBox":"Learn more about [Journal of Cancer Research and Clinical Oncology](https://www.springer.com/journal/432)","snPcode":"432","submissionUrl":"https://submission.nature.com/new-submission/432/3","title":"Journal of Cancer Research and Clinical Oncology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"gastric cancer, esophageal cancer, machine learning, survival analysis, oncological outcome","lastPublishedDoi":"10.21203/rs.3.rs-1318132/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1318132/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003ePurpose\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eSurgical oncologists are frequently confronted with the question of expected long-term prognosis. The aim of this study was to apply machine learning algorithms to optimize survival prediction after oncological resection of gastroesophageal cancers.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eEligible patients underwent oncological resection of gastric or distal esophageal cancer between 2001 and 2020 at Heidelberg University Hospital, Department of General Surgery. Machine learning methods such as multi-task logistic regression and survival forests were compared with usual algorithms to establish an individual estimation.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eThe study included 117 variables with a total of 1,360 patients. The overall missingness was 1.3%. Out of eight machine learning algorithms, the random survival forest (RSF) performed best with a concordance-index of 0.736 and an integrated Brier score of 0.166. The RSF demonstrated a mean area under the curve (AUC) of 0.814 over a time period of 10 years after diagnosis. The most important long-term outcome predictor was lymph node ratio with a mean AUC of 0.730. A numeric risk score was calculated by the RSF for each patient and three risk groups were defined accordingly. Median survival time was 18.8 months in the high-risk group, 44.6 months in the medium-risk group and above 10 years in the low-risk group.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eThe results of this study suggest that RSF is most appropriate to accurately answer the question of long-term prognosis. Furthermore, we could establish a compact risk score model with 20 input parameters and thus provide a clinical tool to improve prediction of oncological outcome after upper gastrointestinal surgery.\u003c/p\u003e","manuscriptTitle":"Machine Learning for Optimal Individual Survival Prediction in Resectable Upper Gastrointestinal Cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-02-23 20:39:05","doi":"10.21203/rs.3.rs-1318132/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2022-02-24T08:10:43+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-02-21T22:21:43+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"Journal of Cancer Research and Clinical Oncology","date":"2022-02-03T12:39:45+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-02-03T12:16:55+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Cancer Research and Clinical Oncology","date":"2022-02-01T10:00:52+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"journal-of-cancer-research-and-clinical-oncology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jocr","sideBox":"Learn more about [Journal of Cancer Research and Clinical Oncology](https://www.springer.com/journal/432)","snPcode":"432","submissionUrl":"https://submission.nature.com/new-submission/432/3","title":"Journal of Cancer Research and Clinical Oncology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"6c47555f-3f70-47ac-851b-8d925e9af08c","owner":[],"postedDate":"February 23rd, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2022-05-10T12:24:09+00:00","versionOfRecord":[],"versionCreatedAt":"2022-02-23 20:39:05","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1318132","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1318132","identity":"rs-1318132","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.