Deep learning models for predicting the survival of patients with medulloblastoma based on a surveillance, epidemiology, and end results analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Deep learning models for predicting the survival of patients with medulloblastoma based on a surveillance, epidemiology, and end results analysis Meng Sun, Jikui Sun, Meng Li This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3975955/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 24 Jun, 2024 Read the published version in Scientific Reports → Version 1 posted 9 You are reading this latest preprint version Abstract Background Medulloblastoma is a malignant neuroepithelial tumor of the central nervous system. Accurate prediction of prognosis is essential for therapeutic decisions in medulloblastoma patients. Several prognostic models have been developed using multivariate Cox regression to predict the1-, 3- and 5-year survival of medulloblastoma patients, but few studies have investigated the results of integrating deep learning algorithms. Compared to simplifying predictions into binary classification tasks, modelling the probability of an event as a function of time by combining it with deep learning may provide greater accuracy and flexibility. Methods Patients diagnosed with medulloblastoma between 2000 and 2019 were extracted from the Surveillance, Epidemiology, and End Results (SEER) registry. Three models—one based on neural networks (DeepSurv), one based on ensemble learning (random survival forest [RSF]), and a typical Cox proportional-hazards (CoxPH) model—were selected for training. The dataset was randomly divided into training and testing datasets in a 7:3 ratio. The model performance was evaluated utilizing the concordance index (C-index), Brier score and integrated Brier score (IBS). The accuracy of predicting 1-, 3- and 5- year survival was assessed using receiver operating characteristic curves (ROC), and the area under the ROC curves (AUC). Results The 2,322 patients with medulloblastoma enrolled in the study were randomly divided into the training cohort (70%, n = 1,625) and the test cohort (30%, n = 697). There was no statistically significant difference in clinical characteristics between the two cohorts ( p > 0.05). We performed Cox proportional hazards regression on the data from the training cohort, which illustrated that age, race, tumour size, histological type, surgery, chemotherapy, and radiotherapy were significant factors influencing survival ( p < 0.05). The Deepsurv outperformed the RSF and classic CoxPH models with C-indexes of 0.763 and 0.751 for the training and test datasets. The DeepSurv model showed better accuracy in predicting 1-, 3- and 5-year survival (AUC: 0.805–0.838). Conclusion The predictive model based on a deep learning algorithm that we have developed can exactly predict the survival rate and duration of medulloblastoma. Biological sciences/Cancer/Cancer models Biological sciences/Cancer/Cns cancer DeepSurv Medulloblastoma Neural network Survival prediction SEER Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Introduction Medulloblastoma is an embryonal tumor that arises from the cerebellum and has the potential to spread throughout the nervous system. It is the most common type of paediatric embryonal tumor, with an incidence ranging from 5 to 11 cases per 1 million individuals 1 , 2 . According to current international consensus, there are four subgroups of medulloblastoma: Wingless (WNT), Sonic Hedgehog (SHH), group 3 (G3), and group 4 (G4) 3 . Multimodal therapy, which includes surgery, external beam irradiation, and/or cytotoxic chemotherapy, can result in survival rates ranging from 50–80% based on clinical staging 4 . Certain prognostic features, such as age at diagnosis, extent of resection, histological subtype, and molecular subgroup classification, have been found to affect survival predictions in individual patients. Previous studies have used the Cox proportional-hazards model (CoxPH) to evaluate the survival rate of medulloblastoma patients 5 6, 7 . This model incorporates survival outcomes and time as target variables, allowing for the simultaneous analysis of multiple factors' impact on survival time. It is extensively used for predicting outcome events when the survival distribution of the analyzed data is unknown 8 . A nomogram is a commonly used method for quantifying and combining important clinical characteristics of patients to calculate the probabilities of outcome events based on the CoxPH model 9 . However, the model assumes that each predictor variable has the same effect throughout the follow-up time, which ignores variations in their impact on individual patients at different times. Therefore, a new method is required to improve the accuracy of predicting the survival rate of cancer patients. In recent years, computer and information technology have shown revolutionary potential for artificial intelligence (AI) in the healthcare industry 10 – 12 . Machine learning models have stronger nonlinear modeling capabilities compared to traditional linear models and can better capture complex relationships among clinical variables. The analysis of these models can provide accurate personalized survival predictions and decision-making support for treatment strategies to improve patient survival rates 13 , 14 . Deep learning, a field in machine learning, explores patterns and representations within data to characterize their distribution 15 , 16 . It is a statistical model that consists of an input layer, hidden layer, and output layer. This model can solve complex, multifactorial, and nonlinear problems. Deep learning-based models have become highly effective predictors of clinical outcomes across various disease domains due to the continuous advancements in deep learning research techniques and the abundance of biomedical big data. Jiang et al. 17 demonstrated the use of an artificial neural network model to predict the survival rate of patients diagnosed with pancreatic neuroendocrine neoplasms, by leveraging clinical information. Katzman et al. 18 integrated deep learning with a multilayer neural network architecture, known as the DeepSurv model, resulting in a personalized treatment recommendation system that showed remarkable performance. To our knowledge, there is a lack of research combining deep learning techniques with the study of medulloblastoma. Therefore, this study aimed to fill this research gap by utilizing data obtained from the Surveillance, Epidemiology, and End Results (SEER) database, which contains information on patients diagnosed with medulloblastoma in the United States. And then the DeepSurv model was used to evaluate their survival rates. Method Data source and patient selection The data of this retrospective cohort study from the SEER database, which encompasses information from 18 cancer registries representing approximately 28% of the entire US population 19 . This database offers extensive and detailed patient data, including demographic characteristics, tumor-related information, cause of death, and survival duration. The SEER*Stat software (version 8.3.6) was used to identify patients with medulloblastoma. The dataset covering the years 2000 to 2019 in the United States was accessed. The patients included in the study had to meet the following criteria: 1) a confirmed pathological diagnosis of medulloblastoma; 2) identification of medulloblastoma cases based on the third edition of the International Classification of Diseases for Oncology (ICD-O3) using specific ICD-O-3 codes for histopathology, including 9,470/3 for medulloblastoma, NOS; 9,471/3 for desmoplastic nodular medulloblastoma; and 9,474/3 for large cell medulloblastoma. Furthermore, patients were required to have a known survival status and time. Afterwards, they were randomly divided into a training group and a testing group at a 7:3 ratio. A flowchart in Fig. 1 illustrates the process of patient selection. Variable’s definitions Several parameters were collected from the samples, including age at diagnosis, sex, race, histological type, tumor size, surgery, chemotherapy, radiation therapy, and survival time. To evaluate the prognostic value of age and tumor size in patients with medulloblastoma objectively, the patients were categorized into two groups based on the optimal cutoff values obtained using the X-tile software ( https://x-tile.software.informer.com , Yale School of Medicine, New Haven, CT, United States). Age cutoff values of ≤ 3 years and > 3 years, and tumor size cutoff values of ≤ 3.4 cm, > 3.4 cm, and/or unknown were utilised. For detailed visual representations, please refer to Fig. 2 . Model development This study selected three models for training: DeepSurv, RSF, and CoxPH. DeepSurv is a deep feedforward neural network used to predict patients' survival time or survival probability. It employs a multi-layer neural network to capture the complex nonlinear relationship between patients' survival probability and input features. This study utilized deep-learning calculations based on the DeepSurv calculation method described by Katzman et al. 18 to predict the survival outcome of patients diagnosed with medulloblastoma. The term RSF refers to Random Survival Forests, which is a survival analysis method based on random forests. When constructing a random survival forest, subsets of samples and features are randomly selected, and multiple decision trees are built using these subsets. Each decision tree splits the samples based on features in the nodes and determines the optimal splitting based on the evaluation of survival time differences. The predictions from multiple decision trees in the random survival forests are combined to obtain the final survival prediction. The CoxPH is a semi-parametric regression model used to analyse survival data and estimate the risk of event occurrence. The Cox proportional-hazards model is used to compare the relative risks of events between different groups and study the impact of various factors on event occurrence. The model functions by modeling the relationship between time and event occurrence as a function of hazard ratios. For the implementation of the algorithms in this research, CoxPH and RSF were implemented using the R packages "survival" and "randomForestSRC", respectively. On the other hand, DeepSurv utilized an open-source Python package. The hyperparameter optimization for DeepSurv was conducted using the Hyperopt package within the TensorFlow framework. Model evaluation The study evaluated the model's performance using several metrics, including C-index, Brier score, integrated Brier score (IBS), Receiver Operating Characteristic (ROC) curves, and Area Under the Curve (AUC) values. The C-index is a commonly used metric for evaluating the accuracy of survival predictions. It measures the concordance or correlation between the predicted survival risk and the actual observed survival time. A C-index of 0.5 indicates random predictions, while a value of 1.0 indicates perfect predictions. The Brier score assesses the mean squared difference between the observed patient statuses (event occurrence or censoring) and the predicted survival probabilities. It ranges from 0 to 1, with 0 indicating a perfect match between predictions and observations. In practice, models with Brier scores less than 0.25 are considered useful. The IBS is a metric that evaluates the overall performance of a survival model across all available time points. It takes into account the model's sensitivity and specificity to time-dependent events, providing a comprehensive measure of predictive accuracy. Receiver Operating Characteristic (ROC) curves are frequently used to assess a model's sensitivity and specificity at various discrimination thresholds. The ROC curve plots the true positive rate against the false positive rate. The Area Under the Curve (AUC) values, which range from 0 to 1, are computed to quantify the overall performance of the model. A higher AUC indicates better discrimination ability. This study calculated AUC values to assess the model's performance at different time points: 1-year, 3-year, and 5-year survival rates. Statistical analysis In the clinical data, continuous variables are expressed as mean ± standard deviation (SD), while categorical variables are described using frequencies and percentages. Statistical tests such as chi-square tests and unpaired t-tests are used to compare variables between groups. Result Basic characteristics This study analysed data from 2,322 medulloblastoma patients registered in the SEER database between 2000 and 2019. Table 1 presents the demographic features of the patients, with 869 cases (37.42%) being female and 1,453 cases (62.58%) being male. The racial distribution was as follows: 185 patients (7.97%) were Black, 1,939 (83.51%) were White, and 198 (8.53%) belonged to other races. Regarding the subtypes of medulloblastoma, 329 patients (14.17%) had desmoplastic/nodular medulloblastoma (DMB), 1,866 (80.36%) had medulloblastoma, not otherwise specified (MB, NOS), and 127 (5.47%) had large-cell/anaplastic medulloblastoma (LC). In terms of surgical interventions, 1,616 patients (69.60%) underwent total resection, 244 (10.51%) underwent subtotal resection, 343 (14.77%) underwent local excision or biopsy, and 119 (5.12%) did not undergo surgery. Of the patients, 1,849 (79.63%) received chemotherapy, 1,766 (76.06%) underwent radiation therapy, and 713 (30.71%) died. The cutoff values for age and tumor size were determined using X-tile analysis ( Fig. 2 ) . Specifically, 324 patients (13.95%) were ≤ 3 years old, and 1,998 patients (86.05%) were older than 3 years. Regarding tumor size, 314 patients (13.52%) had tumors ≤ 3.4 cm, 1,269 patients (54.65%) had tumor size > 3.4 cm, and the tumor size was unknown for 739 patients (31.83%). The predictive model was generated by partitioning the complete dataset into two mutually exclusive subsets. 70% of the dataset was allocated for the training set, while the remaining 30% was used for the testing set. Model generation was performed on 1,625 randomly assigned patients from the training set, while the accuracy of the model was estimated using 697 randomly assigned patients from the validation set. No statistically significant differences in characteristics were found between the two groups (refer to Table 1 ) . Additionally, survival outcomes showed no differences between the two groups (refer to Fig. 3 ) . Table 1 Characteristic distribution of data into raining sets and test sets. Variables Overall N (%) Train cohort N (%) Test cohort N (%) P Patients 2,322 1625 (69.98) 697 (30.02) Age ≤ 3 >3 324 (13.95) 1,998 (86.05) 229 (14.09) 1,396 (85.91) 95 (13.63) 602 (86.37) 0.73 Sex Female Male 869 (37.42) 1,453 (62.58) 625 (38.46) 1,000 (61.54) 244 (35.01) 453 (64.99) 0.12 Race Black White Other 185 (7.97) 1,939 (83.51) 198 (8.53) 135 (8.31) 1,349 (83.02) 141 (8.68) 50 (7.17) 590 (84.65) 57 (8.18) 0.57 Histopathology DMB MB, NOS LC 329 (14.17) 1,866 (80.36) 127 (5.47) 229 (14.09) 1,302 (80.12) 94 (5.78) 100 (14.35) 564 (80.92) 33 (4.73) 0.24 Size (cm) ≤ 3.4 >3.4 Unknown 314 (13.52) 1,269 (54.65) 739 (31.83) 211 (12.98) 896 (55.14) 518 (31.88) 103 (14.78) 373 (53.52) 221 (31.71) 0.17 Surgery Total resection Subtotal resection Local excision/Biopsy No evidence 1,616 (69.60) 244 (10.51) 343 (14.77) 119 (5.12) 1,122 (69.05) 175 (10.77) 238 (14.65) 90 (5.54) 494 (70.88) 69 (9.90) 105 (15.06) 29 (4.16) 0.14 Chemotherapy Yes No evidence 1,849 (79.63) 473 (20.37) 1,274 (78.40) 351 (21.60) 575 (82.50) 122 (17.50) 0.25 Radiotherapy Yes No evidence 1,766 (76.06) 556 (24.94) 1,214 (74.71) 411 (25.29) 552 (79.20) 145 (20.80) 0.33 Status Death Alive 713 (30.71) 1,609 (69.29) 514 (31.63) 1,111 (68.37) 199 (28.55) 498 (71.45) 0.46 Cox proportional-hazard (CoxPH) model The CoxPH model was developed using the training set (refer to Fig. 4 ) . Only variables that showed statistical significance in the univariate analysis were included in the multivariate analysis. The survival of medulloblastoma patients was significantly affected by non-surgical treatment, LC, white race, tumor size ≤ 3.4 cm, total resection, age > 3 years, chemotherapy, and radiotherapy. Furthermore, the survival of the patients was significantly associated with these features in the multivariate analysis. The collinearity analysis also revealed a high correlation between age and radiotherapy, as well as between chemotherapy and radiotherapy (refer to Fig. 5 ) . Ultimately, we included seven features (age, race, tumor size, histological type, surgery, chemotherapy, and radiotherapy) in the model development. Random Survival Forests (RSF) Prediction error is calculated using the out-of-bag (OOB) from the training and the testing set ( Fig. 6 A, B ) . Variable importance (VIMP) can be measured by randomising a specific variable, as shown in Fig. 6 C. A higher VIMP value indicates a greater influence or importance of that variable in accurately predicting the outcome 20 . The interaction between variables in the analyzed data is illustrated and displayed in Fig. 6 D. If one variable's split in a decision tree affects or influences the split of another variable, it suggests an interaction between those variables 21 , 22 . The extent of interactions is assessed based on the minimum depth, which represents the distance from the root node to the node where the variable first splits. In this case, it is observed that chemotherapy and radiotherapy exhibit the lowest minimum depth among the variables considered were expected to be associated with other variables. DeepSurv The loss function curve, which illustrates the relationship between the loss and the number of iterations during the training process, provides valuable insights into the convergence and performance of the model. By examining this curve, we can assess how well the model is optimizing its parameters over iterations 23 . Furthermore, plotting the performance of the training dataset as a function of the number of iterations allows us to evaluate the model's ability to rank the samples accurately. This measure helps us monitor the model's generalization ability and identify any signs of overfitting, where the model may excessively capture details specific to the training dataset but fails to generalize well to unseen samples 24 . The learning process of DeepSurv, a survival prediction model based on deep learning, was visualized ( Fig. 7 ) . The figure demonstrates a good fit of the model, indicating that it is effectively learning and capturing the patterns within the data. Model comparisons The predictive performance of the three models is shown in Table 2 . In the test dataset, the DeepSurv and RSF model exhibited significantly better discrimination abilities (the DeepSurv C-index: 0.751, RSF: 0.750) compared with the CoxPH model (the C-index: 0.748). And in the three models, DeepSurv had the highest C-index of 0.751. The IBS for the three models were as follows: DeepSurv (0.150), RSF (0.160), and CoxPH (0.166). Lower IBS values indicate better model performance. Additionally, the C-index obtained from the train data set (DeepSurv: 0.763, RSF: 0.759, CoxPH: 0.757) differed only slightly with test set, indicating that the models did not exhibit overfitting. Furthermore, in terms of the Brier score ( Fig. 8 ) , DeepSurv outperformed the other two models, indicating its superior accuracy. The AUC for DeepSurv was also higher than the other models (1-year-AUC of DeepSurv: 0.838, RSF: 0.809, CoxPH: 0.808; 3-year-AUCof DeepSurv: 0.820, RSF: 0.791, CoxPH: 0.782;5-year-AUC of DeepSurv: 0.805, RSF: 0.780, CoxPH: 0.773) ( Fig. 9 ) . These results demonstrate that DeepSurv outperforms both RSF and the classical CoxPH model in accurately predicting the prognosis of patients with medulloblastoma. Table 2 Performance of three survival models. Models C index Train Test IBS 1-year ACU 3-year AUC 5-year AUC CoxPH 0.757 0.748 0.166 0.808 (0.77–0.85) 0.782 (0.74–0.82) 0.773 (0.74–0.81) RSF 0.759 0.750 0.160 0.809 (0.77–0.85) 0.791 (0.75–0.83) 0.780 (0.74–0.82) DeepSurv 0.763 0.751 0.150 0.838 (0.80–0.88) 0.820 (0.78–0.86) 0.805 (0.76–0.84) Discussion Medulloblastoma, a malignant brain tumor that mainly impacts children, continues to pose a substantial obstacle in the field of pediatric oncology. Precisely predicting the individual prognosis of patients is crucial for customizing treatment approaches and enhancing survival rates. Prior research has identified several prognostic factors that affect the survival duration of medulloblastoma patients, including age, extent of surgical removal, and the administration of radiotherapy or chemotherapy 7 , 25 , 26 . Moreover, as medical advancements progress, an increasing amount of imaging data 5 and genetic data 27 are being analyzed for survival analysis of medulloblastoma patients. However, classical survival analysis methods, such as the Cox proportional-hazards model, assume a linear relationship between variables, which may be limited in the face of multidimensional data. With the advancement of artificial intelligence, machine learning methods are being applied to clinical, imaging, and genetic data, allowing for the discovery of potential nonlinear relationships within the data 28 – 30 . Within machine learning, deep learning is a specific class of methods that utilizes multilayered neural networks to extract high-order features. Deep learning has gained increasing popularity in the field of cancer survival analysis, and has demonstrated excellent performance 31 – 33 . As far as we know, this approach has not been applied to medulloblastoma. Therefore, we developed a deep learning model (Deepsurv) to predict the overall survival (OS) of medulloblastoma patients and compared its performance to that of a machine learning model (RSF) and a classical model (CoxPH). By extracting potentially significant features from the SEER database, this research developed multiple models to forecast the survival rates of individuals diagnosed with medulloblastoma. Initially, we utilized the X-tile tool to determine the optimal cutoff values for age and tumor size from a cohort of 2,322 medulloblastoma patients. We identified two high-risk factors, age ≤ 3 years old and tumor size > 3.4 cm, that significantly impact the survival duration of patients with medulloblastoma. Subsequently, we employed Cox proportional hazards regression to identify variables associated with the prognosis of medulloblastoma patients. Age, race, tumor size, histological type, surgery, chemotherapy, and radiotherapy were selected for inclusion in the modeling process ( p < 0.05). We established RSF, DeepSurv and CoxPH models and evaluated their performance using metrics such as the C-index, IBS, and ROC curve. The study results demonstrated that the DeepSurv model outperformed both the CoxPH and RSF models, as indicated by its higher C-index in both the training and testing sets. Moreover, the DeepSurv model exhibited the lowest IBS and the largest AUC values when predicting 1-, 3-, and 5-year survival. These findings collectively suggest that the DeepSurv model is more accurate in predicting the survival of patients with medulloblastoma. In previous studies, Guo et al. 7 and Zhou et al. 5 utilized Cox proportional hazard regression for survival analysis of medulloblastoma and developed a nomogram. Compared with their study, the C-index values obtained from the DeepSurv model were higher in both the training cohort, indicating its superior predictive accuracy of the prognosis of patients with medulloblastoma. This finding is consistent with the results reported in several previous studies focusing on cancer prognosis 34 , 35 . The main advantage of the DeepSurv model in its ability to process both linear and nonlinear predictive variables by utilizing a multilevel neural network. This transformation into a linear combination allows the model to uncover associations that may not be readily apparent to the human eye or traditional statistical techniques. Nevertheless, our study encountered several limitations. Firstly, the data collected from the SEER database for medulloblastoma patients contained some missing information that could potentially influence survival outcomes. This includes important details such as molecular subgroups, specific radiotherapy dosages, and chemotherapy regimens. The availability and completeness of these data rely on the ongoing improvements in data collection within the SEER database. Secondly, our model has yet to undergo external validation, and it is necessary to validate its performance on new data. Conducting further validations using independent datasets would enhance the reliability and generalizability of the findings. Another inherent limitation lies within the DeepSurv model itself. Due to its utilization of hidden layers in its architecture, the model operates as a black-box, making it challenging to fully comprehend the computations involved in the model construction process and its associated limitations. Future research should aim to address these concerns and explore the inner workings of the model to improve interpretability. Conclusions This study employed Cox proportional hazards regression analysis to examine the prognostic factors influencing medulloblastoma patients' outcomes, which include age, race, tumor size, histological type, surgery, chemotherapy, and radiotherapy. Subsequently, we developed a groundbreaking DeepSurv prediction model, which exhibited strong predictive capabilities in assessing the prognosis of patients diagnosed with medulloblastoma. This innovative DeepSurv model holds significant potential in accurately predicting the survival duration of medulloblastoma patients. Declarations Funding The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Competing interests The authors have no relevant financial or non-financial interests to disclose. Author Contributions Sun M: conceptualization, data curation, investigation, methodology, software, visualization, writing—original draft; Sun J: conceptualization, data curation, formal analysis; Li M: methodology, project administration, supervision—review and editing. All authors read and approved the final manuscript. Data availability The datasets analyzed during the current study are available in the SEER database repository (https://seer.cancer.gov/). Ethical approval and consent to participate Not applicable. Consent for publication Not applicable. References Gajjar, A. J. & Robinson, G. W. Medulloblastoma-translating discoveries from the bench to the bedside. Nat Rev Clin Oncol 11, 714–722, doi: 10.1038/nrclinonc.2014.181 (2014). Ostrom, Q. T., Cioffi, G., Waite, K., Kruchko, C. & Barnholtz-Sloan, J. S. CBTRUS Statistical Report: Primary Brain and Other Central Nervous System Tumors Diagnosed in the United States in 2014–2018. Neuro Oncol 23, iii1-iii105, doi: 10.1093/neuonc/noab200 (2021). Taylor, M. D. et al. Molecular subgroups of medulloblastoma: the current consensus. Acta Neuropathol 123, 465–472, doi: 10.1007/s00401-011-0922-z (2012). Ramaswamy, V. & Taylor, M. D. Medulloblastoma: From Myth to Molecular. J Clin Oncol 35, 2355–2363, doi: 10.1200/JCO.2017.72.7842 (2017). Zhou, L. et al. Automatic image segmentation and online survival prediction model of medulloblastoma based on machine learning. Eur Radiol, doi: 10.1007/s00330-023-10316-9 (2023). Li, X. & Gong, J. Survival nomogram for medulloblastoma and multi-center external validation cohort. Front Pharmacol 14, 1247812, doi: 10.3389/fphar.2023.1247812 (2023). Guo, C. et al. External Validation of a Nomogram and Risk Grouping System for Predicting Individual Prognosis of Patients With Medulloblastoma. Front Pharmacol 11, 590348, doi: 10.3389/fphar.2020.590348 (2020). Baek, E. T. et al. Survival time prediction by integrating cox proportional hazards network and distribution function network. BMC Bioinformatics 22, 192, doi: 10.1186/s12859-021-04103-w (2021). Iasonos, A., Schrag, D., Raj, G. V. & Panageas, K. S. How to build and interpret a nomogram for cancer prognosis. J Clin Oncol 26, 1364–1370, doi: 10.1200/JCO.2007.12.9791 (2008). Schwalbe, N. & Wahl, B. Artificial intelligence and the future of global health. Lancet 395, 1579–1586, doi: 10.1016/S0140-6736(20)30226-9 (2020). Hamet, P. & Tremblay, J. Artificial intelligence in medicine. Metabolism 69S, S36-S40, doi: 10.1016/j.metabol.2017.01.011 (2017). Hunter, D. J. & Holmes, C. Where Medical Statistics Meets Artificial Intelligence. N Engl J Med 389, 1211–1219, doi: 10.1056/NEJMra2212850 (2023). Connor, C. W. Artificial Intelligence and Machine Learning in Anesthesiology. Anesthesiology 131, 1346–1359, doi: 10.1097/ALN.0000000000002694 (2019). Bhat, M., Rabindranath, M., Chara, B. S. & Simonetto, D. A. Artificial intelligence, machine learning, and deep learning in liver transplantation. J Hepatol 78, 1216–1233, doi: 10.1016/j.jhep.2023.01.006 (2023). Choi, R. Y., Coyner, A. S., Kalpathy-Cramer, J., Chiang, M. F. & Campbell, J. P. Introduction to Machine Learning, Neural Networks, and Deep Learning. Transl Vis Sci Technol 9, 14, doi: 10.1167/tvst.9.2.14 (2020). Greener, J. G., Kandathil, S. M., Moffat, L. & Jones, D. T. A guide to machine learning for biologists. Nat Rev Mol Cell Biol 23, 40–55, doi: 10.1038/s41580-021-00407-0 (2022). Jiang, C. et al. Predicting the survival of patients with pancreatic neuroendocrine neoplasms using deep learning: A study based on Surveillance, Epidemiology, and End Results database. Cancer Med 12, 12413–12424, doi: 10.1002/cam4.5949 (2023). Katzman, J. L. et al. DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol 18, 24, doi: 10.1186/s12874-018-0482-1 (2018). Hankey, B. F., Ries, L. A. & Edwards, B. K. The surveillance, epidemiology, and end results program: a national resource. Cancer Epidemiol Biomarkers Prev 8, 1117–1121 (1999). Taylor, J. M. Random Survival Forests. J Thorac Oncol 6, 1974–1975, doi: 10.1097/JTO.0b013e318233d835 (2011). Gilhodes, J. et al. Comparison of variable selection methods for high-dimensional survival data with competing events. Comput Biol Med 91, 159–167, doi: 10.1016/j.compbiomed.2017.10.021 (2017). Kretowska, M. Tree-based models for survival data with competing risks. Comput Methods Programs Biomed 159, 185–198, doi: 10.1016/j.cmpb.2018.03.017 (2018). Du, J., Zhou, Y., Liu, P., Vong, C. M. & Wang, T. Parameter-Free Loss for Class-Imbalanced Deep Learning in Image Classification. IEEE Trans Neural Netw Learn Syst 34, 3234–3240, doi: 10.1109/TNNLS.2021.3110885 (2023). Serghiou, S. & Rough, K. Deep Learning for Epidemiologists: An Introduction to Neural Networks. Am J Epidemiol 192, 1904–1916, doi: 10.1093/aje/kwad107 (2023). Dasgupta, A. et al. Nomograms based on preoperative multiparametric magnetic resonance imaging for prediction of molecular subgrouping in medulloblastoma: results from a radiogenomics study of 111 patients. Neuro Oncol 21, 115–124, doi: 10.1093/neuonc/noy093 (2019). Liu, H. & Sun, P. A Nomogram Model for Predicting Prognosis of Patients with Medulloblastoma. Turk Neurosurg 34, 38–45, doi: 10.5137/1019-5149.JTN.40397-22.3 (2024). Zhu, S. et al. Identification of a Twelve-Gene Signature and Establishment of a Prognostic Nomogram Predicting Overall Survival for Medulloblastoma. Front Genet 11, 563882, doi: 10.3389/fgene.2020.563882 (2020). Erickson, B. J., Korfiatis, P., Akkus, Z. & Kline, T. L. Machine Learning for Medical Imaging. Radiographics 37, 505–515, doi: 10.1148/rg.2017160130 (2017). Eraslan, G., Avsec, Z., Gagneur, J. & Theis, F. J. Deep learning: new computational modelling techniques for genomics. Nat Rev Genet 20, 389–403, doi: 10.1038/s41576-019-0122-6 (2019). Handelman, G. S. et al. eDoctor: machine learning and the future of medicine. J Intern Med 284, 603–619, doi: 10.1111/joim.12822 (2018). She, Y. et al. Deep learning for predicting major pathological response to neoadjuvant chemoimmunotherapy in non-small cell lung cancer: A multicentre study. EBioMedicine 86, 104364, doi: 10.1016/j.ebiom.2022.104364 (2022). Tran, K. A. et al. Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med 13, 152, doi: 10.1186/s13073-021-00968-x (2021). Foersch, S. et al. Multistain deep learning for prediction of prognosis and therapy response in colorectal cancer. Nat Med 29, 430–439, doi: 10.1038/s41591-022-02134-1 (2023). Huang, B. et al. Deep Learning for the Prediction of the Survival of Midline Diffuse Glioma with an H3K27M Alteration. Brain Sci 13, doi: 10.3390/brainsci13101483 (2023). Zhang, X. et al. Deep learning-based pathology image analysis predicts cancer progression risk in patients with oral leukoplakia. Cancer Med 12, 7508–7518, doi: 10.1002/cam4.5478 (2023). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 24 Jun, 2024 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 25 Mar, 2024 Reviews received at journal 08 Mar, 2024 Reviewers agreed at journal 06 Mar, 2024 Reviewers agreed at journal 06 Mar, 2024 Reviewers invited by journal 06 Mar, 2024 Editor assigned by journal 06 Mar, 2024 Editor invited by journal 05 Mar, 2024 Submission checks completed at journal 05 Mar, 2024 First submitted to journal 21 Feb, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3975955","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":276738385,"identity":"25b86f48-da96-4259-95ec-90b3fc5cd6e5","order_by":0,"name":"Meng Sun","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4ElEQVRIiWNgGAWjYBADOftm5oMPEipqiNdibMDOlmzw4Mwx4rUkbuDnMZN82MJMWKn8jNyjGz7uqGXczsxgVpHYwMbA396dgFeLwY28tJszzxxntmxmSLuRuEOGQeLM2Q34tUjkmN3mbTvGxnCY4diNxDNsQJFc/FrkZwC1/G07xsNwmLGtILGNmbAWhhtALYxtNRIGh5nZGIjSYnDmjdnN3rYDBpLNbMwSCWeO8RD0i3x7jtmNn2119f385z9+/FFRI8ff3kvAYRBwGM7iIUY5CNQRq3AUjIJRMApGIgAANo1LyjsA4c4AAAAASUVORK5CYII=","orcid":"","institution":"The First Affiliated Hospital of Shandong First Medical University","correspondingAuthor":true,"prefix":"","firstName":"Meng","middleName":"","lastName":"Sun","suffix":""},{"id":276738386,"identity":"77fb508a-5cb4-49cf-b446-e9c264cb8de8","order_by":1,"name":"Jikui Sun","email":"","orcid":"","institution":"The First Affiliated Hospital of Shandong First Medical University","correspondingAuthor":false,"prefix":"","firstName":"Jikui","middleName":"","lastName":"Sun","suffix":""},{"id":276738387,"identity":"180784fb-1594-480a-b396-5de0cce6a80f","order_by":2,"name":"Meng Li","email":"","orcid":"","institution":"The First Affiliated Hospital of Shandong First Medical University","correspondingAuthor":false,"prefix":"","firstName":"Meng","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2024-02-21 14:32:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3975955/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3975955/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-024-65367-9","type":"published","date":"2024-06-24T15:32:52+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":52301163,"identity":"d2567619-6e4a-414c-a011-21bcbdf6e5c0","added_by":"auto","created_at":"2024-03-08 18:37:40","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":132275,"visible":true,"origin":"","legend":"\u003cp\u003eStudy profile and analysis pipeline.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/fc1d2980cb442c129bd59095.jpeg"},{"id":52303000,"identity":"1bff1314-eb6a-4334-aafd-c28925dbac59","added_by":"auto","created_at":"2024-03-08 18:53:40","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":117112,"visible":true,"origin":"","legend":"\u003cp\u003eThe X-tile analysis was conducted to determine the best cutoff points for the variables of age and tumor size. (A) X-tile plot of training sets in age. (B) The cutoff point highlighted using a histogram of the entire cohort. (C) Kaplan-Meier plot showing the distinct prognosis determined by the cutoff point. (D) X-tile plot of training sets in tumor size. (E) The cutoff point highlighted using a histogram. (F) Kaplan-Meier plot showing the prognosis determined by the cutoff point. The low subset is depicted in gray, while the high subset is shown in blue.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/cabeef137891d92912a05f06.jpeg"},{"id":52301160,"identity":"7e4ac85b-0b22-462b-8659-67d3a8f3fd1c","added_by":"auto","created_at":"2024-03-08 18:37:40","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":109695,"visible":true,"origin":"","legend":"\u003cp\u003eKaplan-Meier curve of training and testing sets. There was no statistically significant difference between the survival of training and testing sets in log-rank test (\u003cem\u003ep \u003c/em\u003e=0.37).\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/b4284e431f1e0a5e611a026f.jpeg"},{"id":52301164,"identity":"f529b34b-bbb4-48c3-8eba-523ca2b20120","added_by":"auto","created_at":"2024-03-08 18:37:40","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":237024,"visible":true,"origin":"","legend":"\u003cp\u003eUnivariate \u0026amp; Multivariable CoxPH analyses. Variables are sorted in descending order of hazard ratio. *\u003cem\u003ep\u003c/em\u003e<0.05, **\u003cem\u003ep\u003c/em\u003e<0.01, ***\u003cem\u003ep\u003c/em\u003e<0.001\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/69b7ce1ba6839a15035dd95f.jpeg"},{"id":52302293,"identity":"ea5c9774-ab5e-4674-a887-1233e06716f9","added_by":"auto","created_at":"2024-03-08 18:45:40","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":102629,"visible":true,"origin":"","legend":"\u003cp\u003eThe correlation coefficients were calculated for each pair of variables in the dataset. These coefficients represent the strength and direction of the relationship between the variables and range from -1 to +1. The correlation values are displayed using color depth, with values closer to -1 or +1 indicating a stronger negative or positive correlation, respectively.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/a81776ac820655016a8c06ea.jpeg"},{"id":52301167,"identity":"8dfb76a8-696d-44c0-854c-a1b74ba6a3cd","added_by":"auto","created_at":"2024-03-08 18:37:40","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":227211,"visible":true,"origin":"","legend":"\u003cp\u003eRandom survival forest model. 8 features were used to construct the model: Sex, Age, Radiotherapy, Chemotherapy, Histopathology, Race, Surgery, Tumor size. \u003cstrong\u003e(A)\u003c/strong\u003eOut-of-bag (OOB) error rate. \u003cstrong\u003e(B)\u003c/strong\u003e Random survival forest curves. \u003cstrong\u003e(C)\u003c/strong\u003eVariable interaction plot. Higher values of Variable Importance (VIMP) indicate the variable contributes more to predictive accuracy of the model. \u003cstrong\u003e(D)\u003c/strong\u003eVariable interaction plot. Lower values indicate a higher level of interactivity between the variables.\u003c/p\u003e","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/418b1f5dcd67185985fd3766.jpeg"},{"id":52301165,"identity":"7639125d-de33-4cbd-8cdf-ade195dde88d","added_by":"auto","created_at":"2024-03-08 18:37:40","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":184249,"visible":true,"origin":"","legend":"\u003cp\u003eThe training and testing history of DeepSurv. \u003cstrong\u003e(A)\u003c/strong\u003e A plot of loss on the training and testing sets. The error gradually decreases over each iteration during training. \u003cstrong\u003e(B)\u003c/strong\u003e A plot of accuracy that measured as C-index. It is neither fitting the training data too well nor failing to capture important patterns in the data.\u003c/p\u003e","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/50274664d2fb220a5d5fd707.jpeg"},{"id":52301162,"identity":"72f1361b-ba69-44ad-963a-9d1c52a3e9b7","added_by":"auto","created_at":"2024-03-08 18:37:40","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":110298,"visible":true,"origin":"","legend":"\u003cp\u003ePrediction error curve. As a benchmark, a reliable model should aim for a Brier score below 0.25.\u003c/p\u003e","description":"","filename":"floatimage8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/497d182c7cdfd3fe278c4a21.jpeg"},{"id":52302294,"identity":"d8dcf28d-13f0-4f1d-9e41-2dedae86b048","added_by":"auto","created_at":"2024-03-08 18:45:40","extension":"jpeg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":154078,"visible":true,"origin":"","legend":"\u003cp\u003eThe receiver operating characteristic (ROC) curves for 1-year, 3-year, and 5-year survival predictions.\u003c/p\u003e","description":"","filename":"floatimage9.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/3ac4c927742cfca09d82c020.jpeg"},{"id":59202845,"identity":"57d08827-5085-4de2-b485-03b6bc630138","added_by":"auto","created_at":"2024-06-27 15:32:57","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1942270,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3975955/v1/3ad10c94-f317-49a5-a25f-25c43150c7bd.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Deep learning models for predicting the survival of patients with medulloblastoma based on a surveillance, epidemiology, and end results analysis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMedulloblastoma is an embryonal tumor that arises from the cerebellum and has the potential to spread throughout the nervous system. It is the most common type of paediatric embryonal tumor, with an incidence ranging from 5 to 11 cases per 1\u0026nbsp;million individuals\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. According to current international consensus, there are four subgroups of medulloblastoma: Wingless (WNT), Sonic Hedgehog (SHH), group 3 (G3), and group 4 (G4)\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Multimodal therapy, which includes surgery, external beam irradiation, and/or cytotoxic chemotherapy, can result in survival rates ranging from 50\u0026ndash;80% based on clinical staging\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Certain prognostic features, such as age at diagnosis, extent of resection, histological subtype, and molecular subgroup classification, have been found to affect survival predictions in individual patients.\u003c/p\u003e \u003cp\u003ePrevious studies have used the Cox proportional-hazards model (CoxPH) to evaluate the survival rate of medulloblastoma patients\u003csup\u003e5 6,\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. This model incorporates survival outcomes and time as target variables, allowing for the simultaneous analysis of multiple factors' impact on survival time. It is extensively used for predicting outcome events when the survival distribution of the analyzed data is unknown\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. A nomogram is a commonly used method for quantifying and combining important clinical characteristics of patients to calculate the probabilities of outcome events based on the CoxPH model \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. However, the model assumes that each predictor variable has the same effect throughout the follow-up time, which ignores variations in their impact on individual patients at different times. Therefore, a new method is required to improve the accuracy of predicting the survival rate of cancer patients.\u003c/p\u003e \u003cp\u003eIn recent years, computer and information technology have shown revolutionary potential for artificial intelligence (AI) in the healthcare industry \u003csup\u003e\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Machine learning models have stronger nonlinear modeling capabilities compared to traditional linear models and can better capture complex relationships among clinical variables. The analysis of these models can provide accurate personalized survival predictions and decision-making support for treatment strategies to improve patient survival rates \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. Deep learning, a field in machine learning, explores patterns and representations within data to characterize their distribution \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. It is a statistical model that consists of an input layer, hidden layer, and output layer. This model can solve complex, multifactorial, and nonlinear problems. Deep learning-based models have become highly effective predictors of clinical outcomes across various disease domains due to the continuous advancements in deep learning research techniques and the abundance of biomedical big data. Jiang et al. \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e demonstrated the use of an artificial neural network model to predict the survival rate of patients diagnosed with pancreatic neuroendocrine neoplasms, by leveraging clinical information. Katzman et al. \u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e integrated deep learning with a multilayer neural network architecture, known as the DeepSurv model, resulting in a personalized treatment recommendation system that showed remarkable performance.\u003c/p\u003e \u003cp\u003eTo our knowledge, there is a lack of research combining deep learning techniques with the study of medulloblastoma. Therefore, this study aimed to fill this research gap by utilizing data obtained from the Surveillance, Epidemiology, and End Results (SEER) database, which contains information on patients diagnosed with medulloblastoma in the United States. And then the DeepSurv model was used to evaluate their survival rates.\u003c/p\u003e"},{"header":"Method","content":"\u003cp\u003e \u003cb\u003eData source and patient selection\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe data of this retrospective cohort study from the SEER database, which encompasses information from 18 cancer registries representing approximately 28% of the entire US population\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. This database offers extensive and detailed patient data, including demographic characteristics, tumor-related information, cause of death, and survival duration. The SEER*Stat software (version 8.3.6) was used to identify patients with medulloblastoma. The dataset covering the years 2000 to 2019 in the United States was accessed.\u003c/p\u003e \u003cp\u003eThe patients included in the study had to meet the following criteria: 1) a confirmed pathological diagnosis of medulloblastoma; 2) identification of medulloblastoma cases based on the third edition of the International Classification of Diseases for Oncology (ICD-O3) using specific ICD-O-3 codes for histopathology, including 9,470/3 for medulloblastoma, NOS; 9,471/3 for desmoplastic nodular medulloblastoma; and 9,474/3 for large cell medulloblastoma. Furthermore, patients were required to have a known survival status and time. Afterwards, they were randomly divided into a training group and a testing group at a 7:3 ratio. A flowchart in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates the process of patient selection.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eVariable’s definitions\u003c/h3\u003e\n\u003cp\u003eSeveral parameters were collected from the samples, including age at diagnosis, sex, race, histological type, tumor size, surgery, chemotherapy, radiation therapy, and survival time. To evaluate the prognostic value of age and tumor size in patients with medulloblastoma objectively, the patients were categorized into two groups based on the optimal cutoff values obtained using the X-tile software (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://x-tile.software.informer.com\u003c/span\u003e\u003cspan address=\"https://x-tile.software.informer.com\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, Yale School of Medicine, New Haven, CT, United States). Age cutoff values of \u0026le;\u0026thinsp;3 years and \u0026gt;\u0026thinsp;3 years, and tumor size cutoff values of \u0026le;\u0026thinsp;3.4 cm, \u0026gt;\u0026thinsp;3.4 cm, and/or unknown were utilised. For detailed visual representations, please refer to Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eModel development\u003c/h3\u003e\n\u003cp\u003eThis study selected three models for training: DeepSurv, RSF, and CoxPH. DeepSurv is a deep feedforward neural network used to predict patients' survival time or survival probability. It employs a multi-layer neural network to capture the complex nonlinear relationship between patients' survival probability and input features. This study utilized deep-learning calculations based on the DeepSurv calculation method described by Katzman et al.\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e to predict the survival outcome of patients diagnosed with medulloblastoma. The term RSF refers to Random Survival Forests, which is a survival analysis method based on random forests. When constructing a random survival forest, subsets of samples and features are randomly selected, and multiple decision trees are built using these subsets. Each decision tree splits the samples based on features in the nodes and determines the optimal splitting based on the evaluation of survival time differences. The predictions from multiple decision trees in the random survival forests are combined to obtain the final survival prediction. The CoxPH is a semi-parametric regression model used to analyse survival data and estimate the risk of event occurrence. The Cox proportional-hazards model is used to compare the relative risks of events between different groups and study the impact of various factors on event occurrence. The model functions by modeling the relationship between time and event occurrence as a function of hazard ratios.\u003c/p\u003e \u003cp\u003eFor the implementation of the algorithms in this research, CoxPH and RSF were implemented using the R packages \"survival\" and \"randomForestSRC\", respectively. On the other hand, DeepSurv utilized an open-source Python package. The hyperparameter optimization for DeepSurv was conducted using the Hyperopt package within the TensorFlow framework.\u003c/p\u003e\n\u003ch3\u003eModel evaluation\u003c/h3\u003e\n\u003cp\u003eThe study evaluated the model's performance using several metrics, including C-index, Brier score, integrated Brier score (IBS), Receiver Operating Characteristic (ROC) curves, and Area Under the Curve (AUC) values.\u003c/p\u003e \u003cp\u003eThe C-index is a commonly used metric for evaluating the accuracy of survival predictions. It measures the concordance or correlation between the predicted survival risk and the actual observed survival time. A C-index of 0.5 indicates random predictions, while a value of 1.0 indicates perfect predictions. The Brier score assesses the mean squared difference between the observed patient statuses (event occurrence or censoring) and the predicted survival probabilities. It ranges from 0 to 1, with 0 indicating a perfect match between predictions and observations. In practice, models with Brier scores less than 0.25 are considered useful. The IBS is a metric that evaluates the overall performance of a survival model across all available time points. It takes into account the model's sensitivity and specificity to time-dependent events, providing a comprehensive measure of predictive accuracy. Receiver Operating Characteristic (ROC) curves are frequently used to assess a model's sensitivity and specificity at various discrimination thresholds. The ROC curve plots the true positive rate against the false positive rate. The Area Under the Curve (AUC) values, which range from 0 to 1, are computed to quantify the overall performance of the model. A higher AUC indicates better discrimination ability. This study calculated AUC values to assess the model's performance at different time points: 1-year, 3-year, and 5-year survival rates.\u003c/p\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analysis\u003c/h2\u003e \u003cp\u003eIn the clinical data, continuous variables are expressed as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation (SD), while categorical variables are described using frequencies and percentages. Statistical tests such as chi-square tests and unpaired t-tests are used to compare variables between groups.\u003c/p\u003e \u003c/div\u003e"},{"header":"Result","content":"\u003cp\u003e \u003cb\u003eBasic characteristics\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThis study analysed data from 2,322 medulloblastoma patients registered in the SEER database between 2000 and 2019. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents the demographic features of the patients, with 869 cases (37.42%) being female and 1,453 cases (62.58%) being male. The racial distribution was as follows: 185 patients (7.97%) were Black, 1,939 (83.51%) were White, and 198 (8.53%) belonged to other races. Regarding the subtypes of medulloblastoma, 329 patients (14.17%) had desmoplastic/nodular medulloblastoma (DMB), 1,866 (80.36%) had medulloblastoma, not otherwise specified (MB, NOS), and 127 (5.47%) had large-cell/anaplastic medulloblastoma (LC). In terms of surgical interventions, 1,616 patients (69.60%) underwent total resection, 244 (10.51%) underwent subtotal resection, 343 (14.77%) underwent local excision or biopsy, and 119 (5.12%) did not undergo surgery. Of the patients, 1,849 (79.63%) received chemotherapy, 1,766 (76.06%) underwent radiation therapy, and 713 (30.71%) died. The cutoff values for age and tumor size were determined using X-tile analysis \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. Specifically, 324 patients (13.95%) were \u0026le;\u0026thinsp;3 years old, and 1,998 patients (86.05%) were older than 3 years. Regarding tumor size, 314 patients (13.52%) had tumors\u0026thinsp;\u0026le;\u0026thinsp;3.4 cm, 1,269 patients (54.65%) had tumor size\u0026thinsp;\u0026gt;\u0026thinsp;3.4 cm, and the tumor size was unknown for 739 patients (31.83%).\u003c/p\u003e \u003cp\u003eThe predictive model was generated by partitioning the complete dataset into two mutually exclusive subsets. 70% of the dataset was allocated for the training set, while the remaining 30% was used for the testing set. Model generation was performed on 1,625 randomly assigned patients from the training set, while the accuracy of the model was estimated using 697 randomly assigned patients from the validation set. No statistically significant differences in characteristics were found between the two groups \u003cb\u003e(refer to\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. Additionally, survival outcomes showed no differences between the two groups \u003cb\u003e(refer to\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCharacteristic distribution of data into raining sets and test sets.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVariables\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOverall\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTrain cohort\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTest cohort\u003c/p\u003e \u003cp\u003eN (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePatients\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2,322\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1625 (69.98)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e697 (30.02)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003cp\u003e\u0026le;\u0026thinsp;3\u003c/p\u003e \u003cp\u003e\u0026gt;3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e324 (13.95)\u003c/p\u003e \u003cp\u003e1,998 (86.05)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e229 (14.09)\u003c/p\u003e \u003cp\u003e1,396 (85.91)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e95 (13.63)\u003c/p\u003e \u003cp\u003e602 (86.37)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003cp\u003eFemale\u003c/p\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e869 (37.42)\u003c/p\u003e \u003cp\u003e1,453 (62.58)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e625 (38.46)\u003c/p\u003e \u003cp\u003e1,000 (61.54)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e244 (35.01)\u003c/p\u003e \u003cp\u003e453 (64.99)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRace\u003c/p\u003e \u003cp\u003eBlack\u003c/p\u003e \u003cp\u003eWhite\u003c/p\u003e \u003cp\u003eOther\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e185 (7.97)\u003c/p\u003e \u003cp\u003e1,939 (83.51)\u003c/p\u003e \u003cp\u003e198 (8.53)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e135 (8.31)\u003c/p\u003e \u003cp\u003e1,349 (83.02)\u003c/p\u003e \u003cp\u003e141 (8.68)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e50 (7.17)\u003c/p\u003e \u003cp\u003e590 (84.65)\u003c/p\u003e \u003cp\u003e57 (8.18)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.57\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHistopathology\u003c/p\u003e \u003cp\u003eDMB\u003c/p\u003e \u003cp\u003eMB, NOS\u003c/p\u003e \u003cp\u003eLC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e329 (14.17)\u003c/p\u003e \u003cp\u003e1,866 (80.36)\u003c/p\u003e \u003cp\u003e127 (5.47)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e229 (14.09)\u003c/p\u003e \u003cp\u003e1,302 (80.12)\u003c/p\u003e \u003cp\u003e94 (5.78)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e100 (14.35)\u003c/p\u003e \u003cp\u003e564 (80.92)\u003c/p\u003e \u003cp\u003e33 (4.73)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSize (cm)\u003c/p\u003e \u003cp\u003e\u0026le;\u0026thinsp;3.4\u003c/p\u003e \u003cp\u003e\u0026gt;3.4\u003c/p\u003e \u003cp\u003eUnknown\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e314 (13.52)\u003c/p\u003e \u003cp\u003e1,269 (54.65)\u003c/p\u003e \u003cp\u003e739 (31.83)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e211 (12.98)\u003c/p\u003e \u003cp\u003e896 (55.14)\u003c/p\u003e \u003cp\u003e518 (31.88)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e103 (14.78)\u003c/p\u003e \u003cp\u003e373 (53.52)\u003c/p\u003e \u003cp\u003e221 (31.71)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSurgery\u003c/p\u003e \u003cp\u003eTotal resection\u003c/p\u003e \u003cp\u003eSubtotal resection\u003c/p\u003e \u003cp\u003eLocal excision/Biopsy\u003c/p\u003e \u003cp\u003eNo evidence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,616 (69.60)\u003c/p\u003e \u003cp\u003e244 (10.51)\u003c/p\u003e \u003cp\u003e343 (14.77)\u003c/p\u003e \u003cp\u003e119 (5.12)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,122 (69.05)\u003c/p\u003e \u003cp\u003e175 (10.77)\u003c/p\u003e \u003cp\u003e238 (14.65)\u003c/p\u003e \u003cp\u003e90 (5.54)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e494 (70.88)\u003c/p\u003e \u003cp\u003e69 (9.90)\u003c/p\u003e \u003cp\u003e105 (15.06)\u003c/p\u003e \u003cp\u003e29 (4.16)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemotherapy\u003c/p\u003e \u003cp\u003eYes\u003c/p\u003e \u003cp\u003eNo evidence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,849 (79.63)\u003c/p\u003e \u003cp\u003e473 (20.37)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,274 (78.40)\u003c/p\u003e \u003cp\u003e351 (21.60)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e575 (82.50)\u003c/p\u003e \u003cp\u003e122 (17.50)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRadiotherapy\u003c/p\u003e \u003cp\u003eYes\u003c/p\u003e \u003cp\u003eNo evidence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,766 (76.06)\u003c/p\u003e \u003cp\u003e556 (24.94)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,214 (74.71)\u003c/p\u003e \u003cp\u003e411 (25.29)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e552 (79.20)\u003c/p\u003e \u003cp\u003e145 (20.80)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.33\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStatus\u003c/p\u003e \u003cp\u003eDeath\u003c/p\u003e \u003cp\u003eAlive\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e713 (30.71)\u003c/p\u003e \u003cp\u003e1,609 (69.29)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e514 (31.63)\u003c/p\u003e \u003cp\u003e1,111 (68.37)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e199 (28.55)\u003c/p\u003e \u003cp\u003e498 (71.45)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eCox proportional-hazard (CoxPH) model\u003c/h3\u003e\n\u003cp\u003eThe CoxPH model was developed using the training set \u003cb\u003e(refer to\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. Only variables that showed statistical significance in the univariate analysis were included in the multivariate analysis. The survival of medulloblastoma patients was significantly affected by non-surgical treatment, LC, white race, tumor size\u0026thinsp;\u0026le;\u0026thinsp;3.4 cm, total resection, age\u0026thinsp;\u0026gt;\u0026thinsp;3 years, chemotherapy, and radiotherapy. Furthermore, the survival of the patients was significantly associated with these features in the multivariate analysis. The collinearity analysis also revealed a high correlation between age and radiotherapy, as well as between chemotherapy and radiotherapy \u003cb\u003e(refer to\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. Ultimately, we included seven features (age, race, tumor size, histological type, surgery, chemotherapy, and radiotherapy) in the model development.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eRandom Survival Forests (RSF)\u003c/h3\u003e\n\u003cp\u003ePrediction error is calculated using the out-of-bag (OOB) from the training and the testing set \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, B\u003cb\u003e)\u003c/b\u003e. Variable importance (VIMP) can be measured by randomising a specific variable, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eC. A higher VIMP value indicates a greater influence or importance of that variable in accurately predicting the outcome\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. The interaction between variables in the analyzed data is illustrated and displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eD. If one variable's split in a decision tree affects or influences the split of another variable, it suggests an interaction between those variables\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. The extent of interactions is assessed based on the minimum depth, which represents the distance from the root node to the node where the variable first splits. In this case, it is observed that chemotherapy and radiotherapy exhibit the lowest minimum depth among the variables considered were expected to be associated with other variables.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eDeepSurv\u003c/h3\u003e\n\u003cp\u003eThe loss function curve, which illustrates the relationship between the loss and the number of iterations during the training process, provides valuable insights into the convergence and performance of the model. By examining this curve, we can assess how well the model is optimizing its parameters over iterations\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e. Furthermore, plotting the performance of the training dataset as a function of the number of iterations allows us to evaluate the model's ability to rank the samples accurately. This measure helps us monitor the model's generalization ability and identify any signs of overfitting, where the model may excessively capture details specific to the training dataset but fails to generalize well to unseen samples\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. The learning process of DeepSurv, a survival prediction model based on deep learning, was visualized \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. The figure demonstrates a good fit of the model, indicating that it is effectively learning and capturing the patterns within the data.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eModel comparisons\u003c/h3\u003e\n\u003cp\u003eThe predictive performance of the three models is shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. In the test dataset, the DeepSurv and RSF model exhibited significantly better discrimination abilities (the DeepSurv C-index: 0.751, RSF: 0.750) compared with the CoxPH model (the C-index: 0.748). And in the three models, DeepSurv had the highest C-index of 0.751. The IBS for the three models were as follows: DeepSurv (0.150), RSF (0.160), and CoxPH (0.166). Lower IBS values indicate better model performance. Additionally, the C-index obtained from the train data set (DeepSurv: 0.763, RSF: 0.759, CoxPH: 0.757) differed only slightly with test set, indicating that the models did not exhibit overfitting.\u003c/p\u003e \u003cp\u003eFurthermore, in terms of the Brier score \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e, DeepSurv outperformed the other two models, indicating its superior accuracy. The AUC for DeepSurv was also higher than the other models (1-year-AUC of DeepSurv: 0.838, RSF: 0.809, CoxPH: 0.808; 3-year-AUCof DeepSurv: 0.820, RSF: 0.791, CoxPH: 0.782;5-year-AUC of DeepSurv: 0.805, RSF: 0.780, CoxPH: 0.773) \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. These results demonstrate that DeepSurv outperforms both RSF and the classical CoxPH model in accurately predicting the prognosis of patients with medulloblastoma.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of three survival models.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModels\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eC index\u003c/p\u003e \u003cp\u003eTrain Test\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eIBS\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1-year ACU\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e3-year AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5-year AUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCoxPH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.757\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.748\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.166\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.808 (0.77\u0026ndash;0.85)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.782 (0.74\u0026ndash;0.82)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.773 (0.74\u0026ndash;0.81)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRSF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.759\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.750\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.160\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.809 (0.77\u0026ndash;0.85)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.791 (0.75\u0026ndash;0.83)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.780 (0.74\u0026ndash;0.82)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeepSurv\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.763\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.751\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.150\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.838\u003c/b\u003e (0.80\u0026ndash;0.88)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.820\u003c/b\u003e (0.78\u0026ndash;0.86)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.805\u003c/b\u003e (0.76\u0026ndash;0.84)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eMedulloblastoma, a malignant brain tumor that mainly impacts children, continues to pose a substantial obstacle in the field of pediatric oncology. Precisely predicting the individual prognosis of patients is crucial for customizing treatment approaches and enhancing survival rates. Prior research has identified several prognostic factors that affect the survival duration of medulloblastoma patients, including age, extent of surgical removal, and the administration of radiotherapy or chemotherapy \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e,\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. Moreover, as medical advancements progress, an increasing amount of imaging data \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e and genetic data \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e are being analyzed for survival analysis of medulloblastoma patients. However, classical survival analysis methods, such as the Cox proportional-hazards model, assume a linear relationship between variables, which may be limited in the face of multidimensional data. With the advancement of artificial intelligence, machine learning methods are being applied to clinical, imaging, and genetic data, allowing for the discovery of potential nonlinear relationships within the data \u003csup\u003e\u003cspan additionalcitationids=\"CR29\" citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. Within machine learning, deep learning is a specific class of methods that utilizes multilayered neural networks to extract high-order features. Deep learning has gained increasing popularity in the field of cancer survival analysis, and has demonstrated excellent performance \u003csup\u003e\u003cspan additionalcitationids=\"CR32\" citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. As far as we know, this approach has not been applied to medulloblastoma. Therefore, we developed a deep learning model (Deepsurv) to predict the overall survival (OS) of medulloblastoma patients and compared its performance to that of a machine learning model (RSF) and a classical model (CoxPH).\u003c/p\u003e \u003cp\u003eBy extracting potentially significant features from the SEER database, this research developed multiple models to forecast the survival rates of individuals diagnosed with medulloblastoma. Initially, we utilized the X-tile tool to determine the optimal cutoff values for age and tumor size from a cohort of 2,322 medulloblastoma patients. We identified two high-risk factors, age\u0026thinsp;\u0026le;\u0026thinsp;3 years old and tumor size\u0026thinsp;\u0026gt;\u0026thinsp;3.4 cm, that significantly impact the survival duration of patients with medulloblastoma. Subsequently, we employed Cox proportional hazards regression to identify variables associated with the prognosis of medulloblastoma patients. Age, race, tumor size, histological type, surgery, chemotherapy, and radiotherapy were selected for inclusion in the modeling process (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). We established RSF, DeepSurv and CoxPH models and evaluated their performance using metrics such as the C-index, IBS, and ROC curve. The study results demonstrated that the DeepSurv model outperformed both the CoxPH and RSF models, as indicated by its higher C-index in both the training and testing sets. Moreover, the DeepSurv model exhibited the lowest IBS and the largest AUC values when predicting 1-, 3-, and 5-year survival. These findings collectively suggest that the DeepSurv model is more accurate in predicting the survival of patients with medulloblastoma.\u003c/p\u003e \u003cp\u003eIn previous studies, Guo et al. \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e and Zhou et al. \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e utilized Cox proportional hazard regression for survival analysis of medulloblastoma and developed a nomogram. Compared with their study, the C-index values obtained from the DeepSurv model were higher in both the training cohort, indicating its superior predictive accuracy of the prognosis of patients with medulloblastoma. This finding is consistent with the results reported in several previous studies focusing on cancer prognosis\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e,\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e. The main advantage of the DeepSurv model in its ability to process both linear and nonlinear predictive variables by utilizing a multilevel neural network. This transformation into a linear combination allows the model to uncover associations that may not be readily apparent to the human eye or traditional statistical techniques.\u003c/p\u003e \u003cp\u003eNevertheless, our study encountered several limitations. Firstly, the data collected from the SEER database for medulloblastoma patients contained some missing information that could potentially influence survival outcomes. This includes important details such as molecular subgroups, specific radiotherapy dosages, and chemotherapy regimens. The availability and completeness of these data rely on the ongoing improvements in data collection within the SEER database. Secondly, our model has yet to undergo external validation, and it is necessary to validate its performance on new data. Conducting further validations using independent datasets would enhance the reliability and generalizability of the findings. Another inherent limitation lies within the DeepSurv model itself. Due to its utilization of hidden layers in its architecture, the model operates as a black-box, making it challenging to fully comprehend the computations involved in the model construction process and its associated limitations. Future research should aim to address these concerns and explore the inner workings of the model to improve interpretability.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis study employed Cox proportional hazards regression analysis to examine the prognostic factors influencing medulloblastoma patients' outcomes, which include age, race, tumor size, histological type, surgery, chemotherapy, and radiotherapy. Subsequently, we developed a groundbreaking DeepSurv prediction model, which exhibited strong predictive capabilities in assessing the prognosis of patients diagnosed with medulloblastoma. This innovative DeepSurv model holds significant potential in accurately predicting the survival duration of medulloblastoma patients.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that no funds, grants, or other support were received during the preparation of this manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors have no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSun M: conceptualization, data curation, investigation, methodology, software, visualization, writing\u0026mdash;original draft; Sun J: conceptualization, data curation, formal analysis; Li M: methodology, project administration, supervision\u0026mdash;review and editing. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets analyzed during the current study are available in the SEER database repository (https://seer.cancer.gov/).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGajjar, A. J. \u0026amp; Robinson, G. W. Medulloblastoma-translating discoveries from the bench to the bedside. Nat Rev Clin Oncol 11, 714\u0026ndash;722, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nrclinonc.2014.181\u003c/span\u003e\u003cspan address=\"10.1038/nrclinonc.2014.181\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOstrom, Q. T., Cioffi, G., Waite, K., Kruchko, C. \u0026amp; Barnholtz-Sloan, J. S. CBTRUS Statistical Report: Primary Brain and Other Central Nervous System Tumors Diagnosed in the United States in 2014\u0026ndash;2018. Neuro Oncol 23, iii1-iii105, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/neuonc/noab200\u003c/span\u003e\u003cspan address=\"10.1093/neuonc/noab200\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTaylor, M. D. \u003cem\u003eet al.\u003c/em\u003e Molecular subgroups of medulloblastoma: the current consensus. Acta Neuropathol 123, 465\u0026ndash;472, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00401-011-0922-z\u003c/span\u003e\u003cspan address=\"10.1007/s00401-011-0922-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamaswamy, V. \u0026amp; Taylor, M. D. Medulloblastoma: From Myth to Molecular. J Clin Oncol 35, 2355\u0026ndash;2363, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1200/JCO.2017.72.7842\u003c/span\u003e\u003cspan address=\"10.1200/JCO.2017.72.7842\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, L. \u003cem\u003eet al.\u003c/em\u003e Automatic image segmentation and online survival prediction model of medulloblastoma based on machine learning. Eur Radiol, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00330-023-10316-9\u003c/span\u003e\u003cspan address=\"10.1007/s00330-023-10316-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, X. \u0026amp; Gong, J. Survival nomogram for medulloblastoma and multi-center external validation cohort. Front Pharmacol 14, 1247812, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fphar.2023.1247812\u003c/span\u003e\u003cspan address=\"10.3389/fphar.2023.1247812\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, C. \u003cem\u003eet al.\u003c/em\u003e External Validation of a Nomogram and Risk Grouping System for Predicting Individual Prognosis of Patients With Medulloblastoma. Front Pharmacol 11, 590348, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fphar.2020.590348\u003c/span\u003e\u003cspan address=\"10.3389/fphar.2020.590348\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaek, E. T. \u003cem\u003eet al.\u003c/em\u003e Survival time prediction by integrating cox proportional hazards network and distribution function network. BMC Bioinformatics 22, 192, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12859-021-04103-w\u003c/span\u003e\u003cspan address=\"10.1186/s12859-021-04103-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIasonos, A., Schrag, D., Raj, G. V. \u0026amp; Panageas, K. S. How to build and interpret a nomogram for cancer prognosis. J Clin Oncol 26, 1364\u0026ndash;1370, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1200/JCO.2007.12.9791\u003c/span\u003e\u003cspan address=\"10.1200/JCO.2007.12.9791\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchwalbe, N. \u0026amp; Wahl, B. Artificial intelligence and the future of global health. Lancet 395, 1579\u0026ndash;1586, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S0140-6736(20)30226-9\u003c/span\u003e\u003cspan address=\"10.1016/S0140-6736(20)30226-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHamet, P. \u0026amp; Tremblay, J. Artificial intelligence in medicine. Metabolism 69S, S36-S40, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.metabol.2017.01.011\u003c/span\u003e\u003cspan address=\"10.1016/j.metabol.2017.01.011\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHunter, D. J. \u0026amp; Holmes, C. Where Medical Statistics Meets Artificial Intelligence. N Engl J Med 389, 1211\u0026ndash;1219, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1056/NEJMra2212850\u003c/span\u003e\u003cspan address=\"10.1056/NEJMra2212850\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eConnor, C. W. Artificial Intelligence and Machine Learning in Anesthesiology. Anesthesiology 131, 1346\u0026ndash;1359, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/ALN.0000000000002694\u003c/span\u003e\u003cspan address=\"10.1097/ALN.0000000000002694\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBhat, M., Rabindranath, M., Chara, B. S. \u0026amp; Simonetto, D. A. Artificial intelligence, machine learning, and deep learning in liver transplantation. J Hepatol 78, 1216\u0026ndash;1233, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jhep.2023.01.006\u003c/span\u003e\u003cspan address=\"10.1016/j.jhep.2023.01.006\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChoi, R. Y., Coyner, A. S., Kalpathy-Cramer, J., Chiang, M. F. \u0026amp; Campbell, J. P. Introduction to Machine Learning, Neural Networks, and Deep Learning. Transl Vis Sci Technol 9, 14, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1167/tvst.9.2.14\u003c/span\u003e\u003cspan address=\"10.1167/tvst.9.2.14\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGreener, J. G., Kandathil, S. M., Moffat, L. \u0026amp; Jones, D. T. A guide to machine learning for biologists. Nat Rev Mol Cell Biol 23, 40\u0026ndash;55, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41580-021-00407-0\u003c/span\u003e\u003cspan address=\"10.1038/s41580-021-00407-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang, C. \u003cem\u003eet al.\u003c/em\u003e Predicting the survival of patients with pancreatic neuroendocrine neoplasms using deep learning: A study based on Surveillance, Epidemiology, and End Results database. Cancer Med 12, 12413\u0026ndash;12424, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/cam4.5949\u003c/span\u003e\u003cspan address=\"10.1002/cam4.5949\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatzman, J. L. \u003cem\u003eet al.\u003c/em\u003e DeepSurv: personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Med Res Methodol 18, 24, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12874-018-0482-1\u003c/span\u003e\u003cspan address=\"10.1186/s12874-018-0482-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHankey, B. F., Ries, L. A. \u0026amp; Edwards, B. K. The surveillance, epidemiology, and end results program: a national resource. Cancer Epidemiol Biomarkers Prev 8, 1117\u0026ndash;1121 (1999).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTaylor, J. M. Random Survival Forests. J Thorac Oncol 6, 1974\u0026ndash;1975, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/JTO.0b013e318233d835\u003c/span\u003e\u003cspan address=\"10.1097/JTO.0b013e318233d835\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGilhodes, J. \u003cem\u003eet al.\u003c/em\u003e Comparison of variable selection methods for high-dimensional survival data with competing events. Comput Biol Med 91, 159\u0026ndash;167, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.compbiomed.2017.10.021\u003c/span\u003e\u003cspan address=\"10.1016/j.compbiomed.2017.10.021\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKretowska, M. Tree-based models for survival data with competing risks. Comput Methods Programs Biomed 159, 185\u0026ndash;198, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cmpb.2018.03.017\u003c/span\u003e\u003cspan address=\"10.1016/j.cmpb.2018.03.017\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu, J., Zhou, Y., Liu, P., Vong, C. M. \u0026amp; Wang, T. Parameter-Free Loss for Class-Imbalanced Deep Learning in Image Classification. IEEE Trans Neural Netw Learn Syst 34, 3234\u0026ndash;3240, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TNNLS.2021.3110885\u003c/span\u003e\u003cspan address=\"10.1109/TNNLS.2021.3110885\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSerghiou, S. \u0026amp; Rough, K. Deep Learning for Epidemiologists: An Introduction to Neural Networks. Am J Epidemiol 192, 1904\u0026ndash;1916, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/aje/kwad107\u003c/span\u003e\u003cspan address=\"10.1093/aje/kwad107\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDasgupta, A. \u003cem\u003eet al.\u003c/em\u003e Nomograms based on preoperative multiparametric magnetic resonance imaging for prediction of molecular subgrouping in medulloblastoma: results from a radiogenomics study of 111 patients. Neuro Oncol 21, 115\u0026ndash;124, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/neuonc/noy093\u003c/span\u003e\u003cspan address=\"10.1093/neuonc/noy093\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, H. \u0026amp; Sun, P. A Nomogram Model for Predicting Prognosis of Patients with Medulloblastoma. Turk Neurosurg 34, 38\u0026ndash;45, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5137/1019-5149.JTN.40397-22.3\u003c/span\u003e\u003cspan address=\"10.5137/1019-5149.JTN.40397-22.3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu, S. \u003cem\u003eet al.\u003c/em\u003e Identification of a Twelve-Gene Signature and Establishment of a Prognostic Nomogram Predicting Overall Survival for Medulloblastoma. Front Genet 11, 563882, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fgene.2020.563882\u003c/span\u003e\u003cspan address=\"10.3389/fgene.2020.563882\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eErickson, B. J., Korfiatis, P., Akkus, Z. \u0026amp; Kline, T. L. Machine Learning for Medical Imaging. Radiographics 37, 505\u0026ndash;515, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1148/rg.2017160130\u003c/span\u003e\u003cspan address=\"10.1148/rg.2017160130\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEraslan, G., Avsec, Z., Gagneur, J. \u0026amp; Theis, F. J. Deep learning: new computational modelling techniques for genomics. Nat Rev Genet 20, 389\u0026ndash;403, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41576-019-0122-6\u003c/span\u003e\u003cspan address=\"10.1038/s41576-019-0122-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHandelman, G. S. \u003cem\u003eet al.\u003c/em\u003e eDoctor: machine learning and the future of medicine. J Intern Med 284, 603\u0026ndash;619, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/joim.12822\u003c/span\u003e\u003cspan address=\"10.1111/joim.12822\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShe, Y. \u003cem\u003eet al.\u003c/em\u003e Deep learning for predicting major pathological response to neoadjuvant chemoimmunotherapy in non-small cell lung cancer: A multicentre study. EBioMedicine 86, 104364, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ebiom.2022.104364\u003c/span\u003e\u003cspan address=\"10.1016/j.ebiom.2022.104364\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTran, K. A. \u003cem\u003eet al.\u003c/em\u003e Deep learning in cancer diagnosis, prognosis and treatment selection. Genome Med 13, 152, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s13073-021-00968-x\u003c/span\u003e\u003cspan address=\"10.1186/s13073-021-00968-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFoersch, S. \u003cem\u003eet al.\u003c/em\u003e Multistain deep learning for prediction of prognosis and therapy response in colorectal cancer. Nat Med 29, 430\u0026ndash;439, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41591-022-02134-1\u003c/span\u003e\u003cspan address=\"10.1038/s41591-022-02134-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang, B. \u003cem\u003eet al.\u003c/em\u003e Deep Learning for the Prediction of the Survival of Midline Diffuse Glioma with an H3K27M Alteration. Brain Sci 13, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/brainsci13101483\u003c/span\u003e\u003cspan address=\"10.3390/brainsci13101483\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, X. \u003cem\u003eet al.\u003c/em\u003e Deep learning-based pathology image analysis predicts cancer progression risk in patients with oral leukoplakia. Cancer Med 12, 7508\u0026ndash;7518, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/cam4.5478\u003c/span\u003e\u003cspan address=\"10.1002/cam4.5478\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"DeepSurv, Medulloblastoma, Neural network, Survival prediction, SEER","lastPublishedDoi":"10.21203/rs.3.rs-3975955/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3975955/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eMedulloblastoma is a malignant neuroepithelial tumor of the central nervous system. Accurate prediction of prognosis is essential for therapeutic decisions in medulloblastoma patients. Several prognostic models have been developed using multivariate Cox regression to predict the1-, 3- and 5-year survival of medulloblastoma patients, but few studies have investigated the results of integrating deep learning algorithms. Compared to simplifying predictions into binary classification tasks, modelling the probability of an event as a function of time by combining it with deep learning may provide greater accuracy and flexibility.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003ePatients diagnosed with medulloblastoma between 2000 and 2019 were extracted from the Surveillance, Epidemiology, and End Results (SEER) registry. Three models\u0026mdash;one based on neural networks (DeepSurv), one based on ensemble learning (random survival forest [RSF]), and a typical Cox proportional-hazards (CoxPH) model\u0026mdash;were selected for training. The dataset was randomly divided into training and testing datasets in a 7:3 ratio. The model performance was evaluated utilizing the concordance index (C-index), Brier score and integrated Brier score (IBS). The accuracy of predicting 1-, 3- and 5- year survival was assessed using receiver operating characteristic curves (ROC), and the area under the ROC curves (AUC).\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThe 2,322 patients with medulloblastoma enrolled in the study were randomly divided into the training cohort (70%, n\u0026thinsp;=\u0026thinsp;1,625) and the test cohort (30%, n\u0026thinsp;=\u0026thinsp;697). There was no statistically significant difference in clinical characteristics between the two cohorts (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.05). We performed Cox proportional hazards regression on the data from the training cohort, which illustrated that age, race, tumour size, histological type, surgery, chemotherapy, and radiotherapy were significant factors influencing survival (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). The Deepsurv outperformed the RSF and classic CoxPH models with C-indexes of 0.763 and 0.751 for the training and test datasets. The DeepSurv model showed better accuracy in predicting 1-, 3- and 5-year survival (AUC: 0.805\u0026ndash;0.838).\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThe predictive model based on a deep learning algorithm that we have developed can exactly predict the survival rate and duration of medulloblastoma.\u003c/p\u003e","manuscriptTitle":"Deep learning models for predicting the survival of patients with medulloblastoma based on a surveillance, epidemiology, and end results analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-03-08 18:37:35","doi":"10.21203/rs.3.rs-3975955/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-03-25T11:32:44+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-03-08T14:36:10+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"39477ca6-f92e-4c87-b5e9-fdf97b1318a6_SNPRID","date":"2024-03-06T09:32:45+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"e7e74c89-0890-433e-84b7-ed7b9223c4d3","date":"2024-03-06T08:58:05+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-03-06T07:16:18+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-03-06T07:12:42+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-03-05T20:59:41+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-03-05T18:39:29+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2024-02-21T14:28:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4a19ee07-1c52-4db6-a7d1-a238a8746b15","owner":[],"postedDate":"March 8th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":29167544,"name":"Biological sciences/Cancer/Cancer models"},{"id":29167545,"name":"Biological sciences/Cancer/Cns cancer"}],"tags":[],"updatedAt":"2024-06-27T15:32:52+00:00","versionOfRecord":{"articleIdentity":"rs-3975955","link":"https://doi.org/10.1038/s41598-024-65367-9","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2024-06-24 15:32:52","publishedOnDateReadable":"June 24th, 2024"},"versionCreatedAt":"2024-03-08 18:37:35","video":"","vorDoi":"10.1038/s41598-024-65367-9","vorDoiUrl":"https://doi.org/10.1038/s41598-024-65367-9","workflowStages":[]},"version":"v1","identity":"rs-3975955","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3975955","identity":"rs-3975955","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.