Investigation of genes related to oral cancer using time-to-event machine learning approaches | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Investigation of genes related to oral cancer using time-to-event machine learning approaches Niusha Shekari, Payam Amini, Leili Tapak, Mahboobeh Rasouli This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2985174/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Since cancer is one of the most common and deadly diseases, its early diagnosis is very important for treatment and prevents the irreparable physical, mental and social consequences of this disease. Oral cancer is also one of the most common cancers, and factors such as gender, age, and smoking influence the incidence of this disease. One of the most important factors affecting cancer is genetic factors. It is not enough to consider clinical factors for the treatment of this disease, and it is also very important to deal with the genes in people's bodies that are effective in their survival against cancer. Also, the survival of people with oral cancer in the early stages of the disease is 80%, so early detection is very important. Therefore, we are looking for a model to better investigate key and effective genes in this disease. Methods: A publicly available dataset of oral cancer (GSE26549) including information of 29096 genes expression profiles of 86 samples was used. A univariate cox regression was used for each gene’s expression to reduce the number of genes. Cox-Boost, Random Survival Forest and Support survival SVM (Recursive Feature Elimination) were used to identify related genes. Shared genes between three methods were discovered for calculating the prognostic score and the Kaplan-Meier curve. To do validation, common genes were selected from the validation dataset (GSE9844) to provide the ROC curve. Results: The univariate Cox regression models selected 945 significant genes. Four shared genes of RPL24, HTR3B, ASAH2B and TEX29 related to time-to-death in oral cancer patients were then identified by using the Cox-Boost, Random Survival Forest and Support survival SVM (Recursive Feature Elimination). The survival distributions of the high-risk and low-risk groups significantly differed. Conclusion: Common genes between three methods were RPL24, HTR3B, ASAH2B and TEX29 which all of them were significant in multiple Cox. Biological sciences/Genetics Biological sciences/Genetics/Gene expression Biological sciences/Computational biology and bioinformatics/High throughput screening Biological sciences/Computational biology and bioinformatics/Machine learning Biological sciences/Computational biology and bioinformatics/Statistical methods Biological sciences/Cancer Biological sciences/Cancer/Oral cancer Oral cancer Gene expression profiling Random Survival Forest Cox-Boost Survival Support Vector Machines Machine Learning Figures Figure 1 Figure 2 Introduction Around the world, oral cancer has become a significant problem for public health. It is indicated that between 1990 and 2017, there was a nearly one-fold increase in the global incidence, mortality, and disability-adjusted life years of this disease. According to Global Cancer Observatory estimates of incidence and mortality, lip and oral cavity cancer will have 377,713 new cases and 177,757 fatalities in 2020. The majority of oral cancers are squamous cell carcinomas, which are aggressive cancers that frequently spread both locally and to distant sites. It has a significant impact on both the patient's life and society as a whole 1 .OSCC patients' 5-year survival rates range from 63% (for female patients) to 47% (for male patients). The high rate of OSCC recurrence and metastases, as well as the disease's delayed diagnosis, are all associated with mortality. Only one-third of OSCCs are discovered at an early stage (0–I). Therefore, there is great interest in the development of tests that increase our capacity to screen high-risk (e.g., heavy alcohol and tobacco use) and post-therapy patients 2 . Understanding the genetic heterogeneity of cancers can help develop effective biomarkers, raise diagnostic precision, and improve treatment outcomes. The development of next-generation sequencing technologies with high efficiency and accuracy has resulted in the generation of a huge amount of cancer tissue genomic data, the majority of which is stored in the Cancer Genome Atlas (TCGA) database. Studies have looked for a diagnostic model that can quickly and effectively distinguish between different tumors by simultaneously taking into account both genetic and phenotypic features using RNA sequencing expression data 3 .Using the Oncomine database, it was determined that the enhanced expression of CXCL8, DDX60, IFI44L, RSAD2, and RTP44 in oral cancer. The Human Protein Atlas database revealed that DDX60, IFI44L, RSAD2, and RTP44 protein expression levels in tumor tissues were higher than those in normal tissues 4 . Due to the difficulty of analyzing and interpreting genomic datasets, various machine learning methods are used. To this purpose, it has been demonstrated that machine learning techniques (shallow learning) offer improved OSCC prognostication. Notably, it has been claimed that using machine learning can predict outcomes more accurately than using traditional statistical methods. Due to the ability to identify the intricate correlations between the variables in the dataset, machine learning algorithms have demonstrated promising results. Machine learning approaches have garnered a lot of attention recently because of their alleged viability and advantages in the field of cancer prognostication 5 . For early diagnosis of the disease and prevention of irreparable physical, mental and social complications of this type of cancer, in this paper, we aim to use Cox-Boost, Random Survival Forest and Survival Support Vector Machines approaches to investigate related genes with oral cancer. Material and Methods Data : Training dataset (GSE26549) and validation dataset (GSE9844) are available on https://www.ncbi.nlm.nih.gov/geo . Statistical Analysis: The statistical analysis was performed using R programing language version 4.2.0. To begin, we filtered genes by fitting a univariate cox model to 29096 genes expression and extracted significant genes that had a p-value less than 0.001. The genes were reduced to 945. “CoxBoost”, “randomForestSRC” and “sigfeature” libraries were used in R. Each of models suggested many genes but our main goal was to discover common genes between three methods. In the next step, we fitted a multiple cox model to common genes and obtained the prognostic score. Prognostic score was stratified into: low-risk & high-risk groups then Kaplan-Meier curves were generated for the two risk groups for the time-to-event related death. Finally, to do validation, common genes were utilized for prediction probability and ROC curve. Cox PH Model : The Cox regression method is a statistical technique that is frequently used in medical research to forecast the length of survival for various patients. The Cox Regression approach is used to estimate the hazard rate, which is the degree to which certain characteristics affect survival. Semi-parametric models are exemplified by the Cox regression method. The Cox model can be expressed using the hazard function \(h\left(t\right)\) . Shortly, the hazard function gives the likelihood of dying at time t. The estimation is as follows: $$h\left(t\right)={h}_{0}\left(t\right)\times \text{e}\text{x}\text{p}({b}_{1}{x}_{1}+{b}_{2}{x}_{2}+\dots +{b}_{n}{x}_{n})$$ where \("t"\) represents the survival time. In order to calculate the hazard function \(h\left(t\right)\) ," \(n covariates\) " \(\left({x}_{1},{x}_{2},\dots ,{x}_{n}\right)\) are used. With the help of the coefficients \(\left({b}_{1},{b}_{2},\dots ,{b}_{n}\right)\) , the influence of variables is calculated. The baseline hazard is \("h0".\) 6 Cox Boost : These steps are used in the boosting procedure: 1) Set the vector of regression coefficients as 0. 2) Compute the negative gradient vector, \(u=\frac{{\delta }{L}({y},{F}\left({X},{\beta }\right))}{{\delta }{F}({X},{\beta })}|{\beta }=\widehat{\beta }\) 3) Compute the updates: 3.1- Fit the base learner to the negative gradient vector, \(\widehat{h}\left(u,{X}_{j}\right)\) 3.2- Penalize it, \({\widehat{b}}_{j}=v\widehat{h}\left(u,{X}_{j}\right)\) 4) 4- Select the best update j∗ (usually that minimizing the loss function). 5) 5- Update the estimations \({\widehat{\beta }}_{{j}^{*}}={\widehat{\beta }}_{{j}^{*}}+{\widehat{b}}_{{j}^{* }}\) ( \(\widehat{b}\) is called a weak estimator) The stages between 2 and 5 must be performed \({m}_{stops}\) times, where \({m}_{stops}\) represents the number of boosting iterations. A likelihood function is the foundation of the estimation strategy 7 . (Riccardo De Bin provides excellent details on this strategy.) Random Survival Forest : Although clinical professionals may find statistical techniques like classification and regression trees intuitive, they have high variance and poor performance. These are dealt with by random forest, which creates a large number of trees and produces the results through voting. Utilizing all variables gathered and automatically evaluating nonlinear effects and intricate interactions, RSF lowers variance and bias. This approach is fully non-parametric, including the effects of the treatments and predictor variables, whereas traditional methods such as CPH utilize a linear combination of attributes. The “RandomForestSRC” R package was used to train the random survival forest models 8 . Survival Support Vector Machine-Recursive Feature Elimination (SVM-RFE) : As supervised learning techniques, SVMs are typically employed for classification and regression analysis. Support vector machines can also be used to pick out specific features from a batch of data (e.g., SVM-RFE). With some modifications to get the necessary weight values for each feature, we used SVM in "sigFeature" for feature selection. Guyon et al. (2002) presented the SVM-RFE feature selection method for the classification of cancer in 2002. A weight-based approach is the "SVM-RFE." Each stage uses the linear SVM's weight vector coefficients as a feature ranking criterion. The characteristics with the highest weight are the most illuminating. As a result, "SVMRFE" employs a sequential backward feature elimination method to choose the feature with the lowest weight before storing it in a stack. The procedure of iteration is carried out until just one feature variable is left. Both the "SVM-RFE" methodology and our recently developed feature selection algorithm are part of the Wrapper method and choose the feature by recursively removing other features 9 . Significant Feature Selection (sigFeature) : The SVM-RFE package offers a brand-new novel feature selection approach for binary classification that makes use of the t-statistic and support vector machines. The chosen features are differently important across the two classes in this feature selection procedure, and they are also strong classifiers with a greater level of classification accuracy 10 . Prognostic Score : In this study, a prognostic score for each patient is calculated using multiple cox regression coefficients and genes expression (shared genes): $$Prognostic Score=\sum {\beta }_{i}\times {X}_{i} i=1,\dots ,p$$ A higher score was thought to be related to a higher risk of oral cancer. We used the median prognostic score as the cutoff point: Group 1: High-Risk Group: If prognostic score for a patient > Median: patient is at a higher risk of developing cancer. Group 2: Low-Risk Group: If prognostic score for a patient < Median: patient is at a lower risk of developing cancer. Then, the survival distribution of these two groups of patients is compared using the Kaplan-Meier plot. Moreover, we calculated Hazard Ratios for each genes using multiple cox regression. In the final stage, to measure the validity of the results, we find the shared genes from the validation dataset and prepared ROC curve for each of them. Results We analyzed the information of 29096 gene expression of 86 individuals. Dataset is available in GEO repository (GSE26549). After fitting the univariate Cox model, we reached 945 significant genes with p-values less than 0.001. The Cox-Boost, Random Survival Forest and Survival Support Vector Machine methods applied to these 945 genes expression. The AUC for each of these methods obtained in order of 0.921, 0.894 and 0.989. Table 1 shows the multiple cox result with shared genes among these three methods: Table 1 The results of multiple cox regression assessing the impact of shared genes from the RF, SSVM, and Cox-Boost Gene HR (95% Confidence Interval) P-value RPL24 0.0005 (0.0001–0.0265) < 0.001 HTR3B 5.4272 (1.0450- 28.1869) 0.044 ASAH2B 0.0215 (0.0026–0.1763) < 0.001 TEX29 0.2031 (0.0451–0.9150) 0.037 As we can see all shared genes are significant and HTR3B expression are associated with higher hazard ratio and it has higher risk for oral cancer compared with three other genes. We calculated prognostic score with these shared genes and compared it with its median then divided it into two parts: (Prognostic Score > Median: High Risk Group & Prognostic Score < Median: Low Risk Group). Figure 1 shows the Kapelan-Meier curve: As a result, high-risk group has lower survival probability than low risk group. In addition, their survival probability decreases during the time. It can be said that survival distributions for the different levels of risk groups is different. The publicly available dataset from the GEO repository (GSE9844) was applied for validation. we extracted only shared genes from this dataset for obtaining prediction probability. Table 2 shows the AUC for each gene and its 95% confidence interval: Table 2 The area under curve (95% confidence interval) based on the prediction probability of ASAH2B, RPL24, HTR3B, and RNU2_22P for oral cancer occurrence using the Validation data set Genes Area under curve ASAH2B 0.822 (0.667–0.976) RPL24 0.414 (0.208–0.621) HTR3B 0.576 (0.368–0.784) RNU2_22P 0.582 (0.401–0.764) ASAH2B has the highest AUC among the other genes. It means that it has more predictive potency for oral cancer. Figure 2 shows Roc curve for each gene based on prediction probability for oral cancer: Conclusion In conclusion, four key genes including RPL24, HTR3B, ASAH2B and TEX29 were identified for further study. Survival support vector machine was the best model for finding related genes to oral cancer. Common genes between three methods were: RPL24, HTR3B, ASAH2B and TEX29 which all of them were significant in multiple cox and survival distributions for the different levels of risk groups was different. By increasing in the expression of HTR3B, the risk of oral cancer increases 5 times. By using four genes mentioned above from validation dataset, we saw all of them had a high AUC specifically, its value for ASAH2B was more than other genes. Discussion In this paper, we developed "sigFeature" that uses "SVM-RFE" and the t statistic to choose significant features. Furthermore, we compared the "sigFeature" technique to other algorithms (such as Cox-Boost and Random Survival Forest) using AUC. The results shown that the sigFeature method to choose important feature has more accuracy than two other methods. Other machine learning methods, such as neural networks, can also be used to find significant genes with greater accuracy and potency. The ceramidase enzyme encoded with N-acylsphingosine amidohydrolase 2B( ASAH2B ) has a role in process such as sphingosine biosynthetic and ceramide catabolic 11 ( https://www.ncbi.nlm.nih.gov/gene ). Li et al in their study showed that this gene plays an essential role in regulation of Breast Cancer Cell Growth through regulating the signaling pathway of mTOR 12 . The ribosomal protein encoded with ribosomal protein L24 ( RPL24 ) belongs to L24E family of ribosomal proteins and is a component of the large ribosomal subunit. RPL24 was Dysregulated in breast cancer and has a role in the regulation of tumorigenesis 13 . Mutation of ribosomal protein encoding genes associated with cancer development 14 . The protein encoded by 5-hydroxytryptamine receptor 3B ( HTR3B ) as a member of ligand-gated ion channel receptor superfamily is a subunit B of the 5-HT3 receptor 15 . This gene has a role in regulation of inflammasomes and immune response and is associated with prognosis of brain tumors 16 . The Cox model is perhaps the most widely used statistical tool for analyzing survival data, despite the restrictions imposed by the proportional hazards assumption, because of its adaptability and simplicity of interpretation. Due to this, novel statistical/machine learning techniques are typically adjusted to match it, such as boosting, an iterative technique that was initially created in the machine learning community and later extended to the statistical sector. The availability of user-friendly software, such as the R package CoxBoost, which permits the use of boosting in conjunction with the Cox model, has further contributed to the popularity of boosting. One of the best techniques, particularly in classification problems, is boosting, which is one of the ensemble methods that connects a strong learner to a group of several weak learners. Because of this, this method has been expanded in the field of statistics to include regression methods and survival analysis due to its useful and accurate applications and performance. This method's sequential learning process differs from the bagging method's lack of use of independent learners. The main goal of boosting is to minimize the predefined loss function by improving the predictors in a sequential manner, each iteration including the weak predictors from the previous stage. Its simplicity of usage is a key benefit of this model. The algorithm automatically determines how many boosting steps to take and avoids overfitting as a result. This is likely one of the factors contributing to its high performance, and it also means that the user need not set any hyperparameters, making it simple to create a model with high performance 17 . In addition, it is possible to avoid penalizing some variables in CoxBoost package by using the unpen.index option 7 . A generalized version of Bierman’s Random Forest (RF) method for analyzing survival data is the Random Forest Survival (RSF) method. There are two methods for randomization in RF. First, for the growth of each tree, a random sampling with replacement (bootstrap method) is taken from the data set. Second, in each tree node, a number of explanatory variables are randomly selected. The purpose of this two-stage randomization is to make trees independent of each other, which due to the bagging property, reduces the variance for the ensemble. The use of deep trees (to reduce bias), when combined with reduced variance due to bagging (randomization and averaging), enables the random forest method to fit valuable models with low generalizability error. As a result, the RF method's general principles are followed when implementing RSF, and they are as follows: Survival trees are created using bootstrapped data, several explanatory variables are randomly chosen from the set of variables when splitting tree nodes, trees typically grow deep until the cessation condition is met, and the survival forest ensemble is created by averaging the predicted survival measures of the individual trees. The RSF method is advantageous to regression techniques in a number of ways. Since RSF is entirely data-driven, hypothesis testing is not necessary. This approach runs the model that best fits the data rather than testing its goodness of fit. The RSF approach does not require testing any hypotheses, such as those involving the distribution of explanatory variables or proportional hazards. RSF models can be executed directly using raw data. Only a few crucial parameters, such as the number of bootstrap samples or the number of node divisions, must be given because the RSF process is entirely automated. The RSF methodology is thus an appropriate tool for exploratory analysis of survival data where prior knowledge is incomplete. The inability to calculate relative risk and odds ratios is a drawback of the RSF method. Instead, using minimum depth and VIMP indicators, it is possible to determine the significance of each risk factor. To calculate relative risks and odds ratios, Cox regression models can be used to analyze the RSF-selected variables. Another drawback of tree-based approaches is that they prefer to use continuous variables in node splitting. However, if the data set includes both continuous and qualitative factors, this can be overcome by choosing a small number of cut points. The random survival forests method has gained attention as a research topic and is seen as a promising method for high dimensional survival data in many biomedical applications because of its high flexibility, capacity for variable selection, and nonlinear and nonparametric nature. The RSF method performs better or at least better than its competitors in analyzing survival data and creating accurate ensemble predictors, according to evidence from applied medical research papers. RSF is a fantastic tool for figuring out extremely intricate relationships between variables. While older approaches are dependent on various constrained assumptions, and thus are less automated and do not perform well in multicollinearity situation. Additionally, the process of random node dividing allows for the inclusion of highly correlated variables in the model, and the choice of appropriate variables is still feasible even in the presence of multicollinearity. Additionally, the bootstrap sampling method's randomization considerably reduces the overfitting issue. The findings of the review of several articles demonstrate, as anticipated, the usefulness and reliability of ensemble methods like RSF, particularly in the medical sciences 18 . An extension of the standard Support Vector Machine that uses right-censored time-to-event data is called a Survival Support Vector Machine. Its main benefit is that it uses the 'kernel trick' to account for complicated, non-linear relationships between features and survival. In high-dimensional feature spaces where survival can be described by a hyperplane, a kernel function implicitly maps the input features. Because of this, survival support vector machines are very adaptable and can be used with a variety of data. There are two approaches to discuss survival analysis in the context of Support Vector Machines: 1-As a ranking problem: the model learns to assign samples with shorter survival times a lower rank by considering all possible pairs of samples in the training data. 2-As a regression problem: the model learns to directly predict the (log) survival time. The drawback in both situations is that it is difficult to relate predictions to the survival function and cumulative hazard function, which are two fundamental survival analysis concepts. Moreover, they have to retain a copy of the training data to do predictions 19 . Declarations Ethics approval and consent to participate Not applicable Consent for publication Not applicable Availability of data and materials Training dataset (GSE26549) and validation dataset (GSE9844) are available on https://www.ncbi.nlm.nih.gov/geo. Competing Interest The authors declare that they have no competing interests. Funding We didn’t need any funding. Authors’ Contributions NSh: analysis, interpretation, writing and editing. PA: methodology, investigation, analysis, reviewing, critical revision of the manuscript. LT: reviewing, writing. MR: corresponding author, investigation, reviewing, critical revision of the manuscript. All authors read and approved the final manuscript Acknowledgements This article is the result of the first author's master's thesis. The authors sincerely acknowledge the deputy director of research and technology at the IRAN University of Medical Sciences, Department of Public Health, for their academic support. References Stasio D Di, Spagnuolo G, López-Cortés XA, Matamala F, Venegas B, Rivera C. Machine-Learning Applications in Oral Cancer: A Systematic Review. Appl Sci 2022, Vol 12, Page 5715 . 2022;12(11):5715. doi:10.3390/APP12115715 Mentel S, Gallo K, Wagendorf O, et al. Prediction of oral squamous cell carcinoma based on machine learning of breath samples: a prospective controlled study. BMC Oral Health . 2020;21:500. doi:10.1186/s12903-021-01862-z Pratama R, Hwang JJ, Lee JH, Song G, Park HR. Authentication of differential gene expression in oral squamous cell carcinoma using machine learning applications. BMC Oral Health . 2020;21:281. doi:10.1186/s12903-021-01642-9 Reyimu A, Chen Y, Song X, Zhou W, Dai J, Jiang F. Identification of latent biomarkers in connection with progression and prognosis in oral cancer by comprehensive bioinformatics analysis. doi:10.1186/s12957-021-02360-w Piazza C, Marchi F, Martino S, et al. Deep Machine Learning for Oral Cancer: From Precise Diagnosis to Precision Medicine. Oral Heal | www.frontiersin.org . 2022;1:794248. doi:10.3389/froh.2021.794248 Atlam M, Torkey H, El-Fishawy N, Salem H. Coronavirus disease 2019 (COVID-19): survival analysis using deep learning and Cox regression model. Pattern Anal Appl . 2021;24(3):993-1005. doi:10.1007/s10044-021-00958-0 De Bin R. Boosting in Cox regression: a comparison between the likelihood-based and the model-based approaches with focus on the R-packages CoxBoost and mboost. Comput Stat . 2016;31(2):513-531. doi:10.1007/s00180-015-0642-2 Kim DW, Lee S, Kwon S, Nam W, Cha IH, Kim HJ. Deep learning-based survival prediction of oral cancer patients. Sci Rep . 2019;9(1):1-10. doi:10.1038/s41598-019-43372-7 Das P, Roychowdhury A, Das S, Roychoudhury S, Tripathy S. sigFeature: Novel Significant Feature Selection Method for Classification of Gene Expression Data Using Support Vector Machine and t Statistic. Front Genet . 2020;11(April):1-12. doi:10.3389/fgene.2020.00247 Bioconductor - sigFeature. Accessed December 5, 2022. https://www.bioconductor.org/packages/release/bioc/html/sigFeature.html Avramopoulos D, Wang R, Valle D, Fallin MD, Bassett SS. A novel gene derived from a segmental duplication shows perturbed expression in Alzheimer’s disease. Neurogenetics . 2007;8(2):111-120. doi:10.1007/S10048-007-0081-5 Li J, Zhang J, Jin L, Deng H, Wu J. Silencing lnc-ASAH2B-2 Inhibits Breast Cancer Cell Growth via the mTOR Pathway. Anticancer Res . 2018;38(6):3427-3434. doi:10.21873/ANTICANRES.12611 Wilson-Edell KA, Kehasse A, Scott GK, et al. RPL24: a potential therapeutic target whose depletion or acetylation inhibits polysome assembly and cancer cell growth. Oncotarget . 2014;5(13):5165. doi:10.18632/ONCOTARGET.2099 Goudarzi KM, Lindström MS. Role of ribosomal protein mutations in tumor development (Review). Int J Oncol . 2016;48(4):1313-1324. doi:10.3892/IJO.2016.3387 Ma XX, Chen QX, Wu SJ, Hu Y, Fang XM. Polymorphisms of the HTR3B gene are associated with post-surgery emesis in a Chinese Han population. J Clin Pharm Ther . 2013;38(2):150-155. doi:10.1111/JCPT.12033 Belotti Y, Tolomeo S, Yu R, Lim WT, Lim CT. Prognostic Neurotransmitter Receptors Genes Are Associated with Immune Response, Inflammation and Cancer Hallmarks in Brain Tumors. Cancers (Basel) . 2022;14(10). doi:10.3390/CANCERS14102544/S1 Spooner A, Chen E, Sowmya A, et al. A comparison of machine learning methods for survival analysis of high-dimensional clinical data for dementia prediction. Sci Reports 2020 101 . 2020;10(1):1-10. doi:10.1038/s41598-020-77220-w Bozorgnezhad M. Journal of Biostatistics and Epidemiology. J Biostat Epidemiol . 2018;1(1):37-44. Introduction to Survival Support Vector Machine — scikit-survival 0.19.0. Accessed January 19, 2023. https://scikit-survival.readthedocs.io/en/stable/user_guide/survival-svm.html Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2985174","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":205194248,"identity":"d95ec95d-fee4-4b64-b637-487f9396fca0","order_by":0,"name":"Niusha Shekari","email":"","orcid":"","institution":"Iran University of Medical Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Niusha","middleName":"","lastName":"Shekari","suffix":""},{"id":205194249,"identity":"53975680-b6f2-418a-a8d0-238dfd77a8ea","order_by":1,"name":"Payam Amini","email":"","orcid":"","institution":"Iran University of Medical Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Payam","middleName":"","lastName":"Amini","suffix":""},{"id":205194250,"identity":"2421be73-0a75-4f2e-b1d7-e20013cffc2f","order_by":2,"name":"Leili Tapak","email":"","orcid":"","institution":"Hamedan University of Medical Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Leili","middleName":"","lastName":"Tapak","suffix":""},{"id":205194251,"identity":"49dee253-9043-450b-aaab-6d3581691436","order_by":3,"name":"Mahboobeh Rasouli","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1ElEQVRIiWNgGAWjYFACHgjFx8x8AEhJyBCvhY2ZLQGkhYcELQw8BkhcPEC3/eyxDz/+2CS2sfN8fnWjxoKHgf3w0Q34tJidyUue2duWltjGzLvNOucY0GE8aWk38Go5kGPMwNtwGKzFOIcNqEWCxwy/lvNvjBn//AFp4XlmnPOPGC03coyZedjAWpgf57YRpeWNMbNsW5pxGzObGXNunwQPG0G/nM8xZnzzx0a2n//w48853+rk+NkPH8OrBRmwSYBJYpWDAPMHUlSPglEwCkbByAEA6PdBjqXJOcwAAAAASUVORK5CYII=","orcid":"","institution":"Iran University of Medical Sciences","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Mahboobeh","middleName":"","lastName":"Rasouli","suffix":""}],"badges":[],"createdAt":"2023-05-26 10:29:25","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2985174/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2985174/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":37848690,"identity":"7ded16ec-9ea6-4fdb-98f9-41fedd6c4fcf","added_by":"auto","created_at":"2023-06-01 14:41:50","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":16454,"visible":true,"origin":"","legend":"\u003cp\u003eThe Kaplan-Meyer curve assessing the survival probability of low and high-risk groups based on the prognostic scores resulted by the multiple cox regression\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-2985174/v1/a750fc9d04262f838bf0522a.png"},{"id":37848691,"identity":"2bb54d7e-6f18-4545-8534-e2f96a12e012","added_by":"auto","created_at":"2023-06-01 14:41:50","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":15972,"visible":true,"origin":"","legend":"\u003cp\u003eThe Roc curve based on the prediction probability of ASAH2B, RPL24, HTR3B, and RNU2_22P for oral cancer occurrence using the Validation data set\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-2985174/v1/1ac07d06198575f806afc05d.png"},{"id":42120621,"identity":"6a8d51a8-9c8f-4166-b681-f7ef7617a0ee","added_by":"auto","created_at":"2023-08-25 07:07:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":313093,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2985174/v1/713f2164-c4d1-4f3b-8540-22b713998145.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Investigation of genes related to oral cancer using time-to-event machine learning approaches","fulltext":[{"header":"Introduction","content":"\u003cp\u003eAround the world, oral cancer has become a significant problem for public health. It is indicated that between 1990 and 2017, there was a nearly one-fold increase in the global incidence, mortality, and disability-adjusted life years of this disease. According to Global Cancer Observatory estimates of incidence and mortality, lip and oral cavity cancer will have 377,713 new cases and 177,757 fatalities in 2020. The majority of oral cancers are squamous cell carcinomas, which are aggressive cancers that frequently spread both locally and to distant sites. It has a significant impact on both the patient's life and society as a whole\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e.OSCC patients' 5-year survival rates range from 63% (for female patients) to 47% (for male patients). The high rate of OSCC recurrence and metastases, as well as the disease's delayed diagnosis, are all associated with mortality. Only one-third of OSCCs are discovered at an early stage (0\u0026ndash;I). Therefore, there is great interest in the development of tests that increase our capacity to screen high-risk (e.g., heavy alcohol and tobacco use) and post-therapy patients\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eUnderstanding the genetic heterogeneity of cancers can help develop effective biomarkers, raise diagnostic precision, and improve treatment outcomes. The development of next-generation sequencing technologies with high efficiency and accuracy has resulted in the generation of a huge amount of cancer tissue genomic data, the majority of which is stored in the Cancer Genome Atlas (TCGA) database. Studies have looked for a diagnostic model that can quickly and effectively distinguish between different tumors by simultaneously taking into account both genetic and phenotypic features using RNA sequencing expression data\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e.Using the Oncomine database, it was determined that the enhanced expression of CXCL8, DDX60, IFI44L, RSAD2, and RTP44 in oral cancer. The Human Protein Atlas database revealed that DDX60, IFI44L, RSAD2, and RTP44 protein expression levels in tumor tissues were higher than those in normal tissues\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eDue to the difficulty of analyzing and interpreting genomic datasets, various machine learning methods are used. To this purpose, it has been demonstrated that machine learning techniques (shallow learning) offer improved OSCC prognostication. Notably, it has been claimed that using machine learning can predict outcomes more accurately than using traditional statistical methods. Due to the ability to identify the intricate correlations between the variables in the dataset, machine learning algorithms have demonstrated promising results. Machine learning approaches have garnered a lot of attention recently because of their alleged viability and advantages in the field of cancer prognostication\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eFor early diagnosis of the disease and prevention of irreparable physical, mental and social complications of this type of cancer, in this paper, we aim to use Cox-Boost, Random Survival Forest and Survival Support Vector Machines approaches to investigate related genes with oral cancer.\u003c/p\u003e"},{"header":"Material and Methods","content":"\u003ch2\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eData\u003c/span\u003e:\u003c/h2\u003e\n\u003cp\u003eTraining dataset (GSE26549) and validation dataset (GSE9844) are available on \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003eStatistical Analysis:\u003c/h2\u003e\n \u003cp\u003eThe statistical analysis was performed using R programing language version 4.2.0. To begin, we filtered genes by fitting a univariate cox model to 29096 genes expression and extracted significant genes that had a p-value less than 0.001. The genes were reduced to 945. \u0026ldquo;CoxBoost\u0026rdquo;, \u0026ldquo;randomForestSRC\u0026rdquo; and \u0026ldquo;sigfeature\u0026rdquo; libraries were used in R. Each of models suggested many genes but our main goal was to discover common genes between three methods. In the next step, we fitted a multiple cox model to common genes and obtained the prognostic score. Prognostic score was stratified into: low-risk \u0026amp; high-risk groups then Kaplan-Meier curves were generated for the two risk groups for the time-to-event related death. Finally, to do validation, common genes were utilized for prediction probability and ROC curve.\u003c/p\u003e\n \u003ch2\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eCox PH Model\u003c/span\u003e:\u003c/h2\u003e\n \u003cp\u003eThe Cox regression method is a statistical technique that is frequently used in medical research to forecast the length of survival for various patients. The Cox Regression approach is used to estimate the hazard rate, which is the degree to which certain characteristics affect survival. Semi-parametric models are exemplified by the Cox regression method. The Cox model can be expressed using the hazard function\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(h\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e. Shortly, the hazard function gives the likelihood of dying at time t. The estimation is as follows:\u003c/p\u003e\n \u003cdiv id=\"Equa\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e$$h\\left(t\\right)={h}_{0}\\left(t\\right)\\times \\text{e}\\text{x}\\text{p}({b}_{1}{x}_{1}+{b}_{2}{x}_{2}+\\dots +{b}_{n}{x}_{n})$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\u0026quot;t\u0026quot;\\)\u003c/span\u003e\u003c/span\u003erepresents the survival time. In order to calculate the hazard function\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(h\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e,\u0026quot;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n covariates\\)\u003c/span\u003e\u003c/span\u003e\u0026quot;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\left({x}_{1},{x}_{2},\\dots ,{x}_{n}\\right)\\)\u003c/span\u003e\u003c/span\u003eare used. With the help of the coefficients \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\left({b}_{1},{b}_{2},\\dots ,{b}_{n}\\right)\\)\u003c/span\u003e\u003c/span\u003e, the influence of variables is calculated. The baseline hazard is \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\u0026quot;h0\u0026quot;.\\)\u003c/span\u003e\u003c/span\u003e \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eCox Boost\u003c/span\u003e:\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003cp\u003eThese steps are used in the boosting procedure:\u003c/p\u003e\u003cspan\u003e\n \u003cp\u003e1) Set the vector of regression coefficients as 0.\u003c/p\u003e\n \u003c/span\u003e\u003cspan\u003e\n \u003cp\u003e2) Compute the negative gradient vector,\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(u=\\frac{{\\delta }{L}({y},{F}\\left({X},{\\beta }\\right))}{{\\delta }{F}({X},{\\beta })}|{\\beta }=\\widehat{\\beta }\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/span\u003e\u003cspan\u003e\n \u003cp\u003e3) Compute the updates:\u003c/p\u003e\n \u003c/span\u003e\u003cspan\u003e\n \u003cp\u003e3.1- Fit the base learner to the negative gradient vector,\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\widehat{h}\\left(u,{X}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/span\u003e\u003cspan\u003e\n \u003cp\u003e3.2- Penalize it,\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\widehat{b}}_{j}=v\\widehat{h}\\left(u,{X}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/span\u003e\u003cspan\u003e\n \u003cp\u003e4) 4- Select the best update j\u0026lowast; (usually that minimizing the loss function).\u003c/p\u003e\n \u003c/span\u003e\u003cspan\u003e\n \u003cp\u003e5) 5- Update the estimations \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\widehat{\\beta }}_{{j}^{*}}={\\widehat{\\beta }}_{{j}^{*}}+{\\widehat{b}}_{{j}^{* }}\\)\u003c/span\u003e\u003c/span\u003e( \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\widehat{b}\\)\u003c/span\u003e\u003c/span\u003e is called a weak estimator)\u003c/p\u003e\n \u003c/span\u003e\n \u003cp\u003eThe stages between 2 and 5 must be performed \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({m}_{stops}\\)\u003c/span\u003e\u003c/span\u003e times, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({m}_{stops}\\)\u003c/span\u003e\u003c/span\u003e represents the number of boosting iterations. A likelihood function is the foundation of the estimation strategy\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. (Riccardo De Bin provides excellent details on this strategy.)\u003c/p\u003e\n \u003cp\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eRandom Survival Forest\u003c/span\u003e:\u003c/p\u003e\n \u003cp\u003eAlthough clinical professionals may find statistical techniques like classification and regression trees intuitive, they have high variance and poor performance. These are dealt with by random forest, which creates a large number of trees and produces the results through voting. Utilizing all variables gathered and automatically evaluating nonlinear effects and intricate interactions, RSF lowers variance and bias. This approach is fully non-parametric, including the effects of the treatments and predictor variables, whereas traditional methods such as CPH utilize a linear combination of attributes. The \u0026ldquo;RandomForestSRC\u0026rdquo; R package was used to train the random survival forest models\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003cp\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eSurvival Support Vector Machine-Recursive Feature Elimination (SVM-RFE)\u003c/span\u003e:\u003c/p\u003e\n \u003cp\u003eAs supervised learning techniques, SVMs are typically employed for classification and regression analysis. Support vector machines can also be used to pick out specific features from a batch of data (e.g., SVM-RFE). With some modifications to get the necessary weight values for each feature, we used SVM in \u0026quot;sigFeature\u0026quot; for feature selection.\u003c/p\u003e\n \u003cp\u003eGuyon et al. (2002) presented the SVM-RFE feature selection method for the classification of cancer in 2002. A weight-based approach is the \u0026quot;SVM-RFE.\u0026quot; Each stage uses the linear SVM\u0026apos;s weight vector coefficients as a feature ranking criterion. The characteristics with the highest weight are the most illuminating. As a result, \u0026quot;SVMRFE\u0026quot; employs a sequential backward feature elimination method to choose the feature with the lowest weight before storing it in a stack. The procedure of iteration is carried out until just one feature variable is left. Both the \u0026quot;SVM-RFE\u0026quot; methodology and our recently developed feature selection algorithm are part of the Wrapper method and choose the feature by recursively removing other features\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eSignificant Feature Selection (sigFeature)\u003c/span\u003e:\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003cp\u003eThe SVM-RFE package offers a brand-new novel feature selection approach for binary classification that makes use of the t-statistic and support vector machines. The chosen features are differently important across the two classes in this feature selection procedure, and they are also strong classifiers with a greater level of classification accuracy\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n \u003cp\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003ePrognostic Score\u003c/span\u003e:\u003c/p\u003e\n \u003cp\u003eIn this study, a prognostic score for each patient is calculated using multiple cox regression coefficients and genes expression (shared genes):\u003c/p\u003e\n \u003cdiv id=\"Equb\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e$$Prognostic Score=\\sum {\\beta }_{i}\\times {X}_{i} i=1,\\dots ,p$$\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003eA higher score was thought to be related to a higher risk of oral cancer. We used the median prognostic score as the cutoff point:\u003c/p\u003e\n \u003cp\u003eGroup 1: High-Risk Group:\u003c/p\u003e\n \u003cp\u003eIf prognostic score for a patient\u0026thinsp;\u0026gt;\u0026thinsp;Median: patient is at a higher risk of developing cancer.\u003c/p\u003e\n \u003cp\u003eGroup 2: Low-Risk Group:\u003c/p\u003e\n \u003cp\u003eIf prognostic score for a patient\u0026thinsp;\u0026lt;\u0026thinsp;Median: patient is at a lower risk of developing cancer.\u003c/p\u003e\n \u003cp\u003eThen, the survival distribution of these two groups of patients is compared using the Kaplan-Meier plot. Moreover, we calculated Hazard Ratios for each genes using multiple cox regression. In the final stage, to measure the validity of the results, we find the shared genes from the validation dataset and prepared ROC curve for each of them.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eWe analyzed the information of 29096 gene expression of 86 individuals. Dataset is available in GEO repository (GSE26549). After fitting the univariate Cox model, we reached 945 significant genes with p-values less than 0.001. The Cox-Boost, Random Survival Forest and Survival Support Vector Machine methods applied to these 945 genes expression. The AUC for each of these methods obtained in order of 0.921, 0.894 and 0.989. Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e shows the multiple cox result with shared genes among these three methods:\u003c/p\u003e\n\u003cp\u003e\u003c/p\u003e\u0026nbsp;\u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eThe results of multiple cox regression assessing the impact of shared genes from the RF, SSVM, and Cox-Boost\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGene\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eHR (95% Confidence Interval)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eP-value\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRPL24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0005 (0.0001\u0026ndash;0.0265)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHTR3B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5.4272 (1.0450- 28.1869)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.044\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eASAH2B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0215 (0.0026\u0026ndash;0.1763)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eTEX29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.2031 (0.0451\u0026ndash;0.9150)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.037\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003cp\u003eAs we can see all shared genes are significant and HTR3B expression are associated with higher hazard ratio and it has higher risk for oral cancer compared with three other genes.\u003c/p\u003e\n\u003cp\u003eWe calculated prognostic score with these shared genes and compared it with its median then divided it into two parts: (Prognostic Score\u0026thinsp;\u0026gt;\u0026thinsp;Median: High Risk Group \u0026amp; Prognostic Score\u0026thinsp;\u0026lt;\u0026thinsp;Median: Low Risk Group). Figure \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e shows the Kapelan-Meier curve:\u003c/p\u003e\n\u003cp\u003eAs a result, high-risk group has lower survival probability than low risk group. In addition, their survival probability decreases during the time. It can be said that survival distributions for the different levels of risk groups is different.\u003c/p\u003e\n\u003cp\u003eThe publicly available dataset from the GEO repository (GSE9844) was applied for validation. we extracted only shared genes from this dataset for obtaining prediction probability. Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e shows the AUC for each gene and its 95% confidence interval:\u003c/p\u003e\n\u003cp\u003e\u003c/p\u003e\u0026nbsp;\u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eThe area under curve (95% confidence interval) based on the prediction probability of ASAH2B, RPL24, HTR3B, and RNU2_22P for oral cancer occurrence using the Validation data set\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGenes\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eArea under curve\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eASAH2B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.822 (0.667\u0026ndash;0.976)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRPL24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.414 (0.208\u0026ndash;0.621)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHTR3B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.576 (0.368\u0026ndash;0.784)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRNU2_22P\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.582 (0.401\u0026ndash;0.764)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003cp\u003eASAH2B has the highest AUC among the other genes. It means that it has more predictive potency for oral cancer. Figure 2 shows Roc curve for each gene based on prediction probability for oral cancer:\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn conclusion, four key genes including RPL24, HTR3B, ASAH2B and TEX29 were identified for further study. Survival support vector machine was the best model for finding related genes to oral cancer. Common genes between three methods were: RPL24, HTR3B, ASAH2B and TEX29 which all of them were significant in multiple cox and survival distributions for the different levels of risk groups was different. By increasing in the expression of HTR3B, the risk of oral cancer increases 5 times.\u003c/p\u003e \u003cp\u003eBy using four genes mentioned above from validation dataset, we saw all of them had a high AUC specifically, its value for ASAH2B was more than other genes.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this paper, we developed \"sigFeature\" that uses \"SVM-RFE\" and the t statistic to choose significant features. Furthermore, we compared the \"sigFeature\" technique to other algorithms (such as Cox-Boost and Random Survival Forest) using AUC. The results shown that the sigFeature method to choose important feature has more accuracy than two other methods. Other machine learning methods, such as neural networks, can also be used to find significant genes with greater accuracy and potency.\u003c/p\u003e \u003cp\u003eThe ceramidase enzyme encoded with N-acylsphingosine amidohydrolase 2B(\u003cem\u003eASAH2B\u003c/em\u003e) has a role in process such as sphingosine biosynthetic and ceramide catabolic\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/gene\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/gene\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Li et al in their study showed that this gene plays an essential role in regulation of Breast Cancer Cell Growth through regulating the signaling pathway of mTOR\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe ribosomal protein encoded with ribosomal protein L24 (\u003cem\u003eRPL24\u003c/em\u003e) belongs to L24E family of ribosomal proteins and is a component of the large ribosomal subunit. \u003cem\u003eRPL24\u003c/em\u003e was Dysregulated in breast cancer and has a role in the regulation of tumorigenesis\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Mutation of ribosomal protein encoding genes associated with cancer development\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe protein encoded by 5-hydroxytryptamine receptor 3B (\u003cem\u003eHTR3B\u003c/em\u003e) as a member of ligand-gated ion channel receptor superfamily is a subunit B of the 5-HT3 receptor\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. This gene has a role in regulation of inflammasomes and immune response and is associated with prognosis of brain tumors\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe Cox model is perhaps the most widely used statistical tool for analyzing survival data, despite the restrictions imposed by the proportional hazards assumption, because of its adaptability and simplicity of interpretation. Due to this, novel statistical/machine learning techniques are typically adjusted to match it, such as boosting, an iterative technique that was initially created in the machine learning community and later extended to the statistical sector. The availability of user-friendly software, such as the R package CoxBoost, which permits the use of boosting in conjunction with the Cox model, has further contributed to the popularity of boosting. One of the best techniques, particularly in classification problems, is boosting, which is one of the ensemble methods that connects a strong learner to a group of several weak learners. Because of this, this method has been expanded in the field of statistics to include regression methods and survival analysis due to its useful and accurate applications and performance. This method's sequential learning process differs from the bagging method's lack of use of independent learners. The main goal of boosting is to minimize the predefined loss function by improving the predictors in a sequential manner, each iteration including the weak predictors from the previous stage. Its simplicity of usage is a key benefit of this model. The algorithm automatically determines how many boosting steps to take and avoids overfitting as a result. This is likely one of the factors contributing to its high performance, and it also means that the user need not set any hyperparameters, making it simple to create a model with high performance\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. In addition, it is possible to avoid penalizing some variables in CoxBoost package by using the unpen.index option\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eA generalized version of Bierman\u0026rsquo;s Random Forest (RF) method for analyzing survival data is the Random Forest Survival (RSF) method. There are two methods for randomization in RF. First, for the growth of each tree, a random sampling with replacement (bootstrap method) is taken from the data set. Second, in each tree node, a number of explanatory variables are randomly selected. The purpose of this two-stage randomization is to make trees independent of each other, which due to the bagging property, reduces the variance for the ensemble. The use of deep trees (to reduce bias), when combined with reduced variance due to bagging (randomization and averaging), enables the random forest method to fit valuable models with low generalizability error. As a result, the RF method's general principles are followed when implementing RSF, and they are as follows: Survival trees are created using bootstrapped data, several explanatory variables are randomly chosen from the set of variables when splitting tree nodes, trees typically grow deep until the cessation condition is met, and the survival forest ensemble is created by averaging the predicted survival measures of the individual trees. The RSF method is advantageous to regression techniques in a number of ways. Since RSF is entirely data-driven, hypothesis testing is not necessary. This approach runs the model that best fits the data rather than testing its goodness of fit. The RSF approach does not require testing any hypotheses, such as those involving the distribution of explanatory variables or proportional hazards. RSF models can be executed directly using raw data. Only a few crucial parameters, such as the number of bootstrap samples or the number of node divisions, must be given because the RSF process is entirely automated. The RSF methodology is thus an appropriate tool for exploratory analysis of survival data where prior knowledge is incomplete. The inability to calculate relative risk and odds ratios is a drawback of the RSF method. Instead, using minimum depth and VIMP indicators, it is possible to determine the significance of each risk factor. To calculate relative risks and odds ratios, Cox regression models can be used to analyze the RSF-selected variables. Another drawback of tree-based approaches is that they prefer to use continuous variables in node splitting. However, if the data set includes both continuous and qualitative factors, this can be overcome by choosing a small number of cut points. The random survival forests method has gained attention as a research topic and is seen as a promising method for high dimensional survival data in many biomedical applications because of its high flexibility, capacity for variable selection, and nonlinear and nonparametric nature. The RSF method performs better or at least better than its competitors in analyzing survival data and creating accurate ensemble predictors, according to evidence from applied medical research papers. RSF is a fantastic tool for figuring out extremely intricate relationships between variables. While older approaches are dependent on various constrained assumptions, and thus are less automated and do not perform well in multicollinearity situation. Additionally, the process of random node dividing allows for the inclusion of highly correlated variables in the model, and the choice of appropriate variables is still feasible even in the presence of multicollinearity. Additionally, the bootstrap sampling method's randomization considerably reduces the overfitting issue. The findings of the review of several articles demonstrate, as anticipated, the usefulness and reliability of ensemble methods like RSF, particularly in the medical sciences\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAn extension of the standard Support Vector Machine that uses right-censored time-to-event data is called a Survival Support Vector Machine. Its main benefit is that it uses the 'kernel trick' to account for complicated, non-linear relationships between features and survival. In high-dimensional feature spaces where survival can be described by a hyperplane, a kernel function implicitly maps the input features. Because of this, survival support vector machines are very adaptable and can be used with a variety of data. There are two approaches to discuss survival analysis in the context of Support Vector Machines:\u003c/p\u003e \u003cp\u003e1-As a ranking problem: the model learns to assign samples with shorter survival times a lower rank by considering all possible pairs of samples in the training data.\u003c/p\u003e \u003cp\u003e2-As a regression problem: the model learns to directly predict the (log) survival time.\u003c/p\u003e \u003cp\u003eThe drawback in both situations is that it is difficult to relate predictions to the survival function and cumulative hazard function, which are two fundamental survival analysis concepts. Moreover, they have to retain a copy of the training data to do predictions\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTraining dataset (GSE26549) and validation dataset (GSE9844) are available on https://www.ncbi.nlm.nih.gov/geo.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe didn\u0026rsquo;t need any funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNSh: analysis, interpretation, writing and editing. PA: methodology, investigation, analysis, reviewing, critical revision of the manuscript. LT: reviewing, writing. MR: corresponding author, investigation, reviewing, critical revision of the manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAll authors read and approved the final manuscript\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis article is the result of the first author\u0026apos;s master\u0026apos;s thesis. The authors sincerely acknowledge the deputy director of research and technology at the IRAN University of Medical Sciences, Department of Public Health, for their academic support.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eStasio D Di, Spagnuolo G, L\u0026oacute;pez-Cort\u0026eacute;s XA, Matamala F, Venegas B, Rivera C. Machine-Learning Applications in Oral Cancer: A Systematic Review. \u003cem\u003eAppl Sci 2022, Vol 12, Page 5715\u003c/em\u003e. 2022;12(11):5715. doi:10.3390/APP12115715\u003c/li\u003e\n\u003cli\u003eMentel S, Gallo K, Wagendorf O, et al. Prediction of oral squamous cell carcinoma based on machine learning of breath samples: a prospective controlled study. \u003cem\u003eBMC Oral Health\u003c/em\u003e. 2020;21:500. doi:10.1186/s12903-021-01862-z\u003c/li\u003e\n\u003cli\u003ePratama R, Hwang JJ, Lee JH, Song G, Park HR. Authentication of differential gene expression in oral squamous cell carcinoma using machine learning applications. \u003cem\u003eBMC Oral Health\u003c/em\u003e. 2020;21:281. doi:10.1186/s12903-021-01642-9\u003c/li\u003e\n\u003cli\u003eReyimu A, Chen Y, Song X, Zhou W, Dai J, Jiang F. Identification of latent biomarkers in connection with progression and prognosis in oral cancer by comprehensive bioinformatics analysis. doi:10.1186/s12957-021-02360-w\u003c/li\u003e\n\u003cli\u003ePiazza C, Marchi F, Martino S, et al. Deep Machine Learning for Oral Cancer: From Precise Diagnosis to Precision Medicine. \u003cem\u003eOral Heal | www.frontiersin.org\u003c/em\u003e. 2022;1:794248. doi:10.3389/froh.2021.794248\u003c/li\u003e\n\u003cli\u003eAtlam M, Torkey H, El-Fishawy N, Salem H. Coronavirus disease 2019 (COVID-19): survival analysis using deep learning and Cox regression model. \u003cem\u003ePattern Anal Appl\u003c/em\u003e. 2021;24(3):993-1005. doi:10.1007/s10044-021-00958-0\u003c/li\u003e\n\u003cli\u003eDe Bin R. Boosting in Cox regression: a comparison between the likelihood-based and the model-based approaches with focus on the R-packages CoxBoost and mboost. \u003cem\u003eComput Stat\u003c/em\u003e. 2016;31(2):513-531. doi:10.1007/s00180-015-0642-2\u003c/li\u003e\n\u003cli\u003eKim DW, Lee S, Kwon S, Nam W, Cha IH, Kim HJ. Deep learning-based survival prediction of oral cancer patients. \u003cem\u003eSci Rep\u003c/em\u003e. 2019;9(1):1-10. doi:10.1038/s41598-019-43372-7\u003c/li\u003e\n\u003cli\u003eDas P, Roychowdhury A, Das S, Roychoudhury S, Tripathy S. sigFeature: Novel Significant Feature Selection Method for Classification of Gene Expression Data Using Support Vector Machine and t Statistic. \u003cem\u003eFront Genet\u003c/em\u003e. 2020;11(April):1-12. doi:10.3389/fgene.2020.00247\u003c/li\u003e\n\u003cli\u003eBioconductor - sigFeature. Accessed December 5, 2022. https://www.bioconductor.org/packages/release/bioc/html/sigFeature.html\u003c/li\u003e\n\u003cli\u003eAvramopoulos D, Wang R, Valle D, Fallin MD, Bassett SS. A novel gene derived from a segmental duplication shows perturbed expression in Alzheimer\u0026rsquo;s disease. \u003cem\u003eNeurogenetics\u003c/em\u003e. 2007;8(2):111-120. doi:10.1007/S10048-007-0081-5\u003c/li\u003e\n\u003cli\u003eLi J, Zhang J, Jin L, Deng H, Wu J. Silencing lnc-ASAH2B-2 Inhibits Breast Cancer Cell Growth via the mTOR Pathway. \u003cem\u003eAnticancer Res\u003c/em\u003e. 2018;38(6):3427-3434. doi:10.21873/ANTICANRES.12611\u003c/li\u003e\n\u003cli\u003eWilson-Edell KA, Kehasse A, Scott GK, et al. RPL24: a potential therapeutic target whose depletion or acetylation inhibits polysome assembly and cancer cell growth. \u003cem\u003eOncotarget\u003c/em\u003e. 2014;5(13):5165. doi:10.18632/ONCOTARGET.2099\u003c/li\u003e\n\u003cli\u003eGoudarzi KM, Lindstr\u0026ouml;m MS. Role of ribosomal protein mutations in tumor development (Review). \u003cem\u003eInt J Oncol\u003c/em\u003e. 2016;48(4):1313-1324. doi:10.3892/IJO.2016.3387\u003c/li\u003e\n\u003cli\u003eMa XX, Chen QX, Wu SJ, Hu Y, Fang XM. Polymorphisms of the HTR3B gene are associated with post-surgery emesis in a Chinese Han population. \u003cem\u003eJ Clin Pharm Ther\u003c/em\u003e. 2013;38(2):150-155. doi:10.1111/JCPT.12033\u003c/li\u003e\n\u003cli\u003eBelotti Y, Tolomeo S, Yu R, Lim WT, Lim CT. Prognostic Neurotransmitter Receptors Genes Are Associated with Immune Response, Inflammation and Cancer Hallmarks in Brain Tumors. \u003cem\u003eCancers (Basel)\u003c/em\u003e. 2022;14(10). doi:10.3390/CANCERS14102544/S1\u003c/li\u003e\n\u003cli\u003eSpooner A, Chen E, Sowmya A, et al. A comparison of machine learning methods for survival analysis of high-dimensional clinical data for dementia prediction. \u003cem\u003eSci Reports 2020 101\u003c/em\u003e. 2020;10(1):1-10. doi:10.1038/s41598-020-77220-w\u003c/li\u003e\n\u003cli\u003eBozorgnezhad M. Journal of Biostatistics and Epidemiology. \u003cem\u003eJ Biostat Epidemiol\u003c/em\u003e. 2018;1(1):37-44.\u003c/li\u003e\n\u003cli\u003eIntroduction to Survival Support Vector Machine \u0026mdash; scikit-survival 0.19.0. Accessed January 19, 2023. https://scikit-survival.readthedocs.io/en/stable/user_guide/survival-svm.html\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Oral cancer, Gene expression profiling, Random Survival Forest, Cox-Boost, Survival Support Vector Machines, Machine Learning","lastPublishedDoi":"10.21203/rs.3.rs-2985174/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2985174/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground: \u003c/strong\u003eSince cancer is one of the most common and deadly diseases, its early diagnosis is very important for treatment and prevents the irreparable physical, mental and social consequences of this disease. Oral cancer is also one of the most common cancers, and factors such as gender, age, and smoking influence the incidence of this disease. One of the most important factors affecting cancer is genetic factors. It is not enough to consider clinical factors for the treatment of this disease, and it is also very important to deal with the genes in people's bodies that are effective in their survival against cancer. Also, the survival of people with oral cancer in the early stages of the disease is 80%, so early detection is very important. Therefore, we are looking for a model to better investigate key and effective genes in this disease.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods: \u0026nbsp;\u003c/strong\u003eA publicly available dataset of oral cancer (GSE26549) including information of 29096 genes expression profiles of 86 samples was used. A univariate cox regression was used for each gene’s expression to reduce the number of genes. Cox-Boost, Random Survival Forest and Support survival SVM (Recursive Feature Elimination) were used to identify related genes. Shared genes between three methods were discovered for calculating the prognostic score and the Kaplan-Meier curve. To do validation, common genes were selected from the validation dataset (GSE9844) to provide the ROC curve.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eThe univariate Cox regression models selected 945 significant genes. Four shared genes of RPL24, HTR3B, ASAH2B and TEX29 related to time-to-death in oral cancer patients were then identified by using the Cox-Boost, Random Survival Forest and Support survival SVM (Recursive Feature Elimination). The survival distributions of the high-risk and low-risk groups significantly differed.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion: \u003c/strong\u003eCommon genes between three methods were RPL24, HTR3B, ASAH2B and TEX29 which all of them were significant in multiple Cox.\u003c/p\u003e","manuscriptTitle":"Investigation of genes related to oral cancer using time-to-event machine learning approaches","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-06-01 14:41:46","doi":"10.21203/rs.3.rs-2985174/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"30dec9b4-cd88-4620-b3ef-31e0867fbea3","owner":[],"postedDate":"June 1st, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":21945335,"name":"Biological sciences/Genetics"},{"id":21945336,"name":"Biological sciences/Genetics/Gene expression"},{"id":21945337,"name":"Biological sciences/Computational biology and bioinformatics/High throughput screening"},{"id":21945338,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"},{"id":21945339,"name":"Biological sciences/Computational biology and bioinformatics/Statistical methods"},{"id":21945340,"name":"Biological sciences/Cancer"},{"id":21945341,"name":"Biological sciences/Cancer/Oral cancer"}],"tags":[],"updatedAt":"2023-08-25T06:59:14+00:00","versionOfRecord":[],"versionCreatedAt":"2023-06-01 14:41:46","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2985174","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2985174","identity":"rs-2985174","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.