Prediction Model for Lymph Node Metastasis in Papillary Thyroid Carcinoma Based on Electronic Medical Records | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Prediction Model for Lymph Node Metastasis in Papillary Thyroid Carcinoma Based on Electronic Medical Records JingWen Zhang, XiaoWen Zhang, ShuJun Xia, YiJie Dong, Wei Zhou, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3909203/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Purpose This study aimed to establish a novel machine learning model for predicting lymph node metastasis(LNM)of patients with papillary thyroid carcinoma (PTC) by utilizing personal electronic medical records (EMR) data. Methods The study included 5076 PTC patients underwent total thyroidectomy or lobectomy with lymph node dissection. Based on the integrated learning approach, this study designed a predictive model for LNM. The predictive model employs deep neural network (DNN) models to identify features within cases and vectorize clinical data from electronic medical records into feature matrices. Subsequently, a classifier based on machine learning algorithms is designed to analyse the feature matrices for prediction LNM in PTC. To mitigate the risk of overfitting commonly associated with machine learning algorithms processing high-dimensional matrices, multiple DNNS are utilized to distribute the overfitting risk. Five mainstream machine learning algorithms (NB, DT, XGB, GBM, RDF) are tested as classifier algorithms in the predictive model. Model performance is assessed using precision, recall, F1, and AUC. Results Among the patients, 2,261 had lymph node metastasis (LNM), with 2,196 displaying central lymph node metastasis (CLNM) and 472 exhibiting lateral cervical lymph node metastasis (LLNM). The RDF model showcased superior predictive performance compared to other models, achieving a testing AUC of 0.98, precision of 0.98, recall of 0.95, and F1 value of 0.97 in predicting LNM. Moreover, it attained an AUC of 0.98, precision of 0.98, recall of 0.94, and an F1 value of 0.96 in predicting CLNM. Regarding the weighting of the feature matrix for various case data types, gender and multi-focus held higher weights, at 1.24 and 1.23 respectively. Conclusion The LNM predictive model proposed in this study could be used as a cost-effective tool for predicting LNM in PTC patients, by utilizing easily available personal electronic medical data, which can provide valuable support to surgeons in devising a personalized treatment plan. papillary thyroid carcinoma lymph node metastasis machine learning algorithm Electronic Medical Records Figures Figure 1 Figure 2 Introduction Papillary thyroid carcinoma (PTC) is the most common pathological type of thyroid cancer, accounting for more than 90% of thyroid tumors 1 . While the majority of PTC patients experience an indolent clinical course and exhibit a favorable prognosis, lymph node metastases are commonly detected at the time of diagnosis. Surgery remains the primary treatment option, with the extent of surgery being determined by the preoperative assessment of lymph node status 2 . Lymph node metastasis (LNM) of papillary thyroid carcinoma (PTC) is a significant risk factor for regional recurrence 3,4 ., with the incidence rate of 60–70% 5 . At present, LNM is primarily evaluated through preoperative imaging examinations, predominantly ultrasonography. However, the usefulness of pre-operative ultrasonography (US) in assessing LNM is frequently limited 6–9 . Therefore, a more objective and accurate method was needed to predict neck lymph node metastasis in PTC patients. Multiple studies have demonstrated that PTC lymph node metastasis (LNM) is influenced by various factors, including the patient's gender, age, tumor size, location, degree of extrathyroidal extension (ETE), and multifocal disease 3,10–12 . Furthermore, Paes et al. 13 propose that cancers are more aggressive in PTC patients with higher body mass indices (BMIs), leading to increased LNM rates. Some studies indicate that patients with second-generation familial non-medullary thyroid carcinoma may exhibit aggressive tumor behaviors, including cervical lymph node metastasis and recurrence 14 . Additionally, research demonstrates that female patients with a history of breast cancer are more likely to experience cervical lymph node metastasis. Although no relevant reports have shown that LNM of PTC is associated with metabolic disease, some reports have shown that metabolic abnormalities(diabetes, hypertension, high BMI)may significantly increase the risk of thyroid cancer 15 . Unhealthy personal habits, such as smoking and consuming alcohol, may also contribute to cancer metastasis 16–18 . In modern medicine, the electronic medical record (EMR) system has developed rapidly, resulting in a significant increase in computerized medical data. In addition to the patient demographic data, the EMR information contains a vast amount of textual data containing the patients' chief complaints, medical histories, lifestyle habits, familial cancer histories, and marital and reproductive histories, along with tumor-related information in imaging and pathological reports. However, due to technical constraints, these textual data have often been ignored in previous studies. With the development of artificial intelligence (AI), machine learning (ML) has become a potentially powerful tool for mining EMR text data to facilitate disease diagnosis and management. In the medical field, natural language processing (NLP) is widely acknowledged as essential for extracting valuable medical information from textual data 19,20 . Therefore, the goal of this study is to establish a novel machine learning model for predicting LNM of patients with PTC, in a cost-effective way by utilizing easily available personal electronic medical data. Method Materials The study included 5076 PTC patients who underwent total thyroidectomy or lobectomy with lymph node dissection at Ruijin Hospital between January 2013 and December 2016. The PMC patients included 1331 males and 3745 females. All patients underwent pre-operative ultrasonography examinations. All patients had no previous history of of neck surgery or irradiation. The research protocol was approved by the Ethics Committee of Ruijin Hospital, Shanghai Jiao Tong University School of Medicine. Data Collection Electronic clinical and pathologic records were retrospectively reviewed to gather clinicopathological information. The following information is collected: patient age, gender, chief complaint, medical history (hypertension, diabetes, hyperlipidemia and malignant tumor), family history of cancer, BMI, smoking and drinking history; US features were recorded: tumor diameter, tumor location (upper pole, middle, lower pole and isthmus), tumor site (left lobe, right lobe and isthmus), and aspect ratio (< 1 or ≥ 1). Pathological informations were analyzed: Hashimoto's thyroiditis, multifocality (unifocal, multifocal), tumor T stage, and extrathyroidal extension (ETE). Model and Training Firstly, considering the noise introduced by synonyms (such as "elevated blood pressure" and "hypertension", "breast cancer" and "malignant breast tumor", "physical examination finding" and "ultrasonography finding"), as well as colloquial expression in chief complaints, medical histories and family cancer histories, which may impact the prediction of LNM, we employ word embedding techniques to achieve a standardized and normalized representation of the original data. Specifically, we further clean and preprocess the case data, which has undergone standardized collation and desensitization, and vectorize it using a pre-trained word embedding dictionary 21 .Word embedding is a technique that maps words or letters from one-dimensional space to high-dimensional vector space and represent words by multidimensional vectors so that computers can understand and learn 22 . In this study, n represents the number of features in this segment of data, embed_dim refers to the word embedding dimension set during model training, and each lattice of case data is converted into a word embedding matrix of n × embed_dim dimension 23 . The embed_dim is set as 50 during model training, and each word is represented as a 1×50 dimensional matrix vector. Each case data is converted into the corresponding word embedding information for matrix stitching, and the stitched matrix is used as the input of DNN model. This study intends to achieve LNM prediction in PTC by DNN models and ML algorithms. DNN is a complex neural network composed of multiple perceptron models with many hidden layers, which is commonly used to simulate the information transfer process of numerous neurons in the human brain for knowledge learning and reasoning. DNN can fully train the features of long text data to get an auxiliary model of feature matrix with rich semantic information for result prediction, but it often leads to misleading calculation due to overfitting of data in the prediction process. ML is employed to mine the implied laws from mass data for classification prediction 24 . However, traditional ML cannot do a good job in feature mining and learning for complex feature datasets like texts and sequences. To unite the strengths of DNN and ML, based on the integrated learning approach 25 , this study utilizes the DNN models to learn the features of the case vector matrices, to obtain the feature matrices for various kinds of information of case datasets and applies the ML algorithms to learn and calculate the feature matrices, to achieve LNM prediction in PTC. To explore the optimal performance of the proposed LNM prediction model, this study investigates the use of the following ML algorithms as classifiers in the predictive model, including Naive Bayes(NB), Decision Tree (DT), Extreme Gradient Boosting (XGB) and Gradient Boosting Machine (GBM) and Random Forest (RDF) 26–28 . NB is a probability theory-based algorithm that achieves result prediction by probability accumulation of each type of feature in the case dataset. DT introduces the features of datasets into the tree judgment structure layer by layer, judges a certain type of features on each layer of sub-tree in the overall decision tree, and obtains the final result through judging the input data in multiple rounds. GBM algorithm exploits the residual value of the parent node after judgment as the basis of the child node in the tree structure, so that the tree structure can be calculated in parallel, and at the same time, the connection of different features is strengthened during judgment, which can effectively mine the potential connection of features and improve the model as a whole. The XGB algorithm is optimized from the GBM algorithm during engineering implementation by adding the regular term and the residual constant term of last training to the objective function for model training. Although GBM algorithm and XGB algorithm possess generalization and applicability to complex datasets, their performance on high-dimensional sparse vector space and text data features is not ideal. RDF designs a corresponding classification decision tree for each feature, and makes use of the majority voting mechanism to predict the results under the independent and parallel multi-decision subtree structure, which can enhance model performance in high-dimensional space and complex datasets, lower the sensitivity of the model to abnormal values, and promote the generalization ability of the model in extreme cases. Additionally, this study introduces the Stacking multi-model training framework 29 . To address the performance loss caused by the randomness during the model training process, the Stacking framework trains multiple diverse submodels for the same task. Using machine learning algorithms, it assigns credibility weights to each submodel based on its predictive performance. The ultimate prediction results of the model are collectively determined by multiple submodels. The model training flowchart ( Fig. 1 ) is shown as below, illustrating the process from case word embedding to prediction result output by DNN learning and ML algorithms. Patients were divided into two groups based on whether lymph node metastasis was present (positive label) or absent (negative label). The datasets were randomly classified in the ratio of 80:20, 80% of the datasets were randomly selected as the training set for the model and iterative training was performed, and 20% of the datasets were used as the test set for evaluating the trained model. Model stability was assessed, and the optimal threshold probability for a positive outcome was defined by applying cross validation to models trained within the 80% curation training subset; each of 5 cross-validation models was trained with a random 80% sample of the training data (ie, 64% of total curation data) and evaluated using the remaining 20% of the training data (ie, 16% of total curation data). The sigmoid activation layer output score for defining a predicted outcome as positive was defined as half of the mean best F1 scores calculated by cross validation, where the F1 score is defined as the harmonic mean between precision and recall. For each outcome, an ensemble of the cross-validation models was constructed by taking the simple mean of the sigmoid activation layer outputs from each cross-validation model. The ensemble models were applied to the 10% validation subset for evaluation and iterative tuning. In this study, precision, recall and F1 score are used to evaluate the model performance. The definition is as follows: Consider a two-label dataset D with \(\left|\mathbf{D}\right|\) samples \(\left({\varvec{x}}_{\varvec{i}},{\varvec{Y}}_{\varvec{i}}\right), i=1\dots \left|\mathbf{D}\right|, {\varvec{Y}}_{\varvec{i}}\in \varvec{L}\) , where \(\varvec{L}\) denotes a label set. Let \(H\) be a two-label classifier, while \({\varvec{Z}}_{i}=H\left({\varvec{x}}_{\varvec{i}}\right)\) represents the prediction label of \({\varvec{x}}_{\varvec{i}}\) by \(\text{H}\) . Precision, recall and F1 values can be calculated from the following equations, respectively. $$Precision\left(H,D\right)=\frac{{\sum }_{i=1}^{\left|D\right|}\left|{Y}_{i}*{Z}_{i}\right|}{\sum _{i=1}^{\left|D\right|}\left|{Z}_{i}\right|}$$ $$Recall\left(H,D\right)=\frac{\sum _{i=1}^{\left|D\right|}\left|{Y}_{i}*{Z}_{i}\right|}{\sum _{i=1}^{\left|D\right|}\left|{Y}_{i}\right|}$$ $${F}_{1}\left(H,D\right)=\frac{2*Precision\left(H,D\right)*Recall(H,D)}{Precision\left(H,D\right)+Recall(H,D)}$$ In addition, the ROC curve is drawn with the true positive rate (sensitivity) as the y-axis and the false positive rate (1-specificity) as the x-axis, and the area under the curve (AUC) is taken as the classifier evaluation indicator. A higher AUC indicates greater accuracy of the model. Result General Information The PMC patients included 1331 males and 3745 females. The average age of the patients was 44.1 ± 12.4 years. The mean BMI of the patients was 23.6 ± 3.6 (Kg/m 2 ). The mean tumor diameter was 10.9 ± 7.6 mm; there were 2,261 patients with lymph node metastasis (LNM), 2,815 without LNM; Among them, 2,196 with central lymph node metastasis (CLNM), 2,880 without CLNM; 472 with lateral cervical lymph node metastasis (LLNM) and 4,604 without LLNM. Prediction Performance of Model Table 1 , Table 2 and Table 3 respectively show the test results of DNN models and five ML algorithms for predicting LNM, CLNM and LLNM of PTC. The RDF model achieves a testing AUC of 0.98, precision of 0.98, recall of 0.95, F1 value of 0.97 in predicting LNM. AUC of 0.98, precision of 0.98, recall of 0.94, F1 value of 0.96 in predicting CLNM. AUC of 0.97, precision of 0.96, recall of 0.92, F1 value of 0.94 in predicting LLNM. The GBM algorithm shows an AUC of 0.95, 0.91 and 0.89 respectively in predicting CLNM, CLNM and LLNM. However, the indicators of NB, DT and XGB algorithms are significantly lower than those of RDF(Figure 2 ). Table 1 Evaluation Indicators for LNM Prediction in PTC by the Model Model LNM Precision Recall F1 AUC NB 0.66 0.33 0.44 0.64 DT 0.71 0.40 0.51 0.68 XGB 0.80 0.61 0.69 0.84 GBM 0.87 0.84 0.85 0.95 RDF 0.98 0.95 0.97 0.98 Table 2 Evaluation Indicators for CLNM Prediction in PTC by the Model Model CLNM Precision Recall F1 AUC NB 0.60 0.43 0.50 0.64 DT 0.64 0.45 0.53 0.70 XGB 0.75 0.58 0.65 0.82 GBM 0.81 0.80 0.80 0.91 RDF 0.98 0.94 0.96 0.98 Table 3 Evaluation Indicators for LLNM Prediction in PTC by the Model Model LLNM Precision Recall F1 AUC NB 0.52 0.25 0.33 0.59 DT 0.59 0.49 0.53 0.68 XGB 0.75 0.61 0.67 0.83 GBM 0.81 0.72 0.76 0.89 RDF 0.96 0.92 0.94 0.97 Weighting of Feature Matrix for Different Types of Case Data In order to analyze the contribution of different features to the prediction objective in the dataset, this study also calculates the weights on the feature matrix of case data after DNN training. Among them, the weight of gender and multi-focus is higher, being 1.24 and 1.23 respectively, which is followed by the chief complaint (1.04), location of primary focus (0.91) and age (0.82). The three with the lowest weight are tumor diameter (0.18), marital and reproductive history (0.17), and nutritional status (BMI, 0.00), which have essentially no effect on the model's prediction task (Table 4 ). Table 4 Weight of Each Features of Case Data Feature Weight Gender 1.24 Multifocal cancer 1.23 Chief Complaint 1.04 Site (left lobe/right lobe/isthmus) 0.91 Age 0.84 Family history of malignant tumor 0.71 Individual history (smoking and drinking) 0.46 Aspect ratio ≥ 1 (Yes/No) 0.44 Medical history 0.43 Extrathyroidal extension site 0.34 Hashimoto's thyroiditis 0.33 Location (upper/middle/lower pole/isthmus) 0.32 T stage 0.27 Maximum diameter of tumor 0.18 Nutritional status (BMI) 0.00 Discussion There has been a significant increase in the prevalence of PTC worldwide due to pre-operative sonography. Although most patients with PTC have an indolent clinical course and positive prognosis, the incidence of LNM has been identified as a risk factor for recurrence 3,30 . Currently, physicians rely on preoperative neck ultrasound for assessment of lymph node metastasis in the neck region. However, the sensitivity and specificity of preoperative ultrasound in predicting cervical lymph node metastasis demonstrate notable variability, and the incidence of occult LNM was observed to be as high as 55% 6 . That means the utility of preoperative ultrasonography for lymph node assessment is often limited 8,31,32 . Additionally, several mathematical and statistical models, which combine clinical factors with medical imaging examinations, have been developed to predict LNM in PTC. KIM et al. 33 employed multiple regression analysis, considering variables such as patient age, gender, tumor size, multifocality, bilaterality, and Hashimoto’s, to construct a nomogram model prognosticating the risk of CLNM in PTC patients after thyroidectomy, achieving an AUC of 0.70. On the other hand, Jin et al. 10 developed a scoring model for the prediction of CLNM by using multiple regression analysis with correlating factors such as large tumor size, irregular margins and BRAF mutations. This model demonstrated an overall sensitivity of 85.1% and specificity of 75.8%. In summary, these nomograms or scoring systems primarily utilize Multiple Logistic Regression (MLR) based on the amalgamation of various risk factors to predict binary outcomes. However, this approach presents certain limitations. To circumvent overfitting of the dataset, it encompasses only meticulously screened independent variables, enabling the model to anticipate the associated outcome. Furthermore, these models prove inadequate in addressing the prevalent issue of missing values commonly encountered in electronic medical records. Consequently, it is necessary to develop new models that can handle a wider range of variables, reduce the impact of missing data, and provide precision and robustness. Machine Learning is an application of Artificial Intelligence (AI) that mining and learning from historical data to build predictive models, which can overcome or reduce the limitations of MLR 34 . In the era of "big data", ML can help clinicians make appropriate decisions based on large amounts of digital medical information. Previous studies have shown that sophisticated machine learning techniques can construct precise predictive models by utilizing unprocessed data extracted from electronic medical records and medical images 35,36 . In recent years, ML has been broadly applied in the metastasis prediction and prognosis of cancer 37 . For instance, Tseng et al. 38 developed a decision tree model to predict the risk of recurrence in cervical cancer. Singa et al. 39 employed the RDF algorithm to predict the onset of liver cancer in cirrhosis patients. Kim 40 and Liang 41 et al. established a Support Vector Machine (SVM) model to discriminate between recurrent and non-recurrent malignant tumors. In predicting regional lymph node metastasis of cancer, Bollschweiler 37 et al. employed a single-layer neural network to predict LNM in gastric cancer, achieving an accuracy of approximately 79%. Takada 40 et al. employed a decision tree model-based data mining approach to predict the risk of axillary lymph node metastasis in breast cancer patients, achieving an Area Under the Curve (AUC) of 0.77. Andrés et al. 42 using the decision tree algorithm, predicted the presence of occult LNM in patients with early oral squamous cell carcinoma, achieving an AUC of 0.84. Additionally, the SVM algorithm has displayed a robust classification ability in predicting LNM in breast cancer, cutaneous melanoma, and colon cancer 43–45 . In the real world, cancer occurrence and prognosis can be influenced by many relevant factors such as demographics, family history, age, dietary habits, body weight (obesity), poor lifestyle habits (smoking and alcohol consumption) and environmental exposures. This information can be gathered through routine clinical history data present in electronic medical record systems. However, even for the most skilled clinicians, it is not easy to synthesize the above information. Therefore, machine learning is more effective at this task. Advanced machine learning models can extract potential information from input data and consequently become superior predictive tools. For example, hart et al. 46 used machine learning model to effectively predict the risk of endometrial cancer within 5 years using personal health data. The model involved 952 patients with endometrial cancer, of which 57.2% were defined as high risk and 41.8% as intermediate risk. The AUC of the best algorithm was 0.88. Currently, no studies have been found utilizing machine learning algorithms to predict lymph node metastasis in PTC through electronic medical record data. It is hypothesized that machine learning algorithms can enhance the prediction accuracy of lymph node metastasis in patients with PT. This approach was trained and verified using DNN models and machine learning algorithms, with the aim of devising surgical strategies. In this study, an innovative multiparameterized ML-based model was formulated and validated to predict LNM in patients with PTC utilizing readily available clinical data extracted from electronic medical records. The findings of this study indicate that the developed model possesses a remarkable capacity for accurately prognosticating cervical lymph node metastasis in patients with PTC. Notably, the RDF-based prediction model performs better in prediction, attaining an accuracy rate of 0.98,0.98,0.96 in the prediction of LNM, CLNM and LLNM. This probably due to its more sophisticated classification decisions and different weighting ratios compared to other algorithms. Conversely, NB, DT and XGB algorithms manifest suboptimal performance in the predictive task. Studies 47 have demonstrated that RDF stands as one of the most precise machine learning models, surpassing other techniques in its ability to handle large amounts of features and extremely non-linear data. Furthermore, RDF performs well in mitigating data noise and its adaptability makes it easier to adapt and integrate with learning algorithms. It is worth emphasizing that the selection of the most appropriate algorithm is contingent upon numerous parameters, including the type of data collected, the size of the data sample and the prediction results. The advantages of this study lie in the application of deep learning and deep NLP techniques to rationally apply the information recorded in EMR to obtain the classification results of clinical prediction objectives, which are manifested as the following aspects: ( 1 ) In this study, the case dataset is vectorially mapped from a low-dimensional space to a high-dimensional space through word embedding techniques to endow the feature matrix representing the case dataset with rich semantic information, which effectively helps the model to learn and mine the potential features of LNM in PTC. ( 2 ) The training method of multi-DNN model based on the Stacking framework can utilize the characteristics that the parameters are varied when multiple models are trained with the same task under the same training set, by assigning decision weight to each field model, enhance the generalization ability of the model on the test set, promote the accuracy of the model under the condition of small probability, and ensure the overall prediction effect of the model. But this study also has some limitations. First, the model uses DNN model and machine learning algorithms, so clinical interpretation of important features identified by the model may be a challenge. Second, the collected data are all from a single center and the trained model may not perform well in other healthcare institutions. Therefore, multi-center datasets are needed for further training and validation of the model to optimize its diagnostic efficacy and generalizability. Third, not all patients underwent total thyroidectomy and some occult multifocal cancers would be missed. Third, EMR is unstructured data. Disease history and family history were dictated by patients and recorded by doctors, and there may be bias in patients' description of diseases. Further research attempts to augment the sample size, incorporate additional data (e.g., ultrasound images or serological examination), and further improve the model’s performance through external data validation. Conclusion In this study, we constructed a predictive model based on Deep Neural Networks (DNN) and machine learning algorithms, utilizing easily accessible personal electronic medical data to predict Lymph Node Metastasis (LNM) in Papillary Thyroid Carcinoma (PTC). This approach explores the application of artificial intelligence techniques in predicting lymph node metastasis in PTC patients. Declarations Author contributions Concept and design: JQZ, YZS and ZWW. Acquisition of data: JWZ, ZHL, YJD, WZ, LZ and XJX. Analysis and interpretation of data: XWZ, YZS. Manuscript writing: XWZ and JWZ. All the authors gave their final approval for the submission of this version and any revised version for publication. Compliance with ethical standards Conflict of interest The authors declare no competing interests. Ethics approval This study design was approved by the Ethics Committee of Ruijin Hospital, Shanghai Jiao Tong University School of Medicine and the need for informed consent was waived due to the retrospective analysis of this study. Acknowledgments This study was supported by National Natural Science Foundation of China (82071928) References Vaccarella, S. et al. Global patterns and trends in incidence and mortality of thyroid cancer in children and adolescents: a population-based study. Lancet Diabetes Endocrinol 9, 144–152 (2021). https://doi.org:10.1016/S2213-8587(20)30401-0 Haugen, B. R. et al. 2015 American Thyroid Association Management Guidelines for Adult Patients with Thyroid Nodules and Differentiated Thyroid Cancer: The American Thyroid Association Guidelines Task Force on Thyroid Nodules and Differentiated Thyroid Cancer. Thyroid 26, 1-133 (2016). https://doi.org:10.1089/thy.2015.0020 So, Y. K., Kim, M. J., Kim, S. & Son, Y. I. Lateral lymph node metastasis in papillary thyroid carcinoma: A systematic review and meta-analysis for prevalence, risk factors, and location. Int J Surg 50, 94–103 (2018). https://doi.org:10.1016/j.ijsu.2017.12.029 Stack, B. C., Jr. et al. American Thyroid Association consensus review and statement regarding the anatomy, terminology, and rationale for lateral neck dissection in differentiated thyroid cancer. Thyroid 22, 501–508 (2012). https://doi.org:10.1089/thy.2011.0312 Qu, H., Sun, G. R., Liu, Y. & He, Q. S. Clinical risk factors for central lymph node metastasis in papillary thyroid carcinoma: a systematic review and meta-analysis. Clin Endocrinol (Oxf) 83, 124–132 (2015). https://doi.org:10.1111/cen.12583 Lim, Y. S. et al. Lateral cervical lymph node metastases from papillary thyroid carcinoma: predictive factors of nodal metastasis. Surgery 150, 116–121 (2011). https://doi.org:10.1016/j.surg.2011.02.003 Ito, Y. et al. Ultrasonographically and anatomopathologically detectable node metastases in the lateral compartment as indicators of worse relapse-free survival in patients with papillary thyroid carcinoma. World J Surg 29, 917–920 (2005). https://doi.org:10.1007/s00268-005-7789-x Ito, Y. et al. Clinical significance of metastasis to the central compartment from papillary microcarcinoma of the thyroid. World J Surg 30, 91–99 (2006). https://doi.org:10.1007/s00268-005-0113-y Zhang, J., Fei, M., Dong, Y., Xu, S. & Zhan, W. Preoperative Ultrasonographic Staging of Papillary Thyroid Carcinoma With the Eighth American Joint Committee on Cancer Tumor-Node-Metastasis Staging System. Ultrasound Q 36, 158–163 (2020). https://doi.org:10.1097/RUQ.0000000000000469 Jin, W. X. et al. Prediction of central lymph node metastasis in papillary thyroid microcarcinoma according to clinicopathologic factors and thyroid nodule sonographic features: a case-control study. Cancer Manag Res 10, 3237–3243 (2018). https://doi.org:10.2147/CMAR.S169741 Hei, H., Song, Y. & Qin, J. Individual prediction of lateral neck metastasis risk in patients with unifocal papillary thyroid carcinoma. Eur J Surg Oncol 45, 1039–1045 (2019). https://doi.org:10.1016/j.ejso.2019.02.016 Zheng, W., Wang, K., Wu, J., Wang, W. & Shang, J. Multifocality is associated with central neck lymph node metastases in papillary thyroid microcarcinoma. Cancer Manag Res 10, 1527–1533 (2018). https://doi.org:10.2147/CMAR.S163263 Paes, J. E. et al. The relationship between body mass index and thyroid cancer pathology features and outcomes: a clinicopathological cohort study. J Clin Endocrinol Metab 95, 4244–4250 (2010). https://doi.org:10.1210/jc.2010-0440 Wang, X. et al. Endocrine tumours: familial nonmedullary thyroid carcinoma is a more aggressive disease: a systematic review and meta-analysis. Eur J Endocrinol 172, R253-262 (2015). https://doi.org:10.1530/EJE-14-0960 Yin, D.-t. et al. The association between thyroid cancer and insulin resistance, metabolic syndrome and its components: a systematic review and meta-analysis. International Journal of Surgery 57, 66–75 (2018). Yahagi, M. et al. Smoking is a risk factor for pulmonary metastasis in colorectal cancer. Colorectal Dis 19, O322-O328 (2017). https://doi.org:10.1111/codi.13833 Foerster, B. et al. Association of Smoking Status With Recurrence, Metastasis, and Mortality Among Patients With Localized Prostate Cancer Undergoing Prostatectomy or Radiotherapy: A Systematic Review and Meta-analysis. JAMA Oncol 4, 953–961 (2018). https://doi.org:10.1001/jamaoncol.2018.1071 Maeda, M., Nagawa, H., Maeda, T., Koike, H. & Kasai, H. Alcohol consumption enhances liver metastasis in colorectal carcinoma patients. Cancer 83, 1483–1488 (1998). https://doi.org:10.1002/(sici)1097-0142(19981015)83:83.0.co;2-z Shickel, B., Tighe, P. J., Bihorac, A. & Rashidi, P. Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis. IEEE J Biomed Health Inform 22, 1589–1604 (2018). https://doi.org:10.1109/JBHI.2017.2767063 Kehl, K. L. et al. Assessment of Deep Natural Language Processing in Ascertaining Oncologic Outcomes From Radiology Reports. JAMA Oncol 5, 1421–1429 (2019). https://doi.org:10.1001/jamaoncol.2019.1800 Pilehvar, M. T. in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2151–2156. Artetxe, M., Labaka, G. & Agirre, E. in Thirty-Second AAAI Conference on Artificial Intelligence. Song, Y., Shi, S., Li, J. & Zhang, H. in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 175–180. Goodfellow, I., McDaniel, P. & Papernot, N. Making machine learning robust against adversarial inputs. Communications of the ACM 61, 56–66 (2018). Zhang, L., Shah, S. K. & Kakadiaris, I. A. Hierarchical multi-label classification using fully associative ensemble learning. Pattern Recognition 70, 89–103 (2017). Li, Y. Deep reinforcement learning: An overview. arXiv preprint arXiv:1701.07274 (2017). Chu, F., Yuan, S. & Peng, Z. Machine learning techniques. Encyclopedia of Structural Health Monitoring (2009). Chen, T. et al. Xgboost: extreme gradient boosting. R package version 0.4-2 1, 1–4 (2015). Zhou, T., Ishibuchi, H. & Wang, S. Stacked blockwise combination of interpretable TSK fuzzy classifiers by negative correlation learning. IEEE Transactions on Fuzzy Systems 26, 3327–3341 (2018). Ryoo, I. et al. Analysis of postoperative ultrasonography surveillance after total thyroidectomy in patients with papillary thyroid carcinoma: a multicenter study. Acta Radiol 59, 196–203 (2018). https://doi.org:10.1177/0284185117700448 Mulla, M. & Schulte, K. M. Central cervical lymph node metastases in papillary thyroid cancer: a systematic review of imaging-guided and prophylactic removal of the central compartment. Clin Endocrinol (Oxf) 76, 131–136 (2012). https://doi.org:10.1111/j.1365-2265.2011.04162.x Moo, T. A. et al. Impact of prophylactic central neck lymph node dissection on early recurrence in papillary thyroid carcinoma. World J Surg 34, 1187–1191 (2010). https://doi.org:10.1007/s00268-010-0418-3 Kim, S. K. et al. Nomogram for predicting central node metastasis in papillary thyroid carcinoma. J Surg Oncol 115, 266–272 (2017). https://doi.org:10.1002/jso.24512 Obermeyer, Z. & Emanuel, E. J. Predicting the Future - Big Data, Machine Learning, and Clinical Medicine. N Engl J Med 375, 1216–1219 (2016). https://doi.org:10.1056/NEJMp1606181 Rajkomar, A. et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med 1, 18 (2018). https://doi.org:10.1038/s41746-018-0029-1 De Fauw, J. et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat Med 24, 1342–1350 (2018). https://doi.org:10.1038/s41591-018-0107-6 Bollschweiler, E. H. et al. Artificial neural network for prediction of lymph node metastases in gastric cancer: a phase II diagnostic study. Ann Surg Oncol 11, 506–511 (2004). https://doi.org:10.1245/ASO.2004.04.018 Tseng, C.-J., Lu, C.-J., Chang, C.-C. & Chen, G.-D. Application of machine learning to predict the recurrence-proneness for cervical cancer. Neural Computing and Applications 24, 1311–1316 (2014). https://doi.org:10.1007/s00521-013-1359-1 Singal, A. G. et al. Machine learning algorithms outperform conventional regression models in predicting development of hepatocellular carcinoma. Am J Gastroenterol 108, 1723–1730 (2013). https://doi.org:10.1038/ajg.2013.332 Kim, W. et al. Development of novel breast cancer recurrence prediction model using support vector machine. J Breast Cancer 15, 230–238 (2012). https://doi.org:10.4048/jbc.2012.15.2.230 Liang, J. D. et al. Recurrence predictive models for patients with hepatocellular carcinoma after radiofrequency ablation using support vector machines with feature selection methods. Comput Methods Programs Biomed 117, 425–434 (2014). https://doi.org:10.1016/j.cmpb.2014.09.001 Bur, A. M. et al. Machine learning to predict occult nodal metastasis in early oral squamous cell carcinoma. Oral Oncol 92, 20–25 (2019). https://doi.org:10.1016/j.oraloncology.2019.03.011 Sattlecker, M., Bessant, C., Smith, J. & Stone, N. Investigation of support vector machines and Raman spectroscopy for lymph node diagnostics. Analyst 135, 895–901 (2010). Fan, X. J. et al. Epithelial-mesenchymal transition biomarkers and support vector machine guided model in preoperatively predicting regional lymph node metastasis for rectal cancer. Br J Cancer 106, 1735–1741 (2012). https://doi.org:10.1038/bjc.2012.82 Mocellin, S. et al. Support vector machine learning model for the prediction of sentinel node status in patients with cutaneous melanoma. Ann Surg Oncol 13, 1113–1122 (2006). https://doi.org:10.1245/ASO.2006.03.019 Hart, G. R. et al. Population-Based Screening for Endometrial Cancer: Human vs. Machine Intelligence. Front Artif Intell 3, 539879 (2020). https://doi.org:10.3389/frai.2020.539879 Lebedev, A. et al. Random Forest ensembles for detection and prediction of Alzheimer's disease with a good between-cohort robustness. NeuroImage: Clinical 6, 115–125 (2014). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3909203","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":269959692,"identity":"df29f552-9fca-49ec-b1db-09968743c3d4","order_by":0,"name":"JingWen Zhang","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"JingWen","middleName":"","lastName":"Zhang","suffix":""},{"id":269959693,"identity":"086fa3d1-f3e6-4b85-b627-8e054aa43f39","order_by":1,"name":"XiaoWen Zhang","email":"","orcid":"","institution":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100080","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"XiaoWen","middleName":"","lastName":"Zhang","suffix":""},{"id":269959694,"identity":"722f0ccc-60a0-4eb5-ad04-962103d8e759","order_by":2,"name":"ShuJun Xia","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"ShuJun","middleName":"","lastName":"Xia","suffix":""},{"id":269959695,"identity":"810bc6cf-1032-4329-9176-21c43237a1f3","order_by":3,"name":"YiJie Dong","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"YiJie","middleName":"","lastName":"Dong","suffix":""},{"id":269959696,"identity":"7c95949e-1baa-4cfc-b53e-a72f5e0d81e4","order_by":4,"name":"Wei Zhou","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wei","middleName":"","lastName":"Zhou","suffix":""},{"id":269959697,"identity":"b5f54702-2374-4318-82c0-1fb7bd6f485e","order_by":5,"name":"ZhenHua Liu","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"ZhenHua","middleName":"","lastName":"Liu","suffix":""},{"id":269959698,"identity":"a24f7f39-765e-48a2-8d1a-9f259efaef05","order_by":6,"name":"Lu Zhang","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lu","middleName":"","lastName":"Zhang","suffix":""},{"id":269959699,"identity":"36423d05-8d49-4745-86de-cff37f82ebc8","order_by":7,"name":"WeiWei Zhan","email":"","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"WeiWei","middleName":"","lastName":"Zhan","suffix":""},{"id":269959700,"identity":"079228bd-df09-476c-a80e-3f414f9bbe6c","order_by":8,"name":"YuZhong Sun","email":"","orcid":"","institution":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100080","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"YuZhong","middleName":"","lastName":"Sun","suffix":""},{"id":269959701,"identity":"0de8052f-e25c-492f-8ec8-83fd5cade2f9","order_by":9,"name":"JianQiao Zhou","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAu0lEQVRIiWNgGAWjYDACCTBZw8PPzHzwASlajslJtrMlG5CihdnY4DyPmQBROuRnNx+T/NnGlrj5MIMZ0IE20QS1MM45libN2yaTuO0wQ9oDhmNpuQ2EtDBL5JhJMwJtAWo5bsDYcJiwFjagFqDDmBM3NzO2SRClhQeoRYK3Deh9ZmY24rRISKQlW/OcOyYncZiN2SCBGL/Iz0g+ePNHGTAq+89/fPChxoawFlSQQJryUTAKRsEoGAW4AAB7RTdpKnCypQAAAABJRU5ErkJggg==","orcid":"","institution":"Ruijin Hospital, Shanghai Jiaotong University School of Medicine","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"JianQiao","middleName":"","lastName":"Zhou","suffix":""}],"badges":[],"createdAt":"2024-01-29 16:00:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3909203/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3909203/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":50455911,"identity":"219fc4c5-694d-43bf-80d6-d7558b77d471","added_by":"auto","created_at":"2024-01-31 18:49:51","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":13379,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTraining Flowchart for LNM Prediction in PTC by NLP and Deep Learning Models\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-3909203/v1/57e71b8d6842e19f6da6fd3a.png"},{"id":50455912,"identity":"189605df-163a-4cbd-97ae-ef16ba35b5ff","added_by":"auto","created_at":"2024-01-31 18:49:51","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":215636,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe ROC curve of different machine learning models (A: LNM; B: CLNM; C:LLNM)\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-3909203/v1/5d6bb7e9fd32001123f780f0.png"},{"id":52488891,"identity":"5ac80599-0f88-406c-a38c-597b4425adb0","added_by":"auto","created_at":"2024-03-12 08:08:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":574706,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3909203/v1/8d87f070-f835-4ed2-b869-75a86facc446.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Prediction Model for Lymph Node Metastasis in Papillary Thyroid Carcinoma Based on Electronic Medical Records","fulltext":[{"header":"Introduction","content":"\u003cp\u003ePapillary thyroid carcinoma (PTC) is the most common pathological type of thyroid cancer, accounting for more than 90% of thyroid tumors \u003csup\u003e1\u003c/sup\u003e. While the majority of PTC patients experience an indolent clinical course and exhibit a favorable prognosis, lymph node metastases are commonly detected at the time of diagnosis. Surgery remains the primary treatment option, with the extent of surgery being determined by the preoperative assessment of lymph node status \u003csup\u003e2\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eLymph node metastasis (LNM) of papillary thyroid carcinoma (PTC) is a significant risk factor for regional recurrence \u003csup\u003e3,4\u003c/sup\u003e., with the incidence rate of 60\u0026ndash;70%\u003csup\u003e5\u003c/sup\u003e. At present, LNM is primarily evaluated through preoperative imaging examinations, predominantly ultrasonography. However, the usefulness of pre-operative ultrasonography (US) in assessing LNM is frequently limited \u003csup\u003e6\u0026ndash;9\u003c/sup\u003e. Therefore, a more objective and accurate method was needed to predict neck lymph node metastasis in PTC patients.\u003c/p\u003e \u003cp\u003eMultiple studies have demonstrated that PTC lymph node metastasis (LNM) is influenced by various factors, including the patient's gender, age, tumor size, location, degree of extrathyroidal extension (ETE), and multifocal disease \u003csup\u003e3,10\u0026ndash;12\u003c/sup\u003e. Furthermore, Paes et al. \u003csup\u003e13\u003c/sup\u003e propose that cancers are more aggressive in PTC patients with higher body mass indices (BMIs), leading to increased LNM rates. Some studies indicate that patients with second-generation familial non-medullary thyroid carcinoma may exhibit aggressive tumor behaviors, including cervical lymph node metastasis and recurrence \u003csup\u003e14\u003c/sup\u003e. Additionally, research demonstrates that female patients with a history of breast cancer are more likely to experience cervical lymph node metastasis. Although no relevant reports have shown that LNM of PTC is associated with metabolic disease, some reports have shown that metabolic abnormalities(diabetes, hypertension, high BMI)may significantly increase the risk of thyroid cancer \u003csup\u003e15\u003c/sup\u003e. Unhealthy personal habits, such as smoking and consuming alcohol, may also contribute to cancer metastasis \u003csup\u003e16\u0026ndash;18\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn modern medicine, the electronic medical record (EMR) system has developed rapidly, resulting in a significant increase in computerized medical data. In addition to the patient demographic data, the EMR information contains a vast amount of textual data containing the patients' chief complaints, medical histories, lifestyle habits, familial cancer histories, and marital and reproductive histories, along with tumor-related information in imaging and pathological reports. However, due to technical constraints, these textual data have often been ignored in previous studies. With the development of artificial intelligence (AI), machine learning (ML) has become a potentially powerful tool for mining EMR text data to facilitate disease diagnosis and management. In the medical field, natural language processing (NLP) is widely acknowledged as essential for extracting valuable medical information from textual data \u003csup\u003e19,20\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTherefore, the goal of this study is to establish a novel machine learning model for predicting LNM of patients with PTC, in a cost-effective way by utilizing easily available personal electronic medical data.\u003c/p\u003e"},{"header":"Method","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eMaterials\u003c/h2\u003e \u003cp\u003eThe study included 5076 PTC patients who underwent total thyroidectomy or lobectomy with lymph node dissection at Ruijin Hospital between January 2013 and December 2016. The PMC patients included 1331 males and 3745 females. All patients underwent pre-operative ultrasonography examinations. All patients had no previous history of of neck surgery or irradiation. The research protocol was approved by the Ethics Committee of Ruijin Hospital, Shanghai Jiao Tong University School of Medicine.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eData Collection\u003c/h2\u003e \u003cp\u003eElectronic clinical and pathologic records were retrospectively reviewed to gather clinicopathological information. The following information is collected: patient age, gender, chief complaint, medical history (hypertension, diabetes, hyperlipidemia and malignant tumor), family history of cancer, BMI, smoking and drinking history; US features were recorded: tumor diameter, tumor location (upper pole, middle, lower pole and isthmus), tumor site (left lobe, right lobe and isthmus), and aspect ratio (\u0026lt;\u0026thinsp;1 or \u0026ge;\u0026thinsp;1). Pathological informations were analyzed: Hashimoto's thyroiditis, multifocality (unifocal, multifocal), tumor T stage, and extrathyroidal extension (ETE).\u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003eModel and Training\u003c/h2\u003e \u003cp\u003eFirstly, considering the noise introduced by synonyms (such as \"elevated blood pressure\" and \"hypertension\", \"breast cancer\" and \"malignant breast tumor\", \"physical examination finding\" and \"ultrasonography finding\"), as well as colloquial expression in chief complaints, medical histories and family cancer histories, which may impact the prediction of LNM, we employ word embedding techniques to achieve a standardized and normalized representation of the original data. Specifically, we further clean and preprocess the case data, which has undergone standardized collation and desensitization, and vectorize it using a pre-trained word embedding dictionary\u003csup\u003e21\u003c/sup\u003e.Word embedding is a technique that maps words or letters from one-dimensional space to high-dimensional vector space and represent words by multidimensional vectors so that computers can understand and learn\u003csup\u003e22\u003c/sup\u003e. In this study, \u003cb\u003en\u003c/b\u003e represents the number of features in this segment of data, embed_dim refers to the word embedding dimension set during model training, and each lattice of case data is converted into a word embedding matrix of n \u0026times; embed_dim dimension\u003csup\u003e23\u003c/sup\u003e. The embed_dim is set as 50 during model training, and each word is represented as a 1\u0026times;50 dimensional matrix vector. Each case data is converted into the corresponding word embedding information for matrix stitching, and the stitched matrix is used as the input of DNN model.\u003c/p\u003e \u003cp\u003eThis study intends to achieve LNM prediction in PTC by DNN models and ML algorithms. DNN is a complex neural network composed of multiple perceptron models with many hidden layers, which is commonly used to simulate the information transfer process of numerous neurons in the human brain for knowledge learning and reasoning. DNN can fully train the features of long text data to get an auxiliary model of feature matrix with rich semantic information for result prediction, but it often leads to misleading calculation due to overfitting of data in the prediction process. ML is employed to mine the implied laws from mass data for classification prediction\u003csup\u003e24\u003c/sup\u003e. However, traditional ML cannot do a good job in feature mining and learning for complex feature datasets like texts and sequences. To unite the strengths of DNN and ML, based on the integrated learning approach\u003csup\u003e25\u003c/sup\u003e, this study utilizes the DNN models to learn the features of the case vector matrices, to obtain the feature matrices for various kinds of information of case datasets and applies the ML algorithms to learn and calculate the feature matrices, to achieve LNM prediction in PTC.\u003c/p\u003e \u003cp\u003eTo explore the optimal performance of the proposed LNM prediction model, this study investigates the use of the following ML algorithms as classifiers in the predictive model, including Naive Bayes(NB), Decision Tree (DT), Extreme Gradient Boosting (XGB) and Gradient Boosting Machine (GBM) and Random Forest (RDF)\u003csup\u003e26\u0026ndash;28\u003c/sup\u003e. NB is a probability theory-based algorithm that achieves result prediction by probability accumulation of each type of feature in the case dataset. DT introduces the features of datasets into the tree judgment structure layer by layer, judges a certain type of features on each layer of sub-tree in the overall decision tree, and obtains the final result through judging the input data in multiple rounds. GBM algorithm exploits the residual value of the parent node after judgment as the basis of the child node in the tree structure, so that the tree structure can be calculated in parallel, and at the same time, the connection of different features is strengthened during judgment, which can effectively mine the potential connection of features and improve the model as a whole. The XGB algorithm is optimized from the GBM algorithm during engineering implementation by adding the regular term and the residual constant term of last training to the objective function for model training. Although GBM algorithm and XGB algorithm possess generalization and applicability to complex datasets, their performance on high-dimensional sparse vector space and text data features is not ideal. RDF designs a corresponding classification decision tree for each feature, and makes use of the majority voting mechanism to predict the results under the independent and parallel multi-decision subtree structure, which can enhance model performance in high-dimensional space and complex datasets, lower the sensitivity of the model to abnormal values, and promote the generalization ability of the model in extreme cases. Additionally, this study introduces the Stacking multi-model training framework \u003csup\u003e29\u003c/sup\u003e. To address the performance loss caused by the randomness during the model training process, the Stacking framework trains multiple diverse submodels for the same task. Using machine learning algorithms, it assigns credibility weights to each submodel based on its predictive performance. The ultimate prediction results of the model are collectively determined by multiple submodels. The model training flowchart \u003cb\u003e(\u003c/b\u003eFig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) is shown as below, illustrating the process from case word embedding to prediction result output by DNN learning and ML algorithms.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePatients were divided into two groups based on whether lymph node metastasis was present (positive label) or absent (negative label). The datasets were randomly classified in the ratio of 80:20, 80% of the datasets were randomly selected as the training set for the model and iterative training was performed, and 20% of the datasets were used as the test set for evaluating the trained model.\u003c/p\u003e \u003cp\u003eModel stability was assessed, and the optimal threshold probability for a positive outcome was defined by applying cross validation to models trained within the 80% curation training subset; each of 5 cross-validation models was trained with a random 80% sample of the training data (ie, 64% of total curation data) and evaluated using the remaining 20% of the training data (ie, 16% of total curation data). The sigmoid activation layer output score for defining a predicted outcome as positive was defined as half of the mean best F1 scores calculated by cross validation, where the F1 score is defined as the harmonic mean between precision and recall. For each outcome, an ensemble of the cross-validation models was constructed by taking the simple mean of the sigmoid activation layer outputs from each cross-validation model. The ensemble models were applied to the 10% validation subset for evaluation and iterative tuning.\u003c/p\u003e \u003cp\u003eIn this study, precision, recall and F1 score are used to evaluate the model performance. The definition is as follows: Consider a two-label dataset D with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\left|\\mathbf{D}\\right|\\)\u003c/span\u003e\u003c/span\u003e samples \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\left({\\varvec{x}}_{\\varvec{i}},{\\varvec{Y}}_{\\varvec{i}}\\right), i=1\\dots \\left|\\mathbf{D}\\right|, {\\varvec{Y}}_{\\varvec{i}}\\in \\varvec{L}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\varvec{L}\\)\u003c/span\u003e\u003c/span\u003e denotes a label set. Let \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(H\\)\u003c/span\u003e\u003c/span\u003e be a two-label classifier, while \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\varvec{Z}}_{i}=H\\left({\\varvec{x}}_{\\varvec{i}}\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the prediction label of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\varvec{x}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\text{H}\\)\u003c/span\u003e\u003c/span\u003e. Precision, recall and F1 values can be calculated from the following equations, respectively.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$Precision\\left(H,D\\right)=\\frac{{\\sum }_{i=1}^{\\left|D\\right|}\\left|{Y}_{i}*{Z}_{i}\\right|}{\\sum _{i=1}^{\\left|D\\right|}\\left|{Z}_{i}\\right|}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$Recall\\left(H,D\\right)=\\frac{\\sum _{i=1}^{\\left|D\\right|}\\left|{Y}_{i}*{Z}_{i}\\right|}{\\sum _{i=1}^{\\left|D\\right|}\\left|{Y}_{i}\\right|}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$${F}_{1}\\left(H,D\\right)=\\frac{2*Precision\\left(H,D\\right)*Recall(H,D)}{Precision\\left(H,D\\right)+Recall(H,D)}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn addition, the ROC curve is drawn with the true positive rate (sensitivity) as the y-axis and the false positive rate (1-specificity) as the x-axis, and the area under the curve (AUC) is taken as the classifier evaluation indicator. A higher AUC indicates greater accuracy of the model.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Result","content":"\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eGeneral Information\u003c/h2\u003e \u003cp\u003eThe PMC patients included 1331 males and 3745 females. The average age of the patients was 44.1\u0026thinsp;\u0026plusmn;\u0026thinsp;12.4 years. The mean BMI of the patients was 23.6\u0026thinsp;\u0026plusmn;\u0026thinsp;3.6 (Kg/m\u003csup\u003e2\u003c/sup\u003e). The mean tumor diameter was 10.9\u0026thinsp;\u0026plusmn;\u0026thinsp;7.6 mm; there were 2,261 patients with lymph node metastasis (LNM), 2,815 without LNM; Among them, 2,196 with central lymph node metastasis (CLNM), 2,880 without CLNM; 472 with lateral cervical lymph node metastasis (LLNM) and 4,604 without LLNM.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003ePrediction Performance of Model\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e respectively show the test results of DNN models and five ML algorithms for predicting LNM, CLNM and LLNM of PTC. The RDF model achieves a testing AUC of 0.98, precision of 0.98, recall of 0.95, F1 value of 0.97 in predicting LNM. AUC of 0.98, precision of 0.98, recall of 0.94, F1 value of 0.96 in predicting CLNM. AUC of 0.97, precision of 0.96, recall of 0.92, F1 value of 0.94 in predicting LLNM. The GBM algorithm shows an AUC of 0.95, 0.91 and 0.89 respectively in predicting CLNM, CLNM and LLNM. However, the indicators of NB, DT and XGB algorithms are significantly lower than those of RDF(Figure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eEvaluation Indicators for LNM Prediction in PTC by the Model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c5\" namest=\"c2\"\u003e \u003cp\u003eLNM\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eEvaluation Indicators for CLNM Prediction in PTC by the Model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c5\" namest=\"c2\"\u003e \u003cp\u003eCLNM\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eEvaluation Indicators for LLNM Prediction in PTC by the Model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003eLLNM\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.89\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eWeighting of Feature Matrix for Different Types of Case Data\u003c/h2\u003e \u003cp\u003eIn order to analyze the contribution of different features to the prediction objective in the dataset, this study also calculates the weights on the feature matrix of case data after DNN training. Among them, the weight of gender and multi-focus is higher, being 1.24 and 1.23 respectively, which is followed by the chief complaint (1.04), location of primary focus (0.91) and age (0.82). The three with the lowest weight are tumor diameter (0.18), marital and reproductive history (0.17), and nutritional status (BMI, 0.00), which have essentially no effect on the model's prediction task (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eWeight of Each Features of Case\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eData Feature\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWeight\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGender\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMultifocal cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.23\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChief Complaint\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1.04\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSite (left lobe/right lobe/isthmus)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFamily history of malignant tumor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIndividual history (smoking and drinking)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAspect ratio\u0026thinsp;\u0026ge;\u0026thinsp;1 (Yes/No)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.44\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMedical history\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.43\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtrathyroidal extension site\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.34\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHashimoto's thyroiditis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.33\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLocation (upper/middle/lower pole/isthmus)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eT stage\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.27\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMaximum diameter of tumor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNutritional status (BMI)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThere has been a significant increase in the prevalence of PTC worldwide due to pre-operative sonography. Although most patients with PTC have an indolent clinical course and positive prognosis, the incidence of LNM has been identified as a risk factor for recurrence \u003csup\u003e3,30\u003c/sup\u003e. Currently, physicians rely on preoperative neck ultrasound for assessment of lymph node metastasis in the neck region. However, the sensitivity and specificity of preoperative ultrasound in predicting cervical lymph node metastasis demonstrate notable variability, and the incidence of occult LNM was observed to be as high as 55%\u003csup\u003e6\u003c/sup\u003e. That means the utility of preoperative ultrasonography for lymph node assessment is often limited \u003csup\u003e8,31,32\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAdditionally, several mathematical and statistical models, which combine clinical factors with medical imaging examinations, have been developed to predict LNM in PTC. KIM et al. \u003csup\u003e33\u003c/sup\u003e employed multiple regression analysis, considering variables such as patient age, gender, tumor size, multifocality, bilaterality, and Hashimoto\u0026rsquo;s, to construct a nomogram model prognosticating the risk of CLNM in PTC patients after thyroidectomy, achieving an AUC of 0.70. On the other hand, Jin et al. \u003csup\u003e10\u003c/sup\u003e developed a scoring model for the prediction of CLNM by using multiple regression analysis with correlating factors such as large tumor size, irregular margins and BRAF mutations. This model demonstrated an overall sensitivity of 85.1% and specificity of 75.8%. In summary, these nomograms or scoring systems primarily utilize Multiple Logistic Regression (MLR) based on the amalgamation of various risk factors to predict binary outcomes. However, this approach presents certain limitations. To circumvent overfitting of the dataset, it encompasses only meticulously screened independent variables, enabling the model to anticipate the associated outcome. Furthermore, these models prove inadequate in addressing the prevalent issue of missing values commonly encountered in electronic medical records. Consequently, it is necessary to develop new models that can handle a wider range of variables, reduce the impact of missing data, and provide precision and robustness.\u003c/p\u003e \u003cp\u003eMachine Learning is an application of Artificial Intelligence (AI) that mining and learning from historical data to build predictive models, which can overcome or reduce the limitations of MLR\u003csup\u003e34\u003c/sup\u003e. In the era of \"big data\", ML can help clinicians make appropriate decisions based on large amounts of digital medical information. Previous studies have shown that sophisticated machine learning techniques can construct precise predictive models by utilizing unprocessed data extracted from electronic medical records and medical images\u003csup\u003e35,36\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn recent years, ML has been broadly applied in the metastasis prediction and prognosis of cancer\u003csup\u003e37\u003c/sup\u003e. For instance, Tseng et al.\u003csup\u003e38\u003c/sup\u003e developed a decision tree model to predict the risk of recurrence in cervical cancer. Singa et al.\u003csup\u003e39\u003c/sup\u003e employed the RDF algorithm to predict the onset of liver cancer in cirrhosis patients. Kim\u003csup\u003e40\u003c/sup\u003e and Liang\u003csup\u003e41\u003c/sup\u003e et al. established a Support Vector Machine (SVM) model to discriminate between recurrent and non-recurrent malignant tumors. In predicting regional lymph node metastasis of cancer, Bollschweiler\u003csup\u003e37\u003c/sup\u003e et al. employed a single-layer neural network to predict LNM in gastric cancer, achieving an accuracy of approximately 79%. Takada\u003csup\u003e40\u003c/sup\u003e et al. employed a decision tree model-based data mining approach to predict the risk of axillary lymph node metastasis in breast cancer patients, achieving an Area Under the Curve (AUC) of 0.77. Andr\u0026eacute;s et al.\u003csup\u003e42\u003c/sup\u003e using the decision tree algorithm, predicted the presence of occult LNM in patients with early oral squamous cell carcinoma, achieving an AUC of 0.84. Additionally, the SVM algorithm has displayed a robust classification ability in predicting LNM in breast cancer, cutaneous melanoma, and colon cancer\u003csup\u003e43\u0026ndash;45\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn the real world, cancer occurrence and prognosis can be influenced by many relevant factors such as demographics, family history, age, dietary habits, body weight (obesity), poor lifestyle habits (smoking and alcohol consumption) and environmental exposures. This information can be gathered through routine clinical history data present in electronic medical record systems.\u003c/p\u003e \u003cp\u003eHowever, even for the most skilled clinicians, it is not easy to synthesize the above information. Therefore, machine learning is more effective at this task. Advanced machine learning models can extract potential information from input data and consequently become superior predictive tools. For example, hart et al.\u003csup\u003e46\u003c/sup\u003e used machine learning model to effectively predict the risk of endometrial cancer within 5 years using personal health data. The model involved 952 patients with endometrial cancer, of which 57.2% were defined as high risk and 41.8% as intermediate risk. The AUC of the best algorithm was 0.88.\u003c/p\u003e \u003cp\u003eCurrently, no studies have been found utilizing machine learning algorithms to predict lymph node metastasis in PTC through electronic medical record data. It is hypothesized that machine learning algorithms can enhance the prediction accuracy of lymph node metastasis in patients with PT. This approach was trained and verified using DNN models and machine learning algorithms, with the aim of devising surgical strategies.\u003c/p\u003e \u003cp\u003eIn this study, an innovative multiparameterized ML-based model was formulated and validated to predict LNM in patients with PTC utilizing readily available clinical data extracted from electronic medical records. The findings of this study indicate that the developed model possesses a remarkable capacity for accurately prognosticating cervical lymph node metastasis in patients with PTC. Notably, the RDF-based prediction model performs better in prediction, attaining an accuracy rate of 0.98,0.98,0.96 in the prediction of LNM, CLNM and LLNM. This probably due to its more sophisticated classification decisions and different weighting ratios compared to other algorithms. Conversely, NB, DT and XGB algorithms manifest suboptimal performance in the predictive task. Studies \u003csup\u003e47\u003c/sup\u003e have demonstrated that RDF stands as one of the most precise machine learning models, surpassing other techniques in its ability to handle large amounts of features and extremely non-linear data. Furthermore, RDF performs well in mitigating data noise and its adaptability makes it easier to adapt and integrate with learning algorithms. It is worth emphasizing that the selection of the most appropriate algorithm is contingent upon numerous parameters, including the type of data collected, the size of the data sample and the prediction results.\u003c/p\u003e \u003cp\u003eThe advantages of this study lie in the application of deep learning and deep NLP techniques to rationally apply the information recorded in EMR to obtain the classification results of clinical prediction objectives, which are manifested as the following aspects: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) In this study, the case dataset is vectorially mapped from a low-dimensional space to a high-dimensional space through word embedding techniques to endow the feature matrix representing the case dataset with rich semantic information, which effectively helps the model to learn and mine the potential features of LNM in PTC. (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) The training method of multi-DNN model based on the Stacking framework can utilize the characteristics that the parameters are varied when multiple models are trained with the same task under the same training set, by assigning decision weight to each field model, enhance the generalization ability of the model on the test set, promote the accuracy of the model under the condition of small probability, and ensure the overall prediction effect of the model.\u003c/p\u003e \u003cp\u003eBut this study also has some limitations. First, the model uses DNN model and machine learning algorithms, so clinical interpretation of important features identified by the model may be a challenge. Second, the collected data are all from a single center and the trained model may not perform well in other healthcare institutions. Therefore, multi-center datasets are needed for further training and validation of the model to optimize its diagnostic efficacy and generalizability. Third, not all patients underwent total thyroidectomy and some occult multifocal cancers would be missed. Third, EMR is unstructured data. Disease history and family history were dictated by patients and recorded by doctors, and there may be bias in patients' description of diseases. Further research attempts to augment the sample size, incorporate additional data (e.g., ultrasound images or serological examination), and further improve the model\u0026rsquo;s performance through external data validation.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn this study, we constructed a predictive model based on Deep Neural Networks (DNN) and machine learning algorithms, utilizing easily accessible personal electronic medical data to predict Lymph Node Metastasis (LNM) in Papillary Thyroid Carcinoma (PTC). This approach explores the application of artificial intelligence techniques in predicting lymph node metastasis in PTC patients.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor contributions\u0026nbsp;\u003c/strong\u003eConcept and design: JQZ, YZS and ZWW. Acquisition of data: JWZ, ZHL, YJD, WZ, LZ and XJX. Analysis and interpretation of data: XWZ, YZS. Manuscript writing: XWZ and JWZ. All the authors gave their final approval for the submission of this version and any revised version for publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompliance with ethical standards\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest\u003c/strong\u003e The authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval\u003c/strong\u003e This study design was approved by the Ethics Committee of Ruijin Hospital, Shanghai Jiao Tong University School of Medicine and the need for informed consent was waived due to the retrospective analysis of this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was supported by National Natural Science Foundation of China (82071928)\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eVaccarella, S. \u003cem\u003eet al.\u003c/em\u003e Global patterns and trends in incidence and mortality of thyroid cancer in children and adolescents: a population-based study. Lancet Diabetes Endocrinol 9, 144\u0026ndash;152 (2021). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1016/S2213-8587(20)30401-0\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1016/S2213-8587(20)30401-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHaugen, B. R. \u003cem\u003eet al.\u003c/em\u003e 2015 American Thyroid Association Management Guidelines for Adult Patients with Thyroid Nodules and Differentiated Thyroid Cancer: The American Thyroid Association Guidelines Task Force on Thyroid Nodules and Differentiated Thyroid Cancer. \u003cem\u003eThyroid\u003c/em\u003e 26, 1-133 (2016). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1089/thy.2015.0020\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1089/thy.2015.0020\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSo, Y. K., Kim, M. J., Kim, S. \u0026amp; Son, Y. I. Lateral lymph node metastasis in papillary thyroid carcinoma: A systematic review and meta-analysis for prevalence, risk factors, and location. Int J Surg 50, 94\u0026ndash;103 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1016/j.ijsu.2017.12.029\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1016/j.ijsu.2017.12.029\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStack, B. C., Jr. \u003cem\u003eet al.\u003c/em\u003e American Thyroid Association consensus review and statement regarding the anatomy, terminology, and rationale for lateral neck dissection in differentiated thyroid cancer. Thyroid 22, 501\u0026ndash;508 (2012). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1089/thy.2011.0312\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1089/thy.2011.0312\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQu, H., Sun, G. R., Liu, Y. \u0026amp; He, Q. S. Clinical risk factors for central lymph node metastasis in papillary thyroid carcinoma: a systematic review and meta-analysis. Clin Endocrinol (Oxf) 83, 124\u0026ndash;132 (2015). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1111/cen.12583\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1111/cen.12583\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLim, Y. S. \u003cem\u003eet al.\u003c/em\u003e Lateral cervical lymph node metastases from papillary thyroid carcinoma: predictive factors of nodal metastasis. Surgery 150, 116\u0026ndash;121 (2011). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1016/j.surg.2011.02.003\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1016/j.surg.2011.02.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIto, Y. \u003cem\u003eet al.\u003c/em\u003e Ultrasonographically and anatomopathologically detectable node metastases in the lateral compartment as indicators of worse relapse-free survival in patients with papillary thyroid carcinoma. World J Surg 29, 917\u0026ndash;920 (2005). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1007/s00268-005-7789-x\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1007/s00268-005-7789-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIto, Y. \u003cem\u003eet al.\u003c/em\u003e Clinical significance of metastasis to the central compartment from papillary microcarcinoma of the thyroid. World J Surg 30, 91\u0026ndash;99 (2006). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1007/s00268-005-0113-y\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1007/s00268-005-0113-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, J., Fei, M., Dong, Y., Xu, S. \u0026amp; Zhan, W. Preoperative Ultrasonographic Staging of Papillary Thyroid Carcinoma With the Eighth American Joint Committee on Cancer Tumor-Node-Metastasis Staging System. Ultrasound Q 36, 158\u0026ndash;163 (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1097/RUQ.0000000000000469\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1097/RUQ.0000000000000469\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJin, W. X. \u003cem\u003eet al.\u003c/em\u003e Prediction of central lymph node metastasis in papillary thyroid microcarcinoma according to clinicopathologic factors and thyroid nodule sonographic features: a case-control study. Cancer Manag Res 10, 3237\u0026ndash;3243 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.2147/CMAR.S169741\u003c/span\u003e\u003cspan address=\"https://doi.org:10.2147/CMAR.S169741\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHei, H., Song, Y. \u0026amp; Qin, J. Individual prediction of lateral neck metastasis risk in patients with unifocal papillary thyroid carcinoma. Eur J Surg Oncol 45, 1039\u0026ndash;1045 (2019). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1016/j.ejso.2019.02.016\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1016/j.ejso.2019.02.016\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng, W., Wang, K., Wu, J., Wang, W. \u0026amp; Shang, J. Multifocality is associated with central neck lymph node metastases in papillary thyroid microcarcinoma. Cancer Manag Res 10, 1527\u0026ndash;1533 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.2147/CMAR.S163263\u003c/span\u003e\u003cspan address=\"https://doi.org:10.2147/CMAR.S163263\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePaes, J. E. \u003cem\u003eet al.\u003c/em\u003e The relationship between body mass index and thyroid cancer pathology features and outcomes: a clinicopathological cohort study. J Clin Endocrinol Metab 95, 4244\u0026ndash;4250 (2010). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1210/jc.2010-0440\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1210/jc.2010-0440\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, X. \u003cem\u003eet al.\u003c/em\u003e Endocrine tumours: familial nonmedullary thyroid carcinoma is a more aggressive disease: a systematic review and meta-analysis. Eur J Endocrinol 172, R253-262 (2015). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1530/EJE-14-0960\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1530/EJE-14-0960\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYin, D.-t. \u003cem\u003eet al.\u003c/em\u003e The association between thyroid cancer and insulin resistance, metabolic syndrome and its components: a systematic review and meta-analysis. International Journal of Surgery 57, 66\u0026ndash;75 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYahagi, M. \u003cem\u003eet al.\u003c/em\u003e Smoking is a risk factor for pulmonary metastasis in colorectal cancer. Colorectal Dis 19, O322-O328 (2017). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1111/codi.13833\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1111/codi.13833\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFoerster, B. \u003cem\u003eet al.\u003c/em\u003e Association of Smoking Status With Recurrence, Metastasis, and Mortality Among Patients With Localized Prostate Cancer Undergoing Prostatectomy or Radiotherapy: A Systematic Review and Meta-analysis. JAMA Oncol 4, 953\u0026ndash;961 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1001/jamaoncol.2018.1071\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1001/jamaoncol.2018.1071\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaeda, M., Nagawa, H., Maeda, T., Koike, H. \u0026amp; Kasai, H. Alcohol consumption enhances liver metastasis in colorectal carcinoma patients. Cancer 83, 1483\u0026ndash;1488 (1998). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1002/(sici)1097-0142(19981015)83:8\u0026lt;1483::aid-cncr2\u0026gt;3.0.co;2-z\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1002/(sici)1097-0142(19981015)83:8%3C1483::aid-cncr2%3E3.0.co;2-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShickel, B., Tighe, P. J., Bihorac, A. \u0026amp; Rashidi, P. Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis. IEEE J Biomed Health Inform 22, 1589\u0026ndash;1604 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1109/JBHI.2017.2767063\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1109/JBHI.2017.2767063\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKehl, K. L. \u003cem\u003eet al.\u003c/em\u003e Assessment of Deep Natural Language Processing in Ascertaining Oncologic Outcomes From Radiology Reports. JAMA Oncol 5, 1421\u0026ndash;1429 (2019). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1001/jamaoncol.2019.1800\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1001/jamaoncol.2019.1800\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePilehvar, M. T. in \u003cem\u003eProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers).\u003c/em\u003e 2151\u0026ndash;2156.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArtetxe, M., Labaka, G. \u0026amp; Agirre, E. in \u003cem\u003eThirty-Second AAAI Conference on Artificial Intelligence.\u003c/em\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong, Y., Shi, S., Li, J. \u0026amp; Zhang, H. in \u003cem\u003eProceedings of the\u003c/em\u003e 2018 \u003cem\u003eConference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers).\u003c/em\u003e 175\u0026ndash;180.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoodfellow, I., McDaniel, P. \u0026amp; Papernot, N. Making machine learning robust against adversarial inputs. Communications of the ACM 61, 56\u0026ndash;66 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, L., Shah, S. K. \u0026amp; Kakadiaris, I. A. Hierarchical multi-label classification using fully associative ensemble learning. Pattern Recognition 70, 89\u0026ndash;103 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, Y. Deep reinforcement learning: An overview. \u003cem\u003earXiv preprint arXiv:1701.07274\u003c/em\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChu, F., Yuan, S. \u0026amp; Peng, Z. Machine learning techniques. Encyclopedia of Structural Health Monitoring (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, T. \u003cem\u003eet al.\u003c/em\u003e Xgboost: extreme gradient boosting. \u003cem\u003eR package version 0.4-2\u003c/em\u003e 1, 1\u0026ndash;4 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, T., Ishibuchi, H. \u0026amp; Wang, S. Stacked blockwise combination of interpretable TSK fuzzy classifiers by negative correlation learning. IEEE Transactions on Fuzzy Systems 26, 3327\u0026ndash;3341 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRyoo, I. \u003cem\u003eet al.\u003c/em\u003e Analysis of postoperative ultrasonography surveillance after total thyroidectomy in patients with papillary thyroid carcinoma: a multicenter study. Acta Radiol 59, 196\u0026ndash;203 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1177/0284185117700448\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1177/0284185117700448\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMulla, M. \u0026amp; Schulte, K. M. Central cervical lymph node metastases in papillary thyroid cancer: a systematic review of imaging-guided and prophylactic removal of the central compartment. Clin Endocrinol (Oxf) 76, 131\u0026ndash;136 (2012). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1111/j.1365-2265.2011.04162.x\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1111/j.1365-2265.2011.04162.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoo, T. A. \u003cem\u003eet al.\u003c/em\u003e Impact of prophylactic central neck lymph node dissection on early recurrence in papillary thyroid carcinoma. World J Surg 34, 1187\u0026ndash;1191 (2010). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1007/s00268-010-0418-3\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1007/s00268-010-0418-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, S. K. \u003cem\u003eet al.\u003c/em\u003e Nomogram for predicting central node metastasis in papillary thyroid carcinoma. J Surg Oncol 115, 266\u0026ndash;272 (2017). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1002/jso.24512\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1002/jso.24512\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eObermeyer, Z. \u0026amp; Emanuel, E. J. Predicting the Future - Big Data, Machine Learning, and Clinical Medicine. N Engl J Med 375, 1216\u0026ndash;1219 (2016). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1056/NEJMp1606181\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1056/NEJMp1606181\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajkomar, A. \u003cem\u003eet al.\u003c/em\u003e Scalable and accurate deep learning with electronic health records. NPJ Digit Med 1, 18 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1038/s41746-018-0029-1\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1038/s41746-018-0029-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Fauw, J. \u003cem\u003eet al.\u003c/em\u003e Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat Med 24, 1342\u0026ndash;1350 (2018). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1038/s41591-018-0107-6\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1038/s41591-018-0107-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBollschweiler, E. H. \u003cem\u003eet al.\u003c/em\u003e Artificial neural network for prediction of lymph node metastases in gastric cancer: a phase II diagnostic study. Ann Surg Oncol 11, 506\u0026ndash;511 (2004). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1245/ASO.2004.04.018\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1245/ASO.2004.04.018\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTseng, C.-J., Lu, C.-J., Chang, C.-C. \u0026amp; Chen, G.-D. Application of machine learning to predict the recurrence-proneness for cervical cancer. Neural Computing and Applications 24, 1311\u0026ndash;1316 (2014). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1007/s00521-013-1359-1\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1007/s00521-013-1359-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSingal, A. G. \u003cem\u003eet al.\u003c/em\u003e Machine learning algorithms outperform conventional regression models in predicting development of hepatocellular carcinoma. Am J Gastroenterol 108, 1723\u0026ndash;1730 (2013). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1038/ajg.2013.332\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1038/ajg.2013.332\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, W. \u003cem\u003eet al.\u003c/em\u003e Development of novel breast cancer recurrence prediction model using support vector machine. J Breast Cancer 15, 230\u0026ndash;238 (2012). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.4048/jbc.2012.15.2.230\u003c/span\u003e\u003cspan address=\"https://doi.org:10.4048/jbc.2012.15.2.230\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang, J. D. \u003cem\u003eet al.\u003c/em\u003e Recurrence predictive models for patients with hepatocellular carcinoma after radiofrequency ablation using support vector machines with feature selection methods. Comput Methods Programs Biomed 117, 425\u0026ndash;434 (2014). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1016/j.cmpb.2014.09.001\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1016/j.cmpb.2014.09.001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBur, A. M. \u003cem\u003eet al.\u003c/em\u003e Machine learning to predict occult nodal metastasis in early oral squamous cell carcinoma. Oral Oncol 92, 20\u0026ndash;25 (2019). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1016/j.oraloncology.2019.03.011\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1016/j.oraloncology.2019.03.011\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSattlecker, M., Bessant, C., Smith, J. \u0026amp; Stone, N. Investigation of support vector machines and Raman spectroscopy for lymph node diagnostics. Analyst 135, 895\u0026ndash;901 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan, X. J. \u003cem\u003eet al.\u003c/em\u003e Epithelial-mesenchymal transition biomarkers and support vector machine guided model in preoperatively predicting regional lymph node metastasis for rectal cancer. Br J Cancer 106, 1735\u0026ndash;1741 (2012). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1038/bjc.2012.82\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1038/bjc.2012.82\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMocellin, S. \u003cem\u003eet al.\u003c/em\u003e Support vector machine learning model for the prediction of sentinel node status in patients with cutaneous melanoma. Ann Surg Oncol 13, 1113\u0026ndash;1122 (2006). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.1245/ASO.2006.03.019\u003c/span\u003e\u003cspan address=\"https://doi.org:10.1245/ASO.2006.03.019\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHart, G. R. \u003cem\u003eet al.\u003c/em\u003e Population-Based Screening for Endometrial Cancer: Human vs. Machine Intelligence. Front Artif Intell 3, 539879 (2020). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org:10.3389/frai.2020.539879\u003c/span\u003e\u003cspan address=\"https://doi.org:10.3389/frai.2020.539879\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLebedev, A. \u003cem\u003eet al.\u003c/em\u003e Random Forest ensembles for detection and prediction of Alzheimer's disease with a good between-cohort robustness. NeuroImage: Clinical 6, 115\u0026ndash;125 (2014).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"papillary thyroid carcinoma, lymph node metastasis, machine learning algorithm, Electronic Medical Records","lastPublishedDoi":"10.21203/rs.3.rs-3909203/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3909203/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003ePurpose\u003c/h2\u003e \u003cp\u003eThis study aimed to establish a novel machine learning model for predicting lymph node metastasis(LNM)of patients with papillary thyroid carcinoma (PTC) by utilizing personal electronic medical records (EMR) data.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThe study included 5076 PTC patients underwent total thyroidectomy or lobectomy with lymph node dissection. Based on the integrated learning approach, this study designed a predictive model for LNM. The predictive model employs deep neural network (DNN) models to identify features within cases and vectorize clinical data from electronic medical records into feature matrices. Subsequently, a classifier based on machine learning algorithms is designed to analyse the feature matrices for prediction LNM in PTC. To mitigate the risk of overfitting commonly associated with machine learning algorithms processing high-dimensional matrices, multiple DNNS are utilized to distribute the overfitting risk. Five mainstream machine learning algorithms (NB, DT, XGB, GBM, RDF) are tested as classifier algorithms in the predictive model. Model performance is assessed using precision, recall, F1, and AUC.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eAmong the patients, 2,261 had lymph node metastasis (LNM), with 2,196 displaying central lymph node metastasis (CLNM) and 472 exhibiting lateral cervical lymph node metastasis (LLNM). The RDF model showcased superior predictive performance compared to other models, achieving a testing AUC of 0.98, precision of 0.98, recall of 0.95, and F1 value of 0.97 in predicting LNM. Moreover, it attained an AUC of 0.98, precision of 0.98, recall of 0.94, and an F1 value of 0.96 in predicting CLNM. Regarding the weighting of the feature matrix for various case data types, gender and multi-focus held higher weights, at 1.24 and 1.23 respectively.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThe LNM predictive model proposed in this study could be used as a cost-effective tool for predicting LNM in PTC patients, by utilizing easily available personal electronic medical data, which can provide valuable support to surgeons in devising a personalized treatment plan.\u003c/p\u003e","manuscriptTitle":"Prediction Model for Lymph Node Metastasis in Papillary Thyroid Carcinoma Based on Electronic Medical Records","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-01-31 18:49:46","doi":"10.21203/rs.3.rs-3909203/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"29c74deb-d999-4bbc-bdf9-bb75b7e5fc23","owner":[],"postedDate":"January 31st, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-03-12T08:06:54+00:00","versionOfRecord":[],"versionCreatedAt":"2024-01-31 18:49:46","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3909203","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3909203","identity":"rs-3909203","version":["v1"]},"buildId":"cTy_lsJlmDsVRNrSptgXS","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.