Machine learning approaches for prediction of thyroid cancer recurrence using thyroglobulin level, whole body scan

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background- Although thyroid cancer generally has a good prognosis, some patients are prone to recurrence. Multiple factors influence recurrence risk. Machine learning (ML) algorithms offer potential for more accurate and precise prediction models. The aim of the present study is to evaluate recurrence-related factors in thyroid cancer patients using ML algorithms. Methods- This retrospective cohort study included patients with differentiated thyroid cancer referred to a specialized endocrinology clinic, between 2013 and 2023. Demographic data, tumor characteristics, and treatment details were extracted from medical records. Six ML algorithm were employed including logistic regression, Naïve Bayes classifier, decision tree, random forest, XGBoost and LightGBM. Results- A total 355 patients were included (mean age: 41.6914.04 years, 84.22% female). Among ML algorithms, LightGBM demonstrated superior predictive performance, achieving an accuracy of 95.41%, precision of 88.84%, recall of 84.25%, specificity of 97.89%, and an area under the curve of 97.28%. The top five predictors were first-year thyroglobulin level, first response to treatment, age, primary tumor characteristics, and regional lymph nodes involvement, respectively. Conclusion- This study demonstrated that ML algorithms had strong capability to identify thyroid cancer patients at risk of recurrence.
Full text 90,142 characters · extracted from preprint-html · click to expand
Machine learning approaches for prediction of thyroid cancer recurrence using thyroglobulin level, whole body scan | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine learning approaches for prediction of thyroid cancer recurrence using thyroglobulin level, whole body scan Reza Nouri, Sajjad Farashi, Erfan Ayubi, Shiva Borzouei This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7013702/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background- Although thyroid cancer generally has a good prognosis, some patients are prone to recurrence. Multiple factors influence recurrence risk. Machine learning (ML) algorithms offer potential for more accurate and precise prediction models. The aim of the present study is to evaluate recurrence-related factors in thyroid cancer patients using ML algorithms. Methods- This retrospective cohort study included patients with differentiated thyroid cancer referred to a specialized endocrinology clinic, between 2013 and 2023. Demographic data, tumor characteristics, and treatment details were extracted from medical records. Six ML algorithm were employed including logistic regression, Naïve Bayes classifier, decision tree, random forest, XGBoost and LightGBM. Results- A total 355 patients were included (mean age: 41.6914.04 years, 84.22% female). Among ML algorithms, LightGBM demonstrated superior predictive performance, achieving an accuracy of 95.41%, precision of 88.84%, recall of 84.25%, specificity of 97.89%, and an area under the curve of 97.28%. The top five predictors were first-year thyroglobulin level, first response to treatment, age, primary tumor characteristics, and regional lymph nodes involvement, respectively. Conclusion- This study demonstrated that ML algorithms had strong capability to identify thyroid cancer patients at risk of recurrence. thyroid cancer recurrence thyroglobulin machine learning Figures Figure 1 Figure 2 Introduction Thyroid cancer is the most common malignancy of the endocrine system, accounting for approximately 820,000 new cases worldwide and ranking as the 7th most common cancer in 2022 ( 1 ). There is a geographical inequality in the incidence of thyroid cancer around the world, with higher rates observed in high-income countries ( 1 , 2 ). Recent increases in the thyroid cancer incidence rate in low- and middle-income countries have also emerged as a public health challenge. Over the past three decades, thyroid cancer incidence in North Africa and the Middle East has increased by 400%. For example, Iran experienced a sixfold increase in thyroid cancer incidence between 1990 and 2019 ( 3 ). Thyroid cancer generally has a good prognosis, ranking as the 24th most common cause of cancer death globally, with approximately 47,000 deaths annually ( 1 ). However, outcomes such as recurrence are still a significant challenge in some patients. A combination of factors related to treatment approach, tumor characteristics, clinicopathological features, and demographic variables has been identified as predictors of recurrence ( 4 – 8 ). These factors vary in their strength of association and predictive performance across populations ( 4 , 8 ). Stimulated and non-stimulated thyroglobulin tests are commonly used for long-term monitoring of cancer recurrence ( 9 , 10 ). In addition, postoperative stimulated thyroglobulin level is considered a crucial marker for more effective risk stratification in differentiated thyroid cancer ( 11 ). It is hypothesized that combining stimulated and non-stimulated thyroglobulin tests with PET/CT imaging, tumor characteristics, and demographic data can improve the accuracy and precision of recurrence prediction. Accurate assessment of the impact of predictors on outcome prediction requires rigorously developed and validated predictive models. Advances in bioinformatics and machine learning (ML) algorithms have enabled researchers to intelligently develop models that statistically efficiently predict the occurrence of outcomes across various medical fields. ML algorithms identify complex patterns within data, while capturing intricate interactions between variables and the output ( 12 , 13 ). ML has been employed to predict outcomes such as incidence, malignancy, and metastasis in thyroid cancer research ( 14 – 20 ). However, studies focusing on recurrence prediction using these algorithms remain scarce ( 21 ). Given this gap, the present study aims to predict thyroid cancer recurrence by integrating treatment-related factors, tumor characteristics, and demographic variables, applying ML algorithms in a sample of Iranian thyroid cancer patients. Material and Methods Participants and variables In the present study, demographic and clinical data of 355 thyroid cancer patients (56 male and 299 females; mean age of 41.69 years) were recorded. Variables collected included age of participant, gender, history of thyroid cancer in first degree relatives, pathology type, the type of thyroid disease, TNM (Tumor, Node, Metastasis) system of thyroid cancer, thyroid cancer stage, response to treatment, the outcome of treatment, whole body scan, and thyroglobulin value were gathered. Descriptions of these variables, their possible categories, and distributions within the cohort are provided in Table 1 . To categorize age variable, available samples were divided into three groups including adolescent and young adults (≤ 35 years), middle-aged (35 to 55 years), and older adults (> 55 years old). Furthermore, in order to have a more balanced data, following categorization were used. Pathology type was considered to be papillary thyroid carcinoma (n = 283) and others (including follicular carcinoma, Hürthle cell, and papillary thyroid microcarcinoma, n = 71), thyroid disease was considered to be euthyroid (n = 329), hyperthyroid (n = 8), and hypothyroid (n = 18), primary tumor (T) stage was considered to be five categories including (T1a (n = 42), T1b (n = 85), T2(n = 166), T3a(n = 52), T3b( 5 ), and T4(n = 5)), regional lymph node (N) involvement was considered to be two groups including (N0 (n = 248) and N1 (n = 107)), Distant metastasis (M) was divided into two distinct groups including (M0(n = 339), M1(n = 16)), and recurrence status (number of recurrence) was considered to be no recurrence (n = 292) and recurrence happened (n = 81). According to a one-hot-encoding strategy, the categorical variables were transformed into numerical vectors suitable for ML approaches. Study Design This retrospective cohort study included all patients with differentiated thyroid cancer referred to the specialized endocrinology clinic at Hamadan University of Medical Sciences over a 10-year period (2013 to 2023). For each participant, a set of data including age, gender (male, female), family history of thyroid cancer in first-degree (yes, no), the type of thyroid disease (euthyroid, hyperthyroid, hypothyroid), pathology type (MPTC, PTC, FTC, Hürthle cell), cancer stage (I, II,III,IV), classification of tumors according to TNM system (primary tumor, regional lymph nodes, distant metastasis), response to treatment according to the recent guidelines including excellent response, biochemical incomplete response, structural incomplete response, or indeterminate response ( 22 ), the values for stimulated or non-stimulated thyroglobulin (Tg), whole body scan result (positive or negative), tumor recurrence status, the location of recurrence, and the final outcome i.e. alive or dead due to thyroid cancer. Although some patients underwent longitudinal follow-up, only the firs-year thyroglobulin levels and whole-body scan results were used in this study. Machine learning This study aims to investigate the association between thyroid cancer recurrence and demographic characteristics (age, gender, history of thyroid cancer in first-degree) as well as clinical features of thyroid cancer (type of thyroid disease, type of pathology, stage of tumor, recurrence rate, primary tumor, nodes involvement) and thyroglobulin levels, and whole-body scan result. To identify these associations, both statistical and ML approaches were incorporated. In supervised ML methods, a dataset with known labeled samples was used to train and test a predictive model for finding the label of a blind sample. By using thyroglobulin level, whole body scan and ultrasonic test during follow-up period labeled the patient as “recurrence” or “non-recurrence”. This study incorporated multiple ML algorithms including logistic regression, Naïve Bayes classifier, decision tree, random forest, XGBoost ( 23 ), and LightGBM ( 24 ). For all ML methods, twenty-five percent of data was used for classifier optimization using a cross-validated grid search or randomized search method. In such methods, a predefined values for hyperparameters were passed to the GridSearchCV or RandomizedSearchCV functions of Sklearn module of python. This function finds the optimized parameters by implementing a “fit” and a “score” method. The predefined parameters for different classifiers were as follows. For logistic regression ['l1', 'l2'] as penalty, regularization parameter in [-4, 4] range (50 values), for Naïve Bayes, the portion of the biggest variance of all variables that is added to variances for calculation stability was optimized in log (0,-9); for decision tree the criterion method was optimized between ‘gini’ and ‘entropy’, and max tree depth was optimized in [2, 4, 6, 8, 10, 12] vector. For random forest the predefined parameters were number of estimators: [25, 50, 100, 150], 'max_features': ['sqrt', 'log2', None], maximum depth: [3, 6, 15], maximum leaf nodes: [3, 6, 9]; for XGBoost the predefined parameter set were number of estimators: stats.randint(150, 300), learning rate: stats.uniform(0.01, 0.5), 'subsample': stats.uniform(0.3, 0.6), maximum of depth: [3, 6, 9], 'colsample_bytree': stats.uniform(0.5, 0.4), 'min_child_weight': [1, 2, 4], and finally for LightGBM the predefined parameters were number of levels: [5, 20, 31], learning rate: [0.05, 0.1, 0.2], number of estimators: [50, 100, 150]. For random forest, XGBoost, and LightGBM a random search strategy was used. Following hyperparameter optimization, the dataset was split into testing and training subsets. A 5-fold stratified cross-validation without shuffling was implemented to be assure about preventing class-imbalance for training and testing the classifier. Models were trained on the training samples, with performance evaluated through 10-fold cross-validation using accuracy, recall, precision, and area under the receiver operating characteristic curve (AUC-ROC) measures. Finally, the classifier was tested incorporating test samples using stratified 5-fold cross validation. To prevent model overfitting, repeated K-fold cross-validation (K = 5) was implemented and the mean values were reported for all classifier evaluation metrics. ML tools were implemented in python (version 3.12) and different libraries including numpy, pandas, sklearn, scipy, xgboost, and Lightgbm were used. Statistical analysis and performance measures For numerical variables, analysis of variance (ANOVA) was performed to find the significant differences of mean values between groups. Post-hoc analysis using Tukey's HSD was conducted to identify the sources of observed differences. The significance level was adjusted to 0.05. Additionally, to evaluate the performance of discrimination models, measures such as accuracy, precision, recall, sensitivity, and AUC were calculated. The mean (standard deviation) of such values for several iterations were reported accordingly. Results Thyroglobulin value comparison between subgroups Thyroglobulin value (stimulated or non-stimulated) is one the primary biomarkers for thyroid cancer assessment. In the current dataset, thyroglobulin level is the only numerical variable, while other variables are categorical. In this section, statistical analyses for comparing thyroglobulin level between different subgroups (considering other variables) were reported. To mitigate outlier effects, the extreme 10% of values in both distribution tails of thyroglobulin profile were excluded. Furthermore, in the current dataset, variables including “history of thyroid cancer in first-degree relatives”, ‘thyroid disease type’, ‘survival outcome’, ‘TNM stage’, ‘staging’, and ‘gender’ were highly imbalanced. In this regard, subgroup analysis was not conducted for these variables. There is a significant difference between stimulated-Tg value of different age spans (i.e. adolescent and young adults (≤ 35 years of old), middle-aged (36–55), and elderly (> 55 years old)). One-way ANOVA confirmed between-group differences (F (3, 96) = 5.96, p = 0.004). Post-hoc analysis using Tukey's HSD revealed that stimulated-Tg was significantly higher for elderly as compared with other two groups (elderly vs. adolescent and young: mean difference = 25.61, CI: [5.92 45.30], p = 0.007; elderly vs. middle-aged: mean difference = 26.22, CI: [6.62 45.82], p = 0.007). For Non-stimulated-Tg, ANOVA revealed significant differences between groups (F (2,202) = 6.18, p = 0.002). Post-hoc analysis using Tukey's HSD also revealed that Non-stimulated-Tg was significantly higher for elderly as compared with adolescent and young (mean difference = 0.3148, CI: [0.05 0.58], p = 0.02), and middle-aged (mean difference = 0.38, CI: [0.12 0.63], p = 0.002). For both stimulated and non-stimulated-Tg, there was no significant difference between PTC and others thyroid cancer (mean difference=-3.79, CI: [-8.01 0.43], p = 0.07; mean difference=-0.02, CI: [0.71 − 0.16], p = 0.71, respectively). Considering first response to treatment (i.e. ‘excellent response’, ‘biochemical incomplete response’, or ‘structural incomplete response’), a significant difference between subgroups was found by ANOVA for stimulated-Tg value (F(2,149) = 5.62, p = 0.004). The post-hoc Tukey’s HSD revealed that the difference was between ‘excellent response’ and ‘structural incomplete response’ groups (mean difference = 4.1934, CI: [1.01 7.38], p = 0.006), which indicated significantly higher value of thyroglobulin in ‘structural incomplete response’ group. For non-stimulated-Tg, no significant differences were found between subgroups. For TNM_system variable (primary tumor including T1a, T1b, T2, and T3a), ANOVA revealed a significant difference for stimulated-Tg value between groups (F(3,171) = 15.95, p < 0.001). The post-hoc Tukey’s HSD test showed that stimulated-Tg for T3a group was significantly differed with other groups (T3a vs. T1a: mean difference = 18.69, CI: [7.68 29.691], p = 0.0001; T3a vs. T1b: mean difference = 15.6912, CI:[ 8.85 22.53], p < 0.001; T3a vs. T2: mean difference = 19.10, CI:[ 11.82 26.38], p < 0.0001). In all cases, the stimulated-Tg value was higher for T3a subgroup. For non-stimulated-Tg values, no significant differences were found between subgroups. Machine learning A primary objective of this study was to predict the thyroid cancer recurrence using the variables described in Table 1. Models were specially trained and tested according to demographic data and first-year clinical assessments. This means that even though for some patients longitudinal values for data such as thyroglobulin level and whole-body scan were available, only the values for the first year follow-up were used (i.e. thyroglobulin level and whole-body scan for the first year of follow-up). Table 2 reports performance metrics of different ML approaches for predicting the cancer recurrence according to the variables described in Table 1. For some patients, according to the endocrinologist opinion, the value for stimulated-Tg was measured, while for others non-stimulated-Tg level was assessed. For the results in Table 2, stimulated-Tg was used for the thyroglobulin level. When such value was not available, the non-stimulated-Tg was substituted for the thyroglobulin level. In Table 2, the importance ratio column shows the relative weighting of Tg versus whole-body scan obtained by each classifier. Table 2 Performance of machine learning approaches for predicting recurrence of thyroid cancer. Prediction model Accuracy (%) Precision (%) Recall (%) Specificity (%) AUC*100 Importance ratio (Tg/WBS) Logistic regression 90.70 (1.44) 95.00 (6.12) 50.51 (9.37) 99.31 (0.84) 94.94 (3.18) 0.13 Bayes Naïve 88.73 (0.07) 77.78 (0.09) 53.85 (1.00) 96.55 (2.06) 96.49 (0.08) 85390 Decision tree 90.82 (1.96) 88.29 (7.03) 55.95 (7.58) 98.35 (1.14) 91.86 (3.36) 79.75 Random forest 93.75 (2.18) 84.16 (5.96) 79.87 (11.32) 96.71 (1.37) 95.62 (2.86) 24.57 XGBoost 91.55 (2,73) 73.33 (0.80) 84.62 (1.59) 93.10 (1.00) 96.49 (2.03) 0.51 LightGBM 94.16 (2.60) 87.88 (9.32) 81.22 (8.96) 96.92 (2.74) 96.20 (3.44) 24.56 Based on results in Table 2, LightGBM obtained the best classification accuracy. For this reason, the importance of different features for this classifier was plotted in Fig. 1. According to this figure, the most informative features for recurrence prediction were first year thyroglobulin (first year Tg) level, first response to treatment, age, TNM stage, pathology type, gender and, first year whole-body scan. In Fig. 2, the block diagram for LightGBM decision tree is depicted. According to this figure, discriminative levels for thyroid cancer recurrence were thyroglobulin value, first response to treatment, and the Node involvement. The count number shows number of samples in the corresponding leaf. In front of each leaf number are prediction values for no recurrence which can be converted to probability using a sigmoid transformation. Table 3 examined how classification performance was affected when the reduced feature set was used for cancer recurrence prediction. The reduced feature set was selected based on the most-informative features found by LightGBM model (Fig. 1) including age, gender, pathology, TNM stage, first response to treatment, and first year Tg level. Table 3 Prediction performance of a LightGBM classifier for a reduced feature set. Accuracy (%) Precision (%) Recall (%) Specificity (%) AUC*100 95.51 (1.42) 88.84 (9.51) 84.25 (12.19) 97.89 (1.72) 97.28 (1.57) Furthermore, in Table 4, the predictive performance of stimulated-Tg and non-stimulated-Tg values were compared using both reduced and full feature sets. When only non-stimulated-Tg values were used for training and testing the predictor model (according to LightGBM classifier), the prediction performance of the model was not significantly different between full and reduced feature set, while the best performance was achieved for full feature set. However, when stimulated-Tg values were used, the prediction performance was significantly lower for full feature set as compared with reduced feature set (p < 0.05). This result was not dependent to the number of cross-validation repeats. Table 4 Machine learning performance for predicting recurrence based on stimulated or Non-stimulated Tg levels. Accuracy (%) Precision (%) Recall (%) Specificity (%) AUC*100 p-value Non-stimulated Tg Full feature set 94.84 ( 1.58 ) 88.81 ( 9.81 ) 91.59 ( 7.09 ) 96.80 ( 3.00 ) 99.23 ( 1.15 ) AUC: <0.001 Accuracy: <0.001 stimulated Tg Full feature set 88.47 ( 3.62 ) 62.38 ( 23.44 ) 60.00 ( 20.68 ) 93.78 ( 3.52 ) 90.39 ( 5.46 ) Non-stimulated Tg Reduced feature set 92.13 ( 2.84 ) 73.06 ( 11.66 ) 78.41 ( 7.62 ) 94.73 ( 2.51 ) 94.69 ( 2.07 ) AUC: 0.57 Accuracy: 0.08 stimulated Tg Reduced feature set 92.58 ( 1.83 ) 81.56 ( 6.54 ) 79.56 ( 8.20 ) 95.82 ( 1.81 ) 94.27 ( 2.08 ) Discussion The LightGBM classifier achieved the highest accuracy for tumor recurrence prediction (94.16 (2.60) %). Other types of classifiers also attained high classification accuracy (> 88%). This result demonstrates the predictive potential of the variables described in Table 1 for estimating future thyroid tumor recurrence. The predictive performance of the feature set used in the present study was superior compared with similar studies. For example, Schindele et al. proposed an XGBoost model using biomarker and clinical features (including age, primary tumor, and serum Tg level) for thyroid cancer prediction ( 25 ). The area under curve for their model was 0.88 (95% CI: 0.84–0.86). The primary tumor size and Tg levels were the most-informative predictors for thyroid recurrence prediction. Xi et al. used a dataset consisting of demographic information, thyroid test measures, nodules characteristics, and six ML methods including gradient boosting machine, logistic regression, linear discriminant analysis, SVM with radial or linear kernel, and random forest to predict nodule malignancy ( 20 ). According to the results, the random forest model achieved the highest prediction accuracy (78.01%, 95% CI: 0.7670, 0.7930). In another study, Gu et al. proposed a model to predict metastasis of thyroid cancer based on demographic and laboratory test ( 16 ). Using XGBoost model, the best accuracy of that model was 77% for predicting benign/malignant and 70% for lymph node metastasis. Mao et al. utilized a dataset from SEER database and used an XGBoost classifier for suggesting a predicting model for the risk of thyroid cancer recurrence. The model achieved a predictive performance with AUC of 90.4% and identified age, marital status, TNM stage, and surgical methods as the most important risk factors for cancer survival ( 26 ). In Park and Lee study, a dataset was created using demographic, clinicopathological, and ultrasound-extracted parameters and several types of classifiers were used for developing a predictive model. Using a decision tree classifier, maximum F1-score of 28% was achieved, indicating limited prediction performance with the used dataset. Age, tumor size, and lymph node ratio were among the most-informative features ( 21 ). In another study, Kim et al. utilized a dataset compromised of clinical, genetics, laboratory, and pathological data to develop a predictive model using an inductive logic programming ( 27 ). The tumor recurrence prediction accuracy of this model was 71.4%, while the most-informative parameters were Tg level, body mass index, anti-Tg antibody level, TSH, primary tumor, and lymph node involvement ( 27 ). Whole-body scan is a conventional tool for investigating thyroid cancer. However, this approach has several side effects and drawbacks including nausea/vomiting, change in taste and smell, bone marrow suppression ( 28 ), radiation safety concerns, and inability to differentiate benign from malignant tumors ( 29 ). In the present study, the necessity of whole-body scanning for recurrence prediction was evaluated by quantifying its relative importance in the model’s prediction performance. In Table 1 , for each model, the predictive value of thyroglobulin variable was compared with that of whole-body scan, and the importance ratio was reported. According to Table 2 , classifiers including Naïve Bayes, decision tree, random forest, and LightGBM indicated that thyroglobulin had greater predictive capacity than the whole-body scan, whereas other classifiers including logistic regression and XGBoost highlighted the importance of whole-body scan test. Previous studies compared the power of Tg for predicting thyroid cancer recurrence with radionuclide methods. For example, a comparison of the sensitivity of 123I scintigraphy and Tg revealed that whole-body scan did not improve the sensitivity ( 30 ). According to the LightGBM decision tree (Fig. 2 ), a critical level of Tg for tumor recurrence prediction is 4.775. The gain of Tg suggests that it is the primary determinant of recurrence risk. The second level of decision tree indicated the importance of the first response to treatment, where structural incomplete response cases were separated from other subgroups. The third level showed that lymph node involvement was the next determinant risk stratification factor. Based on the terminal leaves, the lowest recurrence risk corresponded to the path to leaf 2 where Tg ≤ 4.775, first response to treatment was excellent or biochemical incomplete response, TNM stage N0(without nodal involvement), and Tg lower than 0.65. The risk for tumor recurrence is higher when Tg > 0.65. The path to leaf 3 (Tg ≤ 4.775 with nodal involvement N1) had lower predictive value for recurrence-free outcome. According to the results of Table 3 , eliminating non-informative features and reducing the size of feature space improved the performance of predictive model in terms of accuracy, precision, recall, and specificity. This result imply that using an eight-dimensional feature space including age, gender, pathology type, TNM stage, first response to treatment, first year Tg and whole-body scan is sufficient to predict the future recurrence of thyroid cancer with high accuracy, specificity, and sensitivity. Conclusion LightGBM demonstrated superior predictive performance for identifying thyroid cancer patients at risk of recurrence, achieving high accuracy and AUC. The most important predictors were first-year thyroglobulin level, first treatment response, age, primary tumor characteristics, and lymph node involvement, respectively. Abbreviations ML Machine learning TNM Tumor, Node, Metastasis Tg Thyroglobulin ANOVA Analysis Of Variance Declarations Acknowledgments This article is the result of a general medicine thesis. We sincerely thank the Hamadan University of Medical Sciences Ethical Research Committee. The work was submitted under the code 140108176791. Ethical Considerations This study was conducted in accordance with the principles of the Helsinki Declaration and was approved by the Ethics Committee of Hamadan University of Medical Sciences (IR.UMSHA.REC.1401.669). Consent to participate - Written informed consent was obtained from all participants prior to data collection. Consent for publication -Not applicable Conflicts of interest/Competing interests -There is nothing to declare. Availability of data and materials - The datasets generated during this study are available from the corresponding author upon reasonable request. Funding- This work was supported by Hamadan University of Medical Sciences, Deputy of research and technology (No. 140108176791). Author contribution- ShB and RN collected the data. SF performed analyses and wrote the initial draft. SF, EA, and ShB discussed the obtained results and finalized the draft. References Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, Jemal A. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229-63. Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71(3):209-49. Nejadghaderi SA, Moghaddam SS, Azadnajafabad S, Rezaei N, Rezaei N, Tavangar SM, et al. Burden of thyroid cancer in North Africa and Middle East 1990-2019. Front Oncol. 2022;12:955358. Guo K, Wang Z. Risk factors influencing the recurrence of papillary thyroid carcinoma: a systematic review and meta-analysis. Int J Clin Exp Pathol. 2014;7(9):5393-403. Hwangbo Y, Kim JM, Park YJ, Lee EK, Lee YJ, Park DJ, et al. Long-Term Recurrence of Small Papillary Thyroid Cancer and Its Risk Factors in a Korean Multicenter Study. J Clin Endocrinol Metab. 2017;102(2):625-33. Ito Y, Kudo T, Kobayashi K, Miya A, Ichihara K, Miyauchi A. Prognostic factors for recurrence of papillary thyroid carcinoma in the lymph nodes, lung, and bone: analysis of 5,768 patients with average 10-year follow-up. World J Surg. 2012;36(6):1274-8. Qi P, Wang Z, Hao X, Ou X, Zhang B, Shi Q, et al. A retrospective study of 17,995 patients investigating the location and recurrence of papillary thyroid cancer. Sci Rep. 2025;15(1):10634. Qu N, Zhang L, Lu ZW, Ji QH, Yang SW, Wei WJ, Zhang Y. Predictive factors for recurrence of differentiated thyroid cancer in patients under 21 years of age and a meta-analysis of the current literature. Tumour Biol. 2016;37(6):7797-808. Jammah AA, Masood A, Akkielah LA, Alhaddad S, Alhaddad MA, Alharbi M, et al. Utility of Stimulated Thyroglobulin in Reclassifying Low Risk Thyroid Cancer Patients' Following Thyroidectomy and Radioactive Iodine Ablation: A 7-Year Prospective Trial. Front Endocrinol (Lausanne). 2020;11:603432. Westbury C, Vini L, Fisher C, Harmer C. Recurrent differentiated thyroid cancer without elevation of serum thyroglobulin. Thyroid. 2000;10(2):171-6. Yang X, Liang J, Li TJ, Yang K, Liang DQ, Yu Z, Lin YS. Postoperative stimulated thyroglobulin level and recurrence risk stratification in differentiated thyroid cancer. Chin Med J (Engl). 2015;128(8):1058-64. Kononenko I. Machine learning for medical diagnosis: history, state of the art and perspective. Artificial Intelligence in Medicine. 2001;23(1):89-109. Shehab M, Abualigah L, Shambour Q, Abu-Hashem MA, Shambour MKY, Alsalibi AI, Gandomi AH. Machine learning in medical applications: A review of state-of-the-art methods. Comput Biol Med. 2022;145:105458. Borzooei S, Briganti G, Golparian M, Lechien JR, Tarokhian A. Machine learning for risk stratification of thyroid cancer patients: a 15-year cohort study. Eur Arch Otorhinolaryngol. 2024;281(4):2095-104. Borzouei S, Safdari A, Ayubi E. Development and Validation of a Clinical Risk Model for Predicting Malignancy in Patients with Thyroid Nodules. Avicenna J Clin Med. 2025;31(4):219-27. Gu J, Xie R, Zhao Y, Zhao Z, Xu D, Ding M, et al. A machine learning-based approach to predicting the malignant and metastasis of thyroid cancer. Front Oncol. 2022;12:938292. Hou F, Zhu Y, Zhao H, Cai H, Wang Y, Peng X, et al. Development and validation of an interpretable machine learning model for predicting the risk of distant metastasis in papillary thyroid cancer: a multicenter study. EClinicalMedicine. 2024;77:102913. Li Y, Wu F, Ge W, Zhang Y, Hu Y, Zhao L, et al. Risk stratification of papillary thyroid cancers using multidimensional machine learning. Int J Surg. 2024;110(1):372-84. Liu W, Wang S, Ye Z, Xu P, Xia X, Guo M. Prediction of lung metastases in thyroid cancer using machine learning based on SEER database. Cancer Med. 2022;11(12):2503-15. Xi NM, Wang L, Yang C. Improving the diagnosis of thyroid cancer by machine learning and clinical data. Sci Rep. 2022;12(1):11143. Park YM, Lee B-J. Machine learning-based prediction model using clinico-pathologic factors for papillary thyroid carcinoma recurrence. Sci Rep. 2021;11(1):4948. Tanaka K, Ozaki T. New TNM classification (AJCC eighth edition) of bone and soft tissue sarcomas: JCOG Bone and Soft Tissue Tumor Study Group. Jpn J Clin Oncol. 2019;49(2):103-7. Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; San Francisco, California, USA: Association for Computing Machinery; 2016. p. 785–94. Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, et al. Lightgbm: A highly efficient gradient boosting decision tree. Adv Neural Inf Process Syst. 2017;30. Schindele A, Krebold A, Heiß U, Nimptsch K, Pfaehler E, Berr C, et al. Interpretable machine learning for thyroid cancer recurrence predicton: Leveraging XGBoost and SHAP analysis. Eur J Radiol. 2025;186:112049. Mao Y, Huang Y, Xu L, Liang J, Lin W, Huang H, et al. Surgical Methods and Social Factors Are Associated With Long-Term Survival in Follicular Thyroid Carcinoma: Construction and Validation of a Prognostic Model Based on Machine Learning Algorithms. Front Oncol. 2022;12:816427. Kim SY, Kim Y-I, Kim HJ, Chang H, Kim S-M, Lee YS, et al. New approach of prediction of recurrence in thyroid cancer patients using machine learning. Medicine. 2021;100(42). Muzahir S, Grady E. Nuclear Imaging and Therapy of Thyroid Disorders. Hall LT editor Molecular Imaging and Therapy. Brisbane (AU): Exon Publications; 2023. Hoang JK, Lee WK, Lee M, Johnson D, Farrell S. US Features of thyroid malignancy: pearls and pitfalls. Radiographics. 2007;27(3):847-60; discussion 61-5. de Geus-Oei LF, Oei HY, Hennemann G, Krenning EP. Sensitivity of 123I whole-body scan and thyroglobulin in the detection of metastases or recurrent differentiated thyroid cancer. Eur J Nucl Med Mol Imaging. 2002;29(6):768-74. Table 1 Table 1 is available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files Table1.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7013702","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":495717999,"identity":"feedcd0b-5c7a-4405-b456-5be5433ec369","order_by":0,"name":"Reza Nouri","email":"","orcid":"","institution":"Hamadan University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Reza","middleName":"","lastName":"Nouri","suffix":""},{"id":495718000,"identity":"b287e0e5-c9fd-4a22-a937-71b535de1d84","order_by":1,"name":"Sajjad Farashi","email":"","orcid":"","institution":"Avicenna Health Research Institute, Hamadan University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Sajjad","middleName":"","lastName":"Farashi","suffix":""},{"id":495718001,"identity":"7684f206-0f14-4fa0-8f46-1fce61367499","order_by":2,"name":"Erfan Ayubi","email":"","orcid":"","institution":"Hamadan University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Erfan","middleName":"","lastName":"Ayubi","suffix":""},{"id":495718002,"identity":"9f7714ba-c469-4305-bec4-3d0e7d64698b","order_by":3,"name":"Shiva Borzouei","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYBACAyjNwyDB2MDAUIEQIVbLGRK0MDBIADFjGxFazNl7H3/4UGEjIz+7uU3i57zD8ubszQcYflRsw6nFsue4geGMM2k8BncOtkn2bjtsuLPnWAJjz5nbuB12I40hmbftMI+BRGKzAe+2w4wbbuQYMDO24ddymLftP4/8jMRmw79zDtsTo4WxmbftAA/DjcTGx7wNhxMJarHsOcbMOONMMo8BSIvMsfTkDWeOJRzE5xdz9jZmYIjZ2cvPSH9w8E2Nte2G480HH/yowK0FHTSDyQNEqweCOlIUj4JRMApGwQgBACuLW4Xu3GlmAAAAAElFTkSuQmCC","orcid":"","institution":"Hamadan University of Medical Sciences","correspondingAuthor":true,"prefix":"","firstName":"Shiva","middleName":"","lastName":"Borzouei","suffix":""}],"badges":[],"createdAt":"2025-06-30 19:23:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7013702/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7013702/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":88506781,"identity":"bcc97db5-c0a7-4400-b297-77d822ea9f3b","added_by":"auto","created_at":"2025-08-07 07:35:06","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":442805,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFeature importance for LighGBM classifier.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7013702/v1/d1097613e97a0b9aefa8df77.png"},{"id":88506850,"identity":"aaa36faf-c5d7-4567-83a9-ff5a3c146342","added_by":"auto","created_at":"2025-08-07 07:35:22","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":129859,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBlock diagram of LightGBM decision tree. Column 12: Thyroglobulin value (stimulated or Non-stimulated), Column 9: first response to treatment (0: biochemical incomplete, 1: Excellent, 2: structural incomplete, 3: indeterminate); Column 6: Nodes involvement (0:N0, 1: N1).\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7013702/v1/0706f3b1487a2e69b3cc5d18.png"},{"id":96808209,"identity":"783a7bb4-2ba2-4af0-b2ff-89607eeda66c","added_by":"auto","created_at":"2025-11-26 09:24:43","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1262456,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7013702/v1/ca0c3c90-2bb6-4f63-8a50-7e828d04bdf8.pdf"},{"id":88506809,"identity":"12c53498-1505-4014-8039-dd8fedad291e","added_by":"auto","created_at":"2025-08-07 07:35:12","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":18570,"visible":true,"origin":"","legend":"","description":"","filename":"Table1.docx","url":"https://assets-eu.researchsquare.com/files/rs-7013702/v1/4dc2c9633829526d34e85fe5.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine learning approaches for prediction of thyroid cancer recurrence using thyroglobulin level, whole body scan","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThyroid cancer is the most common malignancy of the endocrine system, accounting for approximately 820,000 new cases worldwide and ranking as the 7th most common cancer in 2022 (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). There is a geographical inequality in the incidence of thyroid cancer around the world, with higher rates observed in high-income countries (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). Recent increases in the thyroid cancer incidence rate in low- and middle-income countries have also emerged as a public health challenge. Over the past three decades, thyroid cancer incidence in North Africa and the Middle East has increased by 400%. For example, Iran experienced a sixfold increase in thyroid cancer incidence between 1990 and 2019 (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eThyroid cancer generally has a good prognosis, ranking as the 24th most common cause of cancer death globally, with approximately 47,000 deaths annually (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). However, outcomes such as recurrence are still a significant challenge in some patients. A combination of factors related to treatment approach, tumor characteristics, clinicopathological features, and demographic variables has been identified as predictors of recurrence (\u003cspan additionalcitationids=\"CR5 CR6 CR7\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). These factors vary in their strength of association and predictive performance across populations (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eStimulated and non-stimulated thyroglobulin tests are commonly used for long-term monitoring of cancer recurrence (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). In addition, postoperative stimulated thyroglobulin level is considered a crucial marker for more effective risk stratification in differentiated thyroid cancer (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e). It is hypothesized that combining stimulated and non-stimulated thyroglobulin tests with PET/CT imaging, tumor characteristics, and demographic data can improve the accuracy and precision of recurrence prediction.\u003c/p\u003e\u003cp\u003eAccurate assessment of the impact of predictors on outcome prediction requires rigorously developed and validated predictive models. Advances in bioinformatics and machine learning (ML) algorithms have enabled researchers to intelligently develop models that statistically efficiently predict the occurrence of outcomes across various medical fields. ML algorithms identify complex patterns within data, while capturing intricate interactions between variables and the output (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eML has been employed to predict outcomes such as incidence, malignancy, and metastasis in thyroid cancer research (\u003cspan additionalcitationids=\"CR15 CR16 CR17 CR18 CR19\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). However, studies focusing on recurrence prediction using these algorithms remain scarce (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Given this gap, the present study aims to predict thyroid cancer recurrence by integrating treatment-related factors, tumor characteristics, and demographic variables, applying ML algorithms in a sample of Iranian thyroid cancer patients.\u003c/p\u003e"},{"header":"Material and Methods","content":"\u003cp\u003e\u003cstrong\u003eParticipants and variables\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the present study, demographic and clinical data of 355 thyroid cancer patients (56 male and 299 females; mean age of 41.69 years) were recorded. Variables collected included age of participant, gender, history of thyroid cancer in first degree relatives, pathology type, the type of thyroid disease, TNM (Tumor, Node, Metastasis) system of thyroid cancer, thyroid cancer stage, response to treatment, the outcome of treatment, whole body scan, and thyroglobulin value were gathered. Descriptions of these variables, their possible categories, and distributions within the cohort are provided in Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003eTo categorize age variable, available samples were divided into three groups including adolescent and young adults (\u0026le;\u0026thinsp;35 years), middle-aged (35 to 55 years), and older adults (\u0026gt;\u0026thinsp;55 years old). Furthermore, in order to have a more balanced data, following categorization were used. Pathology type was considered to be papillary thyroid carcinoma (n\u0026thinsp;=\u0026thinsp;283) and others (including follicular carcinoma, H\u0026uuml;rthle cell, and papillary thyroid microcarcinoma, n\u0026thinsp;=\u0026thinsp;71), thyroid disease was considered to be euthyroid (n\u0026thinsp;=\u0026thinsp;329), hyperthyroid (n\u0026thinsp;=\u0026thinsp;8), and hypothyroid (n\u0026thinsp;=\u0026thinsp;18), primary tumor (T) stage was considered to be five categories including (T1a (n\u0026thinsp;=\u0026thinsp;42), T1b (n\u0026thinsp;=\u0026thinsp;85), T2(n\u0026thinsp;=\u0026thinsp;166), T3a(n\u0026thinsp;=\u0026thinsp;52), T3b(\u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e), and T4(n\u0026thinsp;=\u0026thinsp;5)), regional lymph node (N) involvement was considered to be two groups including (N0 (n\u0026thinsp;=\u0026thinsp;248) and N1 (n\u0026thinsp;=\u0026thinsp;107)), Distant metastasis (M) was divided into two distinct groups including (M0(n\u0026thinsp;=\u0026thinsp;339), M1(n\u0026thinsp;=\u0026thinsp;16)), and recurrence status (number of recurrence) was considered to be no recurrence (n\u0026thinsp;=\u0026thinsp;292) and recurrence happened (n\u0026thinsp;=\u0026thinsp;81). According to a one-hot-encoding strategy, the categorical variables were transformed into numerical vectors suitable for ML approaches.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStudy Design\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis retrospective cohort study included all patients with differentiated thyroid cancer referred to the specialized endocrinology clinic at Hamadan University of Medical Sciences over a 10-year period (2013 to 2023). For each participant, a set of data including age, gender (male, female), family history of thyroid cancer in first-degree (yes, no), the type of thyroid disease (euthyroid, hyperthyroid, hypothyroid), pathology type (MPTC, PTC, FTC, H\u0026uuml;rthle cell), cancer stage (I, II,III,IV), classification of tumors according to TNM system (primary tumor, regional lymph nodes, distant metastasis), response to treatment according to the recent guidelines including excellent response, biochemical incomplete response, structural incomplete response, or indeterminate response (\u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e), the values for stimulated or non-stimulated thyroglobulin (Tg), whole body scan result (positive or negative), tumor recurrence status, the location of recurrence, and the final outcome i.e. alive or dead due to thyroid cancer. Although some patients underwent longitudinal follow-up, only the firs-year thyroglobulin levels and whole-body scan results were used in this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMachine learning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study aims to investigate the association between thyroid cancer recurrence and demographic characteristics (age, gender, history of thyroid cancer in first-degree) as well as clinical features of thyroid cancer (type of thyroid disease, type of pathology, stage of tumor, recurrence rate, primary tumor, nodes involvement) and thyroglobulin levels, and whole-body scan result. To identify these associations, both statistical and ML approaches were incorporated. In supervised ML methods, a dataset with known labeled samples was used to train and test a predictive model for finding the label of a blind sample. By using thyroglobulin level, whole body scan and ultrasonic test during follow-up period labeled the patient as \u0026ldquo;recurrence\u0026rdquo; or \u0026ldquo;non-recurrence\u0026rdquo;.\u003c/p\u003e\n\u003cp\u003eThis study incorporated multiple ML algorithms including logistic regression, Na\u0026iuml;ve Bayes classifier, decision tree, random forest, XGBoost (\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e), and LightGBM (\u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e). For all ML methods, twenty-five percent of data was used for classifier optimization using a cross-validated grid search or randomized search method. In such methods, a predefined values for hyperparameters were passed to the GridSearchCV or RandomizedSearchCV functions of Sklearn module of python. This function finds the optimized parameters by implementing a \u0026ldquo;fit\u0026rdquo; and a \u0026ldquo;score\u0026rdquo; method. The predefined parameters for different classifiers were as follows. For logistic regression [\u0026apos;l1\u0026apos;, \u0026apos;l2\u0026apos;] as penalty, regularization parameter in [-4, 4] range (50 values), for Na\u0026iuml;ve Bayes, the portion of the biggest variance of all variables that is added to variances for calculation stability was optimized in log (0,-9); for decision tree the criterion method was optimized between \u0026lsquo;gini\u0026rsquo; and \u0026lsquo;entropy\u0026rsquo;, and max tree depth was optimized in [2, 4, 6, 8, 10, 12] vector. For random forest the predefined parameters were number of estimators: [25, 50, 100, 150], \u0026apos;max_features\u0026apos;: [\u0026apos;sqrt\u0026apos;, \u0026apos;log2\u0026apos;, None], maximum depth: [3, 6, 15], maximum leaf nodes: [3, 6, 9]; for XGBoost the predefined parameter set were number of estimators: stats.randint(150, 300), learning rate: stats.uniform(0.01, 0.5), \u0026apos;subsample\u0026apos;: stats.uniform(0.3, 0.6), maximum of depth: [3, 6, 9], \u0026apos;colsample_bytree\u0026apos;: stats.uniform(0.5, 0.4), \u0026apos;min_child_weight\u0026apos;: [1, 2, 4], and finally for LightGBM the predefined parameters were number of levels: [5, 20, 31], learning rate: [0.05, 0.1, 0.2], number of estimators: [50, 100, 150]. For random forest, XGBoost, and LightGBM a random search strategy was used.\u003c/p\u003e\n\u003cp\u003eFollowing hyperparameter optimization, the dataset was split into testing and training subsets. A 5-fold stratified cross-validation without shuffling was implemented to be assure about preventing class-imbalance for training and testing the classifier. Models were trained on the training samples, with performance evaluated through 10-fold cross-validation using accuracy, recall, precision, and area under the receiver operating characteristic curve (AUC-ROC) measures. Finally, the classifier was tested incorporating test samples using stratified 5-fold cross validation. To prevent model overfitting, repeated K-fold cross-validation (K\u0026thinsp;=\u0026thinsp;5) was implemented and the mean values were reported for all classifier evaluation metrics.\u003c/p\u003e\n\u003cp\u003eML tools were implemented in python (version 3.12) and different libraries including numpy, pandas, sklearn, scipy, xgboost, and Lightgbm were used.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical analysis and performance measures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor numerical variables, analysis of variance (ANOVA) was performed to find the significant differences of mean values between groups. Post-hoc analysis using Tukey\u0026apos;s HSD was conducted to identify the sources of observed differences. The significance level was adjusted to 0.05. Additionally, to evaluate the performance of discrimination models, measures such as accuracy, precision, recall, sensitivity, and AUC were calculated. The mean (standard deviation) of such values for several iterations were reported accordingly.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eThyroglobulin value comparison between subgroups\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThyroglobulin value (stimulated or non-stimulated) is one the primary biomarkers for thyroid cancer assessment. In the current dataset, thyroglobulin level is the only numerical variable, while other variables are categorical. In this section, statistical analyses for comparing thyroglobulin level between different subgroups (considering other variables) were reported. To mitigate outlier effects, the extreme 10% of values in both distribution tails of thyroglobulin profile were excluded. Furthermore, in the current dataset, variables including “history of thyroid cancer in first-degree relatives”, ‘thyroid disease type’, ‘survival outcome’, ‘TNM stage’, ‘staging’, and ‘gender’ were highly imbalanced. In this regard, subgroup analysis was not conducted for these variables.\u003c/p\u003e\n\u003cp\u003eThere is a significant difference between stimulated-Tg value of different age spans (i.e. adolescent and young adults (≤ 35 years of old), middle-aged (36–55), and elderly (\u0026gt; 55 years old)). One-way ANOVA confirmed between-group differences (F (3, 96) = 5.96, p = 0.004). Post-hoc analysis using Tukey's HSD revealed that stimulated-Tg was significantly higher for elderly as compared with other two groups (elderly vs. adolescent and young: mean difference = 25.61, CI: [5.92 45.30], p = 0.007; elderly vs. middle-aged: mean difference = 26.22, CI: [6.62 45.82], p = 0.007). For Non-stimulated-Tg, ANOVA revealed significant differences between groups (F (2,202) = 6.18, p = 0.002). Post-hoc analysis using Tukey's HSD also revealed that Non-stimulated-Tg was significantly higher for elderly as compared with adolescent and young (mean difference = 0.3148, CI: [0.05 0.58], p = 0.02), and middle-aged (mean difference = 0.38, CI: [0.12 0.63], p = 0.002).\u003c/p\u003e\n\u003cp\u003eFor both stimulated and non-stimulated-Tg, there was no significant difference between PTC and others thyroid cancer (mean difference=-3.79, CI: [-8.01 0.43], p = 0.07; mean difference=-0.02, CI: [0.71 − 0.16], p = 0.71, respectively). Considering first response to treatment (i.e. ‘excellent response’, ‘biochemical incomplete response’, or ‘structural incomplete response’), a significant difference between subgroups was found by ANOVA for stimulated-Tg value (F(2,149) = 5.62, p = 0.004). The post-hoc Tukey’s HSD revealed that the difference was between ‘excellent response’ and ‘structural incomplete response’ groups (mean difference = 4.1934, CI: [1.01 7.38], p = 0.006), which indicated significantly higher value of thyroglobulin in ‘structural incomplete response’ group. For non-stimulated-Tg, no significant differences were found between subgroups. For TNM_system variable (primary tumor including T1a, T1b, T2, and T3a), ANOVA revealed a significant difference for stimulated-Tg value between groups (F(3,171) = 15.95, p \u0026lt; 0.001). The post-hoc Tukey’s HSD test showed that stimulated-Tg for T3a group was significantly differed with other groups (T3a vs. T1a: mean difference = 18.69, CI: [7.68 29.691], p = 0.0001; T3a vs. T1b: mean difference = 15.6912, CI:[ 8.85 22.53], p \u0026lt; 0.001; T3a vs. T2: mean difference = 19.10, CI:[ 11.82 26.38], p \u0026lt; 0.0001). In all cases, the stimulated-Tg value was higher for T3a subgroup. For non-stimulated-Tg values, no significant differences were found between subgroups.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMachine learning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA primary objective of this study was to predict the thyroid cancer recurrence using the variables described in Table\u0026nbsp;1. Models were specially trained and tested according to demographic data and first-year clinical assessments. This means that even though for some patients longitudinal values for data such as thyroglobulin level and whole-body scan were available, only the values for the first year follow-up were used (i.e. thyroglobulin level and whole-body scan for the first year of follow-up). Table\u0026nbsp;2 reports performance metrics of different ML approaches for predicting the cancer recurrence according to the variables described in Table\u0026nbsp;1. For some patients, according to the endocrinologist opinion, the value for stimulated-Tg was measured, while for others non-stimulated-Tg level was assessed. For the results in Table\u0026nbsp;2, stimulated-Tg was used for the thyroglobulin level. When such value was not available, the non-stimulated-Tg was substituted for the thyroglobulin level. In Table\u0026nbsp;2, the importance ratio column shows the relative weighting of Tg versus whole-body scan obtained by each classifier.\u003c/p\u003e\n\u003cdiv\u003e\n \u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 2\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003ePerformance of machine learning approaches for predicting recurrence of thyroid cancer.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePrediction model\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAccuracy (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePrecision (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRecall (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSpecificity (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAUC*100\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eImportance ratio (Tg/WBS)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLogistic regression\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e90.70 (1.44)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e95.00 (6.12)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e50.51 (9.37)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e99.31 (0.84)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e94.94 (3.18)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.13\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eBayes Naïve\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e88.73 (0.07)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e77.78 (0.09)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e53.85 (1.00)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e96.55 (2.06)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e96.49 (0.08)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e85390\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDecision tree\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e90.82 (1.96)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e88.29 (7.03)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e55.95 (7.58)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e98.35 (1.14)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e91.86 (3.36)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e79.75\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRandom forest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e93.75 (2.18)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e84.16 (5.96)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e79.87 (11.32)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e96.71 (1.37)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e95.62 (2.86)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e24.57\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e91.55 (2,73)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e73.33 (0.80)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e84.62 (1.59)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e93.10 (1.00)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e96.49 (2.03)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.51\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eLightGBM\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e94.16 (2.60)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e87.88 (9.32)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e81.22 (8.96)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e96.92 (2.74)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e\u003cstrong\u003e96.20 (3.44)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e24.56\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eBased on results in Table\u0026nbsp;2, LightGBM obtained the best classification accuracy. For this reason, the importance of different features for this classifier was plotted in Fig.\u0026nbsp;1. According to this figure, the most informative features for recurrence prediction were first year thyroglobulin (first year Tg) level, first response to treatment, age, TNM stage, pathology type, gender and, first year whole-body scan.\u003c/p\u003e\n\u003cp\u003eIn Fig.\u0026nbsp;2, the block diagram for LightGBM decision tree is depicted. According to this figure, discriminative levels for thyroid cancer recurrence were thyroglobulin value, first response to treatment, and the Node involvement. The count number shows number of samples in the corresponding leaf. In front of each leaf number are prediction values for no recurrence which can be converted to probability using a sigmoid transformation.\u003c/p\u003e\n\u003cp\u003eTable\u0026nbsp;3 examined how classification performance was affected when the reduced feature set was used for cancer recurrence prediction. The reduced feature set was selected based on the most-informative features found by LightGBM model (Fig.\u0026nbsp;1) including age, gender, pathology, TNM stage, first response to treatment, and first year Tg level.\u003c/p\u003e\n\u003cdiv\u003e\n \u003ctable id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 3\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003ePrediction performance of a LightGBM classifier for a reduced feature set.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAccuracy (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePrecision (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRecall (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSpecificity (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAUC*100\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e95.51 (1.42)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e88.84 (9.51)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e84.25 (12.19)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e97.89 (1.72)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e97.28 (1.57)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eFurthermore, in Table\u0026nbsp;4, the predictive performance of stimulated-Tg and non-stimulated-Tg values were compared using both reduced and full feature sets. When only non-stimulated-Tg values were used for training and testing the predictor model (according to LightGBM classifier), the prediction performance of the model was not significantly different between full and reduced feature set, while the best performance was achieved for full feature set. However, when stimulated-Tg values were used, the prediction performance was significantly lower for full feature set as compared with reduced feature set (p \u0026lt; 0.05). This result was not dependent to the number of cross-validation repeats.\u003c/p\u003e\n\u003cdiv\u003e\n \u003ctable id=\"Tab4\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 4\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eMachine learning performance for predicting recurrence based on stimulated or Non-stimulated Tg levels.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAccuracy (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePrecision (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRecall (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSpecificity (%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAUC*100\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ep-value\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNon-stimulated Tg Full feature set\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e94.84 ( 1.58 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e88.81 ( 9.81 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e91.59 ( 7.09 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e96.80 ( 3.00 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e99.23 ( 1.15 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" rowspan=\"2\"\u003e\n \u003cp\u003eAUC: \u0026lt;0.001\u003c/p\u003e\n \u003cp\u003eAccuracy: \u0026lt;0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003estimulated Tg Full feature set\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e88.47 ( 3.62 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e62.38 ( 23.44 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e60.00 ( 20.68 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e93.78 ( 3.52 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e90.39 ( 5.46 )\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNon-stimulated Tg Reduced feature set\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e92.13 ( 2.84 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e73.06 ( 11.66 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e78.41 ( 7.62 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e94.73 ( 2.51 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e94.69 ( 2.07 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" rowspan=\"2\"\u003e\n \u003cp\u003eAUC: 0.57\u003c/p\u003e\n \u003cp\u003eAccuracy: 0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003estimulated Tg Reduced feature set\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e92.58 ( 1.83 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e81.56 ( 6.54 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e79.56 ( 8.20 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e95.82 ( 1.81 )\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e94.27 ( 2.08 )\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe LightGBM classifier achieved the highest accuracy for tumor recurrence prediction (94.16 (2.60) %). Other types of classifiers also attained high classification accuracy (\u0026gt;\u0026thinsp;88%). This result demonstrates the predictive potential of the variables described in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e for estimating future thyroid tumor recurrence. The predictive performance of the feature set used in the present study was superior compared with similar studies. For example, Schindele et al. proposed an XGBoost model using biomarker and clinical features (including age, primary tumor, and serum Tg level) for thyroid cancer prediction (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). The area under curve for their model was 0.88 (95% CI: 0.84\u0026ndash;0.86). The primary tumor size and Tg levels were the most-informative predictors for thyroid recurrence prediction. Xi et al. used a dataset consisting of demographic information, thyroid test measures, nodules characteristics, and six ML methods including gradient boosting machine, logistic regression, linear discriminant analysis, SVM with radial or linear kernel, and random forest to predict nodule malignancy (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). According to the results, the random forest model achieved the highest prediction accuracy (78.01%, 95% CI: 0.7670, 0.7930). In another study, Gu et al. proposed a model to predict metastasis of thyroid cancer based on demographic and laboratory test (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). Using XGBoost model, the best accuracy of that model was 77% for predicting benign/malignant and 70% for lymph node metastasis. Mao et al. utilized a dataset from SEER database and used an XGBoost classifier for suggesting a predicting model for the risk of thyroid cancer recurrence. The model achieved a predictive performance with AUC of 90.4% and identified age, marital status, TNM stage, and surgical methods as the most important risk factors for cancer survival (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e). In Park and Lee study, a dataset was created using demographic, clinicopathological, and ultrasound-extracted parameters and several types of classifiers were used for developing a predictive model. Using a decision tree classifier, maximum F1-score of 28% was achieved, indicating limited prediction performance with the used dataset. Age, tumor size, and lymph node ratio were among the most-informative features (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). In another study, Kim et al. utilized a dataset compromised of clinical, genetics, laboratory, and pathological data to develop a predictive model using an inductive logic programming (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e). The tumor recurrence prediction accuracy of this model was 71.4%, while the most-informative parameters were Tg level, body mass index, anti-Tg antibody level, TSH, primary tumor, and lymph node involvement (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eWhole-body scan is a conventional tool for investigating thyroid cancer. However, this approach has several side effects and drawbacks including nausea/vomiting, change in taste and smell, bone marrow suppression (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e), radiation safety concerns, and inability to differentiate benign from malignant tumors (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). In the present study, the necessity of whole-body scanning for recurrence prediction was evaluated by quantifying its relative importance in the model\u0026rsquo;s prediction performance. In Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, for each model, the predictive value of thyroglobulin variable was compared with that of whole-body scan, and the importance ratio was reported. According to Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, classifiers including Na\u0026iuml;ve Bayes, decision tree, random forest, and LightGBM indicated that thyroglobulin had greater predictive capacity than the whole-body scan, whereas other classifiers including logistic regression and XGBoost highlighted the importance of whole-body scan test. Previous studies compared the power of Tg for predicting thyroid cancer recurrence with radionuclide methods. For example, a comparison of the sensitivity of 123I scintigraphy and Tg revealed that whole-body scan did not improve the sensitivity (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eAccording to the LightGBM decision tree (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), a critical level of Tg for tumor recurrence prediction is 4.775. The gain of Tg suggests that it is the primary determinant of recurrence risk. The second level of decision tree indicated the importance of the first response to treatment, where structural incomplete response cases were separated from other subgroups. The third level showed that lymph node involvement was the next determinant risk stratification factor. Based on the terminal leaves, the lowest recurrence risk corresponded to the path to leaf 2 where Tg\u0026thinsp;\u0026le;\u0026thinsp;4.775, first response to treatment was excellent or biochemical incomplete response, TNM stage N0(without nodal involvement), and Tg lower than 0.65. The risk for tumor recurrence is higher when Tg\u0026thinsp;\u0026gt;\u0026thinsp;0.65. The path to leaf 3 (Tg\u0026thinsp;\u0026le;\u0026thinsp;4.775 with nodal involvement N1) had lower predictive value for recurrence-free outcome.\u003c/p\u003e\u003cp\u003eAccording to the results of Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, eliminating non-informative features and reducing the size of feature space improved the performance of predictive model in terms of accuracy, precision, recall, and specificity. This result imply that using an eight-dimensional feature space including age, gender, pathology type, TNM stage, first response to treatment, first year Tg and whole-body scan is sufficient to predict the future recurrence of thyroid cancer with high accuracy, specificity, and sensitivity.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eLightGBM demonstrated superior predictive performance for identifying thyroid cancer patients at risk of recurrence, achieving high accuracy and AUC. The most important predictors were first-year thyroglobulin level, first treatment response, age, primary tumor characteristics, and lymph node involvement, respectively.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eML \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; Machine learning\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTNM \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Tumor, Node, Metastasis\u003c/p\u003e\n\u003cp\u003eTg \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; Thyroglobulin\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eANOVA \u0026nbsp; \u0026nbsp;Analysis Of Variance\u0026nbsp;\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis article is the result of a general medicine thesis. We sincerely thank the Hamadan University of Medical Sciences Ethical Research Committee. The work was submitted under the code 140108176791.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Considerations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was conducted in accordance with the principles of the Helsinki Declaration and was approved by the Ethics Committee of Hamadan University of Medical Sciences (IR.UMSHA.REC.1401.669).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to participate\u003c/strong\u003e-\u0026nbsp;Written informed consent was obtained from all participants prior to data collection.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e-Not applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of interest/Competing interests\u003c/strong\u003e-There is nothing to declare.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e-\u0026nbsp;The datasets generated during this study are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding-\u0026nbsp;\u003c/strong\u003eThis work was supported by Hamadan University of Medical Sciences, Deputy of research and technology (No. 140108176791).\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contribution-\u0026nbsp;\u003c/strong\u003eShB\u003csup\u003e\u0026nbsp;\u003c/sup\u003eand RN collected the data. SF performed analyses and wrote the initial draft. SF, EA, and ShB\u003csup\u003e\u0026nbsp;\u003c/sup\u003e discussed the obtained results and finalized the draft.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, Jemal A. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229-63.\u003c/li\u003e\n\u003cli\u003eSung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71(3):209-49.\u003c/li\u003e\n\u003cli\u003eNejadghaderi SA, Moghaddam SS, Azadnajafabad S, Rezaei N, Rezaei N, Tavangar SM, et al. Burden of thyroid cancer in North Africa and Middle East 1990-2019. Front Oncol. 2022;12:955358.\u003c/li\u003e\n\u003cli\u003eGuo K, Wang Z. Risk factors influencing the recurrence of papillary thyroid carcinoma: a systematic review and meta-analysis. Int J Clin Exp Pathol. 2014;7(9):5393-403.\u003c/li\u003e\n\u003cli\u003eHwangbo Y, Kim JM, Park YJ, Lee EK, Lee YJ, Park DJ, et al. Long-Term Recurrence of Small Papillary Thyroid Cancer and Its Risk Factors in a Korean Multicenter Study. J Clin Endocrinol Metab. 2017;102(2):625-33.\u003c/li\u003e\n\u003cli\u003eIto Y, Kudo T, Kobayashi K, Miya A, Ichihara K, Miyauchi A. Prognostic factors for recurrence of papillary thyroid carcinoma in the lymph nodes, lung, and bone: analysis of 5,768 patients with average 10-year follow-up. World J Surg. 2012;36(6):1274-8.\u003c/li\u003e\n\u003cli\u003eQi P, Wang Z, Hao X, Ou X, Zhang B, Shi Q, et al. A retrospective study of 17,995 patients investigating the location and recurrence of papillary thyroid cancer. Sci Rep. 2025;15(1):10634.\u003c/li\u003e\n\u003cli\u003eQu N, Zhang L, Lu ZW, Ji QH, Yang SW, Wei WJ, Zhang Y. Predictive factors for recurrence of differentiated thyroid cancer in patients under 21 years of age and a meta-analysis of the current literature. Tumour Biol. 2016;37(6):7797-808.\u003c/li\u003e\n\u003cli\u003eJammah AA, Masood A, Akkielah LA, Alhaddad S, Alhaddad MA, Alharbi M, et al. Utility of Stimulated Thyroglobulin in Reclassifying Low Risk Thyroid Cancer Patients\u0026apos; Following Thyroidectomy and Radioactive Iodine Ablation: A 7-Year Prospective Trial. Front Endocrinol (Lausanne). 2020;11:603432.\u003c/li\u003e\n\u003cli\u003eWestbury C, Vini L, Fisher C, Harmer C. Recurrent differentiated thyroid cancer without elevation of serum thyroglobulin. Thyroid. 2000;10(2):171-6.\u003c/li\u003e\n\u003cli\u003eYang X, Liang J, Li TJ, Yang K, Liang DQ, Yu Z, Lin YS. Postoperative stimulated thyroglobulin level and recurrence risk stratification in differentiated thyroid cancer. Chin Med J (Engl). 2015;128(8):1058-64.\u003c/li\u003e\n\u003cli\u003eKononenko I. Machine learning for medical diagnosis: history, state of the art and perspective. Artificial Intelligence in Medicine. 2001;23(1):89-109.\u003c/li\u003e\n\u003cli\u003eShehab M, Abualigah L, Shambour Q, Abu-Hashem MA, Shambour MKY, Alsalibi AI, Gandomi AH. Machine learning in medical applications: A review of state-of-the-art methods. Comput Biol Med. 2022;145:105458.\u003c/li\u003e\n\u003cli\u003eBorzooei S, Briganti G, Golparian M, Lechien JR, Tarokhian A. Machine learning for risk stratification of thyroid cancer patients: a 15-year cohort study. Eur Arch Otorhinolaryngol. 2024;281(4):2095-104.\u003c/li\u003e\n\u003cli\u003eBorzouei S, Safdari A, Ayubi E. Development and Validation of a Clinical Risk Model for Predicting Malignancy in Patients with Thyroid Nodules. Avicenna J Clin Med. 2025;31(4):219-27.\u003c/li\u003e\n\u003cli\u003eGu J, Xie R, Zhao Y, Zhao Z, Xu D, Ding M, et al. A machine learning-based approach to predicting the malignant and metastasis of thyroid cancer. Front Oncol. 2022;12:938292.\u003c/li\u003e\n\u003cli\u003eHou F, Zhu Y, Zhao H, Cai H, Wang Y, Peng X, et al. Development and validation of an interpretable machine learning model for predicting the risk of distant metastasis in papillary thyroid cancer: a multicenter study. EClinicalMedicine. 2024;77:102913.\u003c/li\u003e\n\u003cli\u003eLi Y, Wu F, Ge W, Zhang Y, Hu Y, Zhao L, et al. Risk stratification of papillary thyroid cancers using multidimensional machine learning. Int J Surg. 2024;110(1):372-84.\u003c/li\u003e\n\u003cli\u003eLiu W, Wang S, Ye Z, Xu P, Xia X, Guo M. Prediction of lung metastases in thyroid cancer using machine learning based on SEER database. Cancer Med. 2022;11(12):2503-15.\u003c/li\u003e\n\u003cli\u003eXi NM, Wang L, Yang C. Improving the diagnosis of thyroid cancer by machine learning and clinical data. Sci Rep. 2022;12(1):11143.\u003c/li\u003e\n\u003cli\u003ePark YM, Lee B-J. Machine learning-based prediction model using clinico-pathologic factors for papillary thyroid carcinoma recurrence. Sci Rep. 2021;11(1):4948.\u003c/li\u003e\n\u003cli\u003eTanaka K, Ozaki T. New TNM classification (AJCC eighth edition) of bone and soft tissue sarcomas: JCOG Bone and Soft Tissue Tumor Study Group. Jpn J Clin Oncol. 2019;49(2):103-7.\u003c/li\u003e\n\u003cli\u003eChen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; San Francisco, California, USA: Association for Computing Machinery; 2016. p. 785\u0026ndash;94.\u003c/li\u003e\n\u003cli\u003eKe G, Meng Q, Finley T, Wang T, Chen W, Ma W, et al. Lightgbm: A highly efficient gradient boosting decision tree. Adv Neural Inf Process Syst. 2017;30.\u003c/li\u003e\n\u003cli\u003eSchindele A, Krebold A, Hei\u0026szlig; U, Nimptsch K, Pfaehler E, Berr C, et al. Interpretable machine learning for thyroid cancer recurrence predicton: Leveraging XGBoost and SHAP analysis. Eur J Radiol. 2025;186:112049.\u003c/li\u003e\n\u003cli\u003eMao Y, Huang Y, Xu L, Liang J, Lin W, Huang H, et al. Surgical Methods and Social Factors Are Associated With Long-Term Survival in Follicular Thyroid Carcinoma: Construction and Validation of a Prognostic Model Based on Machine Learning Algorithms. Front Oncol. 2022;12:816427.\u003c/li\u003e\n\u003cli\u003eKim SY, Kim Y-I, Kim HJ, Chang H, Kim S-M, Lee YS, et al. New approach of prediction of recurrence in thyroid cancer patients using machine learning. Medicine. 2021;100(42).\u003c/li\u003e\n\u003cli\u003eMuzahir S, Grady E. Nuclear Imaging and Therapy of Thyroid Disorders. Hall LT editor Molecular Imaging and Therapy. Brisbane (AU): Exon Publications; 2023.\u003c/li\u003e\n\u003cli\u003eHoang JK, Lee WK, Lee M, Johnson D, Farrell S. US Features of thyroid malignancy: pearls and pitfalls. Radiographics. 2007;27(3):847-60; discussion 61-5.\u003c/li\u003e\n\u003cli\u003ede Geus-Oei LF, Oei HY, Hennemann G, Krenning EP. Sensitivity of 123I whole-body scan and thyroglobulin in the detection of metastases or recurrent differentiated thyroid cancer. Eur J Nucl Med Mol Imaging. 2002;29(6):768-74.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Table 1","content":"\u003cp\u003eTable 1 is available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"thyroid cancer, recurrence, thyroglobulin, machine learning","lastPublishedDoi":"10.21203/rs.3.rs-7013702/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7013702/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground-\u003c/h2\u003e\u003cp\u003eAlthough thyroid cancer generally has a good prognosis, some patients are prone to recurrence. Multiple factors influence recurrence risk. Machine learning (ML) algorithms offer potential for more accurate and precise prediction models. The aim of the present study is to evaluate recurrence-related factors in thyroid cancer patients using ML algorithms.\u003c/p\u003e\u003ch2\u003eMethods-\u003c/h2\u003e\u003cp\u003eThis retrospective cohort study included patients with differentiated thyroid cancer referred to a specialized endocrinology clinic, between 2013 and 2023. Demographic data, tumor characteristics, and treatment details were extracted from medical records. Six ML algorithm were employed including logistic regression, Na\u0026iuml;ve Bayes classifier, decision tree, random forest, XGBoost and LightGBM.\u003c/p\u003e\u003ch2\u003eResults-\u003c/h2\u003e\u003cp\u003eA total 355 patients were included (mean age: 41.6914.04 years, 84.22% female). Among ML algorithms, LightGBM demonstrated superior predictive performance, achieving an accuracy of 95.41%, precision of 88.84%, recall of 84.25%, specificity of 97.89%, and an area under the curve of 97.28%. The top five predictors were first-year thyroglobulin level, first response to treatment, age, primary tumor characteristics, and regional lymph nodes involvement, respectively.\u003c/p\u003e\u003ch2\u003eConclusion-\u003c/h2\u003e\u003cp\u003eThis study demonstrated that ML algorithms had strong capability to identify thyroid cancer patients at risk of recurrence.\u003c/p\u003e","manuscriptTitle":"Machine learning approaches for prediction of thyroid cancer recurrence using thyroglobulin level, whole body scan","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-07 07:07:40","doi":"10.21203/rs.3.rs-7013702/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"50e49054-00b4-486b-b6f8-9a1fce6646aa","owner":[],"postedDate":"August 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-11-26T09:24:21+00:00","versionOfRecord":[],"versionCreatedAt":"2025-08-07 07:07:40","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7013702","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7013702","identity":"rs-7013702","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00