A Study on the Feasibility of Optimizing Gastric Cancer Screening to Reduce Screening Costs in China Using a Gradient Boosting Machine: A prospective, large-sample, single-center study

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background and aim: The current cancer screening model in our country involves preliminary screening and identification of individuals who require gastroscopy, in order to control screening costs. The purpose of this study is to optimize the screening process using Gradient Boosting Machines (GBM), a machine learning technique, based on a large-scale prospective gastric cancer screening dataset. The ultimate goal is to further reduce the cost of initial cancer screening. Methods The study constructs a GBM machine learning model based on prospective, large-sample Taizhou City gastric cancer screening data and validates it with data from the Minimum Security Cohort Group (MLGC) in Taizhou City. Both data analysis and machine learning model construction were performed using the R programming language. Results A total of 195,640 cases were used as the training set, and 32,994 cases were used as an external validation set. A GBM was built based on the training set, yielding area under the curve (AUC) and area under the precision-recall curve (AUCPR) values of 0.99938 and 0.99823, respectively. External validation of the model yielded AUC and AUCPR values of 0.99742 and 0.99454, respectively. Through a visual analysis of the model, it was determined that the variable for Helicobacter pylori IgG could be eliminated. The GBM model was then reconstructed without the H. pylori IgG variable. In the training set, the new model achieved an AUC of 0.99817 and an AUCPR of 0.99462, whereas in the external validation set, it achieved an AUC of 0.99742 and an AUCPR of 0.99454. Conclusion This study utilized a dataset of 230,000 samples to train and validate a GBM model, optimizing the initial screening process by excluding the detection of H. pylori IgG antibodies while maintaining satisfactory discriminative performance. This conclusion will contribute to a reduction in the current cost of gastric cancer screening, demonstrating its economic value. Furthermore, the conclusion is derived from a large sample size, giving it clinical significance and generalizability.
Full text 114,057 characters · extracted from preprint-html · click to expand
A Study on the Feasibility of Optimizing Gastric Cancer Screening to Reduce Screening Costs in China Using a Gradient Boosting Machine: A prospective, large-sample, single-center study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Study on the Feasibility of Optimizing Gastric Cancer Screening to Reduce Screening Costs in China Using a Gradient Boosting Machine: A prospective, large-sample, single-center study Xin-yu Fu, Rongbin Qi, Shan-jing Xu, Meng-sha Huang, Cong-ni Zhu, and 12 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3853941/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background and aim: The current cancer screening model in our country involves preliminary screening and identification of individuals who require gastroscopy, in order to control screening costs. The purpose of this study is to optimize the screening process using Gradient Boosting Machines (GBM), a machine learning technique, based on a large-scale prospective gastric cancer screening dataset. The ultimate goal is to further reduce the cost of initial cancer screening. Methods The study constructs a GBM machine learning model based on prospective, large-sample Taizhou City gastric cancer screening data and validates it with data from the Minimum Security Cohort Group (MLGC) in Taizhou City. Both data analysis and machine learning model construction were performed using the R programming language. Results A total of 195,640 cases were used as the training set, and 32,994 cases were used as an external validation set. A GBM was built based on the training set, yielding area under the curve (AUC) and area under the precision-recall curve (AUCPR) values of 0.99938 and 0.99823, respectively. External validation of the model yielded AUC and AUCPR values of 0.99742 and 0.99454, respectively. Through a visual analysis of the model, it was determined that the variable for Helicobacter pylori IgG could be eliminated. The GBM model was then reconstructed without the H. pylori IgG variable. In the training set, the new model achieved an AUC of 0.99817 and an AUCPR of 0.99462, whereas in the external validation set, it achieved an AUC of 0.99742 and an AUCPR of 0.99454. Conclusion This study utilized a dataset of 230,000 samples to train and validate a GBM model, optimizing the initial screening process by excluding the detection of H. pylori IgG antibodies while maintaining satisfactory discriminative performance. This conclusion will contribute to a reduction in the current cost of gastric cancer screening, demonstrating its economic value. Furthermore, the conclusion is derived from a large sample size, giving it clinical significance and generalizability. Gradient boosting machine Helicobacter pylori antibody Gastric cancer Screening Cost reduction Large-sample Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Gastric cancer (GC) remains one of the most common malignancies worldwide, with over one million new cases estimated annually, ranking it as the fifth-most diagnosed malignancy globally [ 1 , 2 ]. Owing to the often late-stage diagnosis of GC, it has a high mortality rate, ranking as the third leading cause of cancer-related deaths [ 3 ]. However, it is worth noting that early GC (EGC) has a 5-year survival rate exceeding 90%, far higher than that of advanced GC (AGC), which has a 30% survival rate [ 4 , 5 ]. Such a high survival rate underscores the importance of GC screening, especially in countries with a high incidence of this disease [ 6 , 7 ]. According to the latest reports from 2020, China has the highest number of registered GC cases, with approximately 820,000 new cases and 580,000 deaths. Therefore, conducting large-scale GC screening programs is imperative in China [ 8 , 9 ]. Currently, China does not have a nationwide GC screening program, and early detection relies solely on opportunistic screening. To address this situation, China's existing national GC screening guidelines recommend screening all high-risk individuals starting at 40 years old [ 10 , 11 ]. Gastric endoscopy and a tissue biopsy for a histological examination are currently the gold standards for GC screening and diagnoses. However, screening the entire "high-risk" population with gastric endoscopy is inefficient and impractical, as it is expected that only 1–3% of this population will actually have GC [ 11 , 12 ]. Furthermore, it is estimated that there are over 300 million "high-risk" individuals in China, and because of the high cost, screening with gastric endoscopy is unlikely to be feasible for all individuals [ 11 ]. Therefore, a risk stratification method is needed as a preliminary screening tool before gastric endoscopy to further identify truly high-risk individuals among the previously defined "high-risk" population [ 13 , 14 ]. In 2018, a nationwide multicenter prospective trial conducted in a Chinese population developed and validated Lee's screening score for this purpose, aiming to distinguish individuals who need further endoscopic evaluations. The screening score demonstrated a good discriminative ability, with an area under the curve (AUC) of 0.76 [ 11 ]. Effective primary screening can significantly reduce the cost of GC screening, but current primary screening still requires the detection of four serological markers: gastrin, Helicobacter pylori IgG antibody, and pepsinogens (PGs) I and II [ 11 ]. The question of whether or not it is possible to optimize the required markers for primary screening to further reduce screening costs is meaningful. In recent years, machine learning has shown strong potential utility in data processing and analyses [ 15 ]. Prediction models constructed using machine learning have demonstrated good predictive performance and stability, thus garnering increasing attention [ 16 , 17 ]. The gradient boosting machine (GBM), a powerful machine-learning algorithm, is capable of efficiently handling high-dimensional sparse data, capturing nonlinear relationships and interactions, and enhancing model predictive performance through iterative optimization [ 18 , 19 ]. In particular, it excels in solving classification problems. This study is based on prospective gastric cancer screening data, aiming to significantly improve and optimize Lee’s screening score using the GBM algorithm in order to further reduce the cost of initial gastric cancer screening. Methods Study design and participants This prospective study on gastric cancer screening is supported by the leadership of the government of Linhai, Taizhou. From January 2021 to July 2023, a 30-month gastric cancer screening project will be carried out targeting the general population aged 45 to 75. The project is led by Taizhou Hospital, with subordinate hospitals participating and Taizhou Hospital conducting comprehensive analysis. This research has been approved by the Ethics Review Board of Wenzhou Medical University, affiliated with Taizhou Hospital in Zhejiang Province (Approval Number: K20210613). To retrospectively collect data on the GC screening of the Minimum Living Guarantee Crowd (MLGC) conducted from January to December 2019. This project, initiated by the Taizhou Municipal Government and led by Taizhou Hospital, was summarized based on the approval of the Ethical Review Committee of Zhejiang Taizhou Hospital affiliated with Wenzhou Medical University (approval number: K20190221). The data of the GC screening on MLGC included a total of 32,994 cases and will serve as the validation dataset for the model. The inclusion exclusion criteria for the general population gastric cancer screening program were as follows: Inclusion criteria: 1) Between 45–75 years of age; 2) Willing to participate in the screening project and sign an informed consent form. Exclusion criteria: 1) Decline to complete the questionnaire; 2) Refuse serological testing; 3) Incomplete or missing data. The inclusion exclusion criteria for the MLGC screening program are as follows Inclusion criteria: 1) All patients participating in the MLGC program. Exclusion criteria: 1) Incomplete or missing data. Questionnaires and serologic tests The questionnaire survey included information about the smoking history, alcohol consumption history, medical history, and other relevant details, which were used to describe the demographic characteristics of the study population. Blood tests included measurement of gastrin-17 (G-17), PG I/II, and H. pylori IgG antibodies. All tests were conducted uniformly by government-designated testing facilities to ensure the stability of blood test results. Definitions of the endoscopy group and room-following group Lee's screening score (See Table 1 of the supplementary document) for GC is based on variables of age, sex, G-17, PG I/II, and H. pylori antibodies. It classifies individuals into low-, intermediate-, and high-risk groups. For individuals classified as intermediate or high risk, further investigations using endoscopy are recommended. In the present study, based on the score table, individuals in the intermediate- and high-risk groups were categorized into the endoscopy group, whereas individuals in the low-risk group were assigned to the follow-up group. Data analyses This study utilizes the "h2o" automatic machine learning package in Java to automatically build and select the most effective machine learning algorithms for model construction. This study uses the general population of Linyi City as the training dataset, employs five-fold cross-validation for internal validation, and uses the MPGC population as the test dataset for external validation. Model evaluation was conducted using the area under the precision-recall curve (AUCPR) and the AUC. Both the AUC and AUCPR involve values between 0.5 and 1, with values closer to 1 indicating better model performance and those closer to 0.5 indicating random guessing. Inter-group comparisons were carried out using various statistical methods, such as the Mann-Whitney U test, Kruskal-Wallis rank-sum test, and Fisher's exact probability method. A significance level of p < 0.05 and q < 0.05 is considered to indicate a significant difference. All statistical analyses and machine-learning procedures were implemented using the R programming language (version 4.2.3). Results Patients From January 2021 to July 2023, a GC screening project was conducted in Linhai City, with 207,005 voluntary participants who completed the survey. Among these, there were 76,784 cases in 2021, 77,330 cases in 2022, and 52,891 cases in 2023. Of these participants, 36,722 had a history of smoking, and 28,586 had a history of alcohol consumption. In addition, 46,473 patients had a history of hypertension, 14,859 had diabetes, 595 had coronary heart disease, 181 had chronic kidney disease, and 7,092 had other diseases. A total of 45,506 (21.98%) patients had a history of undergoing gastroscopy. In 2021, 74,614 patients underwent serological testing, while 2,170 patients were excluded because of refusal or failure to undergo serological testing within the specified time. In 2022, 73,686 patients underwent serological testing and 3,644 patients were excluded for the same reasons. As of July 2023, 52,217 patients had undergone serological testing in 2023, with 674 patients excluded for refusal or failure to adhere to the testing schedule. Ultimately, 200,517 patients underwent serological testing, with 4,877 excluded owing to missing information or incomplete data, leaving 195,640 cases for the analysis and model training. According to the Lee's score table, 48,118 patients were assigned to the endoscopy group, and 147,522 patients were included in the follow-up group. The distribution of indicators for each group is shown in Table 1 . Table 1 Linhai City General Population Baseline Table Variable Overall N = 195,640 1 follow-up group N = 147,522 1 endoscopy group N = 48,118 1 p-value 2 q-value 3 Gender < 0.001 < 0.001 Male 78,539 (40%) 39,700 (27%) 38,839 (81%) Female 117,101 (60%) 107,822 (73%) 9,279 (19%) Age 60 (54, 66) 59 (53, 65) 64 (59, 68) < 0.001 < 0.001 Gastrin G-17 1.65 (1.00, 3.32) 1.24 (1.00, 2.43) 3.30 (2.06, 6.00) < 0.001 < 0.001 Pepsinogen I 87 (70, 111) 86 (70, 108) 92 (70, 120) < 0.001 < 0.001 Pepsinogen II 6.6 (4.7, 9.1) 6.5 (4.8, 8.6) 7.2 (3.3, 10.4) < 0.001 < 0.001 H. pylori IgG antibody < 0.001 < 0.001 Negative 118,057 (60%) 96,856 (66%) 21,201 (44%) Positive 77,583 (40%) 50,666 (34%) 26,917 (56%) 1 n (%); Median (IQR); 2 Pearson's Chi-squared test; Wilcoxon rank sum test; 3 False discovery rate correction for multiple testing From January to December 2019, a GC screening project for the MLGC in Taizhou City included 33,054 voluntary participants who completed the survey. Among these, 10,171 had a history of smoking, 7,771 had a history of alcohol consumption, 8,284 had a history of hypertension, 2,252 had a history of diabetes, 2,230 had hyperlipidemia, and 929 had other diseases. A total of 983 patients (2.97%) had a history of gastroscopy. All of these patients underwent serological testing. Sixty cases were excluded owing to missing information, incomplete data, and other reasons, leaving 32,994 cases for the analysis as an external validation set. In the validation set, 10,608 cases were defined as the endoscopy group, and 22,386 cases were defined as the follow-up group. Please refer to Table 2 for further details. A flowchart is shown in Fig. 1 . Table 2 Taizhou Minimum Living Guarantee Crowd Baseline Table Variable Overall N = 32,994 1 follow-up group N = 22,386 1 endoscopy group N = 10,608 1 p-value 2 q-value 3 Gender < 0.001 < 0.001 Male 20,107 (61%) 10,464 (47%) 9,643 (91%) Female 12,887 (39%) 11,922 (53%) 965 (9%) Age 56 (50, 63) 54 (48, 62) 61 (55, 66) < 0.001 < 0.001 Gastrin G-17 2.0 (1.0, 4.7) 1.3 (1.0, 3.0) 4.0 (2.2, 8.5) < 0.001 < 0.001 Pepsinogen I 136 (102, 165) 129 (99, 165) 154 (114, 165) < 0.001 < 0.001 Pepsinogen II 6.7 (4.4, 11.2) 5.9 (4.0, 9.4) 9.4 (5.9, 15.0) < 0.001 < 0.001 H. pylori IgG antibody < 0.001 < 0.001 Negative 20,489 (62%) 15,035 (67%) 5,454 (51%) Positive 12,505 (38%) 7,351 (33%) 5,154 (49%) 1 n (%); Median (IQR); 2 Pearson's Chi-squared test; Wilcoxon rank sum test; 3 False discovery rate correction for multiple testing Machine-learning model selection Based on the h2o package, automated machine learning was conducted on the training dataset, resulting in the training and validation of the 12 machine learning models. These models encompassed five different algorithms: deep learning, random forest, GBM, generalized linear models, and Stacked Ensemble. Among them, the Stacked Ensemble exhibited the best performance, with an AUC of 0.99884, followed closely by the GBM, with an AUC of 0.99883 (Fig. 2 ). Given that the GBM's predictive performance was comparable to that of the Stacked Ensemble while offering greater model stability, we chose to focus on the GBM for subsequent research. In addition, the Stacked Ensemble, which is an ensemble model composed of multiple sub-models, presents challenges in terms of model interpretability. Therefore, further research in this study is based on the GBM model. Modeling GBM distinctions The GBM classification model was built to distinguish between the endoscopy and follow-up groups using the following variables: age, sex, G-17, PG I, PG II, and IgG. The h2o package for automated machine learning was used, and the best results were obtained when using the GBM_5 learner, which was identified by its learner-type code. The model achieved excellent predictive performance, with an AUC of 0.99938 and an AUCPR of 0.99823 (Fig. 3 ). To assess the robustness of the model, it underwent an internal 5-fold cross-validation, resulting in an AUC of 0.99883 and an AUCPR of 0.99656. External validation was performed using the MLGC dataset, yielding an AUC of 0.99742 and AUCPR of 0.99454 (Fig. 4 ). The GBM model demonstrated nearly 100% predictive accuracy in distinguishing between patients in the endoscopy and follow-up groups. An analysis of the importance of variables A visual analysis of the GBM was conducted to determine the importance of each variable within the model, and the results are shown in Fig. 5 . Among the variables, IgG had the lowest importance score, with a value of 0.07516, followed by PG I, with a score of 0.08827 (Fig. 5 A). In the SHAP plot, it was observed that IgG had a greater negative contribution than a positive contribution (Fig. 5 B). In addition, when creating a partial dependence plot for the IgG variable, it was found that both positive and negative IgG values had a similar effect on the mean response (average response) of the outcome variable (Fig. 5 C). Based on these results, it was concluded that the IgG variable could be excluded from the model, as it had low importance and did not significantly influence the outcome variable. Remodeling of the exclusion variable IgG Based on the h2o package, GBM_2 achieved the best performance. The AUC was 0.99817 and the AUCPR was 0.99462 (Fig. 5 ). During the 5-fold cross-validation, the AUC was 0.99717, and the AUCPR was 0.99154. In the external validation set, the AUC was 0.99605, and the AUCPR was 0.99135 (Fig. 6 ). Even after removing IgG, the model maintained a high level of predictive accuracy. Please refer to Fig. 6 – 7 for visualization of the results. Discussion Successfully conducting large-scale GC screening in China requires dividing the screening efforts into two parts: the initial screening and a follow-up examination [ 11 ]. Differentiating individuals who require a further endoscopic examination remains a challenge. Based on a multi-center prospective study specific to the Chinese population, the feasibility of Lee's score table has been proposed and validated. Currently, Lee's screening score is widely applied throughout China in GC screening programs. The GC screening projects involved in this study were based on Lee’s score table. However, use of this table requires the measurement of four serological markers: G-17, PG I/II, and H. pylori antibodies [ 11 ]. The cost of initial screening remains high for large-scale surveys; therefore, this study attempted to optimize the screening process using machine-learning methods. By training and validating a GBM model, it was found that removing H. pylori antibodies from the GBM model had a negligible effect on its performance for differentiating individuals who require gastroscopy, achieving an approximately 100% discrimination rate, similar to the full model without excluding H. pylori antibodies. In this study, automated machine learning was used to select the GBM algorithm, which had the best fit and was easy to analyze, to train and optimize the initial screening for GC [ 16 ]. The GBM algorithm has strong predictive performance, particularly for handling structured data and regression and classification tasks. It also demonstrates robustness against outliers and noise in data [ 16 , 20 ]. In addition, the GBM algorithm can evaluate feature importance, aiding in understanding the impact of data features on the model. Based on the powerful predictive performance of GBM, we included age, sex, G-17, pepsinogen I and II (PG I/II), and H. pylori antibodies to train the full model, which achieved a nearly 100% discrimination rate (AUC = 0.99938 and AUCPR = 0.99823). Furthermore, the discrimination performance remained close to 100% during both internal and external validations. In the internal validation, the AUC was 0.99938, and the AUCPR was 0.99823, whereas in the external validation, the AUC was 0.99742, and the AUCPR was 0.99454. Therefore, we consider the full model to be highly stable. In the visual analysis of the model, we found that the variable H. pylori antibodies had a very low weight in the model. We also discovered that they had a low weight in Lee's score table, with only 1 point. Therefore, we proposed optimizing the inclusion of IgG antibodies. After removing IgG antibodies, we built a discrimination model based on the GBM algorithm and found that the GBM discrimination model without IgG antibodies performed comparably to the full model, with an AUC of 0.998167 and an AUCPR of 0.9946214. It also demonstrated excellent predictive performance in internal (AUC of 0.9971681 and AUCPR of 0.991536) and external validation (AUC of 0.9960488 and AUCPR of 0.9913469). Based on these results, we conclude that the measurement of H. pylori IgG antibodies was not necessary when using the predictive model constructed with the GBM algorithm to differentiate individuals who require GC screening. H. pylori infection is considered a high-risk factor for GC, especially in China, where GC is prevalent [ 21 , 22 ]. At present, invasive and non-invasive methods are available for detecting H. pylori infections. Invasive methods include the Rapid Urease Test, histological examination, and genetic testing, and non-invasive methods include the Urea Breath Test, serological antibody testing, and stool antigen tests, among others [ 21 , 23 ]. Among the screening options based on Lee’s criteria, serological detection of H. pylori IgG antibodies is the most suitable screening strategy, as both gastrin and PG tests require blood sampling. However, according to numerous studies, IgG antibody testing has a high negative predictive value (NPV). The ability of this test to detect active infection depends on a number of factors, such as the patient's age, clinical condition of the infection, choice of antigen used for antibody preparation in the enzyme-linked immunosorbent assay kit, and prevalence of infection. IgG antibodies are unable to differentiate between current infection and previous exposure and can still be detected several months after treatment, which may interfere with GC screening efforts [ 21 , 24 ]. H. pylori infection can lead to secondary hypergastrinemia and hyperpepsinogenemia, indicating that both G-17 and PG I/II levels can reflect the presence of H. pylori infection to some extent [ 25 – 27 ]. Studies have shown that PG I and PG II levels are associated with H. pylori infection; furthermore, a close relationship between G-17 and H. pylori , as patients with H. pylori infection have significantly higher levels of G-17 than those without infection [ 28 , 29 ]. This suggests that G-17, PG I, and PG II are associated with H. pylori infections. As H. pylori -induced gastritis is the most common condition, these three markers not only reflect the overall gastric function but also indirectly indicate the presence of H. pylori infection in the body [ 23 , 29 ]. There is no denying that H. pylori infection is closely associated with the development of GC and related precancerous lesions. However, the accuracy of IgG antibody testing is questionable, as a positive result for IgG antibodies does not accurately reflect the true infection status [ 21 , 30 ]. A retrospective study conducted in Japan showed that IgG can indicate the presence of gastric mucosal damage to some extent, but its accuracy is influenced by age [ 31 ]. In patients > 65 years old, the response of IgG antibodies to H. pylori is diminished, making it difficult to indicate the presence of H. pylori infection [ 31 ]. The dynamic measurement of antibody titers is reportedly necessary to accurately reflect the status of H. pylori infection and treatment efficacy. However, the implementation of this method requires more manpower, financial resources, and material resources, making it challenging to achieve in the screening process. Based on these findings, we believe that the role of H. pylori IgG antibodies in GC screening is limited. Our trained GBM model based on a dataset of 200,000 samples confirms this viewpoint. Therefore, we believe that H. pylori antibody testing is not necessary when using the GBM model to discriminate individuals who require an endoscopic examination. In China, where GC is highly prevalent, cancer screening is necessary because early GC has a favorable prognosis [ 11 ]. The first step in conducting GC screening in China is to accurately identify individuals who need further gastroscopy examinations at the lowest possible cost [ 11 ]. This study, based on a large sample size, trained a GBM model to distinguish individuals requiring a gastroscopy examination and successfully optimized Lee's screening criteria. Due to the large population in China, a considerable number of individuals need to undergo GC screening, and the cost savings achieved by eliminating one serological marker test are significant. Based on the company's testing fees, the detection cost of H. pylori IgG antibodies was 30 RMB per person. Therefore, based on the currently completed screening population, excluding IgG testing can save 7,201,770 RMB. These savings could cover an additional 38,929 initial screenings. Currently, the detection rate of the GC screening program in Lihai City is 0.22%. By extending the coverage to an additional population of over 30,000 individuals, approximately 84 more patients with GC could be detected. According to relevant literature, GC screening is cost-effective in countries with a high incidence of GC. By improving screening techniques, it is possible to further reduce screening costs and achieve higher quality-adjusted life-years with a lower incremental cost-effectiveness ratio [ 32 , 33 ]. In addition, with approximately 300 million people in China requiring screening, nearly 9 billion RMB in screening costs can be saved, which is quite significant [ 11 ]. The treatment costs for advanced-stage GC are extremely high, and most families cannot afford it [ 34 ]. The funds saved from the screening can help reduce these later-stage costs or expand the screening scope to detect more early-stage cancer cases, which can also help reduce subsequent treatment expenses [ 35 ]. However, there are several limitations worth mentioning regarding this study. First, our data were solely derived from the population in Taizhou, so the conclusions drawn from the experiment may only be applicable to the GC screening program in Taizhou. Whether or not these conclusions also apply to the entire population of China requires prospective studies conducted at multiple centers for validation. In addition, external validation of this study was also conducted using data from the Taizhou population. Although the model demonstrated good stability and predictive performance, further validation using external data from different regions is still necessary to optimize its applicability. Furthermore, although the data used in this study is prospective gastric cancer screening data, the sheer size of the screening population inevitably introduces certain inherent errors at every stage of the screening process. Despite the government's involvement in the project's coordination, it cannot completely negate the aforementioned challenges. Declarations Ethical Approval and consent to participate The ethical approval numbers for this study were K20210613 and K20190221.All participants were willing to provide information on their questionnaires and hematological tests. Consent to Publication Not applicable. Data Availability statement The datasets generated and/or analysed during the current study are not publicly available due but are available from the corresponding author on reasonable request. Conflict of interest The authors declare no conflicts of interest. Funding statement This work was supported in part by Medical Science and Technology Project of Zhejiang Province (2021PY083, ¥15000, Shao-wei Li), Program of Taizhou Science and Technology Grant (20ywb29, ¥15000, Shao-wei Li), Major Research Program of Taizhou Enze Medical Center Grant (19EZZDA2, ¥120000, Shao-wei Li), Research Program of Taizhou Enze Medical Center Grant (22EZC17, ¥10000, Yan-di Lu),Open Project Program of Key Laboratory of Minimally Invasive Techniques & Rapid Rehabilitation of Digestive System Tumor of Zhejiang Province (21SZDSYS01, ¥100000, Shao-wei Li). Open Project Program of Key Laboratory of Minimally Invasive Techniques & Rapid Rehabilitation of Digestive System Tumor of Zhejiang Province (21SZDSYS09, Xiao-kang Li) Acknowledgment Not applicable. Author Contributions XL M, Y C, LL Y, SW L, XY Fand LP Y participated in Gastric Cancer Screening Program. XY F, HW W, ZQ M, YQ S, SP T and ZC L participated in machine learning algorithm analysis. XY F, Y S, SW L, SJ X, RB Q, JW L, JY L and KX L undertook validation, writing, review, and editing. All authors have read and approved the manuscript. References López MJ, Carbajal J, Alfaro AL, Saravia LG, Zanabria D, Araujo JM, Quispe L, Zevallos A, Buleje JL, Cho CE et al : Characteristics of gastric cancer around the world . Crit Rev Oncol Hematol 2023, 181 :103841. Smyth EC, Nilsson M, Grabsch HI, van Grieken NC, Lordick F: Gastric cancer . Lancet 2020, 396 (10251):635-648. Patel TH, Cecchini M: Targeted Therapies in Advanced Gastric Cancer . Curr Treat Options Oncol 2020, 21 (9):70. Sano T, Coit DG, Kim HH, Roviello F, Kassab P, Wittekind C, Yamamoto Y, Ohashi Y: Proposal of a new stage grouping of gastric cancer for TNM classification: International Gastric Cancer Association staging project . Gastric Cancer 2017, 20 (2):217-225. Kim JH: Important considerations when contemplating endoscopic resection of undifferentiated-type early gastric cancer . World J Gastroenterol 2016, 22 (3):1172-1178. Pilonis ND, Tischkowitz M, Fitzgerald RC, di Pietro M: Hereditary Diffuse Gastric Cancer: Approaches to Screening, Surveillance, and Treatment . Annu Rev Med 2021, 72 :263-280. Thrift AP, Wenker TN, El-Serag HB: Global burden of gastric cancer: epidemiological trends, risk factors, screening and prevention . Nat Rev Clin Oncol 2023, 20 (5):338-349. Wei W, Zeng H, Zheng R, Zhang S, An L, Chen R, Wang S, Sun K, Matsuda T, Bray F et al : Cancer registration in China and its role in cancer prevention and control . Lancet Oncol 2020, 21 (7):e342-e349. Ilic M, Ilic I: Epidemiology of stomach cancer . World J Gastroenterol 2022, 28 (12):1187-1203. Zou WB, Yang F, Li ZS: [How to improve the diagnosis rate of early gastric cancer in China] . Zhejiang Da Xue Xue Bao Yi Xue Ban 2015, 44 (1):9-14. Cai Q, Zhu C, Yuan Y, Feng Q, Feng Y, Hao Y, Li J, Zhang K, Ye G, Ye L et al : Development and validation of a prediction rule for estimating gastric cancer risk in the Chinese high-risk population: a nationwide multicentre study . Gut 2019, 68 (9):1576-1587. Liu K, Qin M, Huang J: The prescreening tool for gastric cancer in China . Gut 2020, 69 (9):1. Park CH, Kim EH, Jung DH, Chung H, Park JC, Shin SK, Lee SK, Lee YC: The new modified ABCD method for gastric neoplasm screening . Gastric Cancer 2016, 19 (1):128-135. Terasawa T, Nishida H, Kato K, Miyashiro I, Yoshikawa T, Takaku R, Hamashima C: Prediction of gastric cancer development by serum pepsinogen test and Helicobacter pylori seropositivity in Eastern Asians: a systematic review and meta-analysis . PLoS One 2014, 9 (10):e109783. Silva GFS, Fagundes TP, Teixeira BC, Chiavegatto Filho ADP: Machine Learning for Hypertension Prediction: a Systematic Review . Curr Hypertens Rep 2022, 24 (11):523-533. Dash TK, Chakraborty C, Mahapatra S, Panda G: Gradient Boosting Machine and Efficient Combination of Features for Speech-Based Detection of COVID-19 . IEEE J Biomed Health Inform 2022, 26 (11):5364-5371. Deo RC: Machine Learning in Medicine . Circulation 2015, 132 (20):1920-1930. Asadikia A, Rajabifard A, Kalantari M: Region-income-based prioritisation of Sustainable Development Goals by Gradient Boosting Machine . Sustain Sci 2022, 17 (5):1939-1957. Patel D, Cheetirala SN, Raut G, Tamegue J, Kia A, Glicksberg B, Freeman R, Levin MA, Timsina P, Klang E: Predicting Adult Hospital Admission from Emergency Department Using Machine Learning: An Inclusive Gradient Boosting Model . J Clin Med 2022, 11 (23). Shojaie M, Cabrerizo M, DeKosky ST, Vaillancourt DE, Loewenstein D, Duara R, Adjouadi M: A transfer learning approach based on gradient boosting machine for diagnosis of Alzheimer's disease . Front Aging Neurosci 2022, 14 :966883. Sabbagh P, Mohammadnia-Afrouzi M, Javanian M, Babazadeh A, Koppolu V, Vasigala VR, Nouri HR, Ebrahimpour S: Diagnostic methods for Helicobacter pylori infection: ideals, options, and limitations . Eur J Clin Microbiol Infect Dis 2019, 38 (1):55-66. Song Z, Chen Y, Lu H, Zeng Z, Wang W, Liu X, Zhang G, Du Q, Xia X, Li C et al : Diagnosis and treatment of Helicobacter pylori infection by physicians in China: A nationwide cross-sectional study . Helicobacter 2022, 27 (3):e12889. Crowe SE: Helicobacter pylori Infection . N Engl J Med 2019, 380 (12):1158-1165. Malfertheiner P, Megraud F, O'Morain CA, Atherton J, Axon AT, Bazzoli F, Gensini GF, Gisbert JP, Graham DY, Rokkas T et al : Management of Helicobacter pylori infection--the Maastricht IV/ Florence Consensus Report . Gut 2012, 61 (5):646-664. Massarrat S, Haj-Sheykholeslami A, Mohamadkhani A, Zendehdel N, Aliasgari A, Rakhshani N, Stolte M, Shahidi SM: Pepsinogen II can be a potential surrogate marker of morphological changes in corpus before and after H. pylori eradication . Biomed Res Int 2014, 2014 :481607. Massarrat S, Haj-Sheykholeslami A: Increased Serum Pepsinogen II Level as a Marker of Pangastritis and Corpus-Predominant Gastritis in Gastric Cancer Prevention . Arch Iran Med 2016, 19 (2):137-140. Leung WK, Wu MS, Kakugawa Y, Kim JJ, Yeoh KG, Goh KL, Wu KC, Wu DC, Sollano J, Kachintorn U et al : Screening for gastric cancer in Asia: current evidence and practice . Lancet Oncol 2008, 9 (3):279-287. Shan JH, Bai XJ, Han LL, Yuan Y, Sun XF: Changes with aging in gastric biomarkers levels and in biochemical factors associated with Helicobacter pylori infection in asymptomatic Chinese population . World J Gastroenterol 2017, 23 (32):5945-5953. Zhou JP, Liu CH, Liu BW, Wang YJ, Benghezal M, Marshall BJ, Tang H, Li H: Association of serum pepsinogens and gastrin-17 with Helicobacter pylori infection assessed by urea breath test . Front Cell Infect Microbiol 2022, 12 :980399. Marchildon P, Balaban DH, Sue M, Charles C, Doobay R, Passaretti N, Peacock J, Marshall BJ, Peura DA: Usefulness of serological IgG antibody determinations for confirming eradication of Helicobacter pylori infection . Am J Gastroenterol 1999, 94 (8):2105-2108. Toyoshima O, Nishizawa T, Sakitani K, Yamakawa T, Takahashi Y, Yamamichi N, Hata K, Seto Y, Koike K, Watanabe H et al : Serum anti-Helicobacter pylori antibody titer and its association with gastric nodularity, atrophy, and age: A cross-sectional study . World J Gastroenterol 2018, 24 (35):4061-4068. Ascherman B, Oh A, Hur C: International cost-effectiveness analysis evaluating endoscopic screening for gastric cancer for populations with low and high risk . Gastric Cancer 2021, 24 (4):878-887. Huang HL, Leung CY, Saito E, Katanoda K, Hur C, Kong CY, Nomura S, Shibuya K: Effect and cost-effectiveness of national gastric cancer screening in Japan: a microsimulation modeling study . BMC Med 2020, 18 (1):257. Abraham P, Wang L, Jiang Z, Gricar J, Tan H, Kelly RJ: Healthcare utilization and total costs of care among patients with advanced metastatic gastric and esophageal cancer . Future Oncol 2021, 17 (3):291-299. Yip W, Fu H, Chen AT, Zhai T, Jian W, Xu R, Pan J, Hu M, Zhou Z, Chen Q et al : 10 years of health-care reform in China: progress and gaps in Universal Health Coverage . Lancet 2019, 394 (10204):1192-1204. Supplementary Document Supplementary Document is not available with this version. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3853941","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":273510747,"identity":"96df4d17-d258-4f9e-ad6a-b88f17e8e508","order_by":0,"name":"Xin-yu Fu","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Xin-yu","middleName":"","lastName":"Fu","suffix":""},{"id":273510748,"identity":"33cdd920-f8a0-45dd-b625-bef540cf2ff8","order_by":1,"name":"Rongbin Qi","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Rongbin","middleName":"","lastName":"Qi","suffix":""},{"id":273510749,"identity":"b838316b-6407-4741-a231-177833623acb","order_by":2,"name":"Shan-jing Xu","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Shan-jing","middleName":"","lastName":"Xu","suffix":""},{"id":273510750,"identity":"e0c1d606-d4f2-417f-818d-bbce79650912","order_by":3,"name":"Meng-sha Huang","email":"","orcid":"","institution":"Taizhou First People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Meng-sha","middleName":"","lastName":"Huang","suffix":""},{"id":273510751,"identity":"33f28a34-619a-4912-aad7-fcb0aa403b76","order_by":4,"name":"Cong-ni Zhu","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Cong-ni","middleName":"","lastName":"Zhu","suffix":""},{"id":273510752,"identity":"5870bdb6-df36-44e1-a9d5-f54d6fc62759","order_by":5,"name":"Hao-wen Wu","email":"","orcid":"","institution":"York University","correspondingAuthor":false,"prefix":"","firstName":"Hao-wen","middleName":"","lastName":"Wu","suffix":""},{"id":273510753,"identity":"09e1a6bc-facc-443d-bdc3-87f6ca641bef","order_by":6,"name":"Zong-qing Ma","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Zong-qing","middleName":"","lastName":"Ma","suffix":""},{"id":273510754,"identity":"f112ac44-29ad-498d-8135-7bcced8749be","order_by":7,"name":"Ya-qi Song","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Ya-qi","middleName":"","lastName":"Song","suffix":""},{"id":273510755,"identity":"e6c00e8e-25e3-4e40-b414-df79f9d8bc1b","order_by":8,"name":"Zhi-cheng Liu","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Zhi-cheng","middleName":"","lastName":"Liu","suffix":""},{"id":273510756,"identity":"dbf82102-3a13-4676-9716-eb8509ed63d3","order_by":9,"name":"Shen-Ping Tang","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Shen-Ping","middleName":"","lastName":"Tang","suffix":""},{"id":273510757,"identity":"e9a7749b-6728-430a-b6a3-860344dad8f8","order_by":10,"name":"Yan-di Lu","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yan-di","middleName":"","lastName":"Lu","suffix":""},{"id":273510758,"identity":"34c04f6b-7f2b-41b1-9fd5-fe51b8a73d9d","order_by":11,"name":"Ling-ling Yan","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Ling-ling","middleName":"","lastName":"Yan","suffix":""},{"id":273510759,"identity":"0d885175-d358-476e-b8e2-10a2159d6562","order_by":12,"name":"Xiao-Kang Li","email":"","orcid":"","institution":"National Center for Child Health and Development","correspondingAuthor":false,"prefix":"","firstName":"Xiao-Kang","middleName":"","lastName":"Li","suffix":""},{"id":273510760,"identity":"3651d7b2-3b41-4e7c-a809-170f7b036043","order_by":13,"name":"Jia-wei Liang","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Jia-wei","middleName":"","lastName":"Liang","suffix":""},{"id":273510761,"identity":"7ae720a2-7acc-469a-89e7-e728defb12f0","order_by":14,"name":"Xin-li Mao","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Xin-li","middleName":"","lastName":"Mao","suffix":""},{"id":273510762,"identity":"ee7e5dad-2b5a-4c8e-9c9d-c474f5eb7dfa","order_by":15,"name":"Li-ping Ye","email":"","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Li-ping","middleName":"","lastName":"Ye","suffix":""},{"id":273510763,"identity":"849c679d-e62b-4911-af78-b73f38bf05d2","order_by":16,"name":"Shao-wei Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9klEQVRIie3PsWoCQRCA4TkG1ubOay/cMwgB4UBYnyTNXbOpFvIAR7IinJXa6ltYhXRZGbASbC3jG1wQhDTinCBpdE2ZYv9qWfZjZgF8vn9YGwHwfEK09ts0p2BgXUT8ElEs5wYSJsZN4EIg7FJ0JgBu0gLcv5SymLZETv0P+doZEU8p5ZNjMZHOVqqYD9GSXqskWxdMVkobB8FQkF4QT9EVJZllEhhyEdyHR9KfFD5SryGb3V0CaVTxFGQSNGR7f4pIo4l6m/Fiy3GlHt63PCV3/CWOLS92kN14SlT/VDLONs+7r7qUNwmj+splfvO5z+fz+f7SCbaHW8X+Oaf6AAAAAElFTkSuQmCC","orcid":"","institution":"Taizhou Hospital of Zhejiang Province affiliated to Wenzhou Medical University","correspondingAuthor":true,"prefix":"","firstName":"Shao-wei","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2024-01-11 15:29:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3853941/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3853941/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":51386731,"identity":"01c2ebd8-0c8d-4a54-96ab-a1864df48c56","added_by":"auto","created_at":"2024-02-20 17:51:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":58615,"visible":true,"origin":"","legend":"\u003cp\u003eFlow chart\u003cstrong\u003e.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/c86230571e20bf61bb781eda.png"},{"id":51386733,"identity":"e4441843-e824-49b1-9195-c09c9188d44a","added_by":"auto","created_at":"2024-02-20 17:51:05","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":13843,"visible":true,"origin":"","legend":"\u003cp\u003ePredictive effectiveness of each algorithmic model is plotted. Deep Learning is a machine learning method based on artificial neural networks; Gradient Boosting Machine (GBM) is an integrated learning algorithm based on the gradient boosting method; Distributed Random Forest (DRF) is an integrated learning algorithm based on random forest constructed from decision trees; Generalized Linear Model (GLM) is a generalized linear model for linearly divisible or approximately linearly divisible data; Stacked Ensemble, also known as Stacked Integration, is an integrated learning technique used to further improve model performance. ROC, receiver operating characteristic.\u003c/p\u003e","description":"","filename":"figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/8e7026523cc8e92053652c65.png"},{"id":51386735,"identity":"20819611-c3ac-4ce5-a286-5fd3e16f6040","added_by":"auto","created_at":"2024-02-20 17:51:05","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":412681,"visible":true,"origin":"","legend":"\u003cp\u003eThe\u003cstrong\u003e \u003c/strong\u003eAUC and AUCPR of the original model on the training set. A shows the area under the receiver operating characteristic curve (AUC) of the training set, which measures the performance of the model in terms of its ability to distinguish between positive and negative samples. B shows the area under the precision-recall curve (AUCPR) of the training set, which evaluates the trade-off between precision (the ratio of true positive predictions to all positive predictions) and recall (the ratio of true positive predictions to all actual positive samples).\u003c/p\u003e","description":"","filename":"figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/898fe1f15c418e9de775768e.png"},{"id":51386737,"identity":"019a1c8b-232a-4a66-8ebd-b5dce72aeb81","added_by":"auto","created_at":"2024-02-20 17:51:05","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":386843,"visible":true,"origin":"","legend":"\u003cp\u003eThe\u003cstrong\u003e \u003c/strong\u003eAUC and AUCPR of the original model on the test set. A shows the area under the receiver operating characteristic curve (AUC) of the test set. B shows the area under the precision-recall curve (AUCPR) of the test set.\u003c/p\u003e","description":"","filename":"figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/090c49e5ee9cb411584c62d5.png"},{"id":51387601,"identity":"7b7efcac-803f-48ab-a6ae-a4ea257ce7cf","added_by":"auto","created_at":"2024-02-20 17:59:05","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1703924,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of the importance of model variables. A shows a bar chart indicating the weights of each variable in the model; B shows a SHapley Additive exPlanations (SHAP) plot, displaying the contribution of each feature to the prediction results; C shows a partial dependence plot of \u003cem\u003eHelicobacter pylori\u003c/em\u003e IgG antibodies, illustrating the influence of this variable on the prediction results.\u003c/p\u003e","description":"","filename":"figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/f1a82ef8b44303fd88e56dd8.png"},{"id":51386730,"identity":"ab7f61a7-125c-401f-8056-c27bf5a4831c","added_by":"auto","created_at":"2024-02-20 17:51:04","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":375865,"visible":true,"origin":"","legend":"\u003cp\u003eThe AUC and AUCPR of the model in the training set after removing the variable IgG. A shows the area under the receiver operating characteristic curve (AUC) of the training set. B shows the area under the precision-recall curve (AUCPR) of the training set.\u003c/p\u003e","description":"","filename":"figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/ee0467e2365a1d94532b547f.png"},{"id":51386736,"identity":"fb4ce7a5-58a9-46f2-b66e-ddbce800b44b","added_by":"auto","created_at":"2024-02-20 17:51:05","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":379834,"visible":true,"origin":"","legend":"\u003cp\u003eAUC and AUCPR of the model after removal of the variable IgG in the validation set. A shows the area under the receiver operating characteristic curve (AUC) of the test set. B shows the area under the precision-recall curve (AUCPR) of the test set.\u003c/p\u003e","description":"","filename":"figure7.png","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/a45cff5d93847d2c8b435419.png"},{"id":59407291,"identity":"5bf8ae1c-159c-4af2-9621-f9d90d59f2c3","added_by":"auto","created_at":"2024-07-01 11:40:55","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4716616,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3853941/v1/f7965497-1775-4c7e-badb-8dbc8dc57e71.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Study on the Feasibility of Optimizing Gastric Cancer Screening to Reduce Screening Costs in China Using a Gradient Boosting Machine: A prospective, large-sample, single-center study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eGastric cancer (GC) remains one of the most common malignancies worldwide, with over one million new cases estimated annually, ranking it as the fifth-most diagnosed malignancy globally [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Owing to the often late-stage diagnosis of GC, it has a high mortality rate, ranking as the third leading cause of cancer-related deaths [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. However, it is worth noting that early GC (EGC) has a 5-year survival rate exceeding 90%, far higher than that of advanced GC (AGC), which has a 30% survival rate [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Such a high survival rate underscores the importance of GC screening, especially in countries with a high incidence of this disease [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. According to the latest reports from 2020, China has the highest number of registered GC cases, with approximately 820,000 new cases and 580,000 deaths. Therefore, conducting large-scale GC screening programs is imperative in China [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCurrently, China does not have a nationwide GC screening program, and early detection relies solely on opportunistic screening. To address this situation, China's existing national GC screening guidelines recommend screening all high-risk individuals starting at 40 years old [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Gastric endoscopy and a tissue biopsy for a histological examination are currently the gold standards for GC screening and diagnoses. However, screening the entire \"high-risk\" population with gastric endoscopy is inefficient and impractical, as it is expected that only 1\u0026ndash;3% of this population will actually have GC [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Furthermore, it is estimated that there are over 300\u0026nbsp;million \"high-risk\" individuals in China, and because of the high cost, screening with gastric endoscopy is unlikely to be feasible for all individuals [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTherefore, a risk stratification method is needed as a preliminary screening tool before gastric endoscopy to further identify truly high-risk individuals among the previously defined \"high-risk\" population [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. In 2018, a nationwide multicenter prospective trial conducted in a Chinese population developed and validated Lee's screening score for this purpose, aiming to distinguish individuals who need further endoscopic evaluations. The screening score demonstrated a good discriminative ability, with an area under the curve (AUC) of 0.76 [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eEffective primary screening can significantly reduce the cost of GC screening, but current primary screening still requires the detection of four serological markers: gastrin, \u003cem\u003eHelicobacter pylori\u003c/em\u003e IgG antibody, and pepsinogens (PGs) I and II [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. The question of whether or not it is possible to optimize the required markers for primary screening to further reduce screening costs is meaningful. In recent years, machine learning has shown strong potential utility in data processing and analyses [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Prediction models constructed using machine learning have demonstrated good predictive performance and stability, thus garnering increasing attention [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. The gradient boosting machine (GBM), a powerful machine-learning algorithm, is capable of efficiently handling high-dimensional sparse data, capturing nonlinear relationships and interactions, and enhancing model predictive performance through iterative optimization [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. In particular, it excels in solving classification problems. This study is based on prospective gastric cancer screening data, aiming to significantly improve and optimize Lee\u0026rsquo;s screening score using the GBM algorithm in order to further reduce the cost of initial gastric cancer screening.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003eStudy design and participants\u003c/h2\u003e\n\u003cp\u003eThis prospective study on gastric cancer screening is supported by the leadership of the government of Linhai, Taizhou. From January 2021 to July 2023, a 30-month gastric cancer screening project will be carried out targeting the general population aged 45 to 75. The project is led by Taizhou Hospital, with subordinate hospitals participating and Taizhou Hospital conducting comprehensive analysis. This research has been approved by the Ethics Review Board of Wenzhou Medical University, affiliated with Taizhou Hospital in Zhejiang Province (Approval Number: K20210613). To retrospectively collect data on the GC screening of the Minimum Living Guarantee Crowd (MLGC) conducted from January to December 2019. This project, initiated by the Taizhou Municipal Government and led by Taizhou Hospital, was summarized based on the approval of the Ethical Review Committee of Zhejiang Taizhou Hospital affiliated with Wenzhou Medical University (approval number: K20190221). The data of the GC screening on MLGC included a total of 32,994 cases and will serve as the validation dataset for the model.\u003c/p\u003e\n\u003cp\u003eThe inclusion exclusion criteria for the general population gastric cancer screening program were as follows:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003eInclusion criteria: 1) Between 45\u0026ndash;75 years of age; 2) Willing to participate in the screening project and sign an informed consent form.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExclusion criteria: 1) Decline to complete the questionnaire; 2) Refuse serological testing; 3) Incomplete or missing data.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThe inclusion exclusion criteria for the MLGC screening program are as follows\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003eInclusion criteria: 1) All patients participating in the MLGC program.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eExclusion criteria: 1) Incomplete or missing data.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n\u003ch2\u003eQuestionnaires and serologic tests\u003c/h2\u003e\n\u003cp\u003eThe questionnaire survey included information about the smoking history, alcohol consumption history, medical history, and other relevant details, which were used to describe the demographic characteristics of the study population. Blood tests included measurement of gastrin-17 (G-17), PG I/II, and \u003cem\u003eH. pylori\u003c/em\u003e IgG antibodies. All tests were conducted uniformly by government-designated testing facilities to ensure the stability of blood test results.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n\u003ch2\u003eDefinitions of the endoscopy group and room-following group\u003c/h2\u003e\n\u003cp\u003eLee's screening score (See Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e of the supplementary document) for GC is based on variables of age, sex, G-17, PG I/II, and \u003cem\u003eH. pylori\u003c/em\u003e antibodies. It classifies individuals into low-, intermediate-, and high-risk groups. For individuals classified as intermediate or high risk, further investigations using endoscopy are recommended.\u003c/p\u003e\n\u003cp\u003eIn the present study, based on the score table, individuals in the intermediate- and high-risk groups were categorized into the endoscopy group, whereas individuals in the low-risk group were assigned to the follow-up group.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003eData analyses\u003c/h2\u003e\n\u003cp\u003eThis study utilizes the \"h2o\" automatic machine learning package in Java to automatically build and select the most effective machine learning algorithms for model construction. This study uses the general population of Linyi City as the training dataset, employs five-fold cross-validation for internal validation, and uses the MPGC population as the test dataset for external validation. Model evaluation was conducted using the area under the precision-recall curve (AUCPR) and the AUC. Both the AUC and AUCPR involve values between 0.5 and 1, with values closer to 1 indicating better model performance and those closer to 0.5 indicating random guessing.\u003c/p\u003e\n\u003cp\u003eInter-group comparisons were carried out using various statistical methods, such as the Mann-Whitney U test, Kruskal-Wallis rank-sum test, and Fisher's exact probability method. A significance level of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and q\u0026thinsp;\u0026lt;\u0026thinsp;0.05 is considered to indicate a significant difference. All statistical analyses and machine-learning procedures were implemented using the R programming language (version 4.2.3).\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n\u003ch2\u003ePatients\u003c/h2\u003e\n\u003cp\u003eFrom January 2021 to July 2023, a GC screening project was conducted in Linhai City, with 207,005 voluntary participants who completed the survey. Among these, there were 76,784 cases in 2021, 77,330 cases in 2022, and 52,891 cases in 2023. Of these participants, 36,722 had a history of smoking, and 28,586 had a history of alcohol consumption. In addition, 46,473 patients had a history of hypertension, 14,859 had diabetes, 595 had coronary heart disease, 181 had chronic kidney disease, and 7,092 had other diseases. A total of 45,506 (21.98%) patients had a history of undergoing gastroscopy. In 2021, 74,614 patients underwent serological testing, while 2,170 patients were excluded because of refusal or failure to undergo serological testing within the specified time. In 2022, 73,686 patients underwent serological testing and 3,644 patients were excluded for the same reasons. As of July 2023, 52,217 patients had undergone serological testing in 2023, with 674 patients excluded for refusal or failure to adhere to the testing schedule. Ultimately, 200,517 patients underwent serological testing, with 4,877 excluded owing to missing information or incomplete data, leaving 195,640 cases for the analysis and model training. According to the Lee's score table, 48,118 patients were assigned to the endoscopy group, and 147,522 patients were included in the follow-up group. The distribution of indicators for each group is shown in Table\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eLinhai City General Population Baseline Table\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eVariable\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eOverall\u003c/p\u003e\n\u003cp\u003eN\u0026thinsp;=\u0026thinsp;195,640\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003efollow-up group\u003c/p\u003e\n\u003cp\u003eN\u0026thinsp;=\u0026thinsp;147,522\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eendoscopy group\u003c/p\u003e\n\u003cp\u003eN\u0026thinsp;=\u0026thinsp;48,118\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ep-value\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eq-value\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGender\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMale\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e78,539 (40%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e39,700 (27%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e38,839 (81%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eFemale\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e117,101 (60%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e107,822 (73%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9,279 (19%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eAge\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e60 (54, 66)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e59 (53, 65)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e64 (59, 68)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGastrin G-17\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.65 (1.00, 3.32)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.24 (1.00, 2.43)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.30 (2.06, 6.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePepsinogen I\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e87 (70, 111)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e86 (70, 108)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e92 (70, 120)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePepsinogen II\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6.6 (4.7, 9.1)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6.5 (4.8, 8.6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7.2 (3.3, 10.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eH. pylori IgG antibody\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNegative\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e118,057 (60%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e96,856 (66%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e21,201 (44%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePositive\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e77,583 (40%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50,666 (34%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e26,917 (56%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"6\" align=\"left\"\u003e\n\u003cp\u003e\u003csup\u003e1\u003c/sup\u003en (%); Median (IQR); \u003csup\u003e2\u003c/sup\u003ePearson's Chi-squared test; Wilcoxon rank sum test; \u003csup\u003e3\u003c/sup\u003eFalse discovery rate correction for multiple testing\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eFrom January to December 2019, a GC screening project for the MLGC in Taizhou City included 33,054 voluntary participants who completed the survey. Among these, 10,171 had a history of smoking, 7,771 had a history of alcohol consumption, 8,284 had a history of hypertension, 2,252 had a history of diabetes, 2,230 had hyperlipidemia, and 929 had other diseases. A total of 983 patients (2.97%) had a history of gastroscopy. All of these patients underwent serological testing. Sixty cases were excluded owing to missing information, incomplete data, and other reasons, leaving 32,994 cases for the analysis as an external validation set. In the validation set, 10,608 cases were defined as the endoscopy group, and 22,386 cases were defined as the follow-up group. Please refer to Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e for further details. A flowchart is shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eTaizhou Minimum Living Guarantee Crowd Baseline Table\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eVariable\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eOverall\u003c/p\u003e\n\u003cp\u003eN\u0026thinsp;=\u0026thinsp;32,994\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003efollow-up group\u003c/p\u003e\n\u003cp\u003eN\u0026thinsp;=\u0026thinsp;22,386\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eendoscopy group\u003c/p\u003e\n\u003cp\u003eN\u0026thinsp;=\u0026thinsp;10,608\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ep-value\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eq-value\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGender\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMale\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20,107 (61%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10,464 (47%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9,643 (91%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eFemale\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12,887 (39%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e11,922 (53%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e965 (9%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eAge\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e56 (50, 63)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e54 (48, 62)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e61 (55, 66)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGastrin G-17\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.0 (1.0, 4.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.3 (1.0, 3.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.0 (2.2, 8.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePepsinogen I\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e136 (102, 165)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e129 (99, 165)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e154 (114, 165)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePepsinogen II\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6.7 (4.4, 11.2)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5.9 (4.0, 9.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9.4 (5.9, 15.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eH. pylori IgG antibody\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNegative\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20,489 (62%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e15,035 (67%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5,454 (51%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePositive\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12,505 (38%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7,351 (33%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5,154 (49%)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"6\" align=\"left\"\u003e\n\u003cp\u003e\u003csup\u003e1\u003c/sup\u003en (%); Median (IQR); \u003csup\u003e2\u003c/sup\u003ePearson's Chi-squared test; Wilcoxon rank sum test; \u003csup\u003e3\u003c/sup\u003eFalse discovery rate correction for multiple testing\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n\u003ch2\u003eMachine-learning model selection\u003c/h2\u003e\n\u003cp\u003eBased on the h2o package, automated machine learning was conducted on the training dataset, resulting in the training and validation of the 12 machine learning models. These models encompassed five different algorithms: deep learning, random forest, GBM, generalized linear models, and Stacked Ensemble. Among them, the Stacked Ensemble exhibited the best performance, with an AUC of 0.99884, followed closely by the GBM, with an AUC of 0.99883 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). Given that the GBM's predictive performance was comparable to that of the Stacked Ensemble while offering greater model stability, we chose to focus on the GBM for subsequent research. In addition, the Stacked Ensemble, which is an ensemble model composed of multiple sub-models, presents challenges in terms of model interpretability. Therefore, further research in this study is based on the GBM model.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n\u003ch2\u003eModeling GBM distinctions\u003c/h2\u003e\n\u003cp\u003eThe GBM classification model was built to distinguish between the endoscopy and follow-up groups using the following variables: age, sex, G-17, PG I, PG II, and IgG. The h2o package for automated machine learning was used, and the best results were obtained when using the GBM_5 learner, which was identified by its learner-type code. The model achieved excellent predictive performance, with an AUC of 0.99938 and an AUCPR of 0.99823 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). To assess the robustness of the model, it underwent an internal 5-fold cross-validation, resulting in an AUC of 0.99883 and an AUCPR of 0.99656. External validation was performed using the MLGC dataset, yielding an AUC of 0.99742 and AUCPR of 0.99454 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e). The GBM model demonstrated nearly 100% predictive accuracy in distinguishing between patients in the endoscopy and follow-up groups.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n\u003ch2\u003eAn analysis of the importance of variables\u003c/h2\u003e\n\u003cp\u003eA visual analysis of the GBM was conducted to determine the importance of each variable within the model, and the results are shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. Among the variables, IgG had the lowest importance score, with a value of 0.07516, followed by PG I, with a score of 0.08827 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003eA). In the SHAP plot, it was observed that IgG had a greater negative contribution than a positive contribution (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003eB). In addition, when creating a partial dependence plot for the IgG variable, it was found that both positive and negative IgG values had a similar effect on the mean response (average response) of the outcome variable (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003eC). Based on these results, it was concluded that the IgG variable could be excluded from the model, as it had low importance and did not significantly influence the outcome variable.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n\u003ch2\u003eRemodeling of the exclusion variable IgG\u003c/h2\u003e\n\u003cp\u003eBased on the h2o package, GBM_2 achieved the best performance. The AUC was 0.99817 and the AUCPR was 0.99462 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e). During the 5-fold cross-validation, the AUC was 0.99717, and the AUCPR was 0.99154. In the external validation set, the AUC was 0.99605, and the AUCPR was 0.99135 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e). Even after removing IgG, the model maintained a high level of predictive accuracy. Please refer to Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e for visualization of the results.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eSuccessfully conducting large-scale GC screening in China requires dividing the screening efforts into two parts: the initial screening and a follow-up examination [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e]. Differentiating individuals who require a further endoscopic examination remains a challenge. Based on a multi-center prospective study specific to the Chinese population, the feasibility of Lee\u0026apos;s score table has been proposed and validated. Currently, Lee\u0026apos;s screening score is widely applied throughout China in GC screening programs. The GC screening projects involved in this study were based on Lee\u0026rsquo;s score table. However, use of this table requires the measurement of four serological markers: G-17, PG I/II, and \u003cem\u003eH. pylori\u003c/em\u003e antibodies [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e]. The cost of initial screening remains high for large-scale surveys; therefore, this study attempted to optimize the screening process using machine-learning methods. By training and validating a GBM model, it was found that removing \u003cem\u003eH. pylori\u003c/em\u003e antibodies from the GBM model had a negligible effect on its performance for differentiating individuals who require gastroscopy, achieving an approximately 100% discrimination rate, similar to the full model without excluding \u003cem\u003eH. pylori\u003c/em\u003e antibodies.\u003c/p\u003e\n\u003cp\u003eIn this study, automated machine learning was used to select the GBM algorithm, which had the best fit and was easy to analyze, to train and optimize the initial screening for GC [\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e]. The GBM algorithm has strong predictive performance, particularly for handling structured data and regression and classification tasks. It also demonstrates robustness against outliers and noise in data [\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e20\u003c/span\u003e]. In addition, the GBM algorithm can evaluate feature importance, aiding in understanding the impact of data features on the model. Based on the powerful predictive performance of GBM, we included age, sex, G-17, pepsinogen I and II (PG I/II), and \u003cem\u003eH. pylori\u003c/em\u003e antibodies to train the full model, which achieved a nearly 100% discrimination rate (AUC\u0026thinsp;=\u0026thinsp;0.99938 and AUCPR\u0026thinsp;=\u0026thinsp;0.99823). Furthermore, the discrimination performance remained close to 100% during both internal and external validations. In the internal validation, the AUC was 0.99938, and the AUCPR was 0.99823, whereas in the external validation, the AUC was 0.99742, and the AUCPR was 0.99454. Therefore, we consider the full model to be highly stable.\u003c/p\u003e\n\u003cp\u003eIn the visual analysis of the model, we found that the variable \u003cem\u003eH. pylori\u003c/em\u003e antibodies had a very low weight in the model. We also discovered that they had a low weight in Lee\u0026apos;s score table, with only 1 point. Therefore, we proposed optimizing the inclusion of IgG antibodies. After removing IgG antibodies, we built a discrimination model based on the GBM algorithm and found that the GBM discrimination model without IgG antibodies performed comparably to the full model, with an AUC of 0.998167 and an AUCPR of 0.9946214. It also demonstrated excellent predictive performance in internal (AUC of 0.9971681 and AUCPR of 0.991536) and external validation (AUC of 0.9960488 and AUCPR of 0.9913469). Based on these results, we conclude that the measurement of \u003cem\u003eH. pylori\u003c/em\u003e IgG antibodies was not necessary when using the predictive model constructed with the GBM algorithm to differentiate individuals who require GC screening.\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eH. pylori\u003c/em\u003e infection is considered a high-risk factor for GC, especially in China, where GC is prevalent [\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e]. At present, invasive and non-invasive methods are available for detecting \u003cem\u003eH. pylori\u003c/em\u003e infections. Invasive methods include the Rapid Urease Test, histological examination, and genetic testing, and non-invasive methods include the Urea Breath Test, serological antibody testing, and stool antigen tests, among others [\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e]. Among the screening options based on Lee\u0026rsquo;s criteria, serological detection of \u003cem\u003eH. pylori\u003c/em\u003e IgG antibodies is the most suitable screening strategy, as both gastrin and PG tests require blood sampling. However, according to numerous studies, IgG antibody testing has a high negative predictive value (NPV). The ability of this test to detect active infection depends on a number of factors, such as the patient\u0026apos;s age, clinical condition of the infection, choice of antigen used for antibody preparation in the enzyme-linked immunosorbent assay kit, and prevalence of infection. IgG antibodies are unable to differentiate between current infection and previous exposure and can still be detected several months after treatment, which may interfere with GC screening efforts [\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eH. pylori\u003c/em\u003e infection can lead to secondary hypergastrinemia and hyperpepsinogenemia, indicating that both G-17 and PG I/II levels can reflect the presence of \u003cem\u003eH. pylori\u003c/em\u003e infection to some extent [\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e]. Studies have shown that PG I and PG II levels are associated with \u003cem\u003eH. pylori\u003c/em\u003e infection; furthermore, a close relationship between G-17 and \u003cem\u003eH. pylori\u003c/em\u003e, as patients with \u003cem\u003eH. pylori\u003c/em\u003e infection have significantly higher levels of G-17 than those without infection [\u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e]. This suggests that G-17, PG I, and PG II are associated with \u003cem\u003eH. pylori\u003c/em\u003e infections. As \u003cem\u003eH. pylori\u003c/em\u003e-induced gastritis is the most common condition, these three markers not only reflect the overall gastric function but also indirectly indicate the presence of \u003cem\u003eH. pylori\u003c/em\u003e infection in the body [\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eThere is no denying that \u003cem\u003eH. pylori\u003c/em\u003e infection is closely associated with the development of GC and related precancerous lesions. However, the accuracy of IgG antibody testing is questionable, as a positive result for IgG antibodies does not accurately reflect the true infection status [\u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e]. A retrospective study conducted in Japan showed that IgG can indicate the presence of gastric mucosal damage to some extent, but its accuracy is influenced by age [\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e]. In patients\u0026thinsp;\u0026gt;\u0026thinsp;65 years old, the response of IgG antibodies to \u003cem\u003eH. pylori\u003c/em\u003e is diminished, making it difficult to indicate the presence of \u003cem\u003eH. pylori\u003c/em\u003e infection [\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e]. The dynamic measurement of antibody titers is reportedly necessary to accurately reflect the status of \u003cem\u003eH. pylori\u003c/em\u003e infection and treatment efficacy. However, the implementation of this method requires more manpower, financial resources, and material resources, making it challenging to achieve in the screening process.\u003c/p\u003e\n\u003cp\u003eBased on these findings, we believe that the role of \u003cem\u003eH. pylori\u003c/em\u003e IgG antibodies in GC screening is limited. Our trained GBM model based on a dataset of 200,000 samples confirms this viewpoint. Therefore, we believe that \u003cem\u003eH. pylori\u003c/em\u003e antibody testing is not necessary when using the GBM model to discriminate individuals who require an endoscopic examination.\u003c/p\u003e\n\u003cp\u003eIn China, where GC is highly prevalent, cancer screening is necessary because early GC has a favorable prognosis [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e]. The first step in conducting GC screening in China is to accurately identify individuals who need further gastroscopy examinations at the lowest possible cost [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e]. This study, based on a large sample size, trained a GBM model to distinguish individuals requiring a gastroscopy examination and successfully optimized Lee\u0026apos;s screening criteria. Due to the large population in China, a considerable number of individuals need to undergo GC screening, and the cost savings achieved by eliminating one serological marker test are significant. Based on the company\u0026apos;s testing fees, the detection cost of \u003cem\u003eH. pylori\u003c/em\u003e IgG antibodies was 30 RMB per person. Therefore, based on the currently completed screening population, excluding IgG testing can save 7,201,770 RMB. These savings could cover an additional 38,929 initial screenings. Currently, the detection rate of the GC screening program in Lihai City is 0.22%. By extending the coverage to an additional population of over 30,000 individuals, approximately 84 more patients with GC could be detected. According to relevant literature, GC screening is cost-effective in countries with a high incidence of GC. By improving screening techniques, it is possible to further reduce screening costs and achieve higher quality-adjusted life-years with a lower incremental cost-effectiveness ratio [\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e33\u003c/span\u003e]. In addition, with approximately 300\u0026nbsp;million people in China requiring screening, nearly 9\u0026nbsp;billion RMB in screening costs can be saved, which is quite significant [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e]. The treatment costs for advanced-stage GC are extremely high, and most families cannot afford it [\u003cspan class=\"CitationRef\"\u003e34\u003c/span\u003e]. The funds saved from the screening can help reduce these later-stage costs or expand the screening scope to detect more early-stage cancer cases, which can also help reduce subsequent treatment expenses [\u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eHowever, there are several limitations worth mentioning regarding this study. First, our data were solely derived from the population in Taizhou, so the conclusions drawn from the experiment may only be applicable to the GC screening program in Taizhou. Whether or not these conclusions also apply to the entire population of China requires prospective studies conducted at multiple centers for validation. In addition, external validation of this study was also conducted using data from the Taizhou population. Although the model demonstrated good stability and predictive performance, further validation using external data from different regions is still necessary to optimize its applicability. Furthermore, although the data used in this study is prospective gastric cancer screening data, the sheer size of the screening population inevitably introduces certain inherent errors at every stage of the screening process. Despite the government\u0026apos;s involvement in the project\u0026apos;s coordination, it cannot completely negate the aforementioned challenges.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthical Approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe ethical approval numbers for this study were K20210613 and K20190221.All participants were willing to provide information on their questionnaires and hematological tests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analysed during the current study are not publicly available due but are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported in part by Medical Science and Technology Project of Zhejiang Province (2021PY083, \u0026yen;15000, Shao-wei Li), Program of Taizhou Science and Technology Grant (20ywb29, \u0026yen;15000, Shao-wei Li), Major Research Program of Taizhou Enze Medical Center Grant (19EZZDA2, \u0026yen;120000, Shao-wei Li), Research Program of Taizhou Enze Medical Center Grant (22EZC17, \u0026yen;10000, Yan-di Lu),Open Project Program of Key Laboratory of Minimally Invasive Techniques \u0026amp; Rapid Rehabilitation of Digestive System Tumor of Zhejiang Province (21SZDSYS01, \u0026yen;100000, Shao-wei Li). Open Project Program of Key Laboratory of Minimally Invasive Techniques \u0026amp; Rapid Rehabilitation of Digestive System Tumor of Zhejiang Province (21SZDSYS09, Xiao-kang Li)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgment\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eXL M, Y C, LL Y, SW L, XY Fand LP Y participated in Gastric Cancer Screening Program. XY F, HW W, ZQ M, YQ S, SP T and ZC L participated in machine learning algorithm analysis. XY F, Y S, SW L, SJ X, RB Q, JW L, JY L and KX L undertook validation, writing, review, and editing. All authors have read and approved the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eL\u0026oacute;pez MJ, Carbajal J, Alfaro AL, Saravia LG, Zanabria D, Araujo JM, Quispe L, Zevallos A, Buleje JL, Cho CE\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eCharacteristics of gastric cancer around the world\u003c/strong\u003e. \u003cem\u003eCrit Rev Oncol Hematol \u003c/em\u003e2023, \u003cstrong\u003e181\u003c/strong\u003e:103841.\u003c/li\u003e\n\u003cli\u003eSmyth EC, Nilsson M, Grabsch HI, van Grieken NC, Lordick F: \u003cstrong\u003eGastric cancer\u003c/strong\u003e. \u003cem\u003eLancet \u003c/em\u003e2020, \u003cstrong\u003e396\u003c/strong\u003e(10251):635-648.\u003c/li\u003e\n\u003cli\u003ePatel TH, Cecchini M: \u003cstrong\u003eTargeted Therapies in Advanced Gastric Cancer\u003c/strong\u003e. \u003cem\u003eCurr Treat Options Oncol \u003c/em\u003e2020, \u003cstrong\u003e21\u003c/strong\u003e(9):70.\u003c/li\u003e\n\u003cli\u003eSano T, Coit DG, Kim HH, Roviello F, Kassab P, Wittekind C, Yamamoto Y, Ohashi Y: \u003cstrong\u003eProposal of a new stage grouping of gastric cancer for TNM classification: International Gastric Cancer Association staging project\u003c/strong\u003e. \u003cem\u003eGastric Cancer \u003c/em\u003e2017, \u003cstrong\u003e20\u003c/strong\u003e(2):217-225.\u003c/li\u003e\n\u003cli\u003eKim JH: \u003cstrong\u003eImportant considerations when contemplating endoscopic resection of undifferentiated-type early gastric cancer\u003c/strong\u003e. \u003cem\u003eWorld J Gastroenterol \u003c/em\u003e2016, \u003cstrong\u003e22\u003c/strong\u003e(3):1172-1178.\u003c/li\u003e\n\u003cli\u003ePilonis ND, Tischkowitz M, Fitzgerald RC, di Pietro M: \u003cstrong\u003eHereditary Diffuse Gastric Cancer: Approaches to Screening, Surveillance, and Treatment\u003c/strong\u003e. \u003cem\u003eAnnu Rev Med \u003c/em\u003e2021, \u003cstrong\u003e72\u003c/strong\u003e:263-280.\u003c/li\u003e\n\u003cli\u003eThrift AP, Wenker TN, El-Serag HB: \u003cstrong\u003eGlobal burden of gastric cancer: epidemiological trends, risk factors, screening and prevention\u003c/strong\u003e. \u003cem\u003eNat Rev Clin Oncol \u003c/em\u003e2023, \u003cstrong\u003e20\u003c/strong\u003e(5):338-349.\u003c/li\u003e\n\u003cli\u003eWei W, Zeng H, Zheng R, Zhang S, An L, Chen R, Wang S, Sun K, Matsuda T, Bray F\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eCancer registration in China and its role in cancer prevention and control\u003c/strong\u003e. \u003cem\u003eLancet Oncol \u003c/em\u003e2020, \u003cstrong\u003e21\u003c/strong\u003e(7):e342-e349.\u003c/li\u003e\n\u003cli\u003eIlic M, Ilic I: \u003cstrong\u003eEpidemiology of stomach cancer\u003c/strong\u003e. \u003cem\u003eWorld J Gastroenterol \u003c/em\u003e2022, \u003cstrong\u003e28\u003c/strong\u003e(12):1187-1203.\u003c/li\u003e\n\u003cli\u003eZou WB, Yang F, Li ZS: \u003cstrong\u003e[How to improve the diagnosis rate of early gastric cancer in China]\u003c/strong\u003e. \u003cem\u003eZhejiang Da Xue Xue Bao Yi Xue Ban \u003c/em\u003e2015, \u003cstrong\u003e44\u003c/strong\u003e(1):9-14.\u003c/li\u003e\n\u003cli\u003eCai Q, Zhu C, Yuan Y, Feng Q, Feng Y, Hao Y, Li J, Zhang K, Ye G, Ye L\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eDevelopment and validation of a prediction rule for estimating gastric cancer risk in the Chinese high-risk population: a nationwide multicentre study\u003c/strong\u003e. \u003cem\u003eGut \u003c/em\u003e2019, \u003cstrong\u003e68\u003c/strong\u003e(9):1576-1587.\u003c/li\u003e\n\u003cli\u003eLiu K, Qin M, Huang J: \u003cstrong\u003eThe prescreening tool for gastric cancer in China\u003c/strong\u003e. \u003cem\u003eGut \u003c/em\u003e2020, \u003cstrong\u003e69\u003c/strong\u003e(9):1.\u003c/li\u003e\n\u003cli\u003ePark CH, Kim EH, Jung DH, Chung H, Park JC, Shin SK, Lee SK, Lee YC: \u003cstrong\u003eThe new modified ABCD method for gastric neoplasm screening\u003c/strong\u003e. \u003cem\u003eGastric Cancer \u003c/em\u003e2016, \u003cstrong\u003e19\u003c/strong\u003e(1):128-135.\u003c/li\u003e\n\u003cli\u003eTerasawa T, Nishida H, Kato K, Miyashiro I, Yoshikawa T, Takaku R, Hamashima C: \u003cstrong\u003ePrediction of gastric cancer development by serum pepsinogen test and Helicobacter pylori seropositivity in Eastern Asians: a systematic review and meta-analysis\u003c/strong\u003e. \u003cem\u003ePLoS One \u003c/em\u003e2014, \u003cstrong\u003e9\u003c/strong\u003e(10):e109783.\u003c/li\u003e\n\u003cli\u003eSilva GFS, Fagundes TP, Teixeira BC, Chiavegatto Filho ADP: \u003cstrong\u003eMachine Learning for Hypertension Prediction: a Systematic Review\u003c/strong\u003e. \u003cem\u003eCurr Hypertens Rep \u003c/em\u003e2022, \u003cstrong\u003e24\u003c/strong\u003e(11):523-533.\u003c/li\u003e\n\u003cli\u003eDash TK, Chakraborty C, Mahapatra S, Panda G: \u003cstrong\u003eGradient Boosting Machine and Efficient Combination of Features for Speech-Based Detection of COVID-19\u003c/strong\u003e. \u003cem\u003eIEEE J Biomed Health Inform \u003c/em\u003e2022, \u003cstrong\u003e26\u003c/strong\u003e(11):5364-5371.\u003c/li\u003e\n\u003cli\u003eDeo RC: \u003cstrong\u003eMachine Learning in Medicine\u003c/strong\u003e. \u003cem\u003eCirculation \u003c/em\u003e2015, \u003cstrong\u003e132\u003c/strong\u003e(20):1920-1930.\u003c/li\u003e\n\u003cli\u003eAsadikia A, Rajabifard A, Kalantari M: \u003cstrong\u003eRegion-income-based prioritisation of Sustainable Development Goals by Gradient Boosting Machine\u003c/strong\u003e. \u003cem\u003eSustain Sci \u003c/em\u003e2022, \u003cstrong\u003e17\u003c/strong\u003e(5):1939-1957.\u003c/li\u003e\n\u003cli\u003ePatel D, Cheetirala SN, Raut G, Tamegue J, Kia A, Glicksberg B, Freeman R, Levin MA, Timsina P, Klang E: \u003cstrong\u003ePredicting Adult Hospital Admission from Emergency Department Using Machine Learning: An Inclusive Gradient Boosting Model\u003c/strong\u003e. \u003cem\u003eJ Clin Med \u003c/em\u003e2022, \u003cstrong\u003e11\u003c/strong\u003e(23).\u003c/li\u003e\n\u003cli\u003eShojaie M, Cabrerizo M, DeKosky ST, Vaillancourt DE, Loewenstein D, Duara R, Adjouadi M: \u003cstrong\u003eA transfer learning approach based on gradient boosting machine for diagnosis of Alzheimer\u0026apos;s disease\u003c/strong\u003e. \u003cem\u003eFront Aging Neurosci \u003c/em\u003e2022, \u003cstrong\u003e14\u003c/strong\u003e:966883.\u003c/li\u003e\n\u003cli\u003eSabbagh P, Mohammadnia-Afrouzi M, Javanian M, Babazadeh A, Koppolu V, Vasigala VR, Nouri HR, Ebrahimpour S: \u003cstrong\u003eDiagnostic methods for Helicobacter pylori infection: ideals, options, and limitations\u003c/strong\u003e. \u003cem\u003eEur J Clin Microbiol Infect Dis \u003c/em\u003e2019, \u003cstrong\u003e38\u003c/strong\u003e(1):55-66.\u003c/li\u003e\n\u003cli\u003eSong Z, Chen Y, Lu H, Zeng Z, Wang W, Liu X, Zhang G, Du Q, Xia X, Li C\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eDiagnosis and treatment of Helicobacter pylori infection by physicians in China: A nationwide cross-sectional study\u003c/strong\u003e. \u003cem\u003eHelicobacter \u003c/em\u003e2022, \u003cstrong\u003e27\u003c/strong\u003e(3):e12889.\u003c/li\u003e\n\u003cli\u003eCrowe SE: \u003cstrong\u003eHelicobacter pylori Infection\u003c/strong\u003e. \u003cem\u003eN Engl J Med \u003c/em\u003e2019, \u003cstrong\u003e380\u003c/strong\u003e(12):1158-1165.\u003c/li\u003e\n\u003cli\u003eMalfertheiner P, Megraud F, O\u0026apos;Morain CA, Atherton J, Axon AT, Bazzoli F, Gensini GF, Gisbert JP, Graham DY, Rokkas T\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eManagement of Helicobacter pylori infection--the Maastricht IV/ Florence Consensus Report\u003c/strong\u003e. \u003cem\u003eGut \u003c/em\u003e2012, \u003cstrong\u003e61\u003c/strong\u003e(5):646-664.\u003c/li\u003e\n\u003cli\u003eMassarrat S, Haj-Sheykholeslami A, Mohamadkhani A, Zendehdel N, Aliasgari A, Rakhshani N, Stolte M, Shahidi SM: \u003cstrong\u003ePepsinogen II can be a potential surrogate marker of morphological changes in corpus before and after H. pylori eradication\u003c/strong\u003e. \u003cem\u003eBiomed Res Int \u003c/em\u003e2014, \u003cstrong\u003e2014\u003c/strong\u003e:481607.\u003c/li\u003e\n\u003cli\u003eMassarrat S, Haj-Sheykholeslami A: \u003cstrong\u003eIncreased Serum Pepsinogen II Level as a Marker of Pangastritis and Corpus-Predominant Gastritis in Gastric Cancer Prevention\u003c/strong\u003e. \u003cem\u003eArch Iran Med \u003c/em\u003e2016, \u003cstrong\u003e19\u003c/strong\u003e(2):137-140.\u003c/li\u003e\n\u003cli\u003eLeung WK, Wu MS, Kakugawa Y, Kim JJ, Yeoh KG, Goh KL, Wu KC, Wu DC, Sollano J, Kachintorn U\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eScreening for gastric cancer in Asia: current evidence and practice\u003c/strong\u003e. \u003cem\u003eLancet Oncol \u003c/em\u003e2008, \u003cstrong\u003e9\u003c/strong\u003e(3):279-287.\u003c/li\u003e\n\u003cli\u003eShan JH, Bai XJ, Han LL, Yuan Y, Sun XF: \u003cstrong\u003eChanges with aging in gastric biomarkers levels and in biochemical factors associated with Helicobacter pylori infection in asymptomatic Chinese population\u003c/strong\u003e. \u003cem\u003eWorld J Gastroenterol \u003c/em\u003e2017, \u003cstrong\u003e23\u003c/strong\u003e(32):5945-5953.\u003c/li\u003e\n\u003cli\u003eZhou JP, Liu CH, Liu BW, Wang YJ, Benghezal M, Marshall BJ, Tang H, Li H: \u003cstrong\u003eAssociation of serum pepsinogens and gastrin-17 with Helicobacter pylori infection assessed by urea breath test\u003c/strong\u003e. \u003cem\u003eFront Cell Infect Microbiol \u003c/em\u003e2022, \u003cstrong\u003e12\u003c/strong\u003e:980399.\u003c/li\u003e\n\u003cli\u003eMarchildon P, Balaban DH, Sue M, Charles C, Doobay R, Passaretti N, Peacock J, Marshall BJ, Peura DA: \u003cstrong\u003eUsefulness of serological IgG antibody determinations for confirming eradication of Helicobacter pylori infection\u003c/strong\u003e. \u003cem\u003eAm J Gastroenterol \u003c/em\u003e1999, \u003cstrong\u003e94\u003c/strong\u003e(8):2105-2108.\u003c/li\u003e\n\u003cli\u003eToyoshima O, Nishizawa T, Sakitani K, Yamakawa T, Takahashi Y, Yamamichi N, Hata K, Seto Y, Koike K, Watanabe H\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003eSerum anti-Helicobacter pylori antibody titer and its association with gastric nodularity, atrophy, and age: A cross-sectional study\u003c/strong\u003e. \u003cem\u003eWorld J Gastroenterol \u003c/em\u003e2018, \u003cstrong\u003e24\u003c/strong\u003e(35):4061-4068.\u003c/li\u003e\n\u003cli\u003eAscherman B, Oh A, Hur C: \u003cstrong\u003eInternational cost-effectiveness analysis evaluating endoscopic screening for gastric cancer for populations with low and high risk\u003c/strong\u003e. \u003cem\u003eGastric Cancer \u003c/em\u003e2021, \u003cstrong\u003e24\u003c/strong\u003e(4):878-887.\u003c/li\u003e\n\u003cli\u003eHuang HL, Leung CY, Saito E, Katanoda K, Hur C, Kong CY, Nomura S, Shibuya K: \u003cstrong\u003eEffect and cost-effectiveness of national gastric cancer screening in Japan: a microsimulation modeling study\u003c/strong\u003e. \u003cem\u003eBMC Med \u003c/em\u003e2020, \u003cstrong\u003e18\u003c/strong\u003e(1):257.\u003c/li\u003e\n\u003cli\u003eAbraham P, Wang L, Jiang Z, Gricar J, Tan H, Kelly RJ: \u003cstrong\u003eHealthcare utilization and total costs of care among patients with advanced metastatic gastric and esophageal cancer\u003c/strong\u003e. \u003cem\u003eFuture Oncol \u003c/em\u003e2021, \u003cstrong\u003e17\u003c/strong\u003e(3):291-299.\u003c/li\u003e\n\u003cli\u003eYip W, Fu H, Chen AT, Zhai T, Jian W, Xu R, Pan J, Hu M, Zhou Z, Chen Q\u003cem\u003e et al\u003c/em\u003e: \u003cstrong\u003e10 years of health-care reform in China: progress and gaps in Universal Health Coverage\u003c/strong\u003e. \u003cem\u003eLancet \u003c/em\u003e2019, \u003cstrong\u003e394\u003c/strong\u003e(10204):1192-1204.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Supplementary Document","content":"\u003cp\u003eSupplementary Document is not available with this version.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Gradient boosting machine, Helicobacter pylori antibody, Gastric cancer, Screening, Cost reduction, Large-sample","lastPublishedDoi":"10.21203/rs.3.rs-3853941/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3853941/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground and aim:\u003c/h2\u003e \u003cp\u003eThe current cancer screening model in our country involves preliminary screening and identification of individuals who require gastroscopy, in order to control screening costs. The purpose of this study is to optimize the screening process using Gradient Boosting Machines (GBM), a machine learning technique, based on a large-scale prospective gastric cancer screening dataset. The ultimate goal is to further reduce the cost of initial cancer screening.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThe study constructs a GBM machine learning model based on prospective, large-sample Taizhou City gastric cancer screening data and validates it with data from the Minimum Security Cohort Group (MLGC) in Taizhou City. Both data analysis and machine learning model construction were performed using the R programming language.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eA total of 195,640 cases were used as the training set, and 32,994 cases were used as an external validation set. A GBM was built based on the training set, yielding area under the curve (AUC) and area under the precision-recall curve (AUCPR) values of 0.99938 and 0.99823, respectively. External validation of the model yielded AUC and AUCPR values of 0.99742 and 0.99454, respectively. Through a visual analysis of the model, it was determined that the variable for \u003cem\u003eHelicobacter pylori\u003c/em\u003e IgG could be eliminated. The GBM model was then reconstructed without the \u003cem\u003eH. pylori\u003c/em\u003e IgG variable. In the training set, the new model achieved an AUC of 0.99817 and an AUCPR of 0.99462, whereas in the external validation set, it achieved an AUC of 0.99742 and an AUCPR of 0.99454.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThis study utilized a dataset of 230,000 samples to train and validate a GBM model, optimizing the initial screening process by excluding the detection of \u003cem\u003eH. pylori\u003c/em\u003e IgG antibodies while maintaining satisfactory discriminative performance. This conclusion will contribute to a reduction in the current cost of gastric cancer screening, demonstrating its economic value. Furthermore, the conclusion is derived from a large sample size, giving it clinical significance and generalizability.\u003c/p\u003e","manuscriptTitle":"A Study on the Feasibility of Optimizing Gastric Cancer Screening to Reduce Screening Costs in China Using a Gradient Boosting Machine: A prospective, large-sample, single-center study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-20 17:50:59","doi":"10.21203/rs.3.rs-3853941/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"6b3245ab-0699-45c4-9cb9-901c8082b672","owner":[],"postedDate":"February 20th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-07-01T11:32:43+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-20 17:50:59","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3853941","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3853941","identity":"rs-3853941","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00