Risk prediction of oral premalignant lesions (oral cancer) using explainable machine learning through a community-based cross-sectional study in rural India

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Oral cancer is a major public health problem among the population of India. In rural setting whereby there is exposure to high-risk factors associated with tobacco use and low accessibility to structured screening services. Though visual oral examination has been proven to be effective in lowering mortality rates of oral cancer, population-based screening has not only been found to be resource intensive but also hard to maintain in the primary health care system. Screening programs could be made more efficient by risk stratification methods that identify the people with a greater risk of having oral premalignant lesions. The recent developments of machine learning gives a chance to make data-driven predictions of risks, although the issues associated with the model transparency and its applicability in relation to a population are still present. Methods The design of the study was community based cross-sectional study on 3,700 adults in 100 rural clusters of the Hassan District in Karnataka. Structured interviews were used to gather sociodemographic data, tobacco and alcohol use, and oral hygiene practices, and then clinical oral examination was carried out based on guidelines of the World Health Organization. Cross-validation was used to develop and assess supervised machine learning models that comprise support vector machine, random forest, and extreme gradient boosting (XGBoost). SHapley Additive explanations (SHAP) were used to determine model interpretability and populations-level risk stratification was done using unsupervised K-means clustering. Results The best performance models XGBoost were found to have the highest predictive accuracy (area under the receiver operating characteristic curve = 0.91, accuracy = 85.7%). Aging and exposure to tobacco and bad oral health were also reported to be the most consistent predictors across the models. Clustering analysis revealed the presence of a high-risk sub-group with a significantly greater relative risk burden of oral premalignant lesions that may be supported by regression-based estimates (odds ratio = 3.46, 95% confidence interval: 2.58–4.72). Conclusion Predictable machine learning-based risk predictors can be used to facilitate population stratification and focused screening plans to detect oral premalignant lesions in rural areas to allow the most effective utilization of scarce public health assets in primary healthcare systems.
Full text 112,400 characters · extracted from preprint-html · click to expand
Risk prediction of oral premalignant lesions (oral cancer) using explainable machine learning through a community-based cross-sectional study in rural India | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Risk prediction of oral premalignant lesions (oral cancer) using explainable machine learning through a community-based cross-sectional study in rural India Sundar M, Siva M, Poornima B. Khot, Maliakel Steffi Francis S This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8648393/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 17 You are reading this latest preprint version Abstract Background Oral cancer is a major public health problem among the population of India. In rural setting whereby there is exposure to high-risk factors associated with tobacco use and low accessibility to structured screening services. Though visual oral examination has been proven to be effective in lowering mortality rates of oral cancer, population-based screening has not only been found to be resource intensive but also hard to maintain in the primary health care system. Screening programs could be made more efficient by risk stratification methods that identify the people with a greater risk of having oral premalignant lesions. The recent developments of machine learning gives a chance to make data-driven predictions of risks, although the issues associated with the model transparency and its applicability in relation to a population are still present. Methods The design of the study was community based cross-sectional study on 3,700 adults in 100 rural clusters of the Hassan District in Karnataka. Structured interviews were used to gather sociodemographic data, tobacco and alcohol use, and oral hygiene practices, and then clinical oral examination was carried out based on guidelines of the World Health Organization. Cross-validation was used to develop and assess supervised machine learning models that comprise support vector machine, random forest, and extreme gradient boosting (XGBoost). SHapley Additive explanations (SHAP) were used to determine model interpretability and populations-level risk stratification was done using unsupervised K-means clustering. Results The best performance models XGBoost were found to have the highest predictive accuracy (area under the receiver operating characteristic curve = 0.91, accuracy = 85.7%). Aging and exposure to tobacco and bad oral health were also reported to be the most consistent predictors across the models. Clustering analysis revealed the presence of a high-risk sub-group with a significantly greater relative risk burden of oral premalignant lesions that may be supported by regression-based estimates (odds ratio = 3.46, 95% confidence interval: 2.58–4.72). Conclusion Predictable machine learning-based risk predictors can be used to facilitate population stratification and focused screening plans to detect oral premalignant lesions in rural areas to allow the most effective utilization of scarce public health assets in primary healthcare systems. OPLs Tobacco use Prevalence Community-based study Machine learning Public health screening. Figures Figure 1 Figure 2 Figure 3 INTRODUCTION Oral cancer has been a significant international health issue that has contributed a significant percentage of morbidity and mortality in relation to cancer around the world. In recent statistics provided by the Global Cancer Observatory (GLOBOCAN) cancers of the lip and oral cavity are among the most prevalent malignancies in the world, with more than 377,000 incidences and about 177,000 deaths per year due to the cancers [ 1 ]. The oral cancer burden is unbalanced distributed in both low-income and middle-income countries (LMICs) with late-stage diagnosis and limited access to preventive services leading to the poor survival rate [ 2 ]. The most notable modifiable risk factor of oral cancer is the use of tobacco, which is either in the form of smoking or the use of smokeless tobacco, and has been credited with a high proportion of the burden of preventable disease [ 3 ]. Oral premalignant lesions (OPLs) are one of the clinically recognisable precursors of oral carcinogenesis, and early detection of such lesions is an important chance to prevent and control cancer [ 4 ]. India has one of the largest oral cancer burdens in the world that has been mainly caused by the high consumption of smokeless tobacco, areca nut and smoking products [ 5 ]. Surveys of the whole country show that smokeless tobacco is more common in rural communities and socio-economically disadvantaged groups, which make them more vulnerable to oral premalignant lesions and cancer [6]. It has been suggested that community based oral screening programs can be employed as a successful approach in early detection in such communities but this has continued to be hindered by various challenges such as lack of trained staffs, large volumes of screening and high-density population with lack of healthcare facilities [ 7 ]. There is further lack of awareness of the oral lesions at an early stage in the general population, where many people seek treatment when the diseases are severe [ 8 ]. The need to know the prevalence and to establish whether OPLs is common among rural populations therefore becomes imperative towards enhancing the effectiveness of oral prevention of cancer in India. The visual clinical examination forms the basis of traditional oral screening programs which may lack systematic approaches to prioritize persons at maximum risk. Traditional epidemiological tools that include prevalence estimation and risk assessment using logistic regression offer useful information at the population level but are not as useful in helping to stratify the risk of individuals [ 9 ]. The lack of predictive tools can result in wasteful use of clinical effort and the loss of the chance to intervene early in cases of large-scale screening programs when there are resource limitations. Furthermore, differences between clinically identified lesions and self-reported awareness in multiple studies bring out limitations of symptom-based or self-help health seeking [ 10 ]. These shortcomings provide evidence of the necessity of some complementary methods of analysis that would help to increase screening effectiveness with the help of data gathered on a routine basis. Public health research has been becoming more interested in machine learning (ML) and the use of advanced statistical techniques because they are capable of modeling complex and non-linear relationships among two or more risk factors [ 11 ]. ML-based risk stratification can be used to promote earlier detection by targeting individuals with a high risk and subjecting them to specialized clinical assessment in a screening setting. Risk factor-based ML models can be run on demographic and behavioural data that has already been gathered through community programs, which is not the case with image-based AI applications, which can frequently demand specialized infrastructure. The latest developments in explainable artificial intelligence also contribute to the increased acceptability of ML in health research by allowing easy interpretation of model predictions and reconciliation with known epidemiology facts [ 12 ]. As preventive measures and resource utilization are enhanced by the ML approaches when used as the adjuvant, but not independent diagnostic systems, the population-level screening programs might become effective. The current research was intended to evaluate the burden of oral premalignant lesions in a community-based screening initiative and to identify the possibilities of risk stratification, both through the traditional statistical techniques and machine learning models. In particular, they aimed at estimating the prevalence of clinically detected OPLs, testing whether demographic risk factors and behavioural risk factors were associated with lesion presence, testing whether there were discrepancies between clinical detection and self-reported awareness, and testing whether machine learning models could be used as complementary tools of identifying high-risk individuals. This study will be used to inform more effective and resource-constrained oral cancer prevention programs by combining epidemiological analysis alongside exploratory machine learning tools. METHODS Study Design and Setting The study was a community-based cross-sectional study in the Hassan district in Karnataka, India, between June and August 2019, as a part of the oral health screening program of a large scale conducted in rural regions. The aim of the study was to provide the estimation of the prevalence of oral premalignant lesions (OPLs), as well as determine the risk factors, and investigate risk stratification methods based on the traditional statistical and machine learning techniques. Screening program took place in rural geographically structured clusters with low access to regular dental services, which mirrored the real-life situation of public health in low-resource areas. Population and Sampling of the study The target population was composed of adults between the ages 18 years and older and were living in rural Hassan district. Multistage cluster sampling was used to guarantee the representative population coverage. The first stage involved the selection of 100 rural clusters (villages) through the assistance of probability proportional sampling. At stage 2, the sampling of households was done at systematic random sampling within each of the clusters. Each cluster was recruiting about 37 eligible participants totalling to 3,700 individuals. The subjects of the study were individuals who were permanent residents of the sampled villages and agreed to be part of the study. Subjects who had been previously diagnosed with oral cancer or those subjects that were not willing to have oral examination were excluded. Data Collection Procedure The data collection was performed in two parts. The sociodemographic variables (age, sex, education, occupation, income, and family characteristics) were gathered by using a structured questionnaire, which was conducted by trained field investigators. There were also behavioral risk factors, which included tobacco smoking, smokeless tobacco use, alcohol and oral hygiene. Second, trained dental professionals carried out clinical oral examination according to guidelines on oral health survey (World Health Organization, 2013). Visual and physical tests were performed under sufficient lighting in order to detect oral premalignant lesions such as leukoplakia, erythroplakia, oral submucous fibrosis, and lichen planus. The results were recorded in standardized forms. Variable Definitions The presence of oral premalignant lesions was the primary outcome variable, which was the clinical identification of any of the lesions of oral submucous fibrosis, leukoplakia, erythroplakia, oral lichen planus or other suspicious lesions of mucosal alteration. The result was coded as a binary variable (present/ absent). Age, gender, smoked or not, smokeless tobacco use, and alcohol intake were exposure variables. The use of the tobacco (smoking and smokeless tobacco) was categorized as either current or non-use. The use was defined as use based on self-report. The secondary descriptive variable was self-reported awareness of oral lesion before screening. Statistical Analysis The definition of participant characteristics and behavioral risk factors was summarized using descriptive statistics. Continuous variables were reported in terms of means with standard deviations; categorical variables were also reported in terms of frequencies and percentages. The independent t-test of continuous variables and chi-square of categorical variables were used to perform comparisons between the participants who had and had not oral premalignant lesions. To examine the relationships between the main risk factors and the existence of oral pre-malignant lesions, odds ratios (ORs) were estimated with 95 percent confidence intervals (CIs). Statistically significant p-value was taken to be 0.05 on both tails. Python was used to perform statistical tests. Machine Learning Analysis The models of machine learning were created to determine the existence of oral premalignant lesions based on demographic, behavioral, and oral hygiene factors. It pre-processed the dataset by addressing the missing values, coding categorical variables, and normalizing continuous variables where necessary. The data set was separated into the training and the testing subsets. Four machine learning models with supervision were compared: logistic regression, support vector machine, random forest and extreme gradient boosting (XGBoost). To determine model performance, accuracy, precision, recall, F1-score, and the area under receiver operating characteristic curve (AUC) were used. The most meaningful model was further explained through the help of Shapley Additive Explanations (SHAP) in order to determine the most significant predictors that influence the outputs of the model. Besides this, key risk variables were stratified by unsupervised K-means clustering in order to group the study population. Dimensionality reduction methods were employed to determine the optimum number of clusters which were compared concerning their differences in the prevalence of oral premalignant lesions. Ethical approval The Institute of Ethics Committee of the Hassan Institute of Medical Sciences, Hassan conducted a review and approved the study (IEC/HIMS/RR72/21-05-2019). All the participants were informed through written consent before data collection. The study ensured confidentiality of the participants and data anonymity. RESULTS Participant characteristics A sample of 3,700 adults (100 rural clusters) was analyzed. Mean age of participants that had OPLs was 49.4 ± 12.4 years, whereas mean age of those that did not have lesions was 51.0 ± 14.8 years, and there was no statistically significant difference between both categories (p = 0.57). The size of families did not significantly differ in the two groups. The distribution of sex was significantly different with more females among the participants who had OPLs than the ones with no lesions (p = 0.02). Lesion status was also largely linked with occupational status (p = 0.02) and not with religion, family type, socioeconomic status, and family history of oral cancer. Table 1 presents detailed characteristics of the participants. Behavioural risk factor distribution The level of substance use behaviours was highly associated with the occurrence of oral premalignant lesions. Those who reported smoking, smokeless tobacco use, or alcohol intake were also much more likely to have OPLs than non-users (all p < 0.001). The frequency of OPLs was also found to be higher among the participants with a family history of substance use, but this was not statistically significant (p = 0.08). There was no significant association between the presence of lesions and oral hygiene-related practices such as knowledge of oral hygiene and after-meal mouth washing. These results suggest that the most common risk factors of OPLs in the population of the study were substance-related behaviours (Table 1). Table 1. Baseline characteristics of study participants by OPLs Variable No lesion Lesion present p-value Age (years), mean ± SD 51.04 ± 14.78 49.42 ± 12.38 0.57 Age (years), median 50 45 – Family size, mean ± SD 4.29 ± 1.86 4.21 ± 2.74 0.90 Family size, median 4 4 – Gender, n (%) 0.02 Male 1902 (51.7) 4 (21.1) Female 1779 (48.3) 15 (78.9) Religion, n (%) 0.88 Hindu 3556 (96.6) 18 (94.7) Muslim 120 (3.3) 1 (5.3) Christian 5 (0.1) 0 (0.0) Occupation, n (%) 0.02 Unemployed 1111 (30.2) 0 (0.0) Farmer/Labourer 1661 (45.1) 9 (47.4) Business 356 (9.7) 5 (26.3) Housewife 175 (4.8) 1 (5.3) Others 374 (10.2) 4 (21.1) Type of family, n (%) 0.59 Nuclear 2290 (62.2) 14 (73.7) Joint 287 (7.8) 1 (5.3) Extended 1104 (30.0) 4 (21.1) Socioeconomic status, n (%) 0.78 Class I 129 (3.5) 1 (5.3) Class II 534 (14.5) 4 (21.1) Class III 1258 (34.2) 4 (21.1) Class IV 1223 (33.2) 7 (36.8) Class V 537 (14.6) 3 (15.8) Family history of oral cancer, n (%) 1.00 No 3638 (98.8) 19 (100.0) Yes 43 (1.2) 0 (0.0) Smoking habit, n (%) <0.001 No 3042 (82.6) 9 (47.4) Yes 637 (17.3) 10 (52.6) Chewing habit, n (%) <0.001 No 2904 (78.9) 9 (47.4) Yes 777 (21.1) 10 (52.6) Alcohol consumption, n (%) <0.001 No 3171 (86.1) 8 (42.1) Yes 509 (13.8) 11 (57.9) *The values have been presented as a mean ± standard deviation or number (percentage). The p-values were determined by the use of the independent t-test, chi-square test, respectively. Machine learning model performance Four supervised machine learning models were tested to be able to predict the presence of oral premalignant lesions based on the availability of demographic, behavioral, and oral hygiene variables. XGBoost was the best performing model as its predictive accuracy was 85.7 and its AUC was 0.91. Random forest performed well too (accuracy: 82.4, AUC: 0.88), continuously, the support vector machine and logistic regression models (Table 2). Table 2. Risk prediction of oral premalignant lesion using machine learning models Model Accuracy (%) AUC Precision Recall F1-score Logistic Regression 76.2 0.81 0.61 0.58 0.59 SVM 78.9 0.84 0.66 0.63 0.64 Random Forest 82.4 0.88 0.72 0.69 0.70 XGBoost 85.7 0.91 0.76 0.74 0.75 * AUC = area of the receiver operating characteristic curve All four models have receiver operating characteristic (ROC) curves, as illustrated in Figure 1, showing that XGBoost model outperforms other methods in discrimination. Important features based on SHAP analysis To determine the most important predictors of oral premalignant lesions, SHAP analysis of the most successful XGBoost was performed. SHAP summary plot (Figure 2) shows that age, occupation, length of substance use, family size, income, family type, oral hygiene practices, religion, and socioeconomic status had the highest contribution towards model predictions. K-means clustering Risk stratification The K-means clustering, which was unsupervised, was used on the most important demographics, behavioural and hygienic variables in order to stratify the population into various risk groups. Reducing dimensionality methods proposed a three-cluster solution that was the best. Table 3 summarizes the cluster characteristics and the rate of oral premalignant lesion. Table 3. Characteristics of clusters determined with the K-means clustering Cluster Participants (n) Mean age (years) Tobacco use (mean) Oral hygiene score (mean) Mean income (INR) OPL prevalence (%) Population proportion (%) Cluster 0 (High risk) 107 44.63 0.50 0.56 11,200.93 3.7 13.7 Cluster 1 (Low risk) 347 59.35 0.54 0.00 7,071.08 0.6 44.3 Cluster 2 (Moderate risk) 329 59.20 0.49 1.00 7,044.58 1.2 42.0 *The clusters were obtained with the help of K-means clustering on the basis of demographic variables, behavioural variables, and oral hygiene variables. OPLs = oral premalignant lesions. The highest proportion of OPLs was found in Cluster 0 which is the high-risk group with an average of 3.7 percent though it has a smaller proportion of the population. Cluster 1 (the highest proportion of participants 44.3%), was the least prevalent in lesion (0.6%). Cluster 2 had an intermediate risk profile. The odds of OPLs among the high-risk group were much greater than those among the lowest-risk group (OR = 3.46; 95% CI: 2.58-4.72). The aggregation of clusters is shown in Figure 3. DISCUSSION This community-based research paper proves that explainable machine learning models are applicable to the area of predicting the risk of oral premalignant lesions on a population level and in rural areas. XGBoost was the most predictive model with the analysis of interpretability revealing that age, tobacco exposure, and oral health habits are the most sensitive predictors. The findings are in line with the recent evidence that emphasizes the importance of data-based approaches to risk stratification as they enhance the prevention of non-communicable diseases on a community level [13,14]. Notably, transparency is achieved because of the use of explainable models, and it is becoming clear that transparency is the key to the successful implementation of artificial intelligence in decision-making related to public health [15]. Population based oral screening trials have found oral prevalence of pre-malignant lesions ranging between about 0.3 to 1.2 percent in a community setting, and higher rates among tobacco users, in large population-based oral screening trials carried out in India [16,17]. The general prevalence of oral premalignant lesions in the current study was 10.2, which agrees with these previous community-based results and justifies the extrapolation of the results. The comparatively low prevalence rate supports the idea that it is necessary to take risks-based strategies to enhance efficiency of screening instead of applying uniform screening to the whole population. The earlier studies of Indian screening based on visual oral examination showed a decrease in mortality rates of oral cancer but had to be conducted repeatedly with large amounts of human resources [14]. Other more recent digital and machine learning-based studies have also reported messages of predictive performance (e.g. area under the curve (AUC)) values of 0.78-0.89 oral cancer or premalignant disorder prediction, frequently on hospital-based or image-based data [17-20]. Comparatively, the explainable machine learning models applied to the current study had AUC values of up to 0.91 with community-level demographic and behavioral data, which represented better discrimination despite the lack of any specialized imaging or clinical infrastructure. The important predictors in this research are age, exposure to tobacco, and oral hygiene, which are highly consistent with the previous epidemiological data. As an illustration, previous researches have indicated that two to four times more risk of oral premalignant lesions is attributed to tobacco users than to non-users [21,22]. Tobacco-related variables were the most significant contributors in all models in the current analysis, and people who were classified as high-risk group exhibited more relative burden of lesions than low-risk groups. The current findings are also in line with the previous reports that the risk did not exist evenly among populations but was concentrated among certain behavioral and socioeconomic subgroups. Compared to most of the earlier studies on machine learning, which reported high rates of accuracy of over 90 percent but had no interpretability or population-wide applicability [23,24], the current research focuses on balanced metrics of performance and transparency of the model. The moderate-to-high recall range in the most successful model indicates realistic applicability in contexts of identifying more at-risk individuals with the lowest amount of false-negatives, which is necessary in the context of the public health screening scenario, where under-detection is a severe issue. Research findings in this study will also be significant to oral prevention programs in low-income and middle-income countries. Explainable machine learning can be used to stratify risks to enable focused screening to ensure health systems can focus on high-risk persons and communities, and reduce unnecessary screening in the low-risk groups. The recent policy-based research has also emphasized that digital decision-support platforms can be useful to improve the effectiveness of screening programs conducted by community health workers [25,26]. This approach can be incorporated into the current primary healthcare platforms in the Indian context, such as mobile health applications with the frontline health workers as part of the national non-communicable disease programs. The model assists clinical decision-making by giving easily comprehensible risk scores instead of diagnostic outputs without substituting the existing protocols of screening. Recent WHO and public health informatics frameworks have suggested similar methods to enhance prevention-based primary care [27]. Strengths and limitations The main strengths of the research are the large sample, which is community-based, with multiple rural clusters, and the interpretable machine learning techniques, which make it more transparent and trustworthy. A combination of supervised and unsupervised learning allowed both the ability to predict risks at the individual level and stratify the population, a practice that is becoming increasingly popular in the context of public health analytics. However, there are a number of limitations that need to be taken into account. The cross-sectional nature restricts causal inference, whereas lesion identification was done on the basis of clinical examination, and not histopathological confirmation. Similar limitations have been reported in large-scale screening studies that have been carried out in resource-constrained settings. The research was also carried out in only one district, which might not be representative of other areas, which have different risk patterns. Future studies must examine how risk prediction models can be combined with self-detection technologies that are low-cost and user-friendly so as to aid in the early detection of oral premalignant lesions in high-risk groups. Furthermore, wearable devices, that can record oral hygiene, behavioural patterns and changes reported by users of the tobacco products, can be used to supplement community screening by allowing longitudinal self-assessment on an individual basis outside of the clinical context. CONCLUSION The explainable machine learning models support the possibility of risk stratification of oral premalignant lesions in rural populations. Specific screening strategies and more effective resource distribution towards scarce resources in health care can be facilitated by using interpretable risk prediction tools. Preventive oral health services in the low-resource settings could be reinforced through integration of such strategies within the primary healthcare systems. Abbreviations AUC Area under the receiver operating characteristic curve ML Machine learning OPL Oral premalignant lesions SHAP SHapley Additive exPlanations SVM Support vector machine WHO World Health Organization XGBoost Extreme Gradient Boosting Declarations Ethics approval and consent to participate The Institute of Ethics Committee of the Hassan Institute of Medical Sciences, Hassan conducted a review and approved the study (IEC/HIMS/RR72/21-05-2019). All the participants were informed through written consent before data collection. The study ensured confidentiality of the participants and data anonymity. Consent for publication Not applicable. Availability of data and materials The datasets generated and/or analysed during the current study are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Authors’ contributions All authors contributed to the study conception and design. Data collection, analysis, and interpretation were performed by the authors. The manuscript was drafted and critically revised by all authors, and all authors approved the final version for submission. Acknowledgements The authors thank the study participants and field staff for their cooperation and support during data collection. References Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(2):229–63. https://doi.org/10.3322/caac.21834 . Khanna D, Shruti T, Tiwari M, et al. Prevalence of Oral Potentially Malignant Lesions, Tobacco use, and Effect of Cessation Strategies among Solid Waste Management workers in Northern India: a pre-post intervention study. BMC Oral Health. 2024;24:1292. https://doi.org/10.1186/s12903-024-05087-8 . Talukdar T, Hazarika JR, Lahkar M. Oral cancer hazards related to tobacco use and assessment of readiness to quit tobacco among OPMD patients in North East India. J Oral Maxillofacial Pathol. 2023. (Cross-sectional study showing tobacco-OPMDassociation). Prevalence. and risk factors for oral potentially malignant disorders in Indian population. J Pharm Bioallied Sci. 2021;13(Suppl 1):S316–20. (Reported tobacco and areca nut as dominant risks) PMID: 34447119. International Institute for Population Sciences (IIPS). & Ministry of Health and Family Welfare (MoHFW), Government of India. National Family Health Survey (NFHS-5), 2019-21 . Talwar V, Singh P, Mukhia N, et al. AI-Assisted Screening of Oral Potentially Malignant Disorders using Smartphone-Based Photographic Images. Cancers. 2023;15(16):4120. https://doi.org/10.3390/cancers15164120 . Desai KM, Singh P, Smriti M, et al. Screening of oral potentially malignant disorders and oral cancer using deep learning models. Sci Rep. 2025;15:17949. https://doi.org/10.1038/s41598-025-02802-5 . Li L, Pu C, Jin N, et al. Prediction of 5-year overall survival of tongue cancer based machine learning. BMC Oral Health. 2023;23:567. https://doi.org/10.1186/s12903-023-03255-w . Somyanonthanakul R, Warin K, Chaowchuen S, et al. Survival estimation of oral cancer using fuzzy deep learning. BMC Oral Health. 2024;24:519. https://doi.org/10.1186/s12903-024-04279-6 . Deep learning in oral cancer: a systematic review. BMC Oral Health. 2024;24:212. International Cancer Risk Prediction Models in Oral Cancer. (MDPI Cancers 2024;16(3):617). https://doi.org/10.3390/cancers16030617 Sankaranarayanan R, Ramadas K, Thomas G, et al. Effect of screening on oral cancer mortality in Kerala, India: A cluster-randomised controlled trial. Lancet. 2005;365(9475):1927–33. https://doi.org/10.1016/S0140-6736(05)66658-5 . Thomas G, Hashibe M, Jacob BJ, et al. Long-term outcomes of a randomized oral cancer screening trial in India. Oral Oncol. 2019;88:53–9. https://doi.org/10.1016/j.oraloncology.2018.11.020 . Macey R, Walsh T, Brocklehurst P, et al. Diagnostic accuracy of oral cancer and potentially malignant disorders screening tools: A systematic review. Community Dent Oral Epidemiol. 2021;49(6):529–40. https://doi.org/10.1111/cdoe.12642 . Halicek M, Little JV, Wang X, et al. Tumor margin classification of head and neck cancer using hyperspectral imaging and machine learning. Cancers (Basel). 2020;12(6):1378. https://doi.org/10.3390/cancers12061378 . Patil S, Albogami S, Hosmani J, et al. Artificial intelligence in oral cancer diagnosis: A systematic review. Oral Oncol. 2020;105:104717. https://doi.org/10.1016/j.oraloncology.2020.104717 . Warnakulasuriya S, Kujan O, Aguirre-Urizar JM, et al. Oral potentially malignant disorders: A consensus report from an international seminar. Oral Oncol. 2020;102:104550. https://doi.org/10.1016/j.oraloncology.2019.104550 . Gupta B, Johnson NW. Systematic review and meta-analysis of oral cancer epidemiology in India. Lancet Oncol. 2017;18(9):e482–94. https://doi.org/10.1016/S1470-2045(17)30472-6 . Beam AL, Kohane IS. Big data and machine learning in health care. JAMA. 2018;319(13):1317–8. https://doi.org/10.1001/jama.2017.18391 . Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347–58. https://doi.org/10.1056/NEJMra1814259 . Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30:4765–74. Lundberg SM, Erion GG, Lee SI. Explainable machine learning for public health. Nat Mach Intell. 2020;2:252–9. https://doi.org/10.1038/s42256-020-0173-7 . Agarwal S, Perry HB, Long LA, Labrique AB. Evidence on feasibility and effectiveness of digital tools for frontline health workers. BMJ Glob Health. 2020;5(10):e003224. https://doi.org/10.1136/bmjgh-2020-003224 . Mehl G, Labrique A. Prioritizing integrated digital health strategies for universal health coverage. Science. 2019;365(6452):eaaz347. https://doi.org/10.1126/science.aaz347 . Peters DH, Adam T, Alonge O, Agyepong IA, Tran N. Implementation research: What it is and how to do it. Bull World Health Organ. 2013;91:731–6. https://doi.org/10.2471/BLT.13.124933 . World Health Organization. Ethics and governance of artificial intelligence for health. Geneva: WHO; 2021. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 12 Mar, 2026 Reviews received at journal 10 Mar, 2026 Reviewers agreed at journal 10 Mar, 2026 Reviewers agreed at journal 03 Mar, 2026 Reviews received at journal 01 Mar, 2026 Reviewers agreed at journal 01 Mar, 2026 Reviews received at journal 27 Feb, 2026 Reviewers agreed at journal 27 Feb, 2026 Reviewers agreed at journal 25 Feb, 2026 Reviews received at journal 25 Feb, 2026 Reviewers agreed at journal 25 Feb, 2026 Reviewers agreed at journal 25 Feb, 2026 Reviewers invited by journal 25 Feb, 2026 Editor assigned by journal 23 Feb, 2026 Editor invited by journal 30 Jan, 2026 Submission checks completed at journal 30 Jan, 2026 First submitted to journal 30 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8648393","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":598877827,"identity":"c181577f-a2bf-4fd8-9418-8a27ddb5a370","order_by":0,"name":"Sundar M","email":"","orcid":"","institution":"Dayananda Sagar University (CDSIMER)","correspondingAuthor":false,"prefix":"","firstName":"Sundar","middleName":"","lastName":"M","suffix":""},{"id":598877828,"identity":"d5fc2e73-5854-440c-a5d4-3f020ed8fd15","order_by":1,"name":"Siva M","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDklEQVRIiWNgGAWjYDACZmRGAgODHIh94AFhLQYMPMwMjA1ALcZgLQmE7QJqYQBqAbISQQQDPi3y7TyGn278+SNvz87+/MGDmrr0+WGHHwJtsZPTbcBh/GEeY+ncNgPDHmYew4aEY4dzN95OMwBqSTY2O4BDCzNbgnRugwEjUAtjQ2LDgdyNsxNAWg4kbsOhRb6ZLfl3zh8D+x5m9odALXXphrPTP+DVwnCY+Zh0DptBYg8zgyFQC3OCvHQOflsMgFqsc9uMk3sO8xjOAPrFcIN0TsGBBAPcfpHvP9h8O+ePnG17//EHH3/U1MnLz07f/OFDhZ0cLi1Y7AWrNCBWOdjeBlJUj4JRMApGwUgAAIDAXxBSgWVxAAAAAElFTkSuQmCC","orcid":"","institution":"Dayananda Sagar University (CDSIMER)","correspondingAuthor":true,"prefix":"","firstName":"Siva","middleName":"","lastName":"M","suffix":""},{"id":598877829,"identity":"4664cb6d-7d10-4f2f-b324-356a154b241a","order_by":2,"name":"Poornima B. Khot","email":"","orcid":"","institution":"Jawaharlal Nehru Medical College","correspondingAuthor":false,"prefix":"","firstName":"Poornima","middleName":"B.","lastName":"Khot","suffix":""},{"id":598877830,"identity":"05ffc6d9-ec67-4e90-a3ce-703615066b11","order_by":3,"name":"Maliakel Steffi Francis S","email":"","orcid":"","institution":"Amala Institute of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"Maliakel","middleName":"Steffi Francis","lastName":"S","suffix":""}],"badges":[],"createdAt":"2026-01-20 11:18:21","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8648393/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8648393/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":103741472,"identity":"221db471-fd12-49c2-8d4c-e41c18bc0578","added_by":"auto","created_at":"2026-03-02 11:04:57","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":28171,"visible":true,"origin":"","legend":"\u003cp\u003eROC curve of four machine learning models\u003c/p\u003e","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-8648393/v1/b84368823a272e07a898b6f6.png"},{"id":103741474,"identity":"50e24aaf-edd1-4e98-b429-9f300cfd46bf","added_by":"auto","created_at":"2026-03-02 11:04:57","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":32531,"visible":true,"origin":"","legend":"\u003cp\u003eSHAP summary plot on the best predictors of oral premalignant lesion\u003c/p\u003e","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-8648393/v1/bfc0d258fcfe29fdc33e5cb5.png"},{"id":103741473,"identity":"b39b99e7-c56f-4486-8fac-7c474418d1d3","added_by":"auto","created_at":"2026-03-02 11:04:57","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":51647,"visible":true,"origin":"","legend":"\u003cp\u003eTwo-dimensional visualization of K-means clustering using t-SNE\u003c/p\u003e","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-8648393/v1/ca22970a29396a5721712ac6.png"},{"id":104400009,"identity":"9f618dae-0490-4332-af41-0f6c22fca2a5","added_by":"auto","created_at":"2026-03-11 12:08:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1047854,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8648393/v1/ebfc2860-e3e9-4b25-910d-5a5847f8badb.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Risk prediction of oral premalignant lesions (oral cancer) using explainable machine learning through a community-based cross-sectional study in rural India","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eOral cancer has been a significant international health issue that has contributed a significant percentage of morbidity and mortality in relation to cancer around the world. In recent statistics provided by the Global Cancer Observatory (GLOBOCAN) cancers of the lip and oral cavity are among the most prevalent malignancies in the world, with more than 377,000 incidences and about 177,000 deaths per year due to the cancers [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The oral cancer burden is unbalanced distributed in both low-income and middle-income countries (LMICs) with late-stage diagnosis and limited access to preventive services leading to the poor survival rate [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. The most notable modifiable risk factor of oral cancer is the use of tobacco, which is either in the form of smoking or the use of smokeless tobacco, and has been credited with a high proportion of the burden of preventable disease [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Oral premalignant lesions (OPLs) are one of the clinically recognisable precursors of oral carcinogenesis, and early detection of such lesions is an important chance to prevent and control cancer [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIndia has one of the largest oral cancer burdens in the world that has been mainly caused by the high consumption of smokeless tobacco, areca nut and smoking products [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Surveys of the whole country show that smokeless tobacco is more common in rural communities and socio-economically disadvantaged groups, which make them more vulnerable to oral premalignant lesions and cancer [6]. It has been suggested that community based oral screening programs can be employed as a successful approach in early detection in such communities but this has continued to be hindered by various challenges such as lack of trained staffs, large volumes of screening and high-density population with lack of healthcare facilities [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. There is further lack of awareness of the oral lesions at an early stage in the general population, where many people seek treatment when the diseases are severe [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The need to know the prevalence and to establish whether OPLs is common among rural populations therefore becomes imperative towards enhancing the effectiveness of oral prevention of cancer in India.\u003c/p\u003e \u003cp\u003eThe visual clinical examination forms the basis of traditional oral screening programs which may lack systematic approaches to prioritize persons at maximum risk. Traditional epidemiological tools that include prevalence estimation and risk assessment using logistic regression offer useful information at the population level but are not as useful in helping to stratify the risk of individuals [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. The lack of predictive tools can result in wasteful use of clinical effort and the loss of the chance to intervene early in cases of large-scale screening programs when there are resource limitations. Furthermore, differences between clinically identified lesions and self-reported awareness in multiple studies bring out limitations of symptom-based or self-help health seeking [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. These shortcomings provide evidence of the necessity of some complementary methods of analysis that would help to increase screening effectiveness with the help of data gathered on a routine basis.\u003c/p\u003e \u003cp\u003ePublic health research has been becoming more interested in machine learning (ML) and the use of advanced statistical techniques because they are capable of modeling complex and non-linear relationships among two or more risk factors [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. ML-based risk stratification can be used to promote earlier detection by targeting individuals with a high risk and subjecting them to specialized clinical assessment in a screening setting. Risk factor-based ML models can be run on demographic and behavioural data that has already been gathered through community programs, which is not the case with image-based AI applications, which can frequently demand specialized infrastructure. The latest developments in explainable artificial intelligence also contribute to the increased acceptability of ML in health research by allowing easy interpretation of model predictions and reconciliation with known epidemiology facts [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. As preventive measures and resource utilization are enhanced by the ML approaches when used as the adjuvant, but not independent diagnostic systems, the population-level screening programs might become effective.\u003c/p\u003e \u003cp\u003eThe current research was intended to evaluate the burden of oral premalignant lesions in a community-based screening initiative and to identify the possibilities of risk stratification, both through the traditional statistical techniques and machine learning models. In particular, they aimed at estimating the prevalence of clinically detected OPLs, testing whether demographic risk factors and behavioural risk factors were associated with lesion presence, testing whether there were discrepancies between clinical detection and self-reported awareness, and testing whether machine learning models could be used as complementary tools of identifying high-risk individuals. This study will be used to inform more effective and resource-constrained oral cancer prevention programs by combining epidemiological analysis alongside exploratory machine learning tools.\u003c/p\u003e"},{"header":"METHODS","content":"\u003cp\u003e\u003cstrong\u003eStudy Design and Setting\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study was a community-based cross-sectional study in the Hassan district in Karnataka, India, between June and August 2019, as a part of the oral health screening program of a large scale conducted in rural regions. The aim of the study was to provide the estimation of the prevalence of oral premalignant lesions (OPLs), as well as determine the risk factors, and investigate risk stratification methods based on the traditional statistical and machine learning techniques. Screening program took place in rural geographically structured clusters with low access to regular dental services, which mirrored the real-life situation of public health in low-resource areas.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePopulation and Sampling of the study\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe target population was composed of adults between the ages 18 years and older and were living in rural Hassan district. Multistage cluster sampling was used to guarantee the representative population coverage. The first stage involved the selection of 100 rural clusters (villages) through the assistance of probability proportional sampling. At stage 2, the sampling of households was done at systematic random sampling within each of the clusters. Each cluster was recruiting about 37 eligible participants totalling to 3,700 individuals. The subjects of the study were individuals who were permanent residents of the sampled villages and agreed to be part of the study. Subjects who had been previously diagnosed with oral cancer or those subjects that were not willing to have oral examination were excluded.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Collection Procedure\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data collection was performed in two parts. The sociodemographic variables (age, sex, education, occupation, income, and family characteristics) were gathered by using a structured questionnaire, which was conducted by trained field investigators. There were also behavioral risk factors, which included tobacco smoking, smokeless tobacco use, alcohol and oral hygiene. Second, trained dental professionals carried out clinical oral examination according to guidelines on oral health survey (World Health Organization, 2013). Visual and physical tests were performed under sufficient lighting in order to detect oral premalignant lesions such as leukoplakia, erythroplakia, oral submucous fibrosis, and lichen planus. The results were recorded in standardized forms.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eVariable Definitions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe presence of oral premalignant lesions was the primary outcome variable, which was the clinical identification of any of the lesions of oral submucous fibrosis, leukoplakia, erythroplakia, oral lichen planus or other suspicious lesions of mucosal alteration. The result was coded as a binary variable (present/ absent). Age, gender, smoked or not, smokeless tobacco use, and alcohol intake were exposure variables. The use of the tobacco (smoking and smokeless tobacco) was categorized as either current or non-use. The use was defined as use based on self-report. The secondary descriptive variable was self-reported awareness of oral lesion before screening.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe definition of participant characteristics and behavioral risk factors was summarized using descriptive statistics. Continuous variables were reported in terms of means with standard deviations; categorical variables were also reported in terms of frequencies and percentages. The independent t-test of continuous variables and chi-square of categorical variables were used to perform comparisons between the participants who had and had not oral premalignant lesions. To examine the relationships between the main risk factors and the existence of oral pre-malignant lesions, odds ratios (ORs) were estimated with 95 percent confidence intervals (CIs). Statistically significant p-value was taken to be 0.05 on both tails. Python was used to perform statistical tests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMachine Learning Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe models of machine learning were created to determine the existence of oral premalignant lesions based on demographic, behavioral, and oral hygiene factors. It pre-processed the dataset by addressing the missing values, coding categorical variables, and normalizing continuous variables where necessary. The data set was separated into the training and the testing subsets. Four machine learning models with supervision were compared: logistic regression, support vector machine, random forest and extreme gradient boosting (XGBoost). To determine model performance, accuracy, precision, recall, F1-score, and the area under receiver operating characteristic curve (AUC) were used. The most meaningful model was further explained through the help of Shapley Additive Explanations (SHAP) in order to determine the most significant predictors that influence the outputs of the model. Besides this, key risk variables were stratified by unsupervised K-means clustering in order to group the study population. Dimensionality reduction methods were employed to determine the optimum number of clusters which were compared concerning their differences in the prevalence of oral premalignant lesions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical approval\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Institute of Ethics Committee of the Hassan Institute of Medical Sciences, Hassan conducted a review and approved the study (IEC/HIMS/RR72/21-05-2019). All the participants were informed through written consent before data collection. The study ensured confidentiality of the participants and data anonymity. \u0026nbsp;\u0026nbsp;\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cp\u003e\u003cstrong\u003eParticipant characteristics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA sample of 3,700 adults (100 rural clusters) was analyzed. Mean age of participants that had OPLs was 49.4 \u0026plusmn; 12.4 years, whereas mean age of those that did not have lesions was 51.0 \u0026plusmn; 14.8 years, and there was no statistically significant difference between both categories (p = 0.57). The size of families did not significantly differ in the two groups. The distribution of sex was significantly different with more females among the participants who had OPLs than the ones with no lesions (p = 0.02). Lesion status was also largely linked with occupational status (p = 0.02) and not with religion, family type, socioeconomic status, and family history of oral cancer. Table 1 presents detailed characteristics of the participants.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBehavioural risk factor distribution\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe level of substance use behaviours was highly associated with the occurrence of oral premalignant lesions. Those who reported smoking, smokeless tobacco use, or alcohol intake were also much more likely to have OPLs than non-users (all p \u0026lt; 0.001). The frequency of OPLs was also found to be higher among the participants with a family history of substance use, but this was not statistically significant (p = 0.08). There was no significant association between the presence of lesions and oral hygiene-related practices such as knowledge of oral hygiene and after-meal mouth washing. These results suggest that the most common risk factors of OPLs in the population of the study were substance-related behaviours (Table 1).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1.\u003c/strong\u003e Baseline characteristics of study participants by OPLs\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eVariable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNo lesion\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLesion present\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ep-value\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAge (years), mean \u0026plusmn; SD\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e51.04 \u0026plusmn; 14.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e49.42 \u0026plusmn; 12.38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e0.57\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAge (years), median\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e45\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u0026ndash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFamily size, mean \u0026plusmn; SD\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e4.29 \u0026plusmn; 1.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4.21 \u0026plusmn; 2.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e0.90\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFamily size, median\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u0026ndash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eGender, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.02\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eMale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1902 (51.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4 (21.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eFemale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1779 (48.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e15 (78.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eReligion, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eHindu\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e3556 (96.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e18 (94.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eMuslim\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e120 (3.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e1 (5.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eChristian\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e5 (0.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOccupation, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.02\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eUnemployed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1111 (30.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eFarmer/Labourer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1661 (45.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e9 (47.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eBusiness\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e356 (9.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e5 (26.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eHousewife\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e175 (4.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e1 (5.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eOthers\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e374 (10.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4 (21.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eType of family, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e0.59\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eNuclear\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e2290 (62.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e14 (73.7)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eJoint\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e287 (7.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e1 (5.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eExtended\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1104 (30.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4 (21.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSocioeconomic status, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e0.78\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eClass I\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e129 (3.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e1 (5.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eClass II\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e534 (14.5)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4 (21.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eClass III\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1258 (34.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e4 (21.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eClass IV\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e1223 (33.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e7 (36.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eClass V\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e537 (14.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e3 (15.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFamily history of oral cancer, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e3638 (98.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e19 (100.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e43 (1.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e0 (0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSmoking habit, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u0026lt;0.001\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e3042 (82.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e9 (47.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e637 (17.3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e10 (52.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eChewing habit, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u0026lt;0.001\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e2904 (78.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e9 (47.4)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e777 (21.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e10 (52.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAlcohol consumption, n (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e\u0026lt;0.001\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e3171 (86.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e8 (42.1)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 47px;\"\u003e\n \u003cp\u003eYes\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 19px;\"\u003e\n \u003cp\u003e509 (13.8)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e11 (57.9)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003csup\u003e*The values have been presented as a mean \u0026plusmn; standard deviation or number (percentage). The p-values were determined by the use of the independent t-test, chi-square test, respectively.\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMachine learning model performance\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFour supervised machine learning models were tested to be able to predict the presence of oral premalignant lesions based on the availability of demographic, behavioral, and oral hygiene variables. XGBoost was the best performing model as its predictive accuracy was 85.7 and its AUC was 0.91. Random forest performed well too (accuracy: 82.4, AUC: 0.88), continuously, the support vector machine and logistic regression models (Table 2).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2.\u003c/strong\u003e Risk prediction of oral premalignant lesion using machine learning models\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 27px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAUC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrecision\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRecall\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eF1-score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 27px;\"\u003e\n \u003cp\u003eLogistic Regression\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e76.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15px;\"\u003e\n \u003cp\u003e0.61\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11px;\"\u003e\n \u003cp\u003e0.58\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14px;\"\u003e\n \u003cp\u003e0.59\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 27px;\"\u003e\n \u003cp\u003eSVM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e78.9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003e0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15px;\"\u003e\n \u003cp\u003e0.66\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11px;\"\u003e\n \u003cp\u003e0.63\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14px;\"\u003e\n \u003cp\u003e0.64\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 27px;\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e82.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15px;\"\u003e\n \u003cp\u003e0.72\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11px;\"\u003e\n \u003cp\u003e0.69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14px;\"\u003e\n \u003cp\u003e0.70\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 27px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21px;\"\u003e\n \u003cp\u003e85.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003e0.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15px;\"\u003e\n \u003cp\u003e0.76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11px;\"\u003e\n \u003cp\u003e0.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14px;\"\u003e\n \u003cp\u003e0.75\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003csup\u003e* AUC = area of the receiver operating characteristic curve\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eAll four models have receiver operating characteristic (ROC) curves, as illustrated in Figure 1, showing that XGBoost model outperforms other methods in discrimination.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eImportant features based on SHAP analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo determine the most important predictors of oral premalignant lesions, SHAP analysis of the most successful XGBoost was performed. SHAP summary plot (Figure 2) shows that age, occupation, length of substance use, family size, income, family type, oral hygiene practices, religion, and socioeconomic status had the highest contribution towards model predictions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eK-means clustering Risk stratification\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe K-means clustering, which was unsupervised, was used on the most important demographics, behavioural and hygienic variables in order to stratify the population into various risk groups. Reducing dimensionality methods proposed a three-cluster solution that was the best. Table 3 summarizes the cluster characteristics and the rate of oral premalignant lesion.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3.\u003c/strong\u003e Characteristics of clusters determined with the K-means clustering\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"100%\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCluster\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eParticipants (n)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 9px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMean age (years)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTobacco use (mean)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 10px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOral hygiene score (mean)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMean income (INR)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eOPL prevalence (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePopulation proportion (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003eCluster 0 (High risk)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 15px;\"\u003e\n \u003cp\u003e107\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e44.63\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0.50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 10px;\"\u003e\n \u003cp\u003e0.56\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 12px;\"\u003e\n \u003cp\u003e11,200.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 13px;\"\u003e\n \u003cp\u003e3.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 14px;\"\u003e\n \u003cp\u003e13.7\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003eCluster 1 (Low risk)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 15px;\"\u003e\n \u003cp\u003e347\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e59.35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0.54\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 10px;\"\u003e\n \u003cp\u003e0.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 12px;\"\u003e\n \u003cp\u003e7,071.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 13px;\"\u003e\n \u003cp\u003e0.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 14px;\"\u003e\n \u003cp\u003e44.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 12px;\"\u003e\n \u003cp\u003eCluster 2 (Moderate risk)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 15px;\"\u003e\n \u003cp\u003e329\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 9px;\"\u003e\n \u003cp\u003e59.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 11px;\"\u003e\n \u003cp\u003e0.49\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 10px;\"\u003e\n \u003cp\u003e1.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 12px;\"\u003e\n \u003cp\u003e7,044.58\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 13px;\"\u003e\n \u003cp\u003e1.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 14px;\"\u003e\n \u003cp\u003e42.0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003csup\u003e*The clusters were obtained with the help of K-means clustering on the basis of demographic variables, behavioural variables, and oral hygiene variables. OPLs = oral premalignant lesions.\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eThe highest proportion of OPLs was found in Cluster 0 which is the high-risk group with an average of 3.7 percent though it has a smaller proportion of the population. Cluster 1 (the highest proportion of participants 44.3%), was the least prevalent in lesion (0.6%). Cluster 2 had an intermediate risk profile. The odds of OPLs among the high-risk group were much greater than those among the lowest-risk group (OR = 3.46; 95% CI: 2.58-4.72). The aggregation of clusters is shown in Figure 3.\u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThis community-based research paper proves that explainable machine learning models are applicable to the area of predicting the risk of oral premalignant lesions on a population level and in rural areas. XGBoost was the most predictive model with the analysis of interpretability revealing that age, tobacco exposure, and oral health habits are the most sensitive predictors. The findings are in line with the recent evidence that emphasizes the importance of data-based approaches to risk stratification as they enhance the prevention of non-communicable diseases on a community level [13,14]. Notably, transparency is achieved because of the use of explainable models, and it is becoming clear that transparency is the key to the successful implementation of artificial intelligence in decision-making related to public health [15].\u003c/p\u003e\n\u003cp\u003ePopulation based oral screening trials have found oral prevalence of pre-malignant lesions ranging between about 0.3 to 1.2 percent in a community setting, and higher rates among tobacco users, in large population-based oral screening trials carried out in India [16,17]. The general prevalence of oral premalignant lesions in the current study was 10.2, which agrees with these previous community-based results and justifies the extrapolation of the results. The comparatively low prevalence rate supports the idea that it is necessary to take risks-based strategies to enhance efficiency of screening instead of applying uniform screening to the whole population. The earlier studies of Indian screening based on visual oral examination showed a decrease in mortality rates of oral cancer but had to be conducted repeatedly with large amounts of human resources [14]. Other more recent digital and machine learning-based studies have also reported messages of predictive performance (e.g. area under the curve (AUC)) values of 0.78-0.89 oral cancer or premalignant disorder prediction, frequently on hospital-based or image-based data [17-20]. Comparatively, the explainable machine learning models applied to the current study had AUC values of up to 0.91 with community-level demographic and behavioral data, which represented better discrimination despite the lack of any specialized imaging or clinical infrastructure.\u003c/p\u003e\n\u003cp\u003eThe important predictors in this research are age, exposure to tobacco, and oral hygiene, which are highly consistent with the previous epidemiological data. As an illustration, previous researches have indicated that two to four times more risk of oral premalignant lesions is attributed to tobacco users than to non-users [21,22]. Tobacco-related variables were the most significant contributors in all models in the current analysis, and people who were classified as high-risk group exhibited more relative burden of lesions than low-risk groups. The current findings are also in line with the previous reports that the risk did not exist evenly among populations but was concentrated among certain behavioral and socioeconomic subgroups. Compared to most of the earlier studies on machine learning, which reported high rates of accuracy of over 90 percent but had no interpretability or population-wide applicability [23,24], the current research focuses on balanced metrics of performance and transparency of the model. The moderate-to-high recall range in the most successful model indicates realistic applicability in contexts of identifying more at-risk individuals with the lowest amount of false-negatives, which is necessary in the context of the public health screening scenario, where under-detection is a severe issue.\u003c/p\u003e\n\u003cp\u003eResearch findings in this study will also be significant to oral prevention programs in low-income and middle-income countries. Explainable machine learning can be used to stratify risks to enable focused screening to ensure health systems can focus on high-risk persons and communities, and reduce unnecessary screening in the low-risk groups. The recent policy-based research has also emphasized that digital decision-support platforms can be useful to improve the effectiveness of screening programs conducted by community health workers [25,26]. This approach can be incorporated into the current primary healthcare platforms in the Indian context, such as mobile health applications with the frontline health workers as part of the national non-communicable disease programs. The model assists clinical decision-making by giving easily comprehensible risk scores instead of diagnostic outputs without substituting the existing protocols of screening. Recent WHO and public health informatics frameworks have suggested similar methods to enhance prevention-based primary care [27].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStrengths and limitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe main strengths of the research are the large sample, which is community-based, with multiple rural clusters, and the interpretable machine learning techniques, which make it more transparent and trustworthy. A combination of supervised and unsupervised learning allowed both the ability to predict risks at the individual level and stratify the population, a practice that is becoming increasingly popular in the context of public health analytics.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; However, there are a number of limitations that need to be taken into account. The cross-sectional nature restricts causal inference, whereas lesion identification was done on the basis of clinical examination, and not histopathological confirmation. Similar limitations have been reported in large-scale screening studies that have been carried out in resource-constrained settings. The research was also carried out in only one district, which might not be representative of other areas, which have different risk patterns. Future studies must examine how risk prediction models can be combined with self-detection technologies that are low-cost and user-friendly so as to aid in the early detection of oral premalignant lesions in high-risk groups. Furthermore, wearable devices, that can record oral hygiene, behavioural patterns and changes reported by users of the tobacco products, can be used to supplement community screening by allowing longitudinal self-assessment on an individual basis outside of the clinical context.\u0026nbsp;\u003c/p\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eThe explainable machine learning models support the possibility of risk stratification of oral premalignant lesions in rural populations. Specific screening strategies and more effective resource distribution towards scarce resources in health care can be facilitated by using interpretable risk prediction tools. Preventive oral health services in the low-resource settings could be reinforced through integration of such strategies within the primary healthcare systems.\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAUC \u0026nbsp; Area under the receiver operating characteristic curve\u003c/p\u003e\n\u003cp\u003eML \u0026nbsp; \u0026nbsp;Machine learning\u003c/p\u003e\n\u003cp\u003eOPL \u0026nbsp; Oral premalignant lesions\u003c/p\u003e\n\u003cp\u003eSHAP SHapley Additive exPlanations\u003c/p\u003e\n\u003cp\u003eSVM \u0026nbsp; Support vector machine\u003c/p\u003e\n\u003cp\u003eWHO \u0026nbsp; World Health Organization\u003c/p\u003e\n\u003cp\u003eXGBoost Extreme Gradient Boosting\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Institute of Ethics Committee of the Hassan Institute of Medical Sciences, Hassan conducted a review and approved the study (IEC/HIMS/RR72/21-05-2019). All the participants were informed through written consent before data collection. The study ensured confidentiality of the participants and data anonymity. \u0026nbsp;\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analysed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors’ contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors contributed to the study conception and design. Data collection, analysis, and interpretation were performed by the authors. The manuscript was drafted and critically revised by all authors, and all authors approved the final version for submission.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors thank the study participants and field staff for their cooperation and support during data collection.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBray F, Laversanne M, Sung H, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(2):229\u0026ndash;63. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3322/caac.21834\u003c/span\u003e\u003cspan address=\"10.3322/caac.21834\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhanna D, Shruti T, Tiwari M, et al. Prevalence of Oral Potentially Malignant Lesions, Tobacco use, and Effect of Cessation Strategies among Solid Waste Management workers in Northern India: a pre-post intervention study. BMC Oral Health. 2024;24:1292. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12903-024-05087-8\u003c/span\u003e\u003cspan address=\"10.1186/s12903-024-05087-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTalukdar T, Hazarika JR, Lahkar M. Oral cancer hazards related to tobacco use and assessment of readiness to quit tobacco among OPMD patients in North East India. J Oral Maxillofacial Pathol. 2023. (Cross-sectional study showing tobacco-OPMDassociation).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrevalence. \u003cem\u003eand\u003c/em\u003e risk factors for oral potentially malignant disorders in Indian population. J Pharm Bioallied Sci. 2021;13(Suppl 1):S316\u0026ndash;20. (Reported tobacco and areca nut as dominant risks) PMID: 34447119.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInternational Institute for Population Sciences (IIPS). \u0026amp; Ministry of Health and Family Welfare (MoHFW), Government of India. \u003cem\u003eNational Family Health Survey (NFHS-5), 2019-21\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTalwar V, Singh P, Mukhia N, et al. AI-Assisted Screening of Oral Potentially Malignant Disorders using Smartphone-Based Photographic Images. Cancers. 2023;15(16):4120. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/cancers15164120\u003c/span\u003e\u003cspan address=\"10.3390/cancers15164120\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDesai KM, Singh P, Smriti M, et al. Screening of oral potentially malignant disorders and oral cancer using deep learning models. Sci Rep. 2025;15:17949. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41598-025-02802-5\u003c/span\u003e\u003cspan address=\"10.1038/s41598-025-02802-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi L, Pu C, Jin N, et al. Prediction of 5-year overall survival of tongue cancer based machine learning. BMC Oral Health. 2023;23:567. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12903-023-03255-w\u003c/span\u003e\u003cspan address=\"10.1186/s12903-023-03255-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSomyanonthanakul R, Warin K, Chaowchuen S, et al. Survival estimation of oral cancer using fuzzy deep learning. BMC Oral Health. 2024;24:519. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12903-024-04279-6\u003c/span\u003e\u003cspan address=\"10.1186/s12903-024-04279-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeep learning in oral cancer: a systematic review. BMC Oral Health. 2024;24:212.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInternational Cancer Risk Prediction Models in Oral Cancer. (MDPI Cancers 2024;16(3):617). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/cancers16030617\u003c/span\u003e\u003cspan address=\"10.3390/cancers16030617\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSankaranarayanan R, Ramadas K, Thomas G, et al. Effect of screening on oral cancer mortality in Kerala, India: A cluster-randomised controlled trial. Lancet. 2005;365(9475):1927\u0026ndash;33. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/S0140-6736(05)66658-5\u003c/span\u003e\u003cspan address=\"10.1016/S0140-6736(05)66658-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomas G, Hashibe M, Jacob BJ, et al. Long-term outcomes of a randomized oral cancer screening trial in India. Oral Oncol. 2019;88:53\u0026ndash;9. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.oraloncology.2018.11.020\u003c/span\u003e\u003cspan address=\"10.1016/j.oraloncology.2018.11.020\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMacey R, Walsh T, Brocklehurst P, et al. Diagnostic accuracy of oral cancer and potentially malignant disorders screening tools: A systematic review. Community Dent Oral Epidemiol. 2021;49(6):529\u0026ndash;40. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/cdoe.12642\u003c/span\u003e\u003cspan address=\"10.1111/cdoe.12642\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHalicek M, Little JV, Wang X, et al. Tumor margin classification of head and neck cancer using hyperspectral imaging and machine learning. Cancers (Basel). 2020;12(6):1378. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/cancers12061378\u003c/span\u003e\u003cspan address=\"10.3390/cancers12061378\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePatil S, Albogami S, Hosmani J, et al. Artificial intelligence in oral cancer diagnosis: A systematic review. Oral Oncol. 2020;105:104717. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.oraloncology.2020.104717\u003c/span\u003e\u003cspan address=\"10.1016/j.oraloncology.2020.104717\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWarnakulasuriya S, Kujan O, Aguirre-Urizar JM, et al. Oral potentially malignant disorders: A consensus report from an international seminar. Oral Oncol. 2020;102:104550. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.oraloncology.2019.104550\u003c/span\u003e\u003cspan address=\"10.1016/j.oraloncology.2019.104550\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGupta B, Johnson NW. Systematic review and meta-analysis of oral cancer epidemiology in India. Lancet Oncol. 2017;18(9):e482\u0026ndash;94. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/S1470-2045(17)30472-6\u003c/span\u003e\u003cspan address=\"10.1016/S1470-2045(17)30472-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeam AL, Kohane IS. Big data and machine learning in health care. JAMA. 2018;319(13):1317\u0026ndash;8. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1001/jama.2017.18391\u003c/span\u003e\u003cspan address=\"10.1001/jama.2017.18391\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347\u0026ndash;58. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1056/NEJMra1814259\u003c/span\u003e\u003cspan address=\"10.1056/NEJMra1814259\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30:4765\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundberg SM, Erion GG, Lee SI. Explainable machine learning for public health. Nat Mach Intell. 2020;2:252\u0026ndash;9. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s42256-020-0173-7\u003c/span\u003e\u003cspan address=\"10.1038/s42256-020-0173-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgarwal S, Perry HB, Long LA, Labrique AB. Evidence on feasibility and effectiveness of digital tools for frontline health workers. BMJ Glob Health. 2020;5(10):e003224. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1136/bmjgh-2020-003224\u003c/span\u003e\u003cspan address=\"10.1136/bmjgh-2020-003224\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMehl G, Labrique A. Prioritizing integrated digital health strategies for universal health coverage. Science. 2019;365(6452):eaaz347. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.aaz347\u003c/span\u003e\u003cspan address=\"10.1126/science.aaz347\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeters DH, Adam T, Alonge O, Agyepong IA, Tran N. Implementation research: What it is and how to do it. Bull World Health Organ. 2013;91:731\u0026ndash;6. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2471/BLT.13.124933\u003c/span\u003e\u003cspan address=\"10.2471/BLT.13.124933\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. Ethics and governance of artificial intelligence for health. Geneva: WHO; 2021.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-public-health","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pubh","sideBox":"Learn more about [BMC Public Health](http://bmcpublichealth.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/pubh/default.aspx","title":"BMC Public Health","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"OPLs, Tobacco use, Prevalence, Community-based study, Machine learning, Public health screening.","lastPublishedDoi":"10.21203/rs.3.rs-8648393/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8648393/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eOral cancer is a major public health problem among the population of India. In rural setting whereby there is exposure to high-risk factors associated with tobacco use and low accessibility to structured screening services. Though visual oral examination has been proven to be effective in lowering mortality rates of oral cancer, population-based screening has not only been found to be resource intensive but also hard to maintain in the primary health care system. Screening programs could be made more efficient by risk stratification methods that identify the people with a greater risk of having oral premalignant lesions. The recent developments of machine learning gives a chance to make data-driven predictions of risks, although the issues associated with the model transparency and its applicability in relation to a population are still present.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThe design of the study was community based cross-sectional study on 3,700 adults in 100 rural clusters of the Hassan District in Karnataka. Structured interviews were used to gather sociodemographic data, tobacco and alcohol use, and oral hygiene practices, and then clinical oral examination was carried out based on guidelines of the World Health Organization. Cross-validation was used to develop and assess supervised machine learning models that comprise support vector machine, random forest, and extreme gradient boosting (XGBoost). SHapley Additive explanations (SHAP) were used to determine model interpretability and populations-level risk stratification was done using unsupervised K-means clustering.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThe best performance models XGBoost were found to have the highest predictive accuracy (area under the receiver operating characteristic curve\u0026thinsp;=\u0026thinsp;0.91, accuracy\u0026thinsp;=\u0026thinsp;85.7%). Aging and exposure to tobacco and bad oral health were also reported to be the most consistent predictors across the models. Clustering analysis revealed the presence of a high-risk sub-group with a significantly greater relative risk burden of oral premalignant lesions that may be supported by regression-based estimates (odds ratio\u0026thinsp;=\u0026thinsp;3.46, 95% confidence interval: 2.58\u0026ndash;4.72).\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003ePredictable machine learning-based risk predictors can be used to facilitate population stratification and focused screening plans to detect oral premalignant lesions in rural areas to allow the most effective utilization of scarce public health assets in primary healthcare systems.\u003c/p\u003e","manuscriptTitle":"Risk prediction of oral premalignant lesions (oral cancer) using explainable machine learning through a community-based cross-sectional study in rural India","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-02 11:04:46","doi":"10.21203/rs.3.rs-8648393/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-03-12T05:38:09+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-10T17:17:04+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"257305838664206729753772136748224586113","date":"2026-03-10T16:58:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"154009421854000968612783915996613049167","date":"2026-03-03T11:38:57+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-01T19:25:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"286791572512082683820814991460802146412","date":"2026-03-01T06:23:44+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-28T01:40:11+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"181361264867775987619104261850752538107","date":"2026-02-28T01:30:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"257305838664206729753772136748224586113","date":"2026-02-26T00:07:37+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-25T23:29:04+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"140379672897846594002679395031915484677","date":"2026-02-25T23:16:03+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"252871235012440321136429867398928533108","date":"2026-02-25T19:08:07+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-02-25T19:05:11+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-23T14:00:44+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-01-30T09:18:56+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-30T07:32:48+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Public Health","date":"2026-01-30T07:15:47+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-public-health","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pubh","sideBox":"Learn more about [BMC Public Health](http://bmcpublichealth.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/pubh/default.aspx","title":"BMC Public Health","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d1fa9c14-f85e-4bb9-bc06-ad3e2f1a8060","owner":[],"postedDate":"March 2nd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-01T05:25:39+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-02 11:04:46","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8648393","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8648393","identity":"rs-8648393","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00