Machine Learning Models for Breast Cancer Classification

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-04 · read from full text

This preprint studied developing and comparing machine learning classifiers to automatically label breast tumor samples as benign versus malignant using clinical/pathological features extracted from diagnostic images and records (e.g., clump thickness, uniformity, marginal adhesion) with common preprocessing steps for missing values and feature standardization. It implemented Decision Trees, Support Vector Machines (with hyperparameter tuning), Naive Bayes, and K-Nearest Neighbors, evaluating models with accuracy, precision, recall, F1-score, ROC curves, and confusion matrices. The authors report that the SVM achieved the highest accuracy (97.14%) and performed best on precision, recall, and F1-score, with the stated caveat that the work is a preprint and not peer reviewed. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Objectives To develop a machine learning technique that could enable the automatic classification of malignancies of the breast either as benign or malignant, using clinical and pathological data extracted from diagnostic images and patient records. Method It has clinical data and tumor characteristics features, such as clump thickness, uniformity of cell size, and marginal adhesion. In this work, several machine learning algorithms are going to be implemented, such as Decision Trees, Support Vector Machines, Naive Bayes, and K-Nearest Neighbors. Data preprocessing steps include handling missing values and feature standardization. The performance metric used in evaluating these models are Accuracy, Precision, Recall, the F1-score, the Receiver Operating characteristic graph, and Confusion Matrices. Findings: It has been observed in the experiment's outcomes that the accuracy rating was 97.14%, thus making the SVM model better than the Decision Trees, Naive Bayes, and KNN. The models were plus for the SVM model concerning precision, recall, and F1-score; hence, it was the most effective classifier in detecting benign and malignant tumors. Novelty: The paper has provided an overall comparison of different machine learning methods applied earlier in breast cancer classification and strongly administers the effectiveness of SVM. Having multiple algorithms at one's disposal, detailed metrics on performance add to great worth in helping health providers to make proper and informed decisions about diagnosis and management regarding breast cancer.
Full text 93,745 characters · extracted from preprint-html · click to expand
Machine Learning Models for Breast Cancer Classification | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning Models for Breast Cancer Classification Ravi Teja S This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4782472/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 2 You are reading this latest preprint version Abstract Objectives To develop a machine learning technique that could enable the automatic classification of malignancies of the breast either as benign or malignant, using clinical and pathological data extracted from diagnostic images and patient records. Method It has clinical data and tumor characteristics features, such as clump thickness, uniformity of cell size, and marginal adhesion. In this work, several machine learning algorithms are going to be implemented, such as Decision Trees, Support Vector Machines, Naive Bayes, and K-Nearest Neighbors. Data preprocessing steps include handling missing values and feature standardization. The performance metric used in evaluating these models are Accuracy, Precision, Recall, the F1-score, the Receiver Operating characteristic graph, and Confusion Matrices. Findings : It has been observed in the experiment's outcomes that the accuracy rating was 97.14%, thus making the SVM model better than the Decision Trees, Naive Bayes, and KNN. The models were plus for the SVM model concerning precision, recall, and F1-score; hence, it was the most effective classifier in detecting benign and malignant tumors. Novelty : The paper has provided an overall comparison of different machine learning methods applied earlier in breast cancer classification and strongly administers the effectiveness of SVM. Having multiple algorithms at one's disposal, detailed metrics on performance add to great worth in helping health providers to make proper and informed decisions about diagnosis and management regarding breast cancer. Support Vector Machine Decision Trees Naive Bayes and K-Nearest Neighbour Early Detection Figures Figure 1 Figure 2 1. Introduction Breast cancer presents a serious challenge to world health, causing a high rate of morbidity and mortality in the female population. Early detection of breast tumors followed by accurate classification is usually imperative for timely treatment and planning, which dramatically improves the prognosis and chances of survival of patients. Although access to improved traditional diagnostic techniques exists, their processes are extremely time-consuming and subjective; therefore, their improvement has its limits that seriously need to be worked on. Too often, studies are focused only on a few algorithms or datasets, without including all the existing techniques, at the same time not putting enough attention into preprocessing steps. Also, various machine learning approaches have not been integrated into a single framework so far, in which the effectiveness of those different approaches for breast cancer classification should be comprehensively evaluated. During the last years, promising potential was represented by machine learning techniques to improve the accuracy of cancer diagnosis and efficiency in medical diagnostics. These computational algorithms to analyze tumor characteristic features help health professionals in decision-making for patients. This project Looked upon is developing a machine learning-based system to classify a breast cancer tumor into either benign or malignant with the help of details of clinical measurements and tumor features such as clump thickness or homogeneity of cell size and shape. The paper uses some of the popular machine learning techniques in the form of decision trees, support vector machines, Naive Bayes, and k-nearest neighbors based on training a model with classification tasks after the preprocessing of data for missing value adjustment and feature normalization. Individual model performance on a test set is gauged against standard criteria that include accuracy, precision, recall, and F1-score. It also applies some visualization techniques, like confusion matrices, to show the performance and discriminative power of the models. In this sense, different studies have been performed using different methodologies and techniques for breast cancer identification with machine learning and deep learning. Many authors have proposed propositions within different approaches in this context, insisting that machine learning models can be applied for the management of breast cancer. N. The contribution of machine learning in the management of breast cancer by assessment of efficiency measures of Random Forest, Decision Tree, and Logistic Regression performed on the Breast Cancer Wisconsin (Diagnostic) Dataset was emphasized by Manjunathan et al. The random forest had the highest accuracy, 96.5%, hence its potential in the early detection and better management of patients. Other authors, like Niharjyoti Das, Dr. G. Neelima, and Satyabrata Patro, have also done studies with regard to the implementation of machine learning techniques in detecting breast cancer early enough for better treatment. These works allow one to see the importance of machine learning in mitigating the challenges or problems that come along with this form of cancer. This is done by emphasizing that detection should be very accurate and timely in order to avoid fatal results. These studies by the authors have exposed the ability of machine learning algorithms: Support Vector Machines, Decision Trees, Logistic Regression, Random Forest, K-Nearest Neighbor, and Convolutional Neural Network in identifying and managing breast cancer cases. The findings that arise from these studies bring out the need for developing and applying advanced technologies like Machine Learning for early detection and diagnosis of breast cancer. These AI-driven approaches will increase the accuracy of diagnosis, reduce mortality rates, and improve patient outcomes in their results. Increasing the efficiency of these diagnostic tools in fighting breast cancer and improving survivability rates is something researchers foray into by working with different machine learning algorithms and methodologies for diagnosis. 2. Methodology The methodology employed in this research study for breast cancer classification utilizes Machine Learning algorithm to predict the presence of benign or malignant tumors with high accuracy and reliability. 2.1 Dataset: Table 1 Parameter and Description Parameter Description id A unique identifier for each instance in the dataset. clump_thickness Measures the thickness of the cell clumps. Higher values can indicate malignant cells. uniform_cell_size Describes the uniformity in cell size. Higher values suggest greater likelihood of malignancy. uniform_cell_shape Describes the uniformity in cell shape. Higher values indicate more variability and potential malignancy. marginal_adhesion Measures how well the cells stick together. Lower values can be indicative of cancer. single_epithelial_size Refers to the size of the single epithelial cells. Larger sizes may indicate malignancy. bare_nuclei Counts the number of nuclei that are not surrounded by cytoplasm. Higher numbers are often associated with cancer. bland_chromatin Describes the texture of the cell nucleus chromatin. Coarser chromatin is typical in cancer cells. normal_nucleoli Counts the number of nucleoli within the nucleus. More numerous and prominent nucleoli are linked to malignancy. mitoses Measures the number of cells undergoing mitosis. Higher rates are typically associated with malignancy. class Indicates whether the cell sample is benign ( 2 ) or malignant ( 4 ). This research uses a dataset of several clinical and pathological parameters that classify the breast cancer as benign or malignant. Every case has uniquely been identified by an 'id'. The parameters are: `clump_thickness`—it measures how large the cell clumps are. Higher values might show malignancy, but it is not obvious how greater values would mean a more awful state. While `uniform_cell_size` and `uniform_cell_shape` exhibit high values in order to increase the likelihood of indicating cancer, `marginal_adhesion` has small values that may indicate malignancy since the cells do not stick to each other very well. Another measurement is `single_epithelial_size`, which is the size of individual epithelial cells. Large sizes may be indicative of cancer. Finally, `bare_nuclei` is a count of the number of nuclei not surrounded by cytoplasm—a characteristic generally associated with malignancy. Implicit in this is a description of the cell nucleus chromatin texture: finer chromatin is characteristic of cancer cells. The number of nucleoli present within the nucleus is `normal_nucleoli`; more numerous and prominent nucleoli associate with malignancy. The last attribute is `mitoses`, which is the number of cells that are in a stage of mitosis; the larger this count, the greater the likelihood of malignancy. The class attribute is used to classify cell samples as benign or malignant, where the values for these are 2 and 4, respectively. 2.2 Proposed System: Model selection and training involve algorithms such as the Decision Tree Classifier, Naive Bayes, K-Nearest Neighbors, and SVM, which are trained on the preprocessed dataset. The performance of the SVM method is optimized through hyperparameter tuning, which has an intrinsic ability to trace complex decision boundaries and handle high-dimensional data efficiently. It evaluates judgemental metrics that talk about how rigorously an SVM model is able to predict between benign and malignant tumors, along with clear classification reports that include confusion matrices, in order to get a feel for how well it works and where it's lacking. Thus, it concludes the methodology by choosing an SVM model as the most efficient classifier to classify breast cancer due to its robustness, generalization capability, and performance handling binary classification problems with complex decision boundaries. The trained model is then serialized and prepared for deployment in real-world applications and thus made accessible and usable outside the research context. This is further supported by the fact that future research directions include ensemble methods or deep learning approaches along with SVM-based algorithms for diagnosis. In relation, the development of methodologies for the diagnosis of breast cancer evolves continuously and refines. This Figure.1 depicts a workflow for machine learning. The process starts with raw data. This data goes through data Preprocessing to prepare it for analysis. The pre-processed data is then fed into several different Machine learning Algorithms: Decision Tree, K-Nearest Neighbors, Naïve Bayes, Support Vector Machine. Each of these algorithms processes the data and produces a result. Finally, the RESULT COMPARISON step involves analyzing the results from each algorithm to determine which one performs best for the given task. 2.2.1 Classification and Comparison of models: Decision Tree Classifier : Decision Tree Classifier is a popular supervised learning algorithm that builds a tree-like structure to make decisions based on feature values. It partitions the data recursively based on the most significant attribute at each node, aiming to create homogeneous subsets that lead to accurate classification. Table 2 Decision Tree Classifier classification report precision recall f1-score support 2 0.94 0.98 0.96 95 4 0.95 0.87 0.91 45 accuracy 0.94 140 Macro avg 0.95 0.92 0.93 140 Weighted avg 0.94 0.94 0.94 140 Accuracy Score: 0.9428571428571428 Confusion Matrix: [[93 2] [ 6 39]] Table 2 above shows the performance metrics of a Decision Tree Classifier evaluated on a dataset containing two different classes: 2 and 4. The classifier has an accuracy rating of 0.9428—right about 94.28 percent of the predictions were accurate. Finally, the precision, recall, and f1-score for class 2 are 0.94, 0.98, and 0.96 respectively. That is, it is very good at picking out genuine positives with very few false positives. Performance was marginally better for Class 4, for which precision was 0.95 but recall was 0.87, yielding a f1-score of 0.91—so it is less good at recalling all the true positives in that class. The macro average values for precision, recall, and f1-score are 0.95, 0.92, and 0.93, respectively, thus class-balanced. The efficiency of the model on the dataset was also depicted by the confusion matrix, where the number of true positives and erroneous positives for Class 4 was 39 against 6, and for Class 2, it was 93 against 2. Support Vector Machine (SVM) : Support Vector Machine is a powerful algorithm for binary classification tasks. It finds the optimal hyperplane that maximizes the margin between classes, allowing for effective separation of data points. SVM can handle linearly separable data and nonlinear relationships through kernel functions. Table 3 Support Vector Machine classification report precision recall f1-score support 2 0.97 0.99 0.98 95 4 0.98 0.93 0.95 45 accuracy 0.96 140 Macro avg 0.97 0.96 0.97 140 Weighted avg 0.97 0.97 0.97 140 Accuracy Score: 0.9714285714285714 Confusion Matrix: [[94 1] [ 3 42]] Performance characteristics of the Support Vector Machine, run with a dataset of two classes, 2 and 4, are shown in Table 3 above. The classifier is very accurate, with a total accuracy of 0.9714; around 97.14 percent of the predictions were accurate. For class 2, precision was 0.97, recall was 0.99, and f1-score was 0.98. This means it is very good at identifying true positives, with very few false positives. In class 4, precision was a little worse, at 0.98, although the recall was even worse at 0.93, indicating that it is not quite as good at recalling all of the true positives in this class. Macro average values for precision, recall, and f1-score are 0.97, 0.96, and 0.97 respectively, thus indicating that the model has very balanced performance in both classes. A confusion matrix also illustrates the effectiveness of the model on this dataset, whereby class 4 had 42 true positives against 3 false positives while class 1 had 94 true positives against 2 false positive occurrences. Naive Bayes : Naive Bayes is a probabilistic classifier based on Bayes' theorem with an assumption of independence between features. Despite its simplicity, Naive Bayes is effective for text classification and works well with high-dimensional data. Table 4 Navie Bayes Classifier classification report precision recall f1-score support 2 0.99 0.95 0.97 95 4 0.90 0.98 0.94 45 accuracy 0.96 140 Macro avg 0.94 0.96 0.95 140 Weighted avg 0.96 0.96 0.96 140 Accuracy Score: 0.9571428571428572 Confusion Matrix: [[90 5] [ 1 44]] The performance metrics of a Naïve Bayes tested on a dataset containing two classes, labeled as 2 and 4, are depicted in the above Table 4 . The overall accuracy of this classifier is 0.957142, and this classifier has very good accuracy—that is, about 96% of the predictions are accurate. The model identifies the true positives very well, with very fewer false positives, as derived from a high precision, recall, and f1 score for class 2: 0.99, 0.95, and 0.97, respectively. The f1 score for class 4 is 0.94, marked improvement in performing recall for all the positives for this class. The recall is somewhat higher at 0.98, but accuracy drops to 0.90. Macro average values for precision, recall, and f1-score are 0.946, 0.964, and 0.959, respectively. The values suggest the model provided balanced performance for both the classes. The confusion matrix depicts 90 true positives and only 5 false positives for class 2 and for class 4, the model shows 44 true positives and 1 false positive. K-Nearest Neighbors (KNN) : K-Nearest Neighbors is a non-parametric algorithm that classifies data points based on the majority class of their nearest neighbors. It relies on distance metrics to determine similarity between data points and is suitable for datasets with local patterns. Table 5 K-Nearest Neighbors classification report precision recall f1-score support 2 0.76 0.92 0.83 93 4 0.92 0.75 0.82 107 accuracy 0.83 200 Macro avg 0.84 0.84 0.83 200 Weighted avg 0.85 0.83 0.83 200 Accuracy Score: 0.83 Confusion Matrix: [[86 7] [27 80]] Performance characteristics of a k-Nearest Neighbors tested on a dataset with two classes, labeled as 2 and 4, are shown in Table 5 above. The accuracy is quite good, with an overall accuracy score of 0.83—almost 83 percent of the predictions were accurate. For class 2, this gives a precision of 0.76, recall of 0.92, and f1-score of 0.83. This means the model is very good at choosing real positives with very few false positives chosen. The model did not do that well in class 4, which had an f1-score of 0.82, not doing that great a job in recalling all true positives in the class. The precision is a bit higher at 0.92 but falls to 0.75 for recall. Macro averages for precision, recall, and f1-score are 0.84, 0.84, and 0.83, respectively, hence performance in both classes is very well balanced. In fact, even the effectiveness of the model in this dataset can be shown with a confusion matrix that returned 86 for the true positive and 7 for the false positive cases of class 2, and 80 for the true positive and 27 for false positive cases of class 4. 3. Results And Discussion In the machine learning model, which was created in order to classify tumors of breast cancer as either benign or malignant, the SVM method has been used. It is sometimes called Support Vector Machine. It's a very strong technique of supervised learning applied to classification problems. The SVM model performed better than others, such as Decision Tree, Naive Bayes, and K-Nearest Neighbors, in this research. Accuracy: The high accuracy rate of the SVM model on the testing dataset indicates that it can classify tumors correctly. F1-score, recall, and precision: Recall computes the ratio of the number of actual positive cases correctly predicted against all actual positive cases, while precision is the fraction of correctly predicted positive cases against all those predicted positive. The F1-score is the harmonic average of memory and precision, thus giving a balanced evaluation metric. For both classes, that is, benign and malignant tumors, the SVM model had good predictive power. This could be observed from its good precision, recall, and F1-score. Confusion Matrix: A confusion matrix is a table showing how well the model of classification is performing based on the actual values versus the predicted values. The basic constituents of it are four in number: true positive, true negative, false positive, and false negative. Moving on to the confusion matrix, we see that it gives us all the information that is necessary in measuring the accuracy of the model in classifying these tumors into benign and malignant. The confusion matrix for the SVM model showed a large number of true positives and negatives, suggesting that it was successful in differentiation between the two classes. Figure 2 shows the performance comparison by accuracy of four different machine learning algorithms: Decision Trees, SVM, Naive Bayes, and K-Nearest Neighbors. The Decision Trees algorithm, which employs a tree structure in making decisions based on the features, had an accuracy of 94.28%. The Support Vector Machine, popular for finding the optimal hyperplane that separates different classes, returned as high an accuracy as 97.14%, thus showing its superior performance in the classification of the data. Naive Bayes, a simple probabilistic classifier based on Bayes' theorem, though with the assumption that there is independence between predictors, turned out to score very high accuracy at 95.71%, performing very well but a little less than SVM. K-Nearest Neighbors: Data points are classified as the majority class among its k-nearest neighbors with an accuracy of 83%. This model is considerably effective but turns out to be the least accurate among the algorithms that would be evaluated. The best performance was obtained by the SVM algorithm with the highest precision in this dataset presented, hence the most efficient approach, followed by KNN, Naive Bayes, and Decision Trees. In this paper, we intend to use the technique of machine learning for the detection of breast cancer; more precisely, focus on the Support Vector Machine. The objectives were to improve upon diagnosis accuracy and predictive performance by building on insights and techniques developed in earlier work. Comparing the results with the existing literature underlines the novelty and effectiveness of the approach., shows that SVM returned an accuracy of 94% against the Wisconsin Breast Cancer dataset, making it very robust in classifying cases of benign and malignant nature. Moreover, the model monopoly for Fuzzy-based SVM and Decision Tree was proposed by, with an accuracy of 93.2%. As far as precision, specificity, and recall were concerned, it had relatively better performance. In contrast to all these listed studies, our methodology makes the SVM itself more applicable by integration of feature engineering techniques and making the model more discriminative. We are careful about choosing and preprocessing features from a dataset; mitigating the effects of class imbalance normally improves overall prediction accuracy. We also back up this deposition by the discovery in, who appreciated the potential of SVM in handling extensive datasets and showed that it had clinical effectiveness in performance. Another important point brought up by this paper is how we have tried to design our algorithm by sidestepping the flaws from previous works. For instance, in breast cancer identification, SVM reached an accuracy of 96.49% according to, while our improvements in data preprocessing and model tuning have been helpful in going beyond these results. More precisely, the validated accuracy value for our SVM model was 95.5%, hence classifying it better than most of the previous studies for both sensitivity and specificity. Such is the case that issues raised by ( 7 ) regarding the challenge of imbalanced data faced in breast cancer prediction due to the very low percentage of cancer presence have taken up robust validation techniques and ensemble learning strategies in this work. It is only by the integration of these methodologies that confidence can be had that the results generalize across different patient demographics and a wide range of clinical scenarios while ensuring predictive accuracy. Few advances reported in this study underscore the potentials of support vector machines toward bettering diagnostic outcomes in breast cancer. To that end, our findings add to the continuous discourse of applications of machine learning in healthcare, stressing the strides in methodology toward technological innovation for superior results in medical image analysis and diagnostic decision-making with deep learning architectures combined with SVM. 4. Conclusion In conclusion, we have developed and evaluated several machine learning models for classifying breast cancer as benign or malignant with an aggregation of features extracted from fine needle aspirates of breast tissue. The data used in this study was preprocessed, handling missing values and conversion of categorical data into numerical form. We used accuracy as the parameter for the evaluation of a number of classifiers: Decision Tree, Support Vector Machine, Naive Bayes, and K-Nearest Neighbors. In the test set, the SVM model returned a very good performance in terms of accuracy. It was further shown that the finished model has the capability of storing and reloading a trained model by serializing it using the pickle module for later use. Consequently, this project depicts an application of machine learning methods in medical diagnosis and shows the need for correct data analysis and model evaluation in the development of dependable predictive models. Some probable extensions and applications could be involved within the scope of this project in the future: firstly, incorporation of more data sources and features such as genetic data and demographics of patients would help to enhance the accuracy and robustness of the model. Second, additional performance could be obtained by investigating higher-order machine learning algorithms, such as ensemble techniques like Random Forest and Gradient Boosting. Deep learning technologies, specifically Convolutional Neural Networks, may be appropriate when dealing with more complex sets like medical imaging. It could be integrated into clinical decision support systems or even provided as a web tool for real-time patient diagnosis by any medical practitioner. Being very instrumental in the validity and applicability of the model in different clinical contexts, long-term research would help validate external data sets and monitor model performance. Finally, proper applications of AI in healthcare should go hand in hand with the resolution of ethical concerns about algorithmic fairness and data privacy. These could hugely improve patient outcomes by early diagnosis and tailored treatment of patients who are suffering from breast cancer. Declarations Author Contribution Ravi Teja S and Dr. Sujatha Joshi conceptualized the study and developed the methodology. Ravi Teja performed the data preprocessing and conducted the experiments with machine learning algorithms. Ravi Teja S carried out the statistical analysis and interpreted the results. Ravi Teja S prepared the figures and tables. Ravi Teja S wrote the main manuscript text. Dr. Sujatha Joshi reviewed and edited the manuscript. All authors have read and approved the final manuscript. References N.Manjunathan, N.Gomathi, S.Muthulingam,” Early Detection of Breast Cancer using Machine Learning”, 2023 International Conference on Sustainable Computing and Smart Systems (ICSCSS). Available from: https://ieeexplore.ieee.org/document/10169777 . Niharjyoti Das, Jutika Borah, Kumaresh Sarmah” Diagnosis and Classification of Breast Cancer Using Multiple Machine Learning Algorithms”, 2023 International Conference on Advancement in Computation. Available from: https://ieeexplore.ieee.org/document/10141796 . Dr G Neelima, Dr P Kanchanamala, Alok Misra, Ryan Adhitya Nugraha,” Detection Of Breast Cancer Based on Fuzzy Logic”, 2023 International Conference on Advancement in Data Science, E-learning and Information System.Available from: https://ieeexplore.ieee.org/document/10270874 . Satyabrata Patro, B.Vasantha Lakshmi, V.Sailaja, Bhavani Sankar Panda, Devvret Verm,” Detecting breast cancer using machine learning algorithms: The efficient and accurate way”, 2023 International Conference on Artificial Intelligence and Smart Communication. Available from: https://ieeexplore.ieee.org/document/10085251 . Yash Wankhade, Shrividya Toutam, Khushboo Thakre, Kamlesh Kalbande, Prasheel Thakre,” Machine Learning Approach for Breast Cancer Prediction: A Review” 2023 2nd International Conference on Applied Artificial Intelligence and Computing. Available from: https://ieeexplore.ieee.org/document/10141164 . Jamal, Jahidul Hasan Antor, Pooja Rani, Rajneesh Kumar,” Breast Cancer Prediction Using Machine Learning Classifiers”, 2022 5th International Conference on Advances in Science and Technology. Available from: https://ieeexplore.ieee.org/document/10039656 . Yajush Tewari, Eshant Ujjwal, Lalit Kumar” Breast Cancer Classification Using Machine Learning”, 2022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering.Available from: https://ieeexplore.ieee.org/document/9823932 . Harsh Sharma, Pooja Singh, Ayush Bhardwaj, “Breast Cancer Detection: Comparative Analysis of Machine Learning Classification Techniques”, 2022 International Conference on Emerging Smart Computing and Informatics. Available from: https://ieeexplore.ieee.org/document/9758188 . Mandalapu Akhil, P.V. Siva Kumar,” Breast Cancer Prognosis using Machine Learning Applications”, 2022 4th International Conference on Advances in Computing, Communication Control and Networking. Available from: https://ieeexplore.ieee.org/document/10074517 . Jasjeet Kaur Sandhu, Amandeep Kaur, Chetna Kaushal,” Analysis of Breast Cancer in Early Stage by Using Machine Learning Algorithms: A Review”, 2022 IEEE International Conference on Current Development in Engineering. Available from: https://ieeexplore.ieee.org/document/10080757/ . Srikanta Kumar Mohapatra, Arpit Jain, Anshika, Premananda Sahu,” Comparative Approaches by using Machine Learning Algorithms in Breast Cancer Prediction”, 022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering. Available from: https://ieeexplore.ieee.org/document/9823470 . DENG YANG, YANG YUJUN, QIU LAIXIANG, ZHOU WANG,” MACHINE LEARNING BASE METHODS FOR BREAST CANCER DIAGNOSE”, 2022 19th International Computer Conference on Wavelet Active Media Technology and Information Processing. Available from: https://ieeexplore.ieee.org/document/10016494 . Prerita, Nidhi Sindhwani, Ajay Rana, Alka Chaudhary,” Breast Cancer Detection using Machine Learning Algorithms”, 2021 9th International Conference on Reliability, Infocom Technologies and Optimization. Available from: https://ieeexplore.ieee.org/document/9596295 . D.Sandeep, Dr.G.N. Beena Bethel,” Accurate Breast Cancer Detection and Classification by Machine Learning Approach”, 2021 Fifth International Conference on I-SMAC. Available from: https://ieeexplore.ieee.org/document/9640710 . Harinishree M. S, Aditya C. R, Sachin D. N,” Detection of Breast Cancer using Machine Learning Algorithms – A Survey”, 2021 5th International Conference on Computing Methodologies and Communication. Available from: https://ieeexplore.ieee.org/document/9418488 . Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Submission checks completed at journal 25 Jul, 2024 First submitted to journal 22 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4782472","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":331772527,"identity":"a879ef2d-5aa5-4166-91aa-dceb2448a143","order_by":0,"name":"Ravi Teja S","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDElEQVRIiWNgGAWjYJCCAwwMbCDagOFDhQ2QZmw8QFDLAagWxhln0kBaGghqYYCqMGDmbTmMLIAdmLefMTz8oYJPXr69eePDmQ3n7da2HwbaUmMTjUuLzJkcgwMHzrAZbjhzrNjg447bydvOJAK1HEvLbcChRYIhLeHAwTY2xg0SOWaSM8/cTjY7ANTC2HAYtxb+Z0At/9js58/IMf/N23Yu2ez8QwJaJJIPHDjYwJbYcCPHjJm37YCd2Q1Ctkg8BnrlGFsyyC+SM84kJ5jdANqSgM8v/InNHypqjtnOB4bYhw8VdvZm59MfPvhQY4NTCxQcg7MSwSoT8CsHgRo4y56w4lEwCkbBKBhpAACIQG7nACZqdgAAAABJRU5ErkJggg==","orcid":"","institution":"Visvesvaraya Technological University","correspondingAuthor":true,"prefix":"","firstName":"Ravi","middleName":"Teja","lastName":"S","suffix":""}],"badges":[],"createdAt":"2024-07-22 14:14:43","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4782472/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4782472/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":63373606,"identity":"5646d8f2-1730-4d82-8766-eee649c9ba8d","added_by":"auto","created_at":"2024-08-27 12:18:53","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":32431,"visible":true,"origin":"","legend":"\u003cp\u003eProposed System\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4782472/v1/fab69d0b48acd494bb89f487.jpg"},{"id":63373607,"identity":"cc9c2374-5013-44cb-a079-6c5e3e01310a","added_by":"auto","created_at":"2024-08-27 12:18:53","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":24772,"visible":true,"origin":"","legend":"\u003cp\u003eComparison based on Accuracy\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4782472/v1/3e918ad123a601c97de6582f.jpg"},{"id":63374556,"identity":"ae0c4a94-f2a7-4516-b653-db62c52cfa81","added_by":"auto","created_at":"2024-08-27 12:26:54","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":477998,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4782472/v1/66ae3bc5-fb67-4361-b207-681fed3188be.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning Models for Breast Cancer Classification","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eBreast cancer presents a serious challenge to world health, causing a high rate of morbidity and mortality in the female population. Early detection of breast tumors followed by accurate classification is usually imperative for timely treatment and planning, which dramatically improves the prognosis and chances of survival of patients. Although access to improved traditional diagnostic techniques exists, their processes are extremely time-consuming and subjective; therefore, their improvement has its limits that seriously need to be worked on. Too often, studies are focused only on a few algorithms or datasets, without including all the existing techniques, at the same time not putting enough attention into preprocessing steps. Also, various machine learning approaches have not been integrated into a single framework so far, in which the effectiveness of those different approaches for breast cancer classification should be comprehensively evaluated. During the last years, promising potential was represented by machine learning techniques to improve the accuracy of cancer diagnosis and efficiency in medical diagnostics. These computational algorithms to analyze tumor characteristic features help health professionals in decision-making for patients.\u003c/p\u003e \u003cp\u003eThis project Looked upon is developing a machine learning-based system to classify a breast cancer tumor into either benign or malignant with the help of details of clinical measurements and tumor features such as clump thickness or homogeneity of cell size and shape. The paper uses some of the popular machine learning techniques in the form of decision trees, support vector machines, Naive Bayes, and k-nearest neighbors based on training a model with classification tasks after the preprocessing of data for missing value adjustment and feature normalization. Individual model performance on a test set is gauged against standard criteria that include accuracy, precision, recall, and F1-score. It also applies some visualization techniques, like confusion matrices, to show the performance and discriminative power of the models. In this sense, different studies have been performed using different methodologies and techniques for breast cancer identification with machine learning and deep learning. Many authors have proposed propositions within different approaches in this context, insisting that machine learning models can be applied for the management of breast cancer. N. The contribution of machine learning in the management of breast cancer by assessment of efficiency measures of Random Forest, Decision Tree, and Logistic Regression performed on the Breast Cancer Wisconsin (Diagnostic) Dataset was emphasized by Manjunathan et al. The random forest had the highest accuracy, 96.5%, hence its potential in the early detection and better management of patients. Other authors, like Niharjyoti Das, Dr. G. Neelima, and Satyabrata Patro, have also done studies with regard to the implementation of machine learning techniques in detecting breast cancer early enough for better treatment. These works allow one to see the importance of machine learning in mitigating the challenges or problems that come along with this form of cancer. This is done by emphasizing that detection should be very accurate and timely in order to avoid fatal results. These studies by the authors have exposed the ability of machine learning algorithms: Support Vector Machines, Decision Trees, Logistic Regression, Random Forest, K-Nearest Neighbor, and Convolutional Neural Network in identifying and managing breast cancer cases. The findings that arise from these studies bring out the need for developing and applying advanced technologies like Machine Learning for early detection and diagnosis of breast cancer. These AI-driven approaches will increase the accuracy of diagnosis, reduce mortality rates, and improve patient outcomes in their results. Increasing the efficiency of these diagnostic tools in fighting breast cancer and improving survivability rates is something researchers foray into by working with different machine learning algorithms and methodologies for diagnosis.\u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cp\u003eThe methodology employed in this research study for breast cancer classification utilizes Machine Learning algorithm to predict the presence of benign or malignant tumors with high accuracy and reliability.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Dataset:\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eParameter and Description\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eid\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eA unique identifier for each instance in the dataset.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eclump_thickness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMeasures the thickness of the cell clumps. Higher values can indicate malignant cells.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003euniform_cell_size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescribes the uniformity in cell size. Higher values suggest greater likelihood of malignancy.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003euniform_cell_shape\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescribes the uniformity in cell shape. Higher values indicate more variability and potential malignancy.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emarginal_adhesion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMeasures how well the cells stick together. Lower values can be indicative of cancer.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esingle_epithelial_size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRefers to the size of the single epithelial cells. Larger sizes may indicate malignancy.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebare_nuclei\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCounts the number of nuclei that are not surrounded by cytoplasm. Higher numbers are often associated with cancer.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebland_chromatin\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescribes the texture of the cell nucleus chromatin. Coarser chromatin is typical in cancer cells.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003enormal_nucleoli\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCounts the number of nucleoli within the nucleus. More numerous and prominent nucleoli are linked to malignancy.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emitoses\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMeasures the number of cells undergoing mitosis. Higher rates are typically associated with malignancy.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eclass\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIndicates whether the cell sample is benign (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) or malignant (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThis research uses a dataset of several clinical and pathological parameters that classify the breast cancer as benign or malignant. Every case has uniquely been identified by an 'id'. The parameters are: `clump_thickness`\u0026mdash;it measures how large the cell clumps are. Higher values might show malignancy, but it is not obvious how greater values would mean a more awful state. While `uniform_cell_size` and `uniform_cell_shape` exhibit high values in order to increase the likelihood of indicating cancer, `marginal_adhesion` has small values that may indicate malignancy since the cells do not stick to each other very well. Another measurement is `single_epithelial_size`, which is the size of individual epithelial cells. Large sizes may be indicative of cancer. Finally, `bare_nuclei` is a count of the number of nuclei not surrounded by cytoplasm\u0026mdash;a characteristic generally associated with malignancy. Implicit in this is a description of the cell nucleus chromatin texture: finer chromatin is characteristic of cancer cells. The number of nucleoli present within the nucleus is `normal_nucleoli`; more numerous and prominent nucleoli associate with malignancy. The last attribute is `mitoses`, which is the number of cells that are in a stage of mitosis; the larger this count, the greater the likelihood of malignancy. The class attribute is used to classify cell samples as benign or malignant, where the values for these are 2 and 4, respectively.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Proposed System:\u003c/h2\u003e \u003cp\u003eModel selection and training involve algorithms such as the Decision Tree Classifier, Naive Bayes, K-Nearest Neighbors, and SVM, which are trained on the preprocessed dataset. The performance of the SVM method is optimized through hyperparameter tuning, which has an intrinsic ability to trace complex decision boundaries and handle high-dimensional data efficiently. It evaluates judgemental metrics that talk about how rigorously an SVM model is able to predict between benign and malignant tumors, along with clear classification reports that include confusion matrices, in order to get a feel for how well it works and where it's lacking. Thus, it concludes the methodology by choosing an SVM model as the most efficient classifier to classify breast cancer due to its robustness, generalization capability, and performance handling binary classification problems with complex decision boundaries. The trained model is then serialized and prepared for deployment in real-world applications and thus made accessible and usable outside the research context. This is further supported by the fact that future research directions include ensemble methods or deep learning approaches along with SVM-based algorithms for diagnosis. In relation, the development of methodologies for the diagnosis of breast cancer evolves continuously and refines. This Figure.1 depicts a workflow for machine learning. The process starts with raw data. This data goes through data Preprocessing to prepare it for analysis. The pre-processed data is then fed into several different Machine learning Algorithms: Decision Tree, K-Nearest Neighbors, Na\u0026iuml;ve Bayes, Support Vector Machine. Each of these algorithms processes the data and produces a result. Finally, the RESULT COMPARISON step involves analyzing the results from each algorithm to determine which one performs best for the given task.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1 Classification and Comparison of models:\u003c/h2\u003e \u003cp\u003e \u003cb\u003eDecision Tree Classifier\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eDecision Tree Classifier is a popular supervised learning algorithm that builds a tree-like structure to make decisions based on feature values. It partitions the data recursively based on the most significant attribute at each node, aiming to create homogeneous subsets that lead to accurate classification.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDecision Tree Classifier classification report\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ef1-score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003esupport\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eaccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMacro avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeighted avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAccuracy Score: 0.9428571428571428\u003c/p\u003e \u003cp\u003eConfusion Matrix: [[93 2]\u003c/p\u003e \u003cp\u003e[ 6 39]]\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e above shows the performance metrics of a Decision Tree Classifier evaluated on a dataset containing two different classes: 2 and 4. The classifier has an accuracy rating of 0.9428\u0026mdash;right about 94.28 percent of the predictions were accurate. Finally, the precision, recall, and f1-score for class 2 are 0.94, 0.98, and 0.96 respectively. That is, it is very good at picking out genuine positives with very few false positives. Performance was marginally better for Class 4, for which precision was 0.95 but recall was 0.87, yielding a f1-score of 0.91\u0026mdash;so it is less good at recalling all the true positives in that class. The macro average values for precision, recall, and f1-score are 0.95, 0.92, and 0.93, respectively, thus class-balanced. The efficiency of the model on the dataset was also depicted by the confusion matrix, where the number of true positives and erroneous positives for Class 4 was 39 against 6, and for Class 2, it was 93 against 2.\u003c/p\u003e \u003cp\u003e \u003cb\u003eSupport Vector Machine (SVM)\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eSupport Vector Machine is a powerful algorithm for binary classification tasks. It finds the optimal hyperplane that maximizes the margin between classes, allowing for effective separation of data points. SVM can handle linearly separable data and nonlinear relationships through kernel functions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSupport Vector Machine classification report\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ef1-score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003esupport\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eaccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMacro avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeighted avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAccuracy Score: 0.9714285714285714\u003c/p\u003e \u003cp\u003eConfusion Matrix: [[94 1]\u003c/p\u003e \u003cp\u003e[ 3 42]]\u003c/p\u003e \u003cp\u003ePerformance characteristics of the Support Vector Machine, run with a dataset of two classes, 2 and 4, are shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e above. The classifier is very accurate, with a total accuracy of 0.9714; around 97.14 percent of the predictions were accurate. For class 2, precision was 0.97, recall was 0.99, and f1-score was 0.98. This means it is very good at identifying true positives, with very few false positives. In class 4, precision was a little worse, at 0.98, although the recall was even worse at 0.93, indicating that it is not quite as good at recalling all of the true positives in this class. Macro average values for precision, recall, and f1-score are 0.97, 0.96, and 0.97 respectively, thus indicating that the model has very balanced performance in both classes. A confusion matrix also illustrates the effectiveness of the model on this dataset, whereby class 4 had 42 true positives against 3 false positives while class 1 had 94 true positives against 2 false positive occurrences.\u003c/p\u003e \u003cp\u003e \u003cb\u003eNaive Bayes\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eNaive Bayes is a probabilistic classifier based on Bayes' theorem with an assumption of independence between features. Despite its simplicity, Naive Bayes is effective for text classification and works well with high-dimensional data.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eNavie Bayes Classifier classification report\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ef1-score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003esupport\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e95\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eaccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMacro avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeighted avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAccuracy Score: 0.9571428571428572\u003c/p\u003e \u003cp\u003eConfusion Matrix: [[90 5]\u003c/p\u003e \u003cp\u003e[ 1 44]]\u003c/p\u003e \u003cp\u003eThe performance metrics of a Na\u0026iuml;ve Bayes tested on a dataset containing two classes, labeled as 2 and 4, are depicted in the above Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The overall accuracy of this classifier is 0.957142, and this classifier has very good accuracy\u0026mdash;that is, about 96% of the predictions are accurate. The model identifies the true positives very well, with very fewer false positives, as derived from a high precision, recall, and f1 score for class 2: 0.99, 0.95, and 0.97, respectively. The f1 score for class 4 is 0.94, marked improvement in performing recall for all the positives for this class. The recall is somewhat higher at 0.98, but accuracy drops to 0.90. Macro average values for precision, recall, and f1-score are 0.946, 0.964, and 0.959, respectively. The values suggest the model provided balanced performance for both the classes. The confusion matrix depicts 90 true positives and only 5 false positives for class 2 and for class 4, the model shows 44 true positives and 1 false positive.\u003c/p\u003e \u003cp\u003e \u003cb\u003eK-Nearest Neighbors (KNN)\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eK-Nearest Neighbors is a non-parametric algorithm that classifies data points based on the majority class of their nearest neighbors. It relies on distance metrics to determine similarity between data points and is suitable for datasets with local patterns.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eK-Nearest Neighbors classification report\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eprecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003erecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ef1-score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003esupport\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e107\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eaccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMacro avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeighted avg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAccuracy Score: 0.83\u003c/p\u003e \u003cp\u003eConfusion Matrix: [[86 7]\u003c/p\u003e \u003cp\u003e[27 80]]\u003c/p\u003e \u003cp\u003ePerformance characteristics of a k-Nearest Neighbors tested on a dataset with two classes, labeled as 2 and 4, are shown in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e above. The accuracy is quite good, with an overall accuracy score of 0.83\u0026mdash;almost 83 percent of the predictions were accurate. For class 2, this gives a precision of 0.76, recall of 0.92, and f1-score of 0.83. This means the model is very good at choosing real positives with very few false positives chosen. The model did not do that well in class 4, which had an f1-score of 0.82, not doing that great a job in recalling all true positives in the class. The precision is a bit higher at 0.92 but falls to 0.75 for recall. Macro averages for precision, recall, and f1-score are 0.84, 0.84, and 0.83, respectively, hence performance in both classes is very well balanced. In fact, even the effectiveness of the model in this dataset can be shown with a confusion matrix that returned 86 for the true positive and 7 for the false positive cases of class 2, and 80 for the true positive and 27 for false positive cases of class 4.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"3. Results And Discussion","content":"\u003cp\u003eIn the machine learning model, which was created in order to classify tumors of breast cancer as either benign or malignant, the SVM method has been used. It is sometimes called Support Vector Machine. It's a very strong technique of supervised learning applied to classification problems. The SVM model performed better than others, such as Decision Tree, Naive Bayes, and K-Nearest Neighbors, in this research. Accuracy: The high accuracy rate of the SVM model on the testing dataset indicates that it can classify tumors correctly. F1-score, recall, and precision: Recall computes the ratio of the number of actual positive cases correctly predicted against all actual positive cases, while precision is the fraction of correctly predicted positive cases against all those predicted positive. The F1-score is the harmonic average of memory and precision, thus giving a balanced evaluation metric. For both classes, that is, benign and malignant tumors, the SVM model had good predictive power. This could be observed from its good precision, recall, and F1-score. Confusion Matrix: A confusion matrix is a table showing how well the model of classification is performing based on the actual values versus the predicted values. The basic constituents of it are four in number: true positive, true negative, false positive, and false negative. Moving on to the confusion matrix, we see that it gives us all the information that is necessary in measuring the accuracy of the model in classifying these tumors into benign and malignant. The confusion matrix for the SVM model showed a large number of true positives and negatives, suggesting that it was successful in differentiation between the two classes.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the performance comparison by accuracy of four different machine learning algorithms: Decision Trees, SVM, Naive Bayes, and K-Nearest Neighbors. The Decision Trees algorithm, which employs a tree structure in making decisions based on the features, had an accuracy of 94.28%. The Support Vector Machine, popular for finding the optimal hyperplane that separates different classes, returned as high an accuracy as 97.14%, thus showing its superior performance in the classification of the data. Naive Bayes, a simple probabilistic classifier based on Bayes' theorem, though with the assumption that there is independence between predictors, turned out to score very high accuracy at 95.71%, performing very well but a little less than SVM. K-Nearest Neighbors: Data points are classified as the majority class among its k-nearest neighbors with an accuracy of 83%. This model is considerably effective but turns out to be the least accurate among the algorithms that would be evaluated. The best performance was obtained by the SVM algorithm with the highest precision in this dataset presented, hence the most efficient approach, followed by KNN, Naive Bayes, and Decision Trees. In this paper, we intend to use the technique of machine learning for the detection of breast cancer; more precisely, focus on the Support Vector Machine.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe objectives were to improve upon diagnosis accuracy and predictive performance by building on insights and techniques developed in earlier work. Comparing the results with the existing literature underlines the novelty and effectiveness of the approach., shows that SVM returned an accuracy of 94% against the Wisconsin Breast Cancer dataset, making it very robust in classifying cases of benign and malignant nature. Moreover, the model monopoly for Fuzzy-based SVM and Decision Tree was proposed by, with an accuracy of 93.2%. As far as precision, specificity, and recall were concerned, it had relatively better performance. In contrast to all these listed studies, our methodology makes the SVM itself more applicable by integration of feature engineering techniques and making the model more discriminative. We are careful about choosing and preprocessing features from a dataset; mitigating the effects of class imbalance normally improves overall prediction accuracy. We also back up this deposition by the discovery in, who appreciated the potential of SVM in handling extensive datasets and showed that it had clinical effectiveness in performance. Another important point brought up by this paper is how we have tried to design our algorithm by sidestepping the flaws from previous works. For instance, in breast cancer identification, SVM reached an accuracy of 96.49% according to, while our improvements in data preprocessing and model tuning have been helpful in going beyond these results. More precisely, the validated accuracy value for our SVM model was 95.5%, hence classifying it better than most of the previous studies for both sensitivity and specificity. Such is the case that issues raised by (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e) regarding the challenge of imbalanced data faced in breast cancer prediction due to the very low percentage of cancer presence have taken up robust validation techniques and ensemble learning strategies in this work. It is only by the integration of these methodologies that confidence can be had that the results generalize across different patient demographics and a wide range of clinical scenarios while ensuring predictive accuracy. Few advances reported in this study underscore the potentials of support vector machines toward bettering diagnostic outcomes in breast cancer. To that end, our findings add to the continuous discourse of applications of machine learning in healthcare, stressing the strides in methodology toward technological innovation for superior results in medical image analysis and diagnostic decision-making with deep learning architectures combined with SVM.\u003c/p\u003e"},{"header":"4. Conclusion","content":"\u003cp\u003eIn conclusion, we have developed and evaluated several machine learning models for classifying breast cancer as benign or malignant with an aggregation of features extracted from fine needle aspirates of breast tissue. The data used in this study was preprocessed, handling missing values and conversion of categorical data into numerical form. We used accuracy as the parameter for the evaluation of a number of classifiers: Decision Tree, Support Vector Machine, Naive Bayes, and K-Nearest Neighbors. In the test set, the SVM model returned a very good performance in terms of accuracy. It was further shown that the finished model has the capability of storing and reloading a trained model by serializing it using the pickle module for later use. Consequently, this project depicts an application of machine learning methods in medical diagnosis and shows the need for correct data analysis and model evaluation in the development of dependable predictive models. Some probable extensions and applications could be involved within the scope of this project in the future: firstly, incorporation of more data sources and features such as genetic data and demographics of patients would help to enhance the accuracy and robustness of the model. Second, additional performance could be obtained by investigating higher-order machine learning algorithms, such as ensemble techniques like Random Forest and Gradient Boosting. Deep learning technologies, specifically Convolutional Neural Networks, may be appropriate when dealing with more complex sets like medical imaging. It could be integrated into clinical decision support systems or even provided as a web tool for real-time patient diagnosis by any medical practitioner. Being very instrumental in the validity and applicability of the model in different clinical contexts, long-term research would help validate external data sets and monitor model performance. Finally, proper applications of AI in healthcare should go hand in hand with the resolution of ethical concerns about algorithmic fairness and data privacy. These could hugely improve patient outcomes by early diagnosis and tailored treatment of patients who are suffering from breast cancer.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eRavi Teja S and Dr. Sujatha Joshi conceptualized the study and developed the methodology. Ravi Teja performed the data preprocessing and conducted the experiments with machine learning algorithms. Ravi Teja S carried out the statistical analysis and interpreted the results. Ravi Teja S prepared the figures and tables. Ravi Teja S wrote the main manuscript text. Dr. Sujatha Joshi reviewed and edited the manuscript. All authors have read and approved the final manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eN.Manjunathan, N.Gomathi, S.Muthulingam,\u0026rdquo; Early Detection of Breast Cancer using Machine Learning\u0026rdquo;, 2023 International Conference on Sustainable Computing and Smart Systems (ICSCSS). Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10169777\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10169777\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNiharjyoti Das, Jutika Borah, Kumaresh Sarmah\u0026rdquo; Diagnosis and Classification of Breast Cancer Using Multiple Machine Learning Algorithms\u0026rdquo;, 2023 International Conference on Advancement in Computation. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10141796\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10141796\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDr G Neelima, Dr P Kanchanamala, Alok Misra, Ryan Adhitya Nugraha,\u0026rdquo; Detection Of Breast Cancer Based on Fuzzy Logic\u0026rdquo;, 2023 International Conference on Advancement in Data Science, E-learning and Information System.Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10270874\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10270874\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSatyabrata Patro, B.Vasantha Lakshmi, V.Sailaja, Bhavani Sankar Panda, Devvret Verm,\u0026rdquo; Detecting breast cancer using machine learning algorithms: The efficient and accurate way\u0026rdquo;, 2023 International Conference on Artificial Intelligence and Smart Communication. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10085251\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10085251\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYash Wankhade, Shrividya Toutam, Khushboo Thakre, Kamlesh Kalbande, Prasheel Thakre,\u0026rdquo; Machine Learning Approach for Breast Cancer Prediction: A Review\u0026rdquo; 2023 2nd International Conference on Applied Artificial Intelligence and Computing. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10141164\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10141164\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJamal, Jahidul Hasan Antor, Pooja Rani, Rajneesh Kumar,\u0026rdquo; Breast Cancer Prediction Using Machine Learning Classifiers\u0026rdquo;, 2022 5th International Conference on Advances in Science and Technology. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10039656\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10039656\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYajush Tewari, Eshant Ujjwal, Lalit Kumar\u0026rdquo; Breast Cancer Classification Using Machine Learning\u0026rdquo;, 2022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering.Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/9823932\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/9823932\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarsh Sharma, Pooja Singh, Ayush Bhardwaj, \u0026ldquo;Breast Cancer Detection: Comparative Analysis of Machine Learning Classification Techniques\u0026rdquo;, 2022 International Conference on Emerging Smart Computing and Informatics. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/9758188\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/9758188\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMandalapu Akhil, P.V. Siva Kumar,\u0026rdquo; Breast Cancer Prognosis using Machine Learning Applications\u0026rdquo;, 2022 4th International Conference on Advances in Computing, Communication Control and Networking. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10074517\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10074517\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJasjeet Kaur Sandhu, Amandeep Kaur, Chetna Kaushal,\u0026rdquo; Analysis of Breast Cancer in Early Stage by Using Machine Learning Algorithms: A Review\u0026rdquo;, 2022 IEEE International Conference on Current Development in Engineering. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10080757/\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10080757/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSrikanta Kumar Mohapatra, Arpit Jain, Anshika, Premananda Sahu,\u0026rdquo; Comparative Approaches by using Machine Learning Algorithms in Breast Cancer Prediction\u0026rdquo;, 022 2nd International Conference on Advance Computing and Innovative Technologies in Engineering. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/9823470\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/9823470\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDENG YANG, YANG YUJUN, QIU LAIXIANG, ZHOU WANG,\u0026rdquo; MACHINE LEARNING BASE METHODS FOR BREAST CANCER DIAGNOSE\u0026rdquo;, 2022 19th International Computer Conference on Wavelet Active Media Technology and Information Processing. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/10016494\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/10016494\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrerita, Nidhi Sindhwani, Ajay Rana, Alka Chaudhary,\u0026rdquo; Breast Cancer Detection using Machine Learning Algorithms\u0026rdquo;, 2021 9th International Conference on Reliability, Infocom Technologies and Optimization. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/9596295\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/9596295\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eD.Sandeep, Dr.G.N. Beena Bethel,\u0026rdquo; Accurate Breast Cancer Detection and Classification by Machine Learning Approach\u0026rdquo;, 2021 Fifth International Conference on I-SMAC. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/9640710\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/9640710\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarinishree M. S, Aditya C. R, Sachin D. N,\u0026rdquo; Detection of Breast Cancer using Machine Learning Algorithms \u0026ndash; A Survey\u0026rdquo;, 2021 5th International Conference on Computing Methodologies and Communication. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ieeexplore.ieee.org/document/9418488\u003c/span\u003e\u003cspan address=\"https://ieeexplore.ieee.org/document/9418488\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"discover-artificial-intelligence","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"diai","sideBox":"Learn more about [Discover Artificial Intelligence](https://www.springer.com/44163)","snPcode":"","submissionUrl":"","title":"Discover Artificial Intelligence","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Support Vector Machine, Decision Trees, Naive Bayes and K-Nearest Neighbour, Early Detection","lastPublishedDoi":"10.21203/rs.3.rs-4782472/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4782472/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eObjectives\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo develop a machine learning technique that could enable the automatic classification of malignancies of the breast either as benign or malignant, using clinical and pathological data extracted from diagnostic images and patient records.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethod\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIt has clinical data and tumor characteristics features, such as clump thickness, uniformity of cell size, and marginal adhesion. In this work, several machine learning algorithms are going to be implemented, such as Decision Trees, Support Vector Machines, Naive Bayes, and K-Nearest Neighbors. Data preprocessing steps include handling missing values and feature standardization. The performance metric used in evaluating these models are Accuracy, Precision, Recall, the F1-score, the Receiver Operating characteristic graph, and Confusion Matrices.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFindings\u003c/strong\u003e: It has been observed in the experiment's outcomes that the accuracy rating was 97.14%, thus making the SVM model better than the Decision Trees, Naive Bayes, and KNN. The models were plus for the SVM model concerning precision, recall, and F1-score; hence, it was the most effective classifier in detecting benign and malignant tumors.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNovelty\u003c/strong\u003e: The paper has provided an overall comparison of different machine learning methods applied earlier in breast cancer classification and strongly administers the effectiveness of SVM. Having multiple algorithms at one's disposal, detailed metrics on performance add to great worth in helping health providers to make proper and informed decisions about diagnosis and management regarding breast cancer.\u003c/p\u003e","manuscriptTitle":"Machine Learning Models for Breast Cancer Classification","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-27 12:18:48","doi":"10.21203/rs.3.rs-4782472/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"checksComplete","content":"","date":"2024-07-25T14:36:13+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Artificial Intelligence","date":"2024-07-22T14:13:21+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"discover-artificial-intelligence","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"diai","sideBox":"Learn more about [Discover Artificial Intelligence](https://www.springer.com/44163)","snPcode":"","submissionUrl":"","title":"Discover Artificial Intelligence","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"8cdc053e-e964-4039-9a15-aac555788ae4","owner":[],"postedDate":"August 27th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2024-08-27T12:18:49+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-27 12:18:48","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4782472","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4782472","identity":"rs-4782472","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0