Machine Learning Fake News Classification with Optimal Feature Selection

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Nowadays, information is published in newspapers and social media while transmitted on radio and television about current events and specific fields of interest nationwide and abroad. It becomes difficult to explicit what is real and what is fake due to the explosive growth of online content. As a result, fake news has become epidemic and immensely challenging to analyze fake news to be verified by the producers in the form of data process outlets not to mislead the people. Indeed, it is a big challenge to the government and public to debate the situation depending on case to case. For the purpose several websites were developed for this purpose to classify the news as either real or fake depending on the website logic and algorithm. A mechanism has to be taken on fact-checking rumors and statements, particularly those that get thousands of views and likes before being debunked and refuted by expert sources. Various machine learning techniques have been used to detect and correctly classified of fake news. However, these approaches are restricted in terms of accuracy. This study has applied a Random Forest (RF) classifier to predict fake or real news. For this prpose, twenty-three (23) textual features are extracted from ISOT Fake News Dataset. Four best feature selection techniques like Chi2, Univariate, information gain and Feature importance are used for selecting fourteen best features out of twenty-three. The proposed model and other benchmark techniques are evaluated on the dataset by using best features. Experimental findings show that, the proposed model outperformed state-of-the-art machine learning techniques such as GBM, XGBoost and Ada Boost Regression Model in terms of classification accuracy.
Full text 126,669 characters · extracted from preprint-html · click to expand
Machine Learning Fake News Classification with Optimal Feature Selection | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning Fake News Classification with Optimal Feature Selection Muhammad Fayaz, Atif Khan, Muhammad Bilal, Sanaullah Khan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-835344/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 29 Jan, 2022 Read the published version in Soft Computing → Version 1 posted 5 You are reading this latest preprint version Abstract Nowadays, information is published in newspapers and social media while transmitted on radio and television about current events and specific fields of interest nationwide and abroad. It becomes difficult to explicit what is real and what is fake due to the explosive growth of online content. As a result, fake news has become epidemic and immensely challenging to analyze fake news to be verified by the producers in the form of data process outlets not to mislead the people. Indeed, it is a big challenge to the government and public to debate the situation depending on case to case. For the purpose several websites were developed for this purpose to classify the news as either real or fake depending on the website logic and algorithm. A mechanism has to be taken on fact-checking rumors and statements, particularly those that get thousands of views and likes before being debunked and refuted by expert sources. Various machine learning techniques have been used to detect and correctly classified of fake news. However, these approaches are restricted in terms of accuracy. This study has applied a Random Forest (RF) classifier to predict fake or real news. For this prpose, twenty-three (23) textual features are extracted from ISOT Fake News Dataset. Four best feature selection techniques like Chi2, Univariate, information gain and Feature importance are used for selecting fourteen best features out of twenty-three. The proposed model and other benchmark techniques are evaluated on the dataset by using best features. Experimental findings show that, the proposed model outperformed state-of-the-art machine learning techniques such as GBM, XGBoost and Ada Boost Regression Model in terms of classification accuracy. Geometry Topology Theoretical Computer Science Machine Learning Random Forest Fake news Feature Selection. Figures Figure 1 Figure 2 Figure 3 1. Introduction The internet offers many possibilities along with many challenges when it comes to reporting the news. The number of communication channels is growing over time. In addition to conventional channels such as newspapers and TV channels, news communication channels such as blogs and social networks have arisen since the internet became a spreading source. It has become simpler for customers to receive the latest news on their fingertips. In the current state, these social media sites are highly effective and valuable if this change has a positive aspect on one-hand while negative aspect on the other hand such as fake and inaccurate news because editorial boards do not necessarily determine the trustworthiness of the information posted. In the current state, these social media platforms are useful to share ideas and discuss issues such as governance, education and health. Organizations widely use most sites for their monetary benefits for their objectives. A study placed by Twitter reveals that fake news is 100 times speedily spread rather than real news. The phenomenon of fake news can produce effects that are not significant, and this can lead to results that affect millions of people in certain countries (Vogel and Meghana 2020 ). Such a propanganda and rumors fluctutate stock prices, stock purchases, investment plans and even reaction to natural disasters. The contents of fake news are frame in such a way that it may create mass opinions and fully win over the reader to make them completely confused and their attention divert from real news (Hakak et al. 2021 ). Detection of fake news is a challenging task, as that requires rationalism. Many fact checking websites are deployed to reveal the fake news to counter the growing misinformation. Such websites play a critical role in clarifying false news, but they need time-consuming and expertise. It is quite challenging to detect and analyze the data authenticity(Napoli 2018 ). This study uses benchmark and other machine learning approaches for fake news classification. This research uses Random Forest (RF), as proposed machine learning algorithm, to improve the accuracy of fake news classification. The proposed machine learning model operates as follows: first we extract twenty-three (23) textual features from the ISOT Fake News Dataset publicly available on the ISOT website and describe the news dataset as a features vector. Too many features influence model efficiency and performance, and not all features have the same predictive model contribution. So it is important to strip out non-valuable and less important features to reduce model complexity and increase model accuracy. Therefore, four different feature selection techniques are used, such as Chi-Square, Feature Importance, Information Gain and Univariate, to pick the fourteen (14) best features out of 23 features. Based on the above, we will then abstract the real information from the fake news dataset. For this analysis, the efficiency of the proposed model is contrasted with the benchmark techniques. This study's contributions are as follows: To propose a Model (Random Forest) for the classification of fake and real news. To test the proposed classification model for all textual features (twenty-three) derived from ISOT fake news dataset. To determine the efficacy of the proposed model with respect to the best features collected through four different feature selection techniques (Chi-Square, Univariate, Feature importance, Information Gain). Remaining portions of the paper are designed as follows: Sect. 2 , comprehensive literature is presented. Next Sect. 3 , the proposed machine learning model is illuminated. Section 4 , various experimental findings and a complete discussion is shown. Final Sect. 5 , the conclusion and future recommendation. 2. Related Work The news are very important because it keeps the public informed about activities and events around their premises and beyond their premises. Reports showed that most adults use digital forms such as social media and web/search engines to access their news instead of using traditional media. Fake news detection has got considerably attention (De Choudhury et al. 2014 ). In this section, numerous methods have been suggested to detect fake news in various types of features and datasets. Authors in (Okoro et al. 2018 ) used a Machine-Human (MH) model for detection of fake news on social media. The study (Khan et al. 2021 ) focused on detectition of fabricated opinions and compared the Glove Embedding and Character Embedding features of fake and real news dataset using three datasets. Amongt them, two are standard datasets and one is combination news of distributed topic on social media through Naïve Bayes, CNN, LSTM, Bi-LSTM, C-LSTM, Heterogeneous Graph Neural Network (HAN), Cov-HAS, Char-level C-LSTM model. It is observed that n-gram features show great promising results in fake news on Naive Bayes model which is almost equivalent to the performances of CNN based model. The authors in (Ozbay and Alatas 2020 ) used mixure of text classification techniques and supervised artificial intelligence classifiers. The proposed model was tested on three different real word datasets. The model was evaluated using accuracy, precision, recall and F-measures values. The performance of best mean values was obtained from the Decision Tree Algorithm. Zero, CV parameter selection (CVPS) with 1000 value, seems the best recall metric algorithm. Authors in (Gravanis et al. 2019 ) used an enhanced set of linguistic features for the detection of fake news by evaluated several classification models using five different datasets containing fake and reals news. Adaboost obtained 95% accuracy over all datasets and next is ranking Support Vector Machine (SVM) and Bagging algorithms. The study (Ahmad et al. 2020 ) presented work of machine learning model and ensemble techniques for detection of fake and real social News. Data collected from web contains fake and real news covering different domain. They extracted different textual features from the dataset and used it as input to different machine learning models like Logistic Regression, SVM, MLP, KNN, Random forest(RF) and ensemble models like voting classifier(RF, LR, KNN), voting classifier(LR, LSVM, CART), Bagging classifier (decision tree) and boosting classifier (AdaBoost and XG-Boost), Ensemble model XGBoost performance better than other classifiers and ensemble model in terms of accuracy. The author in (Ahmed et al. 2019 ) have used n-grams and Part of Speech (POS) tagging, they suggested Deep Syntax Analysis using Probabilistic Context-Free Grammars (PCFG). The author in (Ruchansky et al. 2017 ) proposes the CSI hybrid model used for fake news detection. The CSI model is comprised of three modules. The first model captures the pattern of the user's temporal engagement with an article. The second modules capture the characteristic source present in the behavior of users and the third modules are used as integrated of both modules first and second experiment on two datasets, check robustness of CSI model when labeled data is limited. It also inspects suspicious users’ behaviors. The CSI model doest not make assumptions regarding distribution of user behavior specially textual context of the data or the structure of data underlying. In the study of (Mansouri et al. 2020 ), a combined method based on semi-supervised LDA (Linear Discriminant Analysis) and convolutional neural network are used to detect fake news using an unlabeled dataset for the convolutional neural network the unlabeled dataset is labeled. The result of the proposed method of precision is 95.6% and 96.7% recall, which outperforms existing methods for detecting fake news. Another study (Najar et al. 2019 ) Fake news detection using Bayesian interference used Bag of Words using Multinomial Model (MM), Dirichlet Compound Multinomial (DCM) and Deterministic Annealing Expectation-Maximization(EDCM-DAEM) and EDCM-Bayesian, EDCM-Bayesian better accuracy than other classifiers, classification accuracy 87.85 on BS-Detector dataset. In study (Jain et al. 2019 ; Reis et al. 2019 ) used different textual features like language features (syntax) such as n-gram and part of speech tagging, lexical features (character and word-level signals), psycholinguistic features, semantic features and subjectivity and sentiment scores of a text using classification of K-Nearest neighbors (KNN), Naïve Bayes(NB), Random forests(RF), Support Vector Machine (SVM) with RBF kernel (SVM), and XGBoost (XGB). Random forest and XGB performed best using handcraft features, web-based networking media. In study [14] using Naïve Bayes classifier, SVM with comparison Naïve Bayes and CNN. Results show that Naïve Bayes, SVM, NLP are performed better than other machine classifier. The accuracy of proposed model 93.50% at the other machine learning model. Another study (Faustini and Covões 2019) conduct on fake social media news used three datasets of social media (Twitter, WhatsApp and Fake BR Corpus), by extracting of fourteen textual features such as proportion of uppercase characters, exclamation marks, question marks, number of unique words, number of sentence, number of characters, words per sentence, proportion of adjective, adverb, nouns, sentiment of message, proportion of swear words and proportion of spell errors as features for classifiers. In study (Hlaing and Kham 2020 ) presents multidimensional fake news (news content, social engagement and news stance) used synonym-based features using three different classifier Decision Tree classifiers, AdaBoost classifier and Random forest classifier for detection of fake news. Experimental result show that Random Forest perform better than other two classifiers on social media dataset. The study (Mahir et al. 2019 ) reported that SVM performed better than other classifiers including Naïve Bayesian, RNN/LSTM, Logistic Regression in recognizing fake news extracted from twitter. The study (Al-Ash et al. 2019 ) used Indonesian news dataset consist of fake and real news documents to show the classification performance of Random Forest, SVM and Naïve Bayesian Classifiers over this dataset with associate classification approach. In this study (Katsaros et al. 2019 ) eight models were evaluated for classification purpose. These models include Linear Regres. sion, SVM, MLP, Gaussian and Multinomial naïve Bayes, Random Forests, Decision Trees and CNN on three publicly available datasets. The result showed that the CNN is the best performing algorithm. The study of (Choudhary et al. 2021 ) proposed a deep learning architecture called BerConvoNet for classification of Fake news and Real news with marginal error. The proposed architecture was composed of two main blocks, a New Embedding Block (NEB)and a Multi-scale Features Block (MSFB). The NEB used BERT for extracting word embeddings from news articles and then fed it to MSFB as input. In the study (Vogel and Meghana 2020 ) reported that SVM achieved the highest accuracy of 92% in classification of fake news. He used hand crafted features extracted from news dataset like total word(tokens), Unique words, Unique words (types), type/token ratio, Number of sentences, average sentence length, Number of Characters, Average word length, nouns, prepositions, adjectives.The classification models include XG Boost, Random Forest, Naïve Bayesian, KNN, Decision Tree and SVM were used for classification of fake news. 3. Proposed Methodology The methodology section presents the architecture of proposed model for classification of fake news as shown in Figure-1. The methodology section is consists of three Main phases: 1st phase: pre-processing (Sentence segmentation, tokenization, stopword removal and word stemming), 2nd phases include features extraction and best features selection, best feature selection using famous feature selection techniques (Chi-Square, Feature Importance, Information Gain and Univariate) and final phase classification of fake news using the proposed model and other machine learning models. 3.1. Dataset The proposed machine learning model is tested on the ISOT News Dataset, a publicly available dataset containing false and real news. This dataset is commonly used in the false news identification problem. A total of 44,919 fake and real news was used in this research, including 23502 fake news and 2147 real news assessments using multiple machine learning models. Regarding performance metrics, i.e. classification precision, for the assignment of fake news classification, we also tested the suggested Machine Learning classifier with the benchmark models. 3.2. Pre-processing: In order to avoid overfitting, the data is preprocessed before fetching it to the Natural Languange Processing (NLP) system. The preprocessing involves various steps like sentence segmentation, tokenization, stopwords removal and word stemming, as discussed below in detail: 3.2.1. Sentence Segmentation: Sentence segmentation is establish text borders and break the text into sentences. Exclamation (!), interrogation (?), and utter stop (.) signs are widely used as markers to segment the paragraph into sentences. 3.2.2. Tokenization At this stage, the phrases and sentences are divided into separate words by dividing them into white spaces such as tabs, blanks, and signs of punctuation, i.e. dot (.), comma (,), semicolon (;), colon (:), etc. These are the key indications for dividing the sentences into tokens. 3.2.3. Stopwords removal. Words which have occurred repeatedly are called stopwords in a sentence. These consist of prepositions (in, on, at, etc.), conjunctions (and thus, too, etc.), articles (a, an, a), etc. These words have little meaning in text documents and are more weighted, and removing them will help improve the system's performance. 3.2.3. Word Stemming: Stemming is used to bring the word to its basic form. Word stemming plays a significant role in preprocessing. In order to normalize the word token to a standard form, this step changes the derived words to its base or stem word. The famous stemming algorithm, Porter’s stemming (Porter, 1980), is adopted to remove suffixes like –ing, -es, -ers from the text words. For example, the words ‘looking’ and ‘looks’ will be modified to its base type ‘look’ after stemming. 3.3 Features Extraction: In text classification problems, features play a major role. This step aims to mine ISOT fake news Dataset features for the problem of text classification. In this study, we extracted twenty-three (23) features form ISOT fake news dataset. Almost all of these features are textual and can be accurately extracted through text as seen in Table 1 . Table 1 List of all features extracted from ISOT News dataset S.No Features Description 1 Counts-words Count total numbers of words 2 Upper-case-character Count total number of upper case characters 3 Lowe-case-character Count total number of Lower case character 4 Character-count Count total number of characters 5 Count-sentence Count total number of sentences in fake News 6 Automated-Readability-Index Automated Readability Index is readability test 7 Coleman-Liau-index Coleman–Liau index is text readability test 8 Flesch-reading-ease Flesch reading ease is text readability test 9 Flesch-Kincaid-grade Flesch Kincaid grade level is text readability test 10 Dale-Chall-formula Dale–Chall formula is text readability test 11 Topic Topic related words 12 Count-spaces Count number of spaces 13 Part-of-Speech Count numbers of Part of speech 14 Principal-Component-Analysis Principal Component Analysis 15 Gunning-Fox-index Gunning fog index is readability test 16 SMOG-formula Mc-Laughlin's SMOG formula is readability test 17 Sentiment-analysis Sentiment analysis show text positivity and negativity 18 Bag-of-word Bag of words is a text representation representing the presence of words 19 Tf-idf (unigram) TF-IDF with unigram 20 Tf-idf (bi- gram) TF-IDF with Bi-gram 21 Tf-idf (tri- gram) TF-IDF with Tri-Gram 22 Stopwords Count total numbers of stop words 23 Negative-words Count total numbers of negative word 3.3.1 Features Selection It is normally not good to use all twenty-three (23) features to classify the ISOT News dataset as fake and real News. All features do not have the same significance and weight when developing a consistent and effective statistical model. Some features are useful and add more to the model prediction and play a vital role in classification accuracy, while others are less valuable and have a less significant impact on the performance of the model. In addition, the appropriate and useful features eliminate over-fitting, increase precision and reduce the predictive model training time. We used the four best features selection techniques to resolve this issue, which are Chi2, Univariate, information gain and Feature importance to decrease the space size of the features to achieve optimum features and significant features. In Table 2 , Column-2 shows fourteen (14) most important features for the News dataset were chosen by the Chi-square technique and same numbers of feature selected using Univariate feature selection techniques as seen in Column-3 of Table 2 . Similarly, Column-4 picked the fourteen best features from the same dataset using feature importance and Column-5 shows the fourteen best features selected using information gain. Next section evaluated performance of proposed model and other machine learning model on all textual features and best features selected by best features selection techniques as mentioned in Table 1 and Column 2–5 of Table 2 . Table 2 List of fourteen best features chosen by different Features Selection Techniques for ISOT News dataset S.No Features Selected by Chi2 Univariate Feature Importance Information gain 1 Counts-words Counts-words Counts-words Count-unique-words 2 Upper-case-character Upper-case-character Upper-case-character Upper case character 3 Lowe-case-character Lowe-case-character Count-sentence Count sentence 4 Character-count Character-count Flesch-Kincaid-grade-level Flesch Kincaid grade level 5 Count-sentence Count-sentence Flesch-reading-ease Flesch reading ease 6 Automated-readability-index Automated-readability-index Dale-Chall-formula Dale–Chall formula 7 Coleman-liau-index Coleman-liau-index Gunning-fog-index Gunning fog index is readability 8 Flesch-kincaid-grade-level Flesch-kincaid-grade-level Somgs-Formula SMOG formula 9 Flesch-reading-ease Dale–Chall-formula Topic Count lowercase characters 10 Principal-component-analysis Topic Tf-idf (unigram) Tf_idf (unigram) 11 Tf-idf (unigram) Count-spaces Stopwords Tf-idf(bigram) 12 Tf-idf (bi-gram) Part-of-Speech Negative-words Tf-idf(trigram) 13 Tf-idf (tri- gram) Tf-idf (unigram) Part-of-Speech Stopwords 14 Dale–chall-formula Sentiment-analysis Sentiment-analysis Negative words 3.4 Classification Model for Fake News dataset This section aims to classify ISOT news as fake and real news using Random Forest and other Machine Learning Model. We train and evaluate the classifiers initially using 10 fold-cross validations on the ISOT False News dataset to validate the impact of the individual models and the proposed model on all textual features and best features selected by features selection techniques. Random forest is a a traditional machine learning algorithm which is used for classification as well as regression problems by using ensembling of many Decision Trees to solve a complex problems. It uses bagging and bootstrap methods for the prediction of model accuracy. The prediction of each decision trees combines for final prediction using a majority of vote as shown in Fig. 2 . For Decision Trees, Gini Impurity and Entropy are calculated by using following Eq. 1–2, respectively. $$\text{G}\text{i}\text{n}\text{i} \text{I}\text{m}\text{p}\text{u}\text{r}\text{i}\text{t}\text{y}= {\sum }_{\text{i}=1}^{\text{c}}\text{f}\text{i}\left(1-\text{f}\text{i}\right) \left(1\right)$$ $$\text{E}\text{n}\text{t}\text{r}\text{o}\text{p}\text{y}= {\sum }_{\text{i}=1}^{\text{c}}-\text{f}\text{log}\left(\text{f}\text{i}\right) \left(2\right)$$ 4. Experimental Setting The proposed machine learning model is tested on the ISOT News Dataset containing a total of 44,919 fake and real news including 23502 fake news and 2147 real news. The dataset is pre-processed by breaking the news text into sentences. The sentences are tokenized into terms and stopwords are removed. Initially, twenty three (23) textual features are selected from the ISOT News dataset for fake news classification. The proposed model and other machine learning models are evaluated on all twenty three features. As all features do not have the same importance in creating a consistent and accurate predictive model. Some features are meaningful and contribute more to model accuracy, while others are less important and adversely affect the model's performance. In addition, the appropriate and useful features eliminate over-fitting, increase precision and reduce the predictive model training time. We used four best features selection technique to resolve this issue, which are Chi2, Univariate, information gain and Feature importance are used for selecting fourteen best features out of twenty-three and then evaluated proposed model and benchmark techniques in terms of accuracy which is used as performance metrics. 5. Results And Discussions In the first step, the performance of individual models and other machine learning models are evaluated by using all twenty three (23) textual features extracted from ISOT News dataset. The results of this experiment is recorded in Table 3 . It is revealed from the results that the proposed model has the highest score of 97.25% compared to other classifiers on all features. Table 3 Results of classification using the all textual features Classifier Accuracy MLP 96.67 Logistic regression 45.71 Naïve Bayes (G) 45.58 Nave Bayes(M) 66.39 Naïve Bayes(B) 78.17 Decision Tree 92.28 KNN 95.25 Random Forest 97.25 Gradient Boost 95.88 Extra Gradient Boost 94.97 Ada Boost 88.65 In second step, top 14 best features are selected by using Chi-squar features selection technique.All models including proposed model was evaluated using best selected features and results are recorded in Table 4 . Experimental results show that proposed model achieved 97.33% accuracy and performed better than individual models on fourteen (14) best features for the task of fake news classification on ISOT dataset. Table 4 Results of classification using the top 14 features selected by Chi2 Classifier Accuracy MLP 92.64 Logistic regression 45.54 Naïve Bayes (G) 65.87 Nave Bayes(M) 66.36 Naïve Bayes(B) 70.17 Decision Tree 93.07 KNN 95.25 Random Forest 97.33 Gradient Boost 96.27 Extra Gradient Boost 95.73 Ada Boost 86.70 The above results shows that Random forest attained the highest accuracy score of 97.33%, Gradient Boost obtained the second highest accuracy of 96.27 percent, Extra Gradient Boost accuracy is 95.73 percent, KNN accuracy is 95.25 percent, Decision Tree accuracy is 93.07, MLP accuracy is 92.64 percent and Logistic Regression got the lowest accuracy of 45.54 percent. In third step, top 14 best features are selected by using the Univariate feature selection technqiue. All models including proposed model are tested using 14 best features and the results are recorded in Table 5 . The results demonstrate that the proposed model performed better than individual other models by attaining accuracy score of 97.27%. on best features selected using Univariate features selection technique. Table 5 Results of classification using the top 14 features selected by Univariate Classifier Accuracy MLP 87.96 Logistic regression 45.90 Naïve Bayes (G) 43.48 Nave Bayes(M) 66.28 Naïve Bayes(B) 78.17 Decision Tree 85.16 KNN 89.99 Random Forest 97.27 Gradient Boost 86.69 Extra Gradient Boost 86.40 Ada Boost 84.37 Referring to the results presented in above table, the fourteen best features selected using Univariate feature selection techniques, The proposed model attained highest accuracy of 97.27%, KNN achieved the second highest accuracy of 89.99%, MLP accuracy is 87.96%, Gradient Boost accuracy is 86.69 percent. In this case, Random Forest again remained on top in term of accuracy. In fourth step, fourteen (14) best features are selected by using the feature importance technique. All models alongwith proposed model are evaluated on top 14 best features and results are stored in Table 6 . The results show that proposed model performance is better than other classifier. The accuracy of proposed model is 96.60%. Table 6 Results of classification using the top 14 features selected by feature importance Classifier Accuracy MLP 96.30 Logistic regression 48.90 Naïve Bayes (G) 48.48 Nave Bayes(M) 70.28 Naïve Bayes(B) 79.17 Decision Tree 92.53 KNN 95.43 Random Forest 96.60 Gradient Boost 94.89 Extra Gradient Boost 94.85 Ada Boost 88.78 With reference to the results given in Table 6 , the fourteen best features selected using feature importance features selection technique, our proposed model performed better than others by securing highest accuracy of 96.60%. The MLP got the second highest accuracy of 96.30%, KNN accuracy is 95.43%, Gradient Boost accuracy is 94.89 percent. In fifth step,the best features are selected using information gain method of feature selection. All models including our proposed model are evaluated on top 14 best features selected from ISOT dataset and the result in each case are recorded in Table 7 , The proposed model achieved the highest accuracy of 96.42%, MLP achieved the second highest accuracy of 96.02%, KNN accuracy is 95.91%, Gradient Boost accuracy is 94.18 percent, and RF achieved the highest accuracy for individual models and Naive Bayes got lowest accuracy. Table 7 Results of classification using the top 14 features selected by information gain Classifier Accuracy MLP 96.02 Logistic regression 49.90 Naïve Bayes (G) 56.48 Nave Bayes(M) 65.28 Naïve Bayes(B) 79.17 Decision Tree 92.36 KNN 95.91 Random Forest 96.42 Gradient Boost 94.18 Extra Gradient Boost 94.10 Ada Boost 88.50 The results in Fig. 3 show that the accuracy of classifiers moved up and down while using all features, best selected features by Chi2,Univariate, Features importance and information gain. Table 8 Comparative Results of all classifications over all features and top 14 features selected by Four Features Selection Techniques Classifier All features Chi2 Univariate Feature Importance Information gain MLP 96.67 92.64 87.96 96.30 96.02 Logistic regression 45.71 45.54 45.90 48.90 49.90 Naïve Bayes (G) 45.58 65.87 43.48 48.48 56.48 Nave Bayes(M) 66.39 66.36 66.28 70.28 65.28 Naïve Bayes(B) 78.17 70.17 78.17 79.17 79.17 Decision Tree 92.28 93.07 85.16 92.53 92.36 KNN 95.25 95.25 89.99 95.43 95.91 Random Forest 97.25 97.33 97.27 96.60 96.42 Gradient Boost 95.88 96.27 86.69 94.89 94.18 Extra Gradient Boost 94.97 95.73 86.40 94.85 94.10 Ada Boost 88.65 86.70 84.37 88.78 88.50 From experimental results shown in Fig. 3 and Table 8 , the following conclusion are drawn: The accuracy of the proposed model (Random Forest) improved with best features selected by using Chi-square and Univariate. However, the accuracy of proposed model decreased when features selection are perfomed using future importance and information gain features selection techniques. The accuracy of boosting technique like XGBoost, GBM and Ada Boost does not improved as a whole on best features . Perfomance of Gaussian Naïve Bayesian improved significantly on best features selected by Chi-squar as compared to other features selection techniques. Overall, the classification accuracy of the proposed model is superior than all individuals’ models as well as other boosting approaches. 5. Conclusion And Future Work Online fake news detection is a challenging task in the area of text classification. Many attempts have been made by various researchers to address this issue. This study proposed Random Forest as machine learning classifier to classify the news as fake news and real news. For this purpose, twenty three (23) features were extracted from the text of the dataset. Four feature selection techniques like chi-square, Univariate, features importance and information gain, were used to select fourteen (14) best features out of the twenty three (23) extracted features. Proposed model as well as other models were used for classification of fake news and real news using fourteen (14) best features. The experimental results show that proposed model outperformed all other classifiers in term of better classification accuracy in fake news prediction. In future, Deep Ensembling models may be used for fake news detection. 6. Declarations Conflict of interest The authors declare that they have no conflict of interest. Acknowledgements The authors did not receive support from any organization for the submitted work. 7. References Ahmad I, Yousaf M, Yousaf S, Ahmad MO (2020) Fake news detection using machine learning ensemble methods Complexity 2020 Ahmed S, Hinkelmann K, Corradini F Combining machine learning with knowledge engineering to detect fake news in social networks-a survey. In: Proceedings of the AAAI 2019 Spring Symposium, 2019. p 8 Al-Ash HS, Putri MF, Mursanto P, Bustamam A Ensemble learning approach on indonesian fake news classification. In: 2019 3rd International Conference on Informatics and Computational Sciences (ICICoS), 2019. IEEE, pp 1-6 Choudhary M, Chouhan SS, Pilli ES, Vipparthi SK (2021) BerConvoNet: A deep learning framework for fake news classification Applied Soft Computing 110:107614 De Choudhury M, Morris MR, White RW Seeking and sharing health information online: comparing search engines and social media. In: Proceedings of the SIGCHI conference on human factors in computing systems, 2014. pp 1365-1376 Faustini P, Covões T Fake news detection using one-class classification. In: 2019 8th Brazilian Conference on Intelligent Systems (BRACIS), 2019. IEEE, pp 592-597 Gravanis G, Vakali A, Diamantaras K, Karadais P (2019) Behind the cues: A benchmarking study for fake news detection Expert Systems with Applications 128:201-213 Hakak S, Alazab M, Khan S, Gadekallu TR, Maddikunta PKR, Khan WZ (2021) An ensemble machine learning approach through effective feature extraction to classify fake news Future Generation Computer Systems 117:47-58 Hlaing MMM, Kham NSM Defining news authenticity on social media using machine learning approach. In: 2020 IEEE Conference on Computer Applications (ICCA), 2020. IEEE, pp 1-6 Jain A, Shakya A, Khatter H, Gupta AK A smart system for fake news detection using machine learning. In: 2019 International Conference on Issues and Challenges in Intelligent Computing Techniques (ICICT), 2019. IEEE, pp 1-4 Katsaros D, Stavropoulos G, Papakostas D Which machine learning paradigm for fake news detection? In: 2019 IEEE/WIC/ACM International Conference on Web Intelligence (WI), 2019. IEEE, pp 383-387 Khan JY, Khondaker MTI, Afroz S, Uddin G, Iqbal A (2021) A benchmark study of machine learning models for online fake news detection Machine Learning with Applications 4:100032 Mahir EM, Akhter S, Huq MR Detecting fake news using machine learning and deep learning algorithms. In: 2019 7th International Conference on Smart Computing & Communications (ICSCC), 2019. IEEE, pp 1-5 Mansouri R, Naderan-Tahan M, Rashti MJ A Semi-supervised Learning Method for Fake News Detection in Social Media. In: 2020 28th Iranian Conference on Electrical Engineering (ICEE), 2020. IEEE, pp 1-5 Najar F, Zamzami N, Bouguila N Fake news detection using bayesian inference. In: 2019 IEEE 20th International Conference on Information Reuse and Integration for Data Science (IRI), 2019. IEEE, pp 389-394 Napoli PM (2018) What if more speech is no longer the solution: First Amendment theory meets fake news and the filter bubble Fed Comm LJ 70:55 Okoro E, Abara B, Umagba A, Ajonye A, Isa Z (2018) A hybrid approach to fake news detection on social media Nigerian Journal of Technology 37:454-462 Ozbay FA, Alatas B (2020) Fake news detection within online social media using supervised artificial intelligence algorithms Physica A: Statistical Mechanics and its Applications 540:123174 Reis JC, Correia A, Murai F, Veloso A, Benevenuto F (2019) Supervised learning for fake news detection IEEE Intelligent Systems 34:76-81 Ruchansky N, Seo S, Liu Y Csi: A hybrid deep model for fake news detection. In: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, 2017. pp 797-806 Vogel I, Meghana M Detecting Fake News Spreaders on Twitter from a Multilingual Perspective. In: 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA), 2020. IEEE, pp 599-606 Cite Share Download PDF Status: Published Journal Publication published 29 Jan, 2022 Read the published version in Soft Computing → Version 1 posted Editorial decision: Major Revision 31 Oct, 2021 Reviews received at journal 15 Sep, 2021 Reviewers invited by journal 15 Sep, 2021 Editor assigned by journal 28 Aug, 2021 First submitted to journal 21 Aug, 2021 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-835344","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":52526776,"identity":"246cb4fb-393b-43e5-ba16-8b963d59b406","order_by":0,"name":"Muhammad Fayaz","email":"","orcid":"","institution":"University of Peshawar","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"","lastName":"Fayaz","suffix":""},{"id":52526777,"identity":"0cea345e-ebc7-4a57-a2ee-0e5592e0063b","order_by":1,"name":"Atif Khan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4klEQVRIiWNgGAWjYFACxoaPjQ0MDPzMYB6IZCOgg42xcSZIi2Qz8VoYGMFaDA4Qq4V/fnNj48wddnnGx5mfPWCosE5skG5LwKtF4hhjY+PGM8nFZofZzA0YzqQnNsgcO4DfmmOM7Q8ftjEnbjvMwybB2HY4sUEivQGvDnmQLQ/b6hM3N4O0/CNCiwHYYUDDNzCDtDSAtKThd5jhsUSg99uOJ84A+SXhWLpxm8yxBLxa5A4ff9jY21ad2N9/+NmDDzXWsv3SbQZ4tSADNgaQ8WwSRGuARyEpWkbBKBgFo2BEAADIlkur377C/QAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-3628-0262","institution":"Islamia College Peshawar","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Atif","middleName":"","lastName":"Khan","suffix":""},{"id":52526778,"identity":"85f4a31c-c218-4dd5-bc8d-83e0ccf72e88","order_by":2,"name":"Muhammad Bilal","email":"","orcid":"","institution":"Islamia College Peshawar","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"","lastName":"Bilal","suffix":""},{"id":52526779,"identity":"daf7e75f-2df9-4bd1-a223-b87c30482cb8","order_by":3,"name":"Sanaullah Khan","email":"","orcid":"","institution":"Kohat University of Science and Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sanaullah","middleName":"","lastName":"Khan","suffix":""}],"badges":[],"createdAt":"2021-08-22 04:29:50","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-835344/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-835344/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00500-022-06773-x","type":"published","date":"2022-01-29T13:26:32+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":13794923,"identity":"5b149017-9880-4bea-b686-c2a06c02643a","added_by":"auto","created_at":"2021-09-20 20:22:09","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":33015,"visible":true,"origin":"","legend":"Proposed approach for fake news classification","description":"","filename":"fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-835344/v1/d995389c4bbed58c33296857.png"},{"id":13795623,"identity":"5656ec2d-c036-43d1-b2e2-36cb086a3aae","added_by":"auto","created_at":"2021-09-20 20:25:09","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":18845,"visible":true,"origin":"","legend":"Simple Random Forest classifier structure","description":"","filename":"fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-835344/v1/eb540885288705db12b5e599.png"},{"id":13794924,"identity":"d7697100-82c7-4a57-ae14-2537bb2122bc","added_by":"auto","created_at":"2021-09-20 20:22:09","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":59924,"visible":true,"origin":"","legend":"Accuracy of classifiers on all textual features and top 14 features collection obtained for the ISOT News dataset using four feature selection techniques (Chi2, Univariate, feature importance and information gain).","description":"","filename":"fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-835344/v1/526e86e3581d2db5268ed215.png"},{"id":18161263,"identity":"0c5a3732-4182-4594-a314-a483e53b8d60","added_by":"auto","created_at":"2022-02-12 13:27:53","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":421425,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-835344/v1/282c86b2-0013-4bac-815e-4d084eed7c43.pdf"}],"financialInterests":"","formattedTitle":"\u003cp\u003eMachine Learning Fake News Classification with Optimal Feature Selection\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe internet offers many possibilities along with many challenges when it comes to reporting the news. The number of communication channels is growing over time. In addition to conventional channels such as newspapers and TV channels, news communication channels such as blogs and social networks have arisen since the internet became a spreading source. It has become simpler for customers to receive the latest news on their fingertips. In the current state, these social media sites are highly effective and valuable if this change has a positive aspect on one-hand while negative aspect on the other hand such as fake and inaccurate news because editorial boards do not necessarily determine the trustworthiness of the information posted. In the current state, these social media platforms are useful to share ideas and discuss issues such as governance, education and health. Organizations widely use most sites for their monetary benefits for their objectives.\u003c/p\u003e \u003cp\u003eA study placed by Twitter reveals that fake news is 100 times speedily spread rather than real news. The phenomenon of fake news can produce effects that are not significant, and this can lead to results that affect millions of people in certain countries (Vogel and Meghana \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Such a propanganda and rumors fluctutate stock prices, stock purchases, investment plans and even reaction to natural disasters. The contents of fake news are frame in such a way that it may create mass opinions and fully win over the reader to make them completely confused and their attention divert from real news (Hakak et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Detection of fake news is a challenging task, as that requires rationalism. Many fact checking websites are deployed to reveal the fake news to counter the growing misinformation. Such websites play a critical role in clarifying false news, but they need time-consuming and expertise. It is quite challenging to detect and analyze the data authenticity(Napoli \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThis study uses benchmark and other machine learning approaches for fake news classification. This research uses Random Forest (RF), as proposed machine learning algorithm, to improve the accuracy of fake news classification. The proposed machine learning model operates as follows: first we extract twenty-three (23) textual features from the ISOT Fake News Dataset publicly available on the ISOT website and describe the news dataset as a features vector. Too many features influence model efficiency and performance, and not all features have the same predictive model contribution. So it is important to strip out non-valuable and less important features to reduce model complexity and increase model accuracy. Therefore, four different feature selection techniques are used, such as Chi-Square, Feature Importance, Information Gain and Univariate, to pick the fourteen (14) best features out of 23 features. Based on the above, we will then abstract the real information from the fake news dataset. For this analysis, the efficiency of the proposed model is contrasted with the benchmark techniques.\u003c/p\u003e \u003cp\u003eThis study's contributions are as follows:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eTo propose a Model (Random Forest) for the classification of fake and real news.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTo test the proposed classification model for all textual features (twenty-three) derived from ISOT fake news dataset.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTo determine the efficacy of the proposed model with respect to the best features collected through four different feature selection techniques (Chi-Square, Univariate, Feature importance, Information Gain).\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eRemaining portions of the paper are designed as follows: Sect.\u0026nbsp;\u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, comprehensive literature is presented. Next Sect.\u0026nbsp;\u003cspan refid=\"Sec3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, the proposed machine learning model is illuminated. Section \u003cspan refid=\"Sec13\" class=\"InternalRef\"\u003e4\u003c/span\u003e, various experimental findings and a complete discussion is shown. Final Sect.\u0026nbsp;\u003cspan refid=\"Sec14\" class=\"InternalRef\"\u003e5\u003c/span\u003e, the conclusion and future recommendation.\u003c/p\u003e"},{"header":"2. Related Work","content":"\u003cp\u003eThe news are very important because it keeps the public informed about activities and events around their premises and beyond their premises. Reports showed that most adults use digital forms such as social media and web/search engines to access their news instead of using traditional media. Fake news detection has got considerably attention (De Choudhury et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). In this section, numerous methods have been suggested to detect fake news in various types of features and datasets. Authors in (Okoro et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) used a Machine-Human (MH) model for detection of fake news on social media. The study (Khan et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) focused on detectition of fabricated opinions and compared the Glove Embedding and Character Embedding features of fake and real news dataset using three datasets. Amongt them, two are standard datasets and one is combination news of distributed topic on social media through Na\u0026iuml;ve Bayes, CNN, LSTM, Bi-LSTM, C-LSTM, Heterogeneous Graph Neural Network (HAN), Cov-HAS, Char-level C-LSTM model. It is observed that n-gram features show great promising results in fake news on Naive Bayes model which is almost equivalent to the performances of CNN based model. The authors in (Ozbay and Alatas \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) used mixure of text classification techniques and supervised artificial intelligence classifiers. The proposed model was tested on three different real word datasets. The model was evaluated using accuracy, precision, recall and F-measures values. The performance of best mean values was obtained from the Decision Tree Algorithm. Zero, CV parameter selection (CVPS) with 1000 value, seems the best recall metric algorithm. Authors in (Gravanis et al. \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) used an enhanced set of linguistic features for the detection of fake news by evaluated several classification models using five different datasets containing fake and reals news. Adaboost obtained 95% accuracy over all datasets and next is ranking Support Vector Machine (SVM) and Bagging algorithms.\u003c/p\u003e \u003cp\u003eThe study (Ahmad et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) presented work of machine learning model and ensemble techniques for detection of fake and real social News. Data collected from web contains fake and real news covering different domain. They extracted different textual features from the dataset and used it as input to different machine learning models like Logistic Regression, SVM, MLP, KNN, Random forest(RF) and ensemble models like voting classifier(RF, LR, KNN), voting classifier(LR, LSVM, CART), Bagging classifier (decision tree) and boosting classifier (AdaBoost and XG-Boost), Ensemble model XGBoost performance better than other classifiers and ensemble model in terms of accuracy. The author in (Ahmed et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) have used n-grams and Part of Speech (POS) tagging, they suggested Deep Syntax Analysis using Probabilistic Context-Free Grammars (PCFG). The author in (Ruchansky et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2017\u003c/span\u003e) proposes the CSI hybrid model used for fake news detection. The CSI model is comprised of three modules. The first model captures the pattern of the user's temporal engagement with an article. The second modules capture the characteristic source present in the behavior of users and the third modules are used as integrated of both modules first and second experiment on two datasets, check robustness of CSI model when labeled data is limited. It also inspects suspicious users\u0026rsquo; behaviors. The CSI model doest not make assumptions regarding distribution of user behavior specially textual context of the data or the structure of data underlying. In the study of (Mansouri et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), a combined method based on semi-supervised LDA (Linear Discriminant Analysis) and convolutional neural network are used to detect fake news using an unlabeled dataset for the convolutional neural network the unlabeled dataset is labeled. The result of the proposed method of precision is 95.6% and 96.7% recall, which outperforms existing methods for detecting fake news.\u003c/p\u003e \u003cp\u003eAnother study (Najar et al. \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) Fake news detection using Bayesian interference used Bag of Words using Multinomial Model (MM), Dirichlet Compound Multinomial (DCM) and Deterministic Annealing Expectation-Maximization(EDCM-DAEM) and EDCM-Bayesian, EDCM-Bayesian better accuracy than other classifiers, classification accuracy 87.85 on BS-Detector dataset. In study (Jain et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Reis et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) used different textual features like language features (syntax) such as n-gram and part of speech tagging, lexical features (character and word-level signals), psycholinguistic features, semantic features and subjectivity and sentiment scores of a text using classification of K-Nearest neighbors (KNN), Na\u0026iuml;ve Bayes(NB), Random forests(RF), Support Vector Machine (SVM) with RBF kernel (SVM), and XGBoost (XGB). Random forest and XGB performed best using handcraft features, web-based networking media. In study [14] using Na\u0026iuml;ve Bayes classifier, SVM with comparison Na\u0026iuml;ve Bayes and CNN. Results show that Na\u0026iuml;ve Bayes, SVM, NLP are performed better than other machine classifier.\u003c/p\u003e \u003cp\u003eThe accuracy of proposed model 93.50% at the other machine learning model.\u003c/p\u003e \u003cp\u003eAnother study (Faustini and Cov\u0026otilde;es 2019) conduct on fake social media news used three datasets of social media (Twitter, WhatsApp and Fake BR Corpus), by extracting of fourteen textual features such as proportion of uppercase characters, exclamation marks, question marks, number of unique words, number of sentence, number of characters, words per sentence, proportion of adjective, adverb, nouns, sentiment of message, proportion of swear words and proportion of spell errors as features for classifiers. In study (Hlaing and Kham \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) presents multidimensional fake news (news content, social engagement and news stance) used synonym-based features using three different classifier Decision Tree classifiers, AdaBoost classifier and Random forest classifier for detection of fake news. Experimental result show that Random Forest perform better than other two classifiers on social media dataset.\u003c/p\u003e \u003cp\u003eThe study (Mahir et al. \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) reported that SVM performed better than other classifiers including Na\u0026iuml;ve Bayesian, RNN/LSTM, Logistic Regression in recognizing fake news extracted from twitter. The study (Al-Ash et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) used Indonesian news dataset consist of fake and real news documents to show the classification performance of Random Forest, SVM and Na\u0026iuml;ve Bayesian Classifiers over this dataset with associate classification approach. In this study (Katsaros et al. \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) eight models were evaluated for classification purpose. These models include Linear Regres. sion, SVM, MLP, Gaussian and Multinomial na\u0026iuml;ve Bayes, Random Forests, Decision Trees and CNN on three publicly available datasets. The result showed that the CNN is the best performing algorithm.\u003c/p\u003e \u003cp\u003eThe study of (Choudhary et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) proposed a deep learning architecture called BerConvoNet for classification of Fake news and Real news with marginal error. The proposed architecture was composed of two main blocks, a New Embedding Block (NEB)and a Multi-scale Features Block (MSFB). The NEB used BERT for extracting word embeddings from news articles and then fed it to MSFB as input.\u003c/p\u003e \u003cp\u003eIn the study (Vogel and Meghana \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) reported that SVM achieved the highest accuracy of 92% in classification of fake news. He used hand crafted features extracted from news dataset like total word(tokens), Unique words, Unique words (types), type/token ratio, Number of sentences, average sentence length, Number of Characters, Average word length, nouns, prepositions, adjectives.The classification models include XG Boost, Random Forest, Na\u0026iuml;ve Bayesian, KNN, Decision Tree and SVM were used for classification of fake news.\u003c/p\u003e"},{"header":"3. Proposed Methodology","content":"\u003cp\u003eThe methodology section presents the architecture of proposed model for classification of fake news as shown in Figure-1. The methodology section is consists of three Main phases: 1st phase: pre-processing (Sentence segmentation, tokenization, stopword removal and word stemming), 2nd phases include features extraction and best features selection, best feature selection using famous feature selection techniques (Chi-Square, Feature Importance, Information Gain and Univariate) and final phase classification of fake news using the proposed model and other machine learning models.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Dataset\u003c/h2\u003e \u003cp\u003eThe proposed machine learning model is tested on the ISOT News Dataset, a publicly available dataset containing false and real news. This dataset is commonly used in the false news identification problem. A total of 44,919 fake and real news was used in this research, including 23502 fake news and 2147 real news assessments using multiple machine learning models. Regarding performance metrics, i.e. classification precision, for the assignment of fake news classification, we also tested the suggested Machine Learning classifier with the benchmark models.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Pre-processing:\u003c/h2\u003e \u003cp\u003eIn order to avoid overfitting, the data is preprocessed before fetching it to the Natural Languange Processing (NLP) system. The preprocessing involves various steps like sentence segmentation, tokenization, stopwords removal and word stemming, as discussed below in detail:\u003c/p\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1. Sentence Segmentation:\u003c/h2\u003e \u003cp\u003eSentence segmentation is establish text borders and break the text into sentences. Exclamation (!), interrogation (?), and utter stop (.) signs are widely used as markers to segment the paragraph into sentences.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2. Tokenization\u003c/h2\u003e \u003cp\u003eAt this stage, the phrases and sentences are divided into separate words by dividing them into white spaces such as tabs, blanks, and signs of punctuation, i.e. dot (.), comma (,), semicolon (;), colon (:), etc. These are the key indications for dividing the sentences into tokens.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3. Stopwords removal.\u003c/h2\u003e \u003cp\u003eWords which have occurred repeatedly are called stopwords in a sentence. These consist of prepositions (in, on, at, etc.), conjunctions (and thus, too, etc.), articles (a, an, a), etc. These words have little meaning in text documents and are more weighted, and removing them will help improve the system's performance.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3. Word Stemming:\u003c/h2\u003e \u003cp\u003eStemming is used to bring the word to its basic form. Word stemming plays a significant role in preprocessing. In order to normalize the word token to a standard form, this step changes the derived words to its base or stem word. The famous stemming algorithm, Porter\u0026rsquo;s stemming (Porter, 1980), is adopted to remove suffixes like \u0026ndash;ing, -es, -ers from the text words. For example, the words \u0026lsquo;looking\u0026rsquo; and \u0026lsquo;looks\u0026rsquo; will be modified to its base type \u0026lsquo;look\u0026rsquo; after stemming.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Features Extraction:\u003c/h2\u003e \u003cp\u003eIn text classification problems, features play a major role. This step aims to mine ISOT fake news Dataset features for the problem of text classification. In this study, we extracted twenty-three (23) features form ISOT fake news dataset. Almost all of these features are textual and can be accurately extracted through text as seen in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of all features extracted from ISOT News dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eS.No\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFeatures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCounts-words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total numbers of words\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUpper-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total number of upper case characters\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLowe-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total number of Lower case character\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCharacter-count\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total number of characters\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCount-sentence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total number of sentences in fake News\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAutomated-Readability-Index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAutomated Readability Index is readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eColeman-Liau-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eColeman\u0026ndash;Liau index is text readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFlesch-reading-ease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFlesch reading ease is text readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFlesch-Kincaid-grade\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFlesch\u0026nbsp;Kincaid\u0026nbsp;grade\u0026nbsp;level is text readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDale-Chall-formula\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDale\u0026ndash;Chall formula is text readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTopic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTopic related words\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCount-spaces\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount number of spaces\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePart-of-Speech\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount numbers of Part of speech\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrincipal-Component-Analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrincipal Component Analysis\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGunning-Fox-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGunning fog index is readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMOG-formula\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMc-Laughlin's SMOG formula is readability test\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSentiment-analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSentiment analysis show text positivity and negativity\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBag-of-word\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBag of words is a text representation representing the presence of words\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTf-idf (unigram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTF-IDF with unigram\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTf-idf (bi- gram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTF-IDF with Bi-gram\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTf-idf (tri- gram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTF-IDF with Tri-Gram\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStopwords\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total numbers of stop words\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNegative-words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount total numbers of negative word\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e3.3.1 Features Selection\u003c/h2\u003e \u003cp\u003eIt is normally not good to use all twenty-three (23) features to classify the ISOT News dataset as fake and real News. All features do not have the same significance and weight when developing a consistent and effective statistical model. Some features are useful and add more to the model prediction and play a vital role in classification accuracy, while others are less valuable and have a less significant impact on the performance of the model. In addition, the appropriate and useful features eliminate over-fitting, increase precision and reduce the predictive model training time. We used the four best features selection techniques to resolve this issue, which are Chi2, Univariate, information gain and Feature importance to decrease the space size of the features to achieve optimum features and significant features. In Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, Column-2 shows fourteen (14) most important features for the News dataset were chosen by the Chi-square technique and same numbers of feature selected using Univariate feature selection techniques as seen in Column-3 of Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Similarly, Column-4 picked the fourteen best features from the same dataset using feature importance and Column-5 shows the fourteen best features selected using information gain. Next section evaluated performance of proposed model and other machine learning model on all textual features and best features selected by best features selection techniques as mentioned in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Column 2\u0026ndash;5 of Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of fourteen best features chosen by different Features Selection Techniques for ISOT News dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eS.No\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c5\" namest=\"c2\"\u003e \u003cp\u003eFeatures Selected by\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eChi2\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eUnivariate\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eFeature Importance\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eInformation gain\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCounts-words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCounts-words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCounts-words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCount-unique-words\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUpper-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUpper-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUpper-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUpper case character\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLowe-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLowe-case-character\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCount-sentence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCount sentence\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCharacter-count\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCharacter-count\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFlesch-Kincaid-grade-level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFlesch\u0026nbsp;Kincaid\u0026nbsp;grade\u0026nbsp;level\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCount-sentence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount-sentence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFlesch-reading-ease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFlesch reading ease\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAutomated-readability-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAutomated-readability-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDale-Chall-formula\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDale\u0026ndash;Chall formula\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eColeman-liau-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eColeman-liau-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eGunning-fog-index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eGunning fog index is readability\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFlesch-kincaid-grade-level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFlesch-kincaid-grade-level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSomgs-Formula\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eSMOG formula\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFlesch-reading-ease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDale\u0026ndash;Chall-formula\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTopic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCount lowercase characters\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrincipal-component-analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTopic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTf-idf (unigram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTf_idf (unigram)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTf-idf (unigram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCount-spaces\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eStopwords\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTf-idf(bigram)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTf-idf (bi-gram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePart-of-Speech\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNegative-words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTf-idf(trigram)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTf-idf (tri- gram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTf-idf (unigram)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePart-of-Speech\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eStopwords\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDale\u0026ndash;chall-formula\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSentiment-analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSentiment-analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNegative words\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Classification Model for Fake News dataset\u003c/h2\u003e \u003cp\u003eThis section aims to classify ISOT news as fake and real news using Random Forest and other Machine Learning Model. We train and evaluate the classifiers initially using 10 fold-cross validations on the ISOT False News dataset to validate the impact of the individual models and the proposed model on all textual features and best features selected by features selection techniques. Random forest is a a traditional machine learning algorithm which is used for classification as well as regression problems by using ensembling of many Decision Trees to solve a complex problems. It uses bagging and bootstrap methods for the prediction of model accuracy. The prediction of each decision trees combines for final prediction using a majority of vote as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFor Decision Trees, Gini Impurity and Entropy are calculated by using following Eq.\u0026nbsp;1\u0026ndash;2, respectively.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n\n$$\\text{G}\\text{i}\\text{n}\\text{i} \\text{I}\\text{m}\\text{p}\\text{u}\\text{r}\\text{i}\\text{t}\\text{y}= {\\sum }_{\\text{i}=1}^{\\text{c}}\\text{f}\\text{i}\\left(1-\\text{f}\\text{i}\\right) \\left(1\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n\n$$\\text{E}\\text{n}\\text{t}\\text{r}\\text{o}\\text{p}\\text{y}= {\\sum }_{\\text{i}=1}^{\\text{c}}-\\text{f}\\text{log}\\left(\\text{f}\\text{i}\\right) \\left(2\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Experimental Setting","content":"\u003cp\u003eThe proposed machine learning model is tested on the ISOT News Dataset containing a total of 44,919 fake and real news including 23502 fake news and 2147 real news. The dataset is pre-processed by breaking the news text into sentences. The sentences are tokenized into terms and stopwords are removed. Initially, twenty three (23) textual features are selected from the ISOT News dataset for fake news classification. The proposed model and other machine learning models are evaluated on all twenty three features. As all features do not have the same importance in creating a consistent and accurate predictive model. Some features are meaningful and contribute more to model accuracy, while others are less important and adversely affect the model's performance. In addition, the appropriate and useful features eliminate over-fitting, increase precision and reduce the predictive model training time. We used four best features selection technique to resolve this issue, which are Chi2, Univariate, information gain and Feature importance are used for selecting fourteen best features out of twenty-three and then evaluated proposed model and benchmark techniques in terms of accuracy which is used as\u003c/p\u003e \u003cp\u003eperformance metrics.\u003c/p\u003e"},{"header":"5. Results And Discussions","content":"\u003cp\u003eIn the first step, the performance of individual models and other machine learning models are evaluated by using all twenty three (23) textual features extracted from ISOT News dataset. The results of this experiment is recorded in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. It is revealed from the results that the proposed model has the highest score of 97.25% compared to other classifiers on all features.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of classification using the all textual features\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassifier\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.67\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45.71\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes (G)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45.58\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNave Bayes(M)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e66.39\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes(B)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e78.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e97.25\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.88\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Gradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.97\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAda Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e88.65\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn second step, top 14 best features are selected by using Chi-squar features selection technique.All models including proposed model was evaluated using best selected features and results are recorded in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. Experimental results show that proposed model achieved 97.33% accuracy and performed better than individual models on fourteen (14) best features for the task of fake news classification on ISOT dataset.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of classification using the top 14 features selected by Chi2\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassifier\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.64\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45.54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes (G)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e65.87\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNave Bayes(M)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e66.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes(B)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e70.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e93.07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.25\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e97.33\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.27\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Gradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.73\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAda Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e86.70\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe above results shows that Random forest attained the highest accuracy score of 97.33%, Gradient Boost obtained the second highest accuracy of 96.27 percent, Extra Gradient Boost accuracy is 95.73 percent, KNN accuracy is 95.25 percent, Decision Tree accuracy is 93.07, MLP accuracy is 92.64 percent and Logistic Regression got the lowest accuracy of 45.54 percent.\u003c/p\u003e \u003cp\u003eIn third step, top 14 best features are selected by using the Univariate feature selection technqiue. All models including proposed model are tested using 14 best features and the results are recorded in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. The results demonstrate that the proposed model performed better than individual other models by attaining accuracy score of 97.27%. on best features selected using Univariate features selection technique.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of classification using the top 14 features selected by Univariate\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassifier\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e87.96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes (G)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e43.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNave Bayes(M)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e66.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes(B)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e78.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e85.16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e89.99\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e97.27\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e86.69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Gradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e86.40\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAda Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e84.37\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eReferring to the results presented in above table, the fourteen best features selected using Univariate feature selection techniques, The proposed model attained highest accuracy of 97.27%, KNN achieved the second highest accuracy of 89.99%, MLP accuracy is 87.96%, Gradient Boost accuracy is 86.69 percent. In this case, Random Forest again remained on top in term of accuracy.\u003c/p\u003e \u003cp\u003eIn fourth step, fourteen (14) best features are selected by using the feature importance technique. All models alongwith proposed model are evaluated on top 14 best features and results are stored in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. The results show that proposed model performance is better than other classifier. The accuracy of proposed model is 96.60%.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of classification using the top 14 features selected by feature importance\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassifier\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.30\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e48.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes (G)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e48.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNave Bayes(M)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e70.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes(B)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e79.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.53\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.43\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e96.60\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.89\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Gradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.85\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAda Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e88.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWith reference to the results given in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e, the fourteen best features selected using feature importance features selection technique, our proposed model performed better than others by securing highest accuracy of 96.60%. The MLP got the second highest accuracy of 96.30%, KNN accuracy is 95.43%, Gradient Boost accuracy is 94.89 percent.\u003c/p\u003e \u003cp\u003eIn fifth step,the best features are selected using information gain method of feature selection. All models including our proposed model are evaluated on top 14 best features selected from ISOT dataset and the result in each case are recorded in Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e, The proposed model achieved the highest accuracy of 96.42%, MLP achieved the second highest accuracy of 96.02%, KNN accuracy is 95.91%, Gradient Boost accuracy is 94.18 percent, and RF achieved the highest accuracy for individual models and Naive Bayes got lowest accuracy.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of classification using the top 14 features selected by information gain\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassifier\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.02\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e49.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes (G)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e56.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNave Bayes(M)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e65.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes(B)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e79.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e96.42\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Gradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAda Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e88.50\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe results in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e show that the accuracy of classifiers moved up and down while using all features, best selected features by Chi2,Univariate, Features importance and information gain.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparative Results of all classifications over all features and top 14 features selected by Four Features Selection Techniques\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClassifier\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAll features\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eChi2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUnivariate\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFeature Importance\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eInformation gain\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e92.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e87.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e96.30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e96.02\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogistic regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e45.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e45.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e48.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e49.90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes (G)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e45.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e65.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e43.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e48.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e56.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNave Bayes(M)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e66.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e66.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e66.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e70.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e65.28\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNa\u0026iuml;ve Bayes(B)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e78.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e70.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e78.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e79.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e79.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e93.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e85.16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e92.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e92.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e89.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e95.43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e95.91\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e97.25\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e97.33\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e97.27\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e96.60\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e96.42\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e96.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e86.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e94.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e94.18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Gradient Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e86.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e94.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e94.10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAda Boost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e88.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e86.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e84.37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e88.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e88.50\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eFrom experimental results shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and Table \u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, the following conclusion are drawn:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eThe accuracy of the proposed model (Random Forest) improved with best features selected by using Chi-square and Univariate. However, the accuracy of proposed model decreased when features selection are perfomed using future importance and information gain features selection techniques.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe accuracy of boosting technique like XGBoost, GBM and Ada Boost does not improved as a whole on best features .\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePerfomance of Gaussian Na\u0026iuml;ve Bayesian improved significantly on best features selected by Chi-squar as compared to other features selection techniques.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eOverall, the classification accuracy of the proposed model is superior than all individuals\u0026rsquo; models as well as other boosting approaches.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e"},{"header":"5. Conclusion And Future Work","content":"\u003cp\u003eOnline fake news detection is a challenging task in the area of text classification. Many attempts have been made by various researchers to address this issue. This study proposed Random Forest as machine learning classifier to classify the news as fake news and real news. For this purpose, twenty three (23) features were extracted from the text of the dataset. Four feature selection techniques like chi-square, Univariate, features importance and information gain, were used to select fourteen (14) best features out of the twenty three (23) extracted features. Proposed model as well as other models were used for classification of fake news and real news using fourteen (14) best features. The experimental results show that proposed model outperformed all other classifiers in term of better classification accuracy in fake news prediction. In future, Deep Ensembling models may be used for fake news detection.\u003c/p\u003e"},{"header":"6. Declarations","content":"\u003ch2\u003eConflict of interest\u003c/h2\u003e \u003cp\u003eThe authors declare that they have no conflict of interest.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e \u003cp\u003eThe authors did not receive support from any organization for the submitted work.\u003c/p\u003e"},{"header":"7. References","content":"\u003col\u003e\n\u003cli\u003eAhmad I, Yousaf M, Yousaf S, Ahmad MO (2020) Fake news detection using machine learning ensemble methods Complexity 2020\u003c/li\u003e\n\u003cli\u003eAhmed S, Hinkelmann K, Corradini F Combining machine learning with knowledge engineering to detect fake news in social networks-a survey. In: Proceedings of the AAAI 2019 Spring Symposium, 2019. p 8\u003c/li\u003e\n\u003cli\u003eAl-Ash HS, Putri MF, Mursanto P, Bustamam A Ensemble learning approach on indonesian fake news classification. In: 2019 3rd International Conference on Informatics and Computational Sciences (ICICoS), 2019. IEEE, pp 1-6\u003c/li\u003e\n\u003cli\u003eChoudhary M, Chouhan SS, Pilli ES, Vipparthi SK (2021) BerConvoNet: A deep learning framework for fake news classification Applied Soft Computing 110:107614\u003c/li\u003e\n\u003cli\u003eDe Choudhury M, Morris MR, White RW Seeking and sharing health information online: comparing search engines and social media. In: Proceedings of the SIGCHI conference on human factors in computing systems, 2014. pp 1365-1376\u003c/li\u003e\n\u003cli\u003eFaustini P, Cov\u0026otilde;es T Fake news detection using one-class classification. In: 2019 8th Brazilian Conference on Intelligent Systems (BRACIS), 2019. IEEE, pp 592-597\u003c/li\u003e\n\u003cli\u003eGravanis G, Vakali A, Diamantaras K, Karadais P (2019) Behind the cues: A benchmarking study for fake news detection Expert Systems with Applications 128:201-213\u003c/li\u003e\n\u003cli\u003eHakak S, Alazab M, Khan S, Gadekallu TR, Maddikunta PKR, Khan WZ (2021) An ensemble machine learning approach through effective feature extraction to classify fake news Future Generation Computer Systems 117:47-58\u003c/li\u003e\n\u003cli\u003eHlaing MMM, Kham NSM Defining news authenticity on social media using machine learning approach. In: 2020 IEEE Conference on Computer Applications (ICCA), 2020. IEEE, pp 1-6\u003c/li\u003e\n\u003cli\u003eJain A, Shakya A, Khatter H, Gupta AK A smart system for fake news detection using machine learning. In: 2019 International Conference on Issues and Challenges in Intelligent Computing Techniques (ICICT), 2019. IEEE, pp 1-4\u003c/li\u003e\n\u003cli\u003eKatsaros D, Stavropoulos G, Papakostas D Which machine learning paradigm for fake news detection? In: 2019 IEEE/WIC/ACM International Conference on Web Intelligence (WI), 2019. IEEE, pp 383-387\u003c/li\u003e\n\u003cli\u003eKhan JY, Khondaker MTI, Afroz S, Uddin G, Iqbal A (2021) A benchmark study of machine learning models for online fake news detection Machine Learning with Applications 4:100032\u003c/li\u003e\n\u003cli\u003eMahir EM, Akhter S, Huq MR Detecting fake news using machine learning and deep learning algorithms. In: 2019 7th International Conference on Smart Computing \u0026amp; Communications (ICSCC), 2019. IEEE, pp 1-5\u003c/li\u003e\n\u003cli\u003eMansouri R, Naderan-Tahan M, Rashti MJ A Semi-supervised Learning Method for Fake News Detection in Social Media. In: 2020 28th Iranian Conference on Electrical Engineering (ICEE), 2020. IEEE, pp 1-5\u003c/li\u003e\n\u003cli\u003eNajar F, Zamzami N, Bouguila N Fake news detection using bayesian inference. In: 2019 IEEE 20th International Conference on Information Reuse and Integration for Data Science (IRI), 2019. IEEE, pp 389-394\u003c/li\u003e\n\u003cli\u003eNapoli PM (2018) What if more speech is no longer the solution: First Amendment theory meets fake news and the filter bubble Fed Comm LJ 70:55\u003c/li\u003e\n\u003cli\u003eOkoro E, Abara B, Umagba A, Ajonye A, Isa Z (2018) A hybrid approach to fake news detection on social media Nigerian Journal of Technology 37:454-462\u003c/li\u003e\n\u003cli\u003eOzbay FA, Alatas B (2020) Fake news detection within online social media using supervised artificial intelligence algorithms Physica A: Statistical Mechanics and its Applications 540:123174\u003c/li\u003e\n\u003cli\u003eReis JC, Correia A, Murai F, Veloso A, Benevenuto F (2019) Supervised learning for fake news detection IEEE Intelligent Systems 34:76-81\u003c/li\u003e\n\u003cli\u003eRuchansky N, Seo S, Liu Y Csi: A hybrid deep model for fake news detection. In: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, 2017. pp 797-806\u003c/li\u003e\n\u003cli\u003eVogel I, Meghana M Detecting Fake News Spreaders on Twitter from a Multilingual Perspective. In: 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA), 2020. IEEE, pp 599-606\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"soft-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"soco","sideBox":"Learn more about [Soft Computing](https://www.springer.com/journal/500)","snPcode":"500","submissionUrl":"https://submission.nature.com/new-submission/500/3","title":"Soft Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Machine Learning, Random Forest, Fake news, Feature Selection.","lastPublishedDoi":"10.21203/rs.3.rs-835344/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-835344/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eNowadays, information is published in newspapers and social media while transmitted on radio and television about current events and specific fields of interest nationwide and abroad. It becomes difficult to explicit what is real and what is fake due to the explosive growth of online content. As a result, fake news has become epidemic and immensely challenging to analyze fake news to be verified by the producers in the form of data process outlets not to mislead the people. Indeed, it is a big challenge to the government and public to debate the situation depending on case to case. For the purpose several websites were developed for this purpose to classify the news as either real or fake depending on the website logic and algorithm. A mechanism has to be taken on fact-checking rumors and statements, particularly those that get thousands of views and likes before being debunked and refuted by expert sources. Various machine learning techniques have been used to detect and correctly classified of fake news. However, these approaches are restricted in terms of accuracy. This study has applied a Random Forest (RF) classifier to predict fake or real news. For this prpose, twenty-three (23) textual features are extracted from ISOT Fake News Dataset. Four best feature selection techniques like Chi2, Univariate, information gain and Feature importance are used for selecting fourteen best features out of twenty-three. The proposed model and other benchmark techniques are evaluated on the dataset by using best features. Experimental findings show that, the proposed model outperformed state-of-the-art machine learning techniques such as GBM, XGBoost and Ada Boost Regression Model in terms of classification accuracy.\u003c/p\u003e","manuscriptTitle":"Machine Learning Fake News Classification with Optimal Feature Selection","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-09-20 20:22:07","doi":"10.21203/rs.3.rs-835344/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major Revision","date":"2021-10-31T09:30:14+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2021-09-15T17:08:49+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2021-09-15T16:25:13+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2021-08-28T17:46:25+00:00","index":"","fulltext":""},{"type":"submitted","content":"Soft Computing","date":"2021-08-21T07:47:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"soft-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"soco","sideBox":"Learn more about [Soft Computing](https://www.springer.com/journal/500)","snPcode":"500","submissionUrl":"https://submission.nature.com/new-submission/500/3","title":"Soft Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"944e3d5b-d45a-4771-b97c-c49188f885d8","owner":[],"postedDate":"September 20th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":7297456,"name":"Geometry"},{"id":7297457,"name":"Topology"},{"id":7297458,"name":"Theoretical Computer Science"}],"tags":[],"updatedAt":"2022-02-12T13:26:32+00:00","versionOfRecord":{"articleIdentity":"rs-835344","link":"https://doi.org/10.1007/s00500-022-06773-x","journal":{"identity":"soft-computing","isVorOnly":false,"title":"Soft Computing"},"publishedOn":"2022-01-29 13:26:32","publishedOnDateReadable":"January 29th, 2022"},"versionCreatedAt":"2021-09-20 20:22:07","video":"","vorDoi":"10.1007/s00500-022-06773-x","vorDoiUrl":"https://doi.org/10.1007/s00500-022-06773-x","workflowStages":[]},"version":"v1","identity":"rs-835344","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-835344","identity":"rs-835344","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00