Social network textual data classification through a hybrid word embedding approach and Bayesian conditional-based multiple classifiers

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Sentiment analysis (SA) of text holds a pivotal role in today's digital age, particularly within the realm of social media networks. The analysis of textual sentiments emerges as a critical facet of NLP. In social media, individuals extensively engage with a multitude of texts and opinions. SA empowers us to delve into and discover these opinions, sentiments, and viewpoints, thereby extracting valuable insights on a wide array of subjects. The significance of word embeddings for processing textual data lies in their ability to represent words as dense vectors, enabling machines to capture semantic relationships and contextual nuances, thereby enhancing various natural language processing tasks. There are two popular and famous models, BERT and GloVe, for embedding words. Currently, GloVe is considered one of the most precise approaches. However, this method does not take into account the sentiment information present in texts. Consequently, we opted to utilize pre-trained BERT models, which have been trained on extensive text corpora, in combination with the GloVe model to address this limitation. This study leverages a hybrid word embedding model combining BERT and GloVe. Several classifiers are employed to analyze text sentiment. At the decision level, we employ Bayesian Conditional to integrate current results with prior decisions. When combining previous decisions with new ones, the model achieves higher accuracy by refining or adjusting decisions in light of new evidence. Our approach demonstrates notable results, showcasing its practical significance. The results of the experiments on IMDB, Sentiment140, and Twitter US Airline datasets demonstrate that the proposed approach has achieved favorable results, with accuracies of 0.958, 0.925, and 0.946 respectively. These results are considered acceptable when compared to those of other similar studies.
Full text 166,755 characters · extracted from preprint-html · click to expand
Social network textual data classification through a hybrid word embedding approach and Bayesian conditional-based multiple classifiers | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Social network textual data classification through a hybrid word embedding approach and Bayesian conditional-based multiple classifiers Alireza Ghorbanali This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3961336/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Sentiment analysis (SA) of text holds a pivotal role in today's digital age, particularly within the realm of social media networks. The analysis of textual sentiments emerges as a critical facet of NLP. In social media, individuals extensively engage with a multitude of texts and opinions. SA empowers us to delve into and discover these opinions, sentiments, and viewpoints, thereby extracting valuable insights on a wide array of subjects. The significance of word embeddings for processing textual data lies in their ability to represent words as dense vectors, enabling machines to capture semantic relationships and contextual nuances, thereby enhancing various natural language processing tasks. There are two popular and famous models, BERT and GloVe, for embedding words. Currently, GloVe is considered one of the most precise approaches. However, this method does not take into account the sentiment information present in texts. Consequently, we opted to utilize pre-trained BERT models, which have been trained on extensive text corpora, in combination with the GloVe model to address this limitation. This study leverages a hybrid word embedding model combining BERT and GloVe. Several classifiers are employed to analyze text sentiment. At the decision level, we employ Bayesian Conditional to integrate current results with prior decisions. When combining previous decisions with new ones, the model achieves higher accuracy by refining or adjusting decisions in light of new evidence. Our approach demonstrates notable results, showcasing its practical significance. The results of the experiments on IMDB, Sentiment140, and Twitter US Airline datasets demonstrate that the proposed approach has achieved favorable results, with accuracies of 0.958, 0.925, and 0.946 respectively. These results are considered acceptable when compared to those of other similar studies. Sentiment analysis Opinion mining BERT GloVe Ensemble Bayesian Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction In recent years, sentiment analysis of textual data has gained significant attention due to its versatile applications in various fields, ranging from marketing and customer feedback analysis to social media monitoring and political sentiment tracking. Sentiment analysis, also known as opinion mining, involves the extraction and categorization of subjective information from the text to determine the underlying sentiment expressed, which could be positive, negative, or neutral (Ghorbanali & Sohrabi, 2023a ). SA holds great promise in understanding public opinion, consumer preferences, and user sentiments on a large scale. It empowers organizations to make informed decisions, adapt their strategies, and tailor their products or services based on the feedback received from their audience. Moreover, in the realm of social media, SA enables tracking and responding to trends, identifying emerging issues, and even gauging the sentiment surrounding specific events. SA can be done in an unimodal and multimodal way (Ghorbanali et al., 2022 ). It can also be done as a single-domain or multi-domain (Ghorbanali & Sohrabi, 2023b ). In the pursuit of SA, various techniques have been employed to extract and interpret sentiments from textual data. These techniques span from traditional lexicon-based approaches to modern machine learning (ML) and deep learning (DL) methodologies. Lexicon-based methods involve utilizing sentiment lexicons or dictionaries to assign sentiment scores to words or phrases, while machine learning techniques encompass approaches like Support Vector Machines (SVM), Naive Bayes (NB), and Random Forest (RF), which learn sentiment patterns from labeled data. In recent times, deep learning models, particularly those pre-trained on large text corpora, have demonstrated remarkable performance in sentiment analysis tasks. Notably, the use of pre-trained language models like BERT and GloVe has become prevalent. BERT captures contextual information effectively, while GloVe focuses on global semantic relationships among words. The combination of these models in a hybrid architecture can further enhance sentiment analysis accuracy by leveraging their complementary strengths. This paper aims to explore the potential of a hybrid BERT and GloVe model for sentiment analysis and contribute to the advancement of sentiment analysis techniques, enabling a more accurate and nuanced understanding of textual sentiments in diverse applications. For this purpose, at first, we presented a hybrid word embedding model of BERT and GloVe for embedding words. Then we sent its results to several classifiers for classification. Next, to determine the final polarity and emotions, we used the results obtained from the classifiers using previous decisions and probabilities. The primary features and impacts of this study can be outlined as follows: Creating a hybrid word embedding model based on BERT and GloVe Utilizing ensemble learning techniques based on CNNs Utilizing previous classifications and their current results for decision-making Fusing results at the decision level using Bayesian Conditional Using the IMDB, Twitter US Airline, and Sentiment140 datasets for experiments The rest of the article is structured as follows: In Section 2, we will have Research objectives and motivation, In Section 3, we delve into the existing literature covering word embedding, deep neural networks, text sentiment analysis, ensemble learning, and fusion applied to text SA. Section 4 offers a comprehensive explanation of our proposed method's methodology. The outcomes of our experiments are outlined and assessed in Section 5, and we wrap up with the conclusions of our work in Section 6. 2. Research objectives and motivation In today's digital world, textual sentiment analysis plays a vital role, especially in social networks where messages and comments are generated in huge amounts. This analysis enables us to delve deeply into these opinions, feelings, and perspectives and extract valuable information for a variety of topics. The research objectives of this study encompass the development and assessment of a hybrid word embedding model that combines BERT and GloVe for SA within social media text. In addition, the study aims to employ a diverse range of classifiers to effectively analyze and categorize text sentiment. Furthermore, the investigation delves into the integration of current sentiment analysis outcomes with prior decisions using Bayesian Conditional. Finally, the research seeks to quantify the practical impact of the proposed approach by measuring the percentage improvement in accuracy on a selected sample dataset, thereby providing a clear and comprehensive roadmap for the study's goals. 3. Literature and backgrounds In this section, we will review the studies conducted in the field of text SA. We will also have a brief overview of word embedding methods. 3.1 Word embedding Word embedding is a technique used in NLP to represent words as vectors in a continuous, high-dimensional space. The main idea behind word embedding is to convert words, which are discrete symbols, into numerical vectors that capture semantic relationships between words. This enables machines to process and understand language in a way that's more suitable for various computational tasks. Traditional methods of representing words, like one-hot encoding, treat each word as a distinct entity and represent it as a sparse binary vector. However, this approach doesn't capture the relationships and meanings between words. Word embedding, on the other hand, aims to create dense vector representations where similar words are located closer to each other in the vector space, allowing for semantic relationships to be encoded. Word embedding models are trained on large text corpora using techniques like Word2Vec, GloVe, and BERT. These models learn to predict words based on their contexts or vice versa, and in the process, they learn to represent words as continuous vectors. The resulting word embeddings can be used as features for various NLP tasks, such as sentiment analysis, machine translation, text generation, and more. They provide a way for machines to understand the meaning and context of words, even in tasks where direct text matching might not be effective. Ghorbanali et al. (Ghorbanali & Sohrabi, 2023b ; Ghorbanali et al., 2022 ), used a BERT pre-trained model to discover features and word embedding. In the study (Ni & Cao, 2020 ), The GloVe model was used to optimally use global and local information. These two popular models have been used in many studies (Mahto & Yadav, 2023 ; Zouzou & El Azami, 2021 ). In Table 1 , we have examined two popular models (BERT and GloVe). Table 1 Comparison of BERT and GloVE models Model Advantages Disadvantages BERT 1. Whole Text Attention (Bidirectional Context Understanding). 2. Deep Analysis with Hierarchical Information 3. State-of-the-Art Performance 4. Fine-Tuning Capability 5. Multilingual Support 1. Large Model Size 2. Computationally Intensive 3. Lack of Interpretability 4. Limited Scalability GloVe 1. Statistical-Based Embeddings 2. Wide Applicability 3. Effective for Small Datasets 4. Simplicity 1. Lack of Sentence Structure 2. Pre-processing Required 3. Fixed Representations 4. Less Effective in Some Tasks 3.2 Deep neural network A deep neural network (DNN) is a type of ML model inspired by the structure and function of the human brain's interconnected neurons. DNNs excel at tasks such as image and speech recognition, natural language processing, and more, as they can automatically learn relevant features from raw data without the need for explicit feature engineering. DNNs can be categorized into different types based on their architectures and use cases: 1)FNN 1 , 2)CNN 2 , 3)RNN 3 , 4)LSTM 4 , 5)GRU 5 , and 6)GAN 6 . In the following, we will give a brief definition of CNN, RNN, LSTM, and GRU models. CNNs use convolutional layers to automatically learn hierarchical features from the data. They are well-suited for image classification, object detection, and image generation. These networks originally designed for image processing, have also found successful applications in NLP (Such as text classification and sentiment analysis tasks). In studies (Ghorbanali et al., 2022 ; Kim, 2014 ; Zhang & Wallace, 2015 ), convolutional networks were used to classify sentences and texts. Figure 1 also shows the use of the above network for sentence classification. RNN networks are designed for sequence data, such as time series or natural language. RNNs have loops that allow information to persist across different time steps. However, they suffer from vanishing gradient problems for long sequences. LSTM is a type of RNN, LSTMs are designed to overcome the vanishing gradient problem. They have memory cells that can store information over long sequences, making them effective for tasks like language modeling and machine translation. GRU is Similar to LSTMs, GRUs are designed to address the vanishing gradient problem in RNNs. They have a simpler architecture but still provide a level of memory retention. In the paper(Zouzou & El Azami, 2021 ), they proposed to use GloVe as a word embedding and introduce a developed classification using CNN, GRU, and a hybrid model of GRU and CNN applied on IMDB consisting of 50k movie reviews. They achieved 86.34% accuracy. In the study (Tan, Lee, Anbananthen, et al., 2022), they used the combination of RoBERTa (Robustly optimized BERT) and LSTM to deal with the challenge of limitations of sequence for SA. They used their proposed model on the IMDB dataset and achieved an accuracy of 92.96%. Mishra et al.(Mishra & Patil, 2023 ) performed SA on the IMDb dataset using a combined CNN-LSTM approach and achieved an accuracy of 0.85%. They also used CNN and LSTM networks separately, which reached 85% and 87% results. They stated that the combination of these two networks is suitable for long-term dependency. 3.3 Text sentiment analysis Sentiment analysis or SA has developed rapidly as an important task in recent years and has been used for a wide range of applications, including election forecasting, stock market forecasting, product evaluation, medicine, tourism, and others. SA is very important because it helps businesses quickly understand their customers' opinions about their products. In this section, we review several previous studies on text SA. SA of text is one of the problems in the field of NLP, which is done by two methods based on Lexicon and ML. ML methods are divided into supervised, semi-supervised, unsupervised, and reinforcement learning categories. Lexicon-based approaches usually use emotional words or phrases to assess the polarity of sentences. Text SA can be classified into three main levels: 1)Document Level, 2)Sentence Level, and 3)Aspect-Based. There are also other levels. SA of text can be done in single domain and multi domains (Ghorbanali & Sohrabi, 2023b ). In the study (Zhou et al., 2022 ), a hybrid model was presented using doc2vec, CNN, and BiLSTM along with the attention mechanism for the SA of texts. The proposed model significantly diminished the loss of semantic information, improved accuracy, and reduced losses from 22.1% and 19.9%. Zhao et al. (Zhao et al., 2021 ), used collected customer comments from e-commerce sites for SA. This article introduces a novel and enhanced ML technique known as the Elman Neural Network (ENN) with Local Search Improved Bat Algorithm (LSIBA-ENN). After collecting the data and preprocessing it, they performed feature extraction and polarity (emotion or sentiment) classification. The results demonstrate that in terms of sentiment classification, the LSIBA-ENN achieves good performance. In the study (Wang et al., 2023 ), a microblog SA model using TEXTCNN and Bayesian techniques was proposed. Their goal is to provide a model that can be interpreted and accurately analyze emotions. To achieve this purpose, they utilized a combination of word2vec, ELMo, and CNN. Aspect-level SA offers a richer and more comprehensive insight compared to document-level SA. This is due to its focus on forecasting sentiment for different aspects within a similar text. The challenge posed by the mentioned level is that various aspect categories within the same text might exhibit diverse polarities. In the above paper, a multi-task model for aspect-level SA was offered, leveraging RoBERTa. Treating individual aspect categories as subtasks, they utilized RoBERTa, to extract features from both text and aspect tokens. Additionally, they implemented a cross-attention mechanism to direct the model's attention toward the most pertinent features corresponding to the specified aspect category. The results of the implementation of the above method show that the proposed method is effective and has succeeded in obtaining good results in the analysis of aspect emotions compared to other action models. In the study (Truşcǎ et al., 2020 ), the authors presented an innovative hybrid approach named HAABSA (Hybrid Aspect-Based SA). The method is developed in two stages. Initially, they substituted non-textual word embeddings with deep contextual embeddings, effectively handling word semantics within a given text. Subsequently, they introduced hierarchical attention, incorporating an extra attention layer into the high-level HAABSA representations. This addition aims to enhance the method's capacity for modeling input data with greater flexibility. They used SemEval 2015 and SemEval 2016 datasets for their work. Classifying genuine sentiments from textual social media data is a challenging and important task. Mahto et al. (Mahto & Yadav, 2023 ), presented the HeBiLSTM model to solve this challenge. They introduce an extra Bi-LSTM hidden layer, referred to as the Enhanced GloVe-HeBiLSTM model. Furthermore, the researchers introduced a practical framework for analyzing authentic emotional content within textual data. This framework, named HeBi-CuDNNLSTM, employs a Hierarchical Bi-CuDNNLSTM architecture and leverages the NVIDIA CUDA deep neural network library. They evaluated their proposed models and achieved better results than the basic models. Deep learning architectures (such as CNN and LSTM) have demonstrated their effectiveness in text SA. Nevertheless, challenges like class imbalance and the presence of unlabeled corpora continue to constrain the accuracy of text emotion prediction. Jiang et al. (Jiang et al., 2023 ), To surmount these challenges, they introduced a novel classification model named KSCB. KSCB is obtained from the integration of the Bi-LSTM, K-means++, CNN, and SMOTE models and is used for SA of textual data. Twitter posts exhibit a lack of structure, and informal tone, and contain considerable noise. Additionally, they encompass linguistic intricacies such as polysemy, where words possess multiple meanings. To solve this challenge, Naseem et al. (Naseem et al., 2019 ), introduced a combination of word representation and BiLSTM 7 with attention mechanisms, resulting in enhanced tweet quality. This approach not only tackles textual noise but also handles the presence of multiple semantic interpretations. The results of the experiments show that the introduced way can overcome the ongoing limitations and improve classification accuracy. The use of combined methods can lead to an increase in classification accuracy, many studies have also used the advantages of different methods in a combined way to classify the emotions of texts. Alyoubi et al. (Alyoubi & Sharma, 2023 ), presented a hybrid approach using CNN and RNN networks. They utilized the BERT model for producing word embeddings, performed feature extraction from CNN, and employed a bidirectional RNN to harness contextual and temporal features. Their combined method achieved 96% accuracy. In the study (Jain et al., 2021 ), a hybrid method was presented using CNN-LSTM networks to analyze consumer sentiments. The proposed method was employed by batch normalization (BN), max pooling, and dropout to attain outcomes. Their hybrid model achieved 91.3% accuracy in SA on used datasets. Another study (Yue & Li, 2020 ), also used the combination of CNN-BiLSTM to analyze the sentiments of the text and achieved a classification accuracy of 91.48%. They used the Word2vec model to embed the words. This paper (Go et al., 2009 ) introduces an innovative approach to automatically categorize the sentiment of Twitter messages through the utilization of ML algorithms. This method leverages emoticon data as somewhat noisy labels, demonstrating an impressive accuracy rate of over 80% when trained with this particular dataset. Additionally, the paper delves into the preprocessing steps essential for achieving high accuracy and emphasizes the significance of incorporating emoticon-laden tweets for distant supervised learning. Twitter SA is a computational process that categorizes tweets into positive or negative ones. The main challenge is identifying the most suitable sentiment classifier. This paper (Saleena, 2018 ) proposes an ensemble classifier that combines base learning classifiers, improving performance and accuracy. Results show the proposed classifier outperforms stand-alone classifiers and majority voting ensemble classifiers. This paper (Siddiqua et al., 2016 ) proposes a method for SA on Twitter, combining a rule-based classifier with a majority voting ensemble of supervised classifiers. The rule-based classifier is based on emoticons and sentiment-bearing words, while the supervised classifiers are trained using Twitter-specific, textual, parts-of-speech, lexicon-based, and bag-of-words features. The method uses chi-square statistics and information gain to select the best feature combination. Experimental results show its effectiveness. Managing unstructured data from social media can be a formidable challenge. To tackle this issue, deep learning algorithms provide a suitable solution for addressing these complexities. This paper (Tyagi et al., 2020 ) presents a CNN-LSTM-based deep learning method for managing unstructured social media data. The method automatically extracts features for sentiment analysis and classification of reviews, achieving better performance on a benchmark dataset compared to baseline machine learning methods. The paper (Tan, Lee, Lim, et al., 2022) presents an ensemble hybrid DL model for SA, combining three models: RoBERTa, LSTM, BiLSTM, and GRU. The model projected textual input sequences, captured long-range dependencies, and combined predictions using averaging ensemble and majority voting. The model outperformed state-of-the-art methods with accuracy of 94.9%, 91.77%, and 89.81% on IMDb, Twitter US Airline Sentiment dataset, and Sentiment140 dataset. The paper (Hossen et al., 2021 ) introduces a deep learning methodology employing LSTM and GRUs to train hotel review data for the purpose of identifying customer opinions. The study (Jain et al., 2021 ) presents a hybrid CNN-LSTM model for SA using dropout, max pooling, and batch normalization. Experimental analysis on Airlinequality and Twitter sentiment datasets shows the model outperforms classical ML models with 91.3% accuracy in SA. SA is crucial in understanding the sentiment context of social media messages, which can impact strategic decision-making in business and politics. Challenges include lexical diversity, imbalanced datasets, and long-distance dependencies. In study(Tan, Lee, Anbananthen, et al., 2022), a hybrid deep learning method was proposed, leveraging the strengths of both sequence and Transformer models while mitigating their limitations. The model combines a robustly optimized BERT approach with LSTM for SA, effectively mapping words into a concise, semantically meaningful word embedding space. 3.4 Ensemble learning Ensemble learning is an ML technique that involves combining multiple models (often referred to as "base" or "weak" models) to create a stronger, more robust model. The idea behind ensemble learning is that by combining the predictions of multiple models, the overall performance can be improved compared to that of any individual model. Ensembles are often used to increase the accuracy, stability, and generalization of ML algorithms. There are several common techniques for creating ensemble models: 1) Bagging or Bootstrap Aggregating, 2) Boosting, 3) Stacking, and 4) Voting. In various studies, ensemble learning has been used to achieve higher accuracy in the field of SA. In the study (Fersini et al., 2014 ), by using ensemble learning, they seek to increase accuracy and reduce noise and ambiguity of language. Ghorbanali et al. (Ghorbanali et al., 2022 ), performed text classification using several weighted ensemble CNN networks. 3.5 Fusion Fusion in the context of ML and data analysis refers to the process of combining information, features, or predictions from multiple sources or models to create a unified and more informative result. Fusion aims to leverage the strengths of each source or model to improve the overall accuracy, robustness, or understanding of the outcome. Fusion techniques are commonly used in various fields, including computer vision, NLP, remote sensing, sensor networks, and more. There are different types of fusion techniques: 1) Data Fusion, 2) Feature Fusion, 3) Model Fusion, 4) Sensor Fusion, and 5) Information Fusion. Fusion is particularly useful when dealing with complex, noisy, or incomplete data, as well as when aiming to improve the overall performance and reliability of a system. The choice of fusion technique depends on the nature of the data and the specific goals of the analysis or task at hand. In the study(Huang et al., 2019 ), fusion is divided into three levels: early fusion, intermediate fusion, and late fusion. 4. The proposed model As shown in Fig. 2 , we have used three classifiers of CNN for our proposed model. As can be seen, PDO represents the previous decision and NPO represents the new decision. In the proposed model, first, the input text is embedded using the hybrid approach of word embedding, then it is entered into the CNN networks, and in the final stage, the results are fused at the decision level using Bayesian conditional. The details of the proposed model are given below. 3.1 Dataset We have used the IMDB dataset to implement and train our proposed model. The IMDB dataset is one of the most renowned and extensive collections for SA, particularly in the domain of NLP. This dataset is tailored for SA tasks involving textual reviews and critiques of movies and television shows. Each review or critique is accompanied by a numerical rating, representing the writer's overall sentiment towards the content, which can be either positive or negative. This dataset comprises over 50,000 reviews, approximately evenly split between positive and negative sentiments, making it a valuable resource for conducting comprehensive and extensive analyses. Each review is associated with a numerical score, serving as the ground truth label, indicating whether the sentiment is positive or negative. The texts within this dataset represent genuine reviews and critiques authored by individuals for various films and TV programs. These texts faithfully capture the sentiments of the authors toward the content. The number of positive and negative reviews is balanced within the dataset, ensuring statistical validity for empirical analyses. Its details are given in Table 2 . We have preprocessed the data from the above dataset before using it. To access additional details about the dataset, kindly review the provided link 8 . We have used two other datasets named Twitter US Airline 9 (This data originally came from Crowdflower's 10 Data for Everyone library.) and Sentiment140 (Go et al., 2009 ). Twitter US Airline dataset comprises tweets related to major U.S. airlines, collected since February 2015. It encompasses various attributes, including Twitter user IDs, sentiment confidence scores, reasons for negativity and positivity, retweet counts, tweet content, date, time, and location information. Additionally, the sentiment analysis dataset includes tagged reviews for numerous Amazon products, offering both positive and negative sentiments. These reviews are associated with ratings on a scale of 1 to 5 stars, which can be converted into binary format if needed. Next dataset is known as Sentiment140 and has been compiled by extracting 1,600,000 tweets through the Twitter API. These tweets have been carefully labeled and can be employed for SA purposes. Each dataset is divided into three parts: a 70% training set, a 15% validation set, and a 15% testing set. Table 2 Dataset Sentiment IMDB Twitter US Airline Sentiment140 Positive (1) 25000 2363 800000 Negative (0) 25000 9178 800000 Neutral (2) - 3099 - Total 50000 14640 1600000 3.2 Word embedding layer (Hybrid BERT and GloVe) The fusion of BERT and GloVe models in a hybrid architecture involves a synergistic approach to harnessing the strengths of both methodologies. By integrating BERT's contextualized embeddings, which excel in capturing intricate contextual nuances of language, with GloVe's global and statistical word representations, which encode semantic relationships and co-occurrence patterns, the hybrid model achieves a comprehensive understanding of the text. This integration enables the model to not only grasp fine-grained semantic nuances at the token level through BERT's embeddings but also grasp broader semantic similarities and associations between words using GloVe's embeddings. The amalgamation of these embeddings within the hybrid architecture results in an enriched feature space that encapsulates a more holistic view of the text's meaning. This holistic representation empowers the model to perform more effectively in various natural language processing tasks such as sentiment analysis, text classification, and information retrieval. The combination also potentially mitigates the limitations of each model, while BERT might struggle with out-of-vocabulary terms, GloVe comprehensive word vectors can contribute to addressing this challenge. Conversely, BERT's fine-grained context understanding can enhance the interpretability and context awareness of GloVe embeddings. Ultimately, this hybridization capitalizes on the strengths of both BERT and GloVe, leading to improved performance, enhanced semantic representation, and potentially more robust generalization across diverse textual datasets and tasks in the field of natural language processing. The architecture of the proposed hybrid word embedding model is shown in Fig. 3 . The dimensions of this combination are also shown in Fig. 4 . This combination is done in the following five steps: Input text : The initial stage involves feeding the input text into the architecture. Tokenization: In this step, the input text is tokenized, breaking it down into individual tokens or words. BERT embeddings : The tokenized text is then processed by the BERT model, generating contextualized embeddings for each token (We using the BERT model to extract high-dimensional features like "semantic embeddings."). GloVe embeddings : Simultaneously, the tokenized text is also used to obtain GloVe embeddings, which capture global semantic information for each token (We using the GloVe model to extract low-dimensional features.). Combination layer : The BERT and GloVe embeddings are combined using a specific layer that merges their information, allowing the model to capture both local and global context (We then use the combined vectors as input features for a new model, such as an ensemble CNNs.). These stages collectively define the hybrid architecture of combining BERT and GloVe embeddings for various downstream NLP tasksThe concept of this process is given in Eq. 1 . $${E}_{BERT: BERT{\prime }s embedding vector for a specific token.}$$ $${E}_{GloVe: GloVe{\prime }s embedding vector for the same token.}$$ $${E}_{Combined= \frac{{E}_{BERT}+ {E}_{GloVe}}{2}}$$ 1 This averaging approach provides a balanced integration of both BERT's contextually rich embeddings, which capture fine-grained semantics, and GloVe's global semantic embeddings, which represent broader word relationships. The combined embedding \({E}_{Combined}\) effectively captures both local context and global semantics, enhancing the model's ability to interpret and comprehend text comprehensively. The strength of this hybrid approach lies in its ability to enhance the semantic richness of embeddings, leading to improved performance and robustness in various natural language processing tasks. There are other ways to combine these two models. The choice of the best fusion method for combining BERT and GloVe embeddings depends on the specific tasks and data at hand. Each of the combination techniques (such as averaging, concatenation, weighted combination, fine-tuning, etc.) comes with its advantages and limitations, and selecting the optimal approach should consider the following factors: Task nature : The type of task at hand (such as sentiment analysis, text classification, translation, etc.) significantly influences the choice of the fusion method. Data constraints : The volume and nature of our training and testing data are crucial. Specific cases : Certain specialized cases might require unique fusion strategies. For instance, dealing with anomalous data (like fake news) could necessitate incorporating anomaly detection techniques within the fusion. Model complexity : The inherent complexity of our model plays a role as well. If your model is already deep and intricate, adding additional layers for fusion might exacerbate complexity challenges. Empirical results : Experimentation and evaluating models using various metrics are usually essential to determine the best approach for your specific tasks. Table 3 shows other equations. Table 3 Equations of word embedding combination and other method Combining BERT and GloVe Equations and methods Averaging \({E}_{Combined= \frac{{E}_{BERT}+ {E}_{GloVe}}{2}}\) (1) Concatenation \({E}_{Combined= }\left[{E}_{BERT},{E}_{GloVe}\right]\) (2) Weighted combination \({E}_{Combined= }{{W}_{BERT}\times E}_{BERT}+{W}_{GloVe}{\times E}_{GloVe}\) (3) Other method Fine-tuning, feature fusion layers, and adaptive combination In Table 3 , \(\left[ \right]\) denotes vector concatenation. Where \({W}_{BERT}\) and \({W}_{GloVe}\) are weight coefficients. Each of these methods offers a unique way to leverage the complementary strengths of BERT and GloVe, and the choice of approach depends on the specific nature of the task, the available data, and the desired trade-off between contextual understanding and semantic richness. 3.3 Classification To extract features, and classify and predict emotions, we have used the ensemble learning approach using three CNN classifications. CNN networks are generally composed of four parts, in the first layer, which is an input layer, its input is the output of the word embedding layer. The next layer, called the Convolution layer, uses filters to extract features (We used Convolution1D for texts). The next layer is called pooling, this layer is used to select features and reduce their dimensions. Classification of texts is done in the last layer using a fully connected layer (FC). Ensembling multiple CNN classifiers can often lead to improved performance and better generalization. We built three parallel CNN classifiers and integrated their final results at the decision level. This ensemble approach allows each CNN classifier to learn different features from the text data, which can lead to improved performance compared to a single classifier. We have used a different configuration for each classifier because the combination can be useful when the models have different structures or are trained on different datasets. Table 4 summarizes the parameters used in the ensemble of CNN classifiers. Table 4 CNN classifier parameters Parameter Classifier 1 Classifier 2 Classifier 3 Vocabulary Size 10,000 10,000 10,000 Sequence Length 100 100 100 Embedding Dim 100 100 100 Num Filters 128 64 96 Filter Size 3 4 5 Dropout Rate 0.3 0.4 0.5 Hidden Units 64 64 64 Output Classes 2 2 2 3.4 Combined approach using Bayesian conditional probability (Fusion) Initial assumption for previous decisions : Define initial probabilities for your previous decisions. For instance, if your past decisions were binary (positive/negative), you can assign an initial probability of selecting "positive" as 0.7 and "negative" as 0.3. Calculate new probabilities based on new results : For each of the new classifiers, calculate the conditional probabilities of each of the previous decisions given the new results. These probabilities can be computed using Bayesian conditional probability. The calculations can take into account various factors, such as previous decision outcomes and similarities with previous decisions. Combine new and previous probabilities : Now that you have the probabilities of choosing previous decisions and the probabilities calculated based on the new classifier results, you can combine these probabilities using specified weights to strike a balance between the new information and the previous decisions. Final decision : Select the decision with the highest final probability as your ultimate decision. Here is how this process works: Probabilities of previous decisions: P(Decision_A), P(Decision_B), ... Probabilities calculated based on new classifier results: P(Decision_A|NewResults), P(Decision_B|NewResults), ... Combining probabilities with specified weights: Final_P(Decision_A) = alpha * P(Decision_A) + (1 - alpha) * P(Decision_A|NewResults) Final Decision: Choose the decision with the highest final probability. In this approach, alpha represents a weight that you can adjust. This weight determines the influence of the previous decisions on the final decision. 4. Result and implications We have used the following evaluation criteria to evaluate our model (Eq. 4–9). It is noteworthy that, the specific evaluation criteria used for a model will depend on the specific task at hand. For example, if the model is classifying samples into two classes, one might use the accuracy, sensitivity, or F1 score. If the model classifies samples into multiple classes, AUC may be used. The hyperparameters for the model are shown in Table 5 . Also, Fig. 5 and Tables 6 – 8 show the results of implementing the proposed model on the IMDB, Twitter US Airline, and Sentiment140 datasets. We have used the confusion matrix to analyze the proposed model. The confusion matrix is a tool that can be used to evaluate the performance of ML classification models. The confusion matrix is shown in Fig. 6 . As the confusion matrices show, the proposed model has been able to get acceptable results on the test data. It shows how well the model correctly classified the positive and negative samples. The negative, positive, and neutral polarities are represented in the confusion matrix with 0, 1, and 2 respectively. Accuracy (ACC) = (TP + TN) / (TP + TN + FP + FN) (4) Where: TP = True Positives TN = True Negatives FP = False Positives FN = False Negatives Sensitivity = TP / (TP + FN) (5) Specificity = TN / (TN + FP) (6) Precision (P) = TP / (TP + FP) (7) Recall (R) = TP / (TP + FN) (8) F1-Score = 2 \(\times\) (P \(\times\) R) / (P + R) (9) Table 5 Hyperparameters for the proposed model Hyperparameters Value Learning rate: 1e-5 Weight decay: 0.01 Epochs: 10 Batch size: 32 Table 6 Comparing the results of other studies on the IMDB dataset Model ACC F1-Score P R Year Lexicon-based (Giatsoglou et al., 2017 ) 0.830 0.906 - - 2017 Combination of CNN and LSTM (Yenter & Verma, 2017 ) 0.895 - - - 2017 CNN-BiLSTM-Attention(Wang et al., 2019 ) 0.903 - - - 2019 Multiple CNN and LSTM layers (Qaisar, 2020 ) 0.899 - - - 2020 Hybrid lexicon and neural networks (Shaukat et al., 2020 ) 0.919 - - - 2020 CNN_GRU(Zouzou & El Azami, 2021 ) 0.863 - - - 2021 Doc2vec + CNN + BiLSTM + Attention(Zhou et al., 2022 ) 0.913 - - - 2022 Ensemble hybrid deep learning model (Tan, Lee, Lim, et al., 2022) 0.949 0.95 0.95 0.95 2022 CNN-LSTM(Mishra & Patil, 2023 ) 0.850 - - - 2023 BERT-based (Arora et al., 2023 ) 0.924 0.921 0.892 0.964 2023 MBi-GRUMCONV (Başarslan & Kayaalp, 2023 ) 0.953 - - - 2023 Our model 0.958 0.980 0.970 0.990 2023 Table 7 Comparing the results of other studies on the Twitter US Airline dataset Model ACC F1-Score P R Year GRU (Hossen et al., 2021 ) 0.785 0.720 0.730 0.710 2021 LSTM (Hossen et al., 2021 ) 0.775 0.690 0.710 0.690 2021 CNN-LSTM (Jain et al., 2021 ) 0.760 0.690 0.680 0.690 2021 RoBERTa-LSTM (Tan, Lee, Anbananthen, et al., 2022) 0.913 0.910 0.910 0.910 2022 Ensemble hybrid deep learning model (Tan, Lee, Lim, et al., 2022) 0.917 0.92 0.92 0.92 2022 Our model 0.946 0.936 0.932 0.941 2023 Table 8 Comparing the results of other studies on the Sentiment140 dataset Model ACC F1-Score P R Year Naive Bayes (Unigram) (Go et al., 2009 ) 0.813 - - - 2009 MaxEnt (Unigram) (Go et al., 2009 ) 0.805 - - - 2009 SVM (Unigram) (Go et al., 2009 ) 0.822 - - - 2009 Naive Bayes (Unigram + Bigram) (Go et al., 2009 ) 0.827 - - - 2009 MaxEnt (Unigram + Bigram) (Go et al., 2009 ) 0.830 - - - 2009 SVM (Unigram + Bigram) (Go et al., 2009 ) 0.816 - - - 2009 Rule-basedclassifier (Siddiqua et al., 2016 ) 0.877 0.884 0.843 0.923 2016 Ensemble method (Saleena, 2018 ) 0.746 0.733 - - 2018 CNN-LSTM (Tyagi et al., 2020 ) 0.812 - - - 2020 Ensemble hybrid deep learning model (Tan, Lee, Lim, et al., 2022) 0.898 0.90 0.90 0.90 2022 Our model 0.925 0.937 0.940 0.935 2023 4.1 Discussion of results and implications In this section, we delve into the theoretical and practical implications of our research, shedding light on its distinctive features in comparison to existing work. 4.1.1 Theoretical Implications Our study advances the field of SA by introducing a novel approach that combines BERT and GloVe word embedding models. By leveraging the contextual understanding of BERT and the global statistical patterns captured by GloVe, we offer a more comprehensive perspective on SA. This integration of diverse word embedding techniques provides a deeper understanding of text sentiment, contributing to the theoretical foundation of SA methodologies. 4.1.2 Practical implications From a practical standpoint, our research offers tangible benefits. The combination of BERT and GloVe significantly enhances SA accuracy, as demonstrated by our experimental results. This improved accuracy holds practical significance in various applications, including social media monitoring, customer feedback analysis, and product review sentiment assessment. Businesses and organizations can leverage our approach to gain more accurate insights into public sentiment, thereby making informed decisions and improving user experiences. 4.1.3 Distinguishing factors While existing work often relies on single-word embedding models, our research innovatively integrates two prominent models. This distinctive approach yields improved sentiment analysis performance, as evidenced by the percentage improvement in accuracy on our sample dataset. The use of Bayesian Conditional for result integration further sets our work apart, offering a unique perspective on combining sentiment analysis outcomes. The proposed fusion method is not only effective, but it is simpler than other methods such as Dempster Shafer evidence theory, and it is also suitable for incomplete data and systems with uncertainty. 4.1.4 Implementation environment and hardware We have implemented our proposed model using the Python 3.x programming language and the TensorFlow library. The specifications of the computer system used are given in Table 9 . Table 9 Information about the implementation environment System Configuration Operating system Windows 10 CPU Intel® Core™ i7-7700K @ 4.2GHz RAM 32 GB H.D.D 1 TB 5. Conclusions We have covered the challenge of incomplete data, uncertainty, and increasing accuracy and reliability by presenting a hybrid method along with an ensemble learning technique. First, we presented a hybrid model of word embedding to apply the advantages of BERT and GloVe. Combining GloVe and BERT, two popular NLP techniques, could offer several advantages for various NLP tasks. which consists of semantic understanding, feature combination, efficiency, reduced Data requirements, and interpretable semantic features. It's important to note that the effectiveness of combining GloVe and BERT depends on the specific task and dataset. We used 3 classifiers based on ensemble CNN networks for classification. CNNs in NLP are particularly effective for tasks like text classification, sentiment analysis, topic modeling, and even sequence tagging. Ensemble techniques provide numerous benefits compared to individual models. These advantages encompass enhanced precision and accuracy, particularly when tackling intricate and chaotic issues. Furthermore, they mitigate the chances of falling into the traps of overfitting and underfitting by striking a balance between bias and variance. We fused the results obtained from the classifications at the decision level. Fusion at this level allows us to obtain the final result even if part of the data and results were insufficient. We did the fusion of results using Bayesian conditional and based on previous decisions and new decisions. Bayesian conditional fusion offers a flexible and adaptive approach to decision-making that can lead to more informed, accurate, and reliable outcomes, particularly in scenarios where both historical and new data are relevant. Our proposed model was evaluated on 3 datasets and obtained good results. Declarations Data Availability The dataset used in this study can be made available with a reasonable request to the corresponding author. Conflict of Interest Author declare that they have no conflict of interest. Funding Declaration The authors did not receive support from any organization for the submitted work. No funding was received to assist with the preparation of this manuscript. No funding was received for conducting this study. References Alyoubi, K. H., & Sharma, A. (2023). A Deep CRNN-Based Sentiment Analysis System with Hybrid BERT Embedding. International Journal of Pattern Recognition and Artificial Intelligence , 37 (05), 2352006. Arora, K., Gupta, N., & Pathak, S. (2023). Sentimental Analysis on IMDb Movies Review using BERT. 2023 4th International Conference on Electronics and Sustainable Communication Systems (ICESC), Başarslan, M. S., & Kayaalp, F. (2023). MBi-GRUMCONV: A novel Multi Bi-GRU and Multi CNN-Based deep learning model for social media sentiment analysis. Journal of Cloud Computing , 12 (1), 5. Fersini, E., Messina, E., & Pozzi, F. A. (2014). Sentiment analysis: Bayesian ensemble learning. Decision Support Systems , 68 , 26-38. Ghorbanali, A., & Sohrabi, M. K. (2023a). A comprehensive survey on deep learning-based approaches for multimodal sentiment analysis. Artificial Intelligence Review , 1-34. Ghorbanali, A., & Sohrabi, M. K. (2023b). Exploiting bi-directional deep neural networks for multi-domain sentiment analysis using capsule network. Multimedia Tools and Applications , 1-18. Ghorbanali, A., Sohrabi, M. K., & Yaghmaee, F. (2022). Ensemble transfer learning-based multimodal sentiment analysis using weighted convolutional neural networks. Information Processing & Management , 59 (3), 102929. Giatsoglou, M., Vozalis, M. G., Diamantaras, K., Vakali, A., Sarigiannidis, G., & Chatzisavvas, K. C. (2017). Sentiment analysis leveraging emotions and word embeddings. Expert Systems with Applications , 69 , 214-224. Go, A., Bhayani, R., & Huang, L. (2009). Twitter sentiment classification using distant supervision. CS224N project report, Stanford , 1 (12), 2009. Hossen, M. S., Jony, A. H., Tabassum, T., Islam, M. T., Rahman, M. M., & Khatun, T. (2021). Hotel review analysis for the prediction of business using deep learning approach. 2021 International Conference on Artificial Intelligence and Smart Systems (ICAIS), Huang, F., Zhang, X., Zhao, Z., Xu, J., & Li, Z. (2019). Image–text sentiment analysis via deep multimodal attentive fusion. Knowledge-based systems , 167 , 26-37. Jain, P. K., Saravanan, V., & Pamula, R. (2021). A hybrid CNN-LSTM: A deep learning approach for consumer sentiment analysis using qualitative user-generated contents. Transactions on Asian and Low-Resource Language Information Processing , 20 (5), 1-15. Jiang, W., Zhou, K., Xiong, C., Du, G., Ou, C., & Zhang, J. (2023). KSCB: A novel unsupervised method for text sentiment analysis. Applied Intelligence , 53 (1), 301-311. Kim, Y. (2014). Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882 . Mahto, D., & Yadav, S. C. (2023). Emotion prediction for textual data using GloVe based HeBi-CuDNNLSTM model. Multimedia Tools and Applications , 1-26. Mishra, M., & Patil, A. (2023). Sentiment Prediction of IMDb Movie Reviews Using CNN-LSTM Approach. 2023 International Conference on Control, Communication and Computing (ICCC), Naseem, U., Khan, S. K., Razzak, I., & Hameed, I. A. (2019). Hybrid words representation for airlines sentiment analysis. AI 2019: Advances in Artificial Intelligence: 32nd Australasian Joint Conference, Adelaide, SA, Australia, December 2–5, 2019, Proceedings 32, Ni, R., & Cao, H. (2020). Sentiment Analysis based on GloVe and LSTM-GRU. 2020 39th Chinese control conference (CCC), Qaisar, S. M. (2020). Sentiment analysis of IMDb movie reviews using long short-term memory. 2020 2nd International Conference on Computer and Information Sciences (ICCIS), Saleena, N. (2018). An ensemble classification system for twitter sentiment analysis. Procedia Computer Science , 132 , 937-946. Shaukat, Z., Zulfiqar, A. A., Xiao, C., Azeem, M., & Mahmood, T. (2020). Sentiment analysis on IMDB using lexicon and neural networks. SN Applied Sciences , 2 , 1-10. Siddiqua, U. A., Ahsan, T., & Chy, A. N. (2016). Combining a rule-based classifier with ensemble of feature sets and machine learning techniques for sentiment analysis on microblog. 2016 19th international conference on computer and information technology (ICCIT), Tan, K. L., Lee, C. P., Anbananthen, K. S. M., & Lim, K. M. (2022). RoBERTa-LSTM: a hybrid model for sentiment analysis with transformer and recurrent neural network. IEEE Access , 10 , 21517-21525. Tan, K. L., Lee, C. P., Lim, K. M., & Anbananthen, K. S. M. (2022). Sentiment analysis with ensemble hybrid deep learning model. IEEE Access , 10 , 103694-103704. Truşcǎ, M. M., Wassenberg, D., Frasincar, F., & Dekker, R. (2020). A hybrid approach for aspect-based sentiment analysis using deep contextual word embeddings and hierarchical attention. Web Engineering: 20th International Conference, ICWE 2020, Helsinki, Finland, June 9–12, 2020, Proceedings 20, Tyagi, V., Kumar, A., & Das, S. (2020). Sentiment analysis on twitter data using deep learning approach. 2020 2nd international conference on advances in computing, communication control and networking (ICACCCN), Wang, L., Liu, C. H., Cai, D., Zhao, T., & Wang, M. (2019). Text sentiment analysis based on CNN-BiLSTM network and attention model. Journal of Wuhan Institute of Technology , 41 (4), 386-391. Wang, Z., Yao, L., Shao, X., & Wang, H. (2023). A combination of TEXTCNN model and Bayesian classifier for microblog sentiment analysis. Journal of Combinatorial Optimization , 45 (4), 109. Yenter, A., & Verma, A. (2017). Deep CNN-LSTM with combined kernels from multiple branches for IMDb review sentiment analysis. 2017 IEEE 8th annual ubiquitous computing, electronics and mobile communication conference (UEMCON), Yue, W., & Li, L. (2020). Sentiment analysis using Word2vec-CNN-BiLSTM classification. 2020 seventh international conference on social networks analysis, management and security (SNAMS), Zhang, Y., & Wallace, B. (2015). A sensitivity analysis of (and practitioners' guide to) convolutional neural networks for sentence classification. arXiv preprint arXiv:1510.03820 . Zhao, H., Liu, Z., Yao, X., & Yang, Q. (2021). A machine learning-based sentiment analysis of online product reviews with a novel term weighting and feature selection approach. Information Processing & Management , 58 (5), 102656. Zhou, Y., Zhang, Q., Wang, D., & Gu, X. (2022). Text Sentiment Analysis Based on a New Hybrid Network Model. Computational Intelligence and Neuroscience , 2022 . Zouzou, A., & El Azami, I. (2021). Text sentiment analysis with CNN & GRU model using GloVe. 2021 Fifth International Conference On Intelligent Computing in Data Sciences (ICDS), Footnotes Feedforward Neural Networks Convolutional Neural Networks Recurrent Neural Networks Long Short-Term Memory Gated Recurrent Units Generative Adversarial Networks Bi-directional Long Short Term Memory https://ai.stanford.edu/~amaas/data/sentiment/ https://www.kaggle.com/datasets/crowdflower/twitter-airline-sentiment http://www.crowdflower.com/data-for-everyone Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3961336","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":274869380,"identity":"70780cbf-b526-4aed-b9fc-141f65c3e3b1","order_by":0,"name":"Alireza Ghorbanali","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABPElEQVRIie2RMUvDQBSALwQuyytZE6rtX0gp1BaU/pWEQCZFNx1EziUuUdf8jICT4HDlYbsI3SSQDs2SOS6lYBAvUmhIIq4O93HHO97x8e69I0Qi+YfQn0W42CrjeeUGdtGoK/peUdgsJNbfiskqigoV5Vcsrr2vP15W/SNtdovHReEwHV+fLq5XB3poK/mWjM8bClwOwiwbPAcOwzPfcpjheUk4z8CIbdUMiDFhDcXrAkcl4qXCSgVGCVAEEtukK3qpv3W6U6bRMmU4LoSiLzcJfCH0RZXPFkX0Mi8VJ4pFFUKFQk5p0vERrNimbVUsBGqGPHOjOGWze3849A1vlHQeEAZvqT8JrKayuMuMnK9OoqWL+bboHT7qmCWwwWlv4WK8vbppzF2FWoruj4qYVdtHaeuWpEQikUgqfANUZ3qH3V6szQAAAABJRU5ErkJggg==","orcid":"","institution":"Islamic Azad University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Alireza","middleName":"","lastName":"Ghorbanali","suffix":""}],"badges":[],"createdAt":"2024-02-16 13:30:20","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3961336/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3961336/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":51774570,"identity":"b4626218-d691-431a-99ee-2cb30f3649f4","added_by":"auto","created_at":"2024-02-28 20:27:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":49205,"visible":true,"origin":"","legend":"\u003cp\u003eCNN Model for two classes (Using example sentence) (Zhang \u0026amp; Wallace, 2015)\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/6c3c90faff73b86db9bd7d4b.png"},{"id":51774572,"identity":"baf3fbd1-a8a2-4abe-be7d-2f471c0c02a6","added_by":"auto","created_at":"2024-02-28 20:27:30","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":61570,"visible":true,"origin":"","legend":"\u003cp\u003eModel architecture with three CNN classifiers\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/17a19bfb08583a1cd03d9d20.png"},{"id":51774571,"identity":"4b5aeed7-ace7-492a-9a5f-7de115d0151c","added_by":"auto","created_at":"2024-02-28 20:27:30","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":64089,"visible":true,"origin":"","legend":"\u003cp\u003eEmbedding model architecture\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/65bda49d9f7a2508c8d80616.png"},{"id":51774573,"identity":"5b788b0f-f409-4ee3-8bbc-b73b182e154d","added_by":"auto","created_at":"2024-02-28 20:27:30","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":45169,"visible":true,"origin":"","legend":"\u003cp\u003eCombined BERT and GloVe Embedding for IMDB dataset\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/935a9a895a997efb754b7aef.png"},{"id":51774574,"identity":"d27bb04a-a771-459e-997d-2b3319931b6d","added_by":"auto","created_at":"2024-02-28 20:27:30","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":321357,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy and validation of the IMDB, Twitter US Airline, and Sentiment140 datasets\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/b9e95d49c17e1692a2f195ef.png"},{"id":51774575,"identity":"2e794951-d51c-446a-9bf4-c707b1989aa0","added_by":"auto","created_at":"2024-02-28 20:27:30","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":193972,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix for the proposed model on IMDB, Twitter US Airline, and Sentiment140 datasets\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/4361e5a7a44160edcdc79196.png"},{"id":55264418,"identity":"5cbc35a3-52bc-4085-a6c3-468ae825bba9","added_by":"auto","created_at":"2024-04-25 01:42:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1491792,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3961336/v1/988f1cc8-c399-4d65-b6e4-39d29972debb.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Social network textual data classification through a hybrid word embedding approach and Bayesian conditional-based multiple classifiers","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eIn recent years, sentiment analysis of textual data has gained significant attention due to its versatile applications in various fields, ranging from marketing and customer feedback analysis to social media monitoring and political sentiment tracking. Sentiment analysis, also known as opinion mining, involves the extraction and categorization of subjective information from the text to determine the underlying sentiment expressed, which could be positive, negative, or neutral (Ghorbanali \u0026amp; Sohrabi, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2023a\u003c/span\u003e). SA holds great promise in understanding public opinion, consumer preferences, and user sentiments on a large scale. It empowers organizations to make informed decisions, adapt their strategies, and tailor their products or services based on the feedback received from their audience. Moreover, in the realm of social media, SA enables tracking and responding to trends, identifying emerging issues, and even gauging the sentiment surrounding specific events. SA can be done in an unimodal and multimodal way (Ghorbanali et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). It can also be done as a single-domain or multi-domain (Ghorbanali \u0026amp; Sohrabi, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2023b\u003c/span\u003e). In the pursuit of SA, various techniques have been employed to extract and interpret sentiments from textual data. These techniques span from traditional lexicon-based approaches to modern machine learning (ML) and deep learning (DL) methodologies. Lexicon-based methods involve utilizing sentiment lexicons or dictionaries to assign sentiment scores to words or phrases, while machine learning techniques encompass approaches like Support Vector Machines (SVM), Naive Bayes (NB), and Random Forest (RF), which learn sentiment patterns from labeled data. In recent times, deep learning models, particularly those pre-trained on large text corpora, have demonstrated remarkable performance in sentiment analysis tasks. Notably, the use of pre-trained language models like BERT and GloVe has become prevalent. BERT captures contextual information effectively, while GloVe focuses on global semantic relationships among words. The combination of these models in a hybrid architecture can further enhance sentiment analysis accuracy by leveraging their complementary strengths. This paper aims to explore the potential of a hybrid BERT and GloVe model for sentiment analysis and contribute to the advancement of sentiment analysis techniques, enabling a more accurate and nuanced understanding of textual sentiments in diverse applications. For this purpose, at first, we presented a hybrid word embedding model of BERT and GloVe for embedding words. Then we sent its results to several classifiers for classification. Next, to determine the final polarity and emotions, we used the results obtained from the classifiers using previous decisions and probabilities.\u003c/p\u003e \u003cp\u003eThe primary features and impacts of this study can be outlined as follows:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eCreating a hybrid word embedding model based on BERT and GloVe\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eUtilizing ensemble learning techniques based on CNNs\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eUtilizing previous classifications and their current results for decision-making\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFusing results at the decision level using Bayesian Conditional\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eUsing the IMDB, Twitter US Airline, and Sentiment140 datasets for experiments\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe rest of the article is structured as follows: In Section 2, we will have Research objectives and motivation, In Section 3, we delve into the existing literature covering word embedding, deep neural networks, text sentiment analysis, ensemble learning, and fusion applied to text SA. Section 4 offers a comprehensive explanation of our proposed method's methodology. The outcomes of our experiments are outlined and assessed in Section 5, and we wrap up with the conclusions of our work in Section 6.\u003c/p\u003e"},{"header":"2. Research objectives and motivation","content":"\u003cp\u003eIn today's digital world, textual sentiment analysis plays a vital role, especially in social networks where messages and comments are generated in huge amounts. This analysis enables us to delve deeply into these opinions, feelings, and perspectives and extract valuable information for a variety of topics. The research objectives of this study encompass the development and assessment of a hybrid word embedding model that combines BERT and GloVe for SA within social media text. In addition, the study aims to employ a diverse range of classifiers to effectively analyze and categorize text sentiment. Furthermore, the investigation delves into the integration of current sentiment analysis outcomes with prior decisions using Bayesian Conditional. Finally, the research seeks to quantify the practical impact of the proposed approach by measuring the percentage improvement in accuracy on a selected sample dataset, thereby providing a clear and comprehensive roadmap for the study's goals.\u003c/p\u003e"},{"header":"3. Literature and backgrounds","content":"\u003cp\u003eIn this section, we will review the studies conducted in the field of text SA. We will also have a brief overview of word embedding methods.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Word embedding\u003c/h2\u003e \u003cp\u003eWord embedding is a technique used in NLP to represent words as vectors in a continuous, high-dimensional space. The main idea behind word embedding is to convert words, which are discrete symbols, into numerical vectors that capture semantic relationships between words. This enables machines to process and understand language in a way that's more suitable for various computational tasks. Traditional methods of representing words, like one-hot encoding, treat each word as a distinct entity and represent it as a sparse binary vector. However, this approach doesn't capture the relationships and meanings between words. Word embedding, on the other hand, aims to create dense vector representations where similar words are located closer to each other in the vector space, allowing for semantic relationships to be encoded. Word embedding models are trained on large text corpora using techniques like Word2Vec, GloVe, and BERT. These models learn to predict words based on their contexts or vice versa, and in the process, they learn to represent words as continuous vectors. The resulting word embeddings can be used as features for various NLP tasks, such as sentiment analysis, machine translation, text generation, and more. They provide a way for machines to understand the meaning and context of words, even in tasks where direct text matching might not be effective. Ghorbanali et al. (Ghorbanali \u0026amp; Sohrabi, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2023b\u003c/span\u003e; Ghorbanali et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), used a BERT pre-trained model to discover features and word embedding. In the study (Ni \u0026amp; Cao, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), The GloVe model was used to optimally use global and local information. These two popular models have been used in many studies (Mahto \u0026amp; Yadav, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Zouzou \u0026amp; El Azami, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). In Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, we have examined two popular models (BERT and GloVe).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of BERT and GloVE models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdvantages\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDisadvantages\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1. Whole Text Attention (Bidirectional Context Understanding).\u003c/p\u003e \u003cp\u003e2. Deep Analysis with Hierarchical Information\u003c/p\u003e \u003cp\u003e3. State-of-the-Art Performance\u003c/p\u003e \u003cp\u003e4. Fine-Tuning Capability\u003c/p\u003e \u003cp\u003e5. Multilingual Support\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. Large Model Size\u003c/p\u003e \u003cp\u003e2. Computationally Intensive\u003c/p\u003e \u003cp\u003e3. Lack of Interpretability\u003c/p\u003e \u003cp\u003e4. Limited Scalability\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGloVe\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1. Statistical-Based Embeddings\u003c/p\u003e \u003cp\u003e2. Wide Applicability\u003c/p\u003e \u003cp\u003e3. Effective for Small Datasets\u003c/p\u003e \u003cp\u003e4. Simplicity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1. Lack of Sentence Structure\u003c/p\u003e \u003cp\u003e2. Pre-processing Required\u003c/p\u003e \u003cp\u003e3. Fixed Representations\u003c/p\u003e \u003cp\u003e4. Less Effective in Some Tasks\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Deep neural network\u003c/h2\u003e \u003cp\u003eA deep neural network (DNN) is a type of ML model inspired by the structure and function of the human brain's interconnected neurons. DNNs excel at tasks such as image and speech recognition, natural language processing, and more, as they can automatically learn relevant features from raw data without the need for explicit feature engineering. DNNs can be categorized into different types based on their architectures and use cases: 1)FNN\u003csup\u003e1\u003c/sup\u003e, 2)CNN\u003csup\u003e2\u003c/sup\u003e, 3)RNN\u003csup\u003e3\u003c/sup\u003e, 4)LSTM\u003csup\u003e4\u003c/sup\u003e, 5)GRU\u003csup\u003e5\u003c/sup\u003e, and 6)GAN\u003csup\u003e6\u003c/sup\u003e. In the following, we will give a brief definition of CNN, RNN, LSTM, and GRU models. CNNs use convolutional layers to automatically learn hierarchical features from the data. They are well-suited for image classification, object detection, and image generation. These networks originally designed for image processing, have also found successful applications in NLP (Such as text classification and sentiment analysis tasks). In studies (Ghorbanali et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Kim, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2014\u003c/span\u003e; Zhang \u0026amp; Wallace, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2015\u003c/span\u003e), convolutional networks were used to classify sentences and texts. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e also shows the use of the above network for sentence classification. RNN networks are designed for sequence data, such as time series or natural language. RNNs have loops that allow information to persist across different time steps. However, they suffer from vanishing gradient problems for long sequences. LSTM is a type of RNN, LSTMs are designed to overcome the vanishing gradient problem. They have memory cells that can store information over long sequences, making them effective for tasks like language modeling and machine translation. GRU is Similar to LSTMs, GRUs are designed to address the vanishing gradient problem in RNNs. They have a simpler architecture but still provide a level of memory retention. In the paper(Zouzou \u0026amp; El Azami, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), they proposed to use GloVe as a word embedding and introduce a developed classification using CNN, GRU, and a hybrid model of GRU and CNN applied on IMDB consisting of 50k movie reviews. They achieved 86.34% accuracy. In the study (Tan, Lee, Anbananthen, et al., 2022), they used the combination of RoBERTa (Robustly optimized BERT) and LSTM to deal with the challenge of limitations of sequence for SA. They used their proposed model on the IMDB dataset and achieved an accuracy of 92.96%. Mishra et al.(Mishra \u0026amp; Patil, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) performed SA on the IMDb dataset using a combined CNN-LSTM approach and achieved an accuracy of 0.85%. They also used CNN and LSTM networks separately, which reached 85% and 87% results. They stated that the combination of these two networks is suitable for long-term dependency.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Text sentiment analysis\u003c/h2\u003e \u003cp\u003eSentiment analysis or SA has developed rapidly as an important task in recent years and has been used for a wide range of applications, including election forecasting, stock market forecasting, product evaluation, medicine, tourism, and others. SA is very important because it helps businesses quickly understand their customers' opinions about their products. In this section, we review several previous studies on text SA. SA of text is one of the problems in the field of NLP, which is done by two methods based on Lexicon and ML. ML methods are divided into supervised, semi-supervised, unsupervised, and reinforcement learning categories. Lexicon-based approaches usually use emotional words or phrases to assess the polarity of sentences. Text SA can be classified into three main levels: 1)Document Level, 2)Sentence Level, and 3)Aspect-Based. There are also other levels. SA of text can be done in single domain and multi domains (Ghorbanali \u0026amp; Sohrabi, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2023b\u003c/span\u003e). In the study (Zhou et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), a hybrid model was presented using doc2vec, CNN, and BiLSTM along with the attention mechanism for the SA of texts. The proposed model significantly diminished the loss of semantic information, improved accuracy, and reduced losses from 22.1% and 19.9%. Zhao et al. (Zhao et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), used collected customer comments from e-commerce sites for SA. This article introduces a novel and enhanced ML technique known as the Elman Neural Network (ENN) with Local Search Improved Bat Algorithm (LSIBA-ENN). After collecting the data and preprocessing it, they performed feature extraction and polarity (emotion or sentiment) classification. The results demonstrate that in terms of sentiment classification, the LSIBA-ENN achieves good performance. In the study (Wang et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), a microblog SA model using TEXTCNN and Bayesian techniques was proposed. Their goal is to provide a model that can be interpreted and accurately analyze emotions. To achieve this purpose, they utilized a combination of word2vec, ELMo, and CNN. Aspect-level SA offers a richer and more comprehensive insight compared to document-level SA. This is due to its focus on forecasting sentiment for different aspects within a similar text. The challenge posed by the mentioned level is that various aspect categories within the same text might exhibit diverse polarities. In the above paper, a multi-task model for aspect-level SA was offered, leveraging RoBERTa. Treating individual aspect categories as subtasks, they utilized RoBERTa, to extract features from both text and aspect tokens. Additionally, they implemented a cross-attention mechanism to direct the model's attention toward the most pertinent features corresponding to the specified aspect category. The results of the implementation of the above method show that the proposed method is effective and has succeeded in obtaining good results in the analysis of aspect emotions compared to other action models. In the study (Truşcǎ et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), the authors presented an innovative hybrid approach named HAABSA (Hybrid Aspect-Based SA). The method is developed in two stages. Initially, they substituted non-textual word embeddings with deep contextual embeddings, effectively handling word semantics within a given text. Subsequently, they introduced hierarchical attention, incorporating an extra attention layer into the high-level HAABSA representations. This addition aims to enhance the method's capacity for modeling input data with greater flexibility. They used SemEval 2015 and SemEval 2016 datasets for their work. Classifying genuine sentiments from textual social media data is a challenging and important task. Mahto et al. (Mahto \u0026amp; Yadav, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), presented the HeBiLSTM model to solve this challenge. They introduce an extra Bi-LSTM hidden layer, referred to as the Enhanced GloVe-HeBiLSTM model. Furthermore, the researchers introduced a practical framework for analyzing authentic emotional content within textual data. This framework, named HeBi-CuDNNLSTM, employs a Hierarchical Bi-CuDNNLSTM architecture and leverages the NVIDIA CUDA deep neural network library. They evaluated their proposed models and achieved better results than the basic models. Deep learning architectures (such as CNN and LSTM) have demonstrated their effectiveness in text SA. Nevertheless, challenges like class imbalance and the presence of unlabeled corpora continue to constrain the accuracy of text emotion prediction. Jiang et al. (Jiang et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), To surmount these challenges, they introduced a novel classification model named KSCB. KSCB is obtained from the integration of the Bi-LSTM, K-means++, CNN, and SMOTE models and is used for SA of textual data. Twitter posts exhibit a lack of structure, and informal tone, and contain considerable noise. Additionally, they encompass linguistic intricacies such as polysemy, where words possess multiple meanings. To solve this challenge, Naseem et al. (Naseem et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), introduced a combination of word representation and BiLSTM\u003csup\u003e7\u003c/sup\u003e with attention mechanisms, resulting in enhanced tweet quality. This approach not only tackles textual noise but also handles the presence of multiple semantic interpretations. The results of the experiments show that the introduced way can overcome the ongoing limitations and improve classification accuracy. The use of combined methods can lead to an increase in classification accuracy, many studies have also used the advantages of different methods in a combined way to classify the emotions of texts. Alyoubi et al. (Alyoubi \u0026amp; Sharma, \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), presented a hybrid approach using CNN and RNN networks. They utilized the BERT model for producing word embeddings, performed feature extraction from CNN, and employed a bidirectional RNN to harness contextual and temporal features. Their combined method achieved 96% accuracy. In the study (Jain et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), a hybrid method was presented using CNN-LSTM networks to analyze consumer sentiments. The proposed method was employed by batch normalization (BN), max pooling, and dropout to attain outcomes. Their hybrid model achieved 91.3% accuracy in SA on used datasets. Another study (Yue \u0026amp; Li, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), also used the combination of CNN-BiLSTM to analyze the sentiments of the text and achieved a classification accuracy of 91.48%. They used the Word2vec model to embed the words.\u003c/p\u003e \u003cp\u003eThis paper (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) introduces an innovative approach to automatically categorize the sentiment of Twitter messages through the utilization of ML algorithms. This method leverages emoticon data as somewhat noisy labels, demonstrating an impressive accuracy rate of over 80% when trained with this particular dataset. Additionally, the paper delves into the preprocessing steps essential for achieving high accuracy and emphasizes the significance of incorporating emoticon-laden tweets for distant supervised learning. Twitter SA is a computational process that categorizes tweets into positive or negative ones. The main challenge is identifying the most suitable sentiment classifier. This paper (Saleena, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) proposes an ensemble classifier that combines base learning classifiers, improving performance and accuracy. Results show the proposed classifier outperforms stand-alone classifiers and majority voting ensemble classifiers. This paper (Siddiqua et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) proposes a method for SA on Twitter, combining a rule-based classifier with a majority voting ensemble of supervised classifiers. The rule-based classifier is based on emoticons and sentiment-bearing words, while the supervised classifiers are trained using Twitter-specific, textual, parts-of-speech, lexicon-based, and bag-of-words features. The method uses chi-square statistics and information gain to select the best feature combination. Experimental results show its effectiveness. Managing unstructured data from social media can be a formidable challenge. To tackle this issue, deep learning algorithms provide a suitable solution for addressing these complexities. This paper (Tyagi et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) presents a CNN-LSTM-based deep learning method for managing unstructured social media data. The method automatically extracts features for sentiment analysis and classification of reviews, achieving better performance on a benchmark dataset compared to baseline machine learning methods. The paper (Tan, Lee, Lim, et al., 2022) presents an ensemble hybrid DL model for SA, combining three models: RoBERTa, LSTM, BiLSTM, and GRU. The model projected textual input sequences, captured long-range dependencies, and combined predictions using averaging ensemble and majority voting. The model outperformed state-of-the-art methods with accuracy of 94.9%, 91.77%, and 89.81% on IMDb, Twitter US Airline Sentiment dataset, and Sentiment140 dataset. The paper (Hossen et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) introduces a deep learning methodology employing LSTM and GRUs to train hotel review data for the purpose of identifying customer opinions. The study (Jain et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) presents a hybrid CNN-LSTM model for SA using dropout, max pooling, and batch normalization. Experimental analysis on Airlinequality and Twitter sentiment datasets shows the model outperforms classical ML models with 91.3% accuracy in SA. SA is crucial in understanding the sentiment context of social media messages, which can impact strategic decision-making in business and politics. Challenges include lexical diversity, imbalanced datasets, and long-distance dependencies. In study(Tan, Lee, Anbananthen, et al., 2022), a hybrid deep learning method was proposed, leveraging the strengths of both sequence and Transformer models while mitigating their limitations. The model combines a robustly optimized BERT approach with LSTM for SA, effectively mapping words into a concise, semantically meaningful word embedding space.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Ensemble learning\u003c/h2\u003e \u003cp\u003eEnsemble learning is an ML technique that involves combining multiple models (often referred to as \"base\" or \"weak\" models) to create a stronger, more robust model. The idea behind ensemble learning is that by combining the predictions of multiple models, the overall performance can be improved compared to that of any individual model. Ensembles are often used to increase the accuracy, stability, and generalization of ML algorithms. There are several common techniques for creating ensemble models: 1) Bagging or Bootstrap Aggregating, 2) Boosting, 3) Stacking, and 4) Voting. In various studies, ensemble learning has been used to achieve higher accuracy in the field of SA. In the study (Fersini et al., \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2014\u003c/span\u003e), by using ensemble learning, they seek to increase accuracy and reduce noise and ambiguity of language. Ghorbanali et al. (Ghorbanali et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), performed text classification using several weighted ensemble CNN networks.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Fusion\u003c/h2\u003e \u003cp\u003eFusion in the context of ML and data analysis refers to the process of combining information, features, or predictions from multiple sources or models to create a unified and more informative result. Fusion aims to leverage the strengths of each source or model to improve the overall accuracy, robustness, or understanding of the outcome. Fusion techniques are commonly used in various fields, including computer vision, NLP, remote sensing, sensor networks, and more. There are different types of fusion techniques: 1) Data Fusion, 2) Feature Fusion, 3) Model Fusion, 4) Sensor Fusion, and 5) Information Fusion. Fusion is particularly useful when dealing with complex, noisy, or incomplete data, as well as when aiming to improve the overall performance and reliability of a system. The choice of fusion technique depends on the nature of the data and the specific goals of the analysis or task at hand. In the study(Huang et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), fusion is divided into three levels: early fusion, intermediate fusion, and late fusion.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. The proposed model","content":"\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, we have used three classifiers of CNN for our proposed model. As can be seen, \u003cb\u003ePDO\u003c/b\u003e represents the previous decision and \u003cb\u003eNPO\u003c/b\u003e represents the new decision. In the proposed model, first, the input text is embedded using the hybrid approach of word embedding, then it is entered into the CNN networks, and in the final stage, the results are fused at the decision level using Bayesian conditional. The details of the proposed model are given below.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Dataset\u003c/h2\u003e \u003cp\u003eWe have used the IMDB dataset to implement and train our proposed model. The IMDB dataset is one of the most renowned and extensive collections for SA, particularly in the domain of NLP. This dataset is tailored for SA tasks involving textual reviews and critiques of movies and television shows. Each review or critique is accompanied by a numerical rating, representing the writer's overall sentiment towards the content, which can be either positive or negative. This dataset comprises over 50,000 reviews, approximately evenly split between positive and negative sentiments, making it a valuable resource for conducting comprehensive and extensive analyses. Each review is associated with a numerical score, serving as the ground truth label, indicating whether the sentiment is positive or negative. The texts within this dataset represent genuine reviews and critiques authored by individuals for various films and TV programs. These texts faithfully capture the sentiments of the authors toward the content. The number of positive and negative reviews is balanced within the dataset, ensuring statistical validity for empirical analyses. Its details are given in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. We have preprocessed the data from the above dataset before using it. To access additional details about the dataset, kindly review the provided link \u003csup\u003e8\u003c/sup\u003e. We have used two other datasets named Twitter US Airline \u003csup\u003e9\u003c/sup\u003e (This data originally came from Crowdflower's\u003csup\u003e10\u003c/sup\u003e Data for Everyone library.) and Sentiment140 (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). Twitter US Airline dataset comprises tweets related to major U.S. airlines, collected since February 2015. It encompasses various attributes, including Twitter user IDs, sentiment confidence scores, reasons for negativity and positivity, retweet counts, tweet content, date, time, and location information. Additionally, the sentiment analysis dataset includes tagged reviews for numerous Amazon products, offering both positive and negative sentiments. These reviews are associated with ratings on a scale of 1 to 5 stars, which can be converted into binary format if needed. Next dataset is known as Sentiment140 and has been compiled by extracting 1,600,000 tweets through the Twitter API. These tweets have been carefully labeled and can be employed for SA purposes. Each dataset is divided into three parts: a 70% training set, a 15% validation set, and a 15% testing set.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSentiment\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIMDB\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTwitter US Airline\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSentiment140\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePositive (1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2363\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e800000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNegative (0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e25000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9178\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e800000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNeutral (2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3099\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e50000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e14640\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1600000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Word embedding layer (Hybrid BERT and GloVe)\u003c/h2\u003e \u003cp\u003eThe fusion of BERT and GloVe models in a hybrid architecture involves a synergistic approach to harnessing the strengths of both methodologies. By integrating BERT's contextualized embeddings, which excel in capturing intricate contextual nuances of language, with GloVe's global and statistical word representations, which encode semantic relationships and co-occurrence patterns, the hybrid model achieves a comprehensive understanding of the text. This integration enables the model to not only grasp fine-grained semantic nuances at the token level through BERT's embeddings but also grasp broader semantic similarities and associations between words using GloVe's embeddings. The amalgamation of these embeddings within the hybrid architecture results in an enriched feature space that encapsulates a more holistic view of the text's meaning. This holistic representation empowers the model to perform more effectively in various natural language processing tasks such as sentiment analysis, text classification, and information retrieval. The combination also potentially mitigates the limitations of each model, while BERT might struggle with out-of-vocabulary terms, GloVe comprehensive word vectors can contribute to addressing this challenge. Conversely, BERT's fine-grained context understanding can enhance the interpretability and context awareness of GloVe embeddings. Ultimately, this hybridization capitalizes on the strengths of both BERT and GloVe, leading to improved performance, enhanced semantic representation, and potentially more robust generalization across diverse textual datasets and tasks in the field of natural language processing. The architecture of the proposed hybrid word embedding model is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The dimensions of this combination are also shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. This combination is done in the following five steps:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eInput text\u003c/b\u003e: The initial stage involves feeding the input text into the architecture.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eTokenization: In this step, the input text is tokenized, breaking it down into individual tokens or words.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eBERT embeddings\u003c/b\u003e: The tokenized text is then processed by the BERT model, generating contextualized embeddings for each token (We using the BERT model to extract high-dimensional features like \"semantic embeddings.\").\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eGloVe embeddings\u003c/b\u003e: Simultaneously, the tokenized text is also used to obtain GloVe embeddings, which capture global semantic information for each token (We using the GloVe model to extract low-dimensional features.).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eCombination layer\u003c/b\u003e: The BERT and GloVe embeddings are combined using a specific layer that merges their information, allowing the model to capture both local and global context (We then use the combined vectors as input features for a new model, such as an ensemble CNNs.).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eThese stages collectively define the hybrid architecture of combining BERT and GloVe embeddings for various downstream NLP tasksThe concept of this process is given in Eq.\u0026nbsp;\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$${E}_{BERT: BERT{\\prime }s embedding vector for a specific token.}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$${E}_{GloVe: GloVe{\\prime }s embedding vector for the same token.}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$${E}_{Combined= \\frac{{E}_{BERT}+ {E}_{GloVe}}{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThis averaging approach provides a balanced integration of both BERT's contextually rich embeddings, which capture fine-grained semantics, and GloVe's global semantic embeddings, which represent broader word relationships. The combined embedding \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({E}_{Combined}\\)\u003c/span\u003e\u003c/span\u003e effectively captures both local context and global semantics, enhancing the model's ability to interpret and comprehend text comprehensively. The strength of this hybrid approach lies in its ability to enhance the semantic richness of embeddings, leading to improved performance and robustness in various natural language processing tasks. There are other ways to combine these two models. The choice of the best fusion method for combining BERT and GloVe embeddings depends on the specific tasks and data at hand. Each of the combination techniques (such as averaging, concatenation, weighted combination, fine-tuning, etc.) comes with its advantages and limitations, and selecting the optimal approach should consider the following factors:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eTask nature\u003c/b\u003e: The type of task at hand (such as sentiment analysis, text classification, translation, etc.) significantly influences the choice of the fusion method.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData constraints\u003c/b\u003e: The volume and nature of our training and testing data are crucial.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eSpecific cases\u003c/b\u003e: Certain specialized cases might require unique fusion strategies. For instance, dealing with anomalous data (like fake news) could necessitate incorporating anomaly detection techniques within the fusion.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eModel complexity\u003c/b\u003e: The inherent complexity of our model plays a role as well. If your model is already deep and intricate, adding additional layers for fusion might exacerbate complexity challenges.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eEmpirical results\u003c/b\u003e: Experimentation and evaluating models using various metrics are usually essential to determine the best approach for your specific tasks. Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows other equations.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eEquations of word embedding combination and other method\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCombining BERT and GloVe\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEquations and methods\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAveraging\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({E}_{Combined= \\frac{{E}_{BERT}+ {E}_{GloVe}}{2}}\\)\u003c/span\u003e\u003c/span\u003e (1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConcatenation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({E}_{Combined= }\\left[{E}_{BERT},{E}_{GloVe}\\right]\\)\u003c/span\u003e\u003c/span\u003e (2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeighted combination\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({E}_{Combined= }{{W}_{BERT}\\times E}_{BERT}+{W}_{GloVe}{\\times E}_{GloVe}\\)\u003c/span\u003e\u003c/span\u003e (3)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOther method\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFine-tuning, feature fusion layers, and adaptive combination\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\left[ \\right]\\)\u003c/span\u003e\u003c/span\u003e denotes vector concatenation. Where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{BERT}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{GloVe}\\)\u003c/span\u003e\u003c/span\u003e are weight coefficients. Each of these methods offers a unique way to leverage the complementary strengths of BERT and GloVe, and the choice of approach depends on the specific nature of the task, the available data, and the desired trade-off between contextual understanding and semantic richness.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Classification\u003c/h2\u003e \u003cp\u003eTo extract features, and classify and predict emotions, we have used the ensemble learning approach using three CNN classifications. CNN networks are generally composed of four parts, in the first layer, which is an input layer, its input is the output of the word embedding layer. The next layer, called the Convolution layer, uses filters to extract features (We used Convolution1D for texts). The next layer is called pooling, this layer is used to select features and reduce their dimensions. Classification of texts is done in the last layer using a fully connected layer (FC). Ensembling multiple CNN classifiers can often lead to improved performance and better generalization. We built three parallel CNN classifiers and integrated their final results at the decision level. This ensemble approach allows each CNN classifier to learn different features from the text data, which can lead to improved performance compared to a single classifier. We have used a different configuration for each classifier because the combination can be useful when the models have different structures or are trained on different datasets. Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e summarizes the parameters used in the ensemble of CNN classifiers.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eCNN classifier parameters\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClassifier 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eClassifier 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eClassifier 3\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVocabulary Size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10,000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSequence Length\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEmbedding Dim\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNum Filters\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e128\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFilter Size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDropout Rate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHidden Units\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOutput Classes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Combined approach using Bayesian conditional probability (Fusion)\u003c/h2\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eInitial assumption for previous decisions\u003c/b\u003e: Define initial probabilities for your previous decisions. For instance, if your past decisions were binary (positive/negative), you can assign an initial probability of selecting \"positive\" as 0.7 and \"negative\" as 0.3.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eCalculate new probabilities based on new results\u003c/b\u003e: For each of the new classifiers, calculate the conditional probabilities of each of the previous decisions given the new results. These probabilities can be computed using Bayesian conditional probability. The calculations can take into account various factors, such as previous decision outcomes and similarities with previous decisions.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eCombine new and previous probabilities\u003c/b\u003e: Now that you have the probabilities of choosing previous decisions and the probabilities calculated based on the new classifier results, you can combine these probabilities using specified weights to strike a balance between the new information and the previous decisions.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eFinal decision\u003c/b\u003e: Select the decision with the highest final probability as your ultimate decision. Here is how this process works:\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eProbabilities of previous decisions: P(Decision_A), P(Decision_B), ...\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eProbabilities calculated based on new classifier results: P(Decision_A|NewResults), P(Decision_B|NewResults), ...\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCombining probabilities with specified weights: Final_P(Decision_A)\u0026thinsp;=\u0026thinsp;alpha * P(Decision_A) + (1 - alpha) * P(Decision_A|NewResults)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFinal Decision: Choose the decision with the highest final probability.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eIn this approach, alpha represents a weight that you can adjust. This weight determines the influence of the previous decisions on the final decision.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Result and implications","content":"\u003cp\u003eWe have used the following evaluation criteria to evaluate our model (Eq.\u0026nbsp;4\u0026ndash;9). It is noteworthy that, the specific evaluation criteria used for a model will depend on the specific task at hand. For example, if the model is classifying samples into two classes, one might use the accuracy, sensitivity, or F1 score. If the model classifies samples into multiple classes, AUC may be used. The hyperparameters for the model are shown in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. Also, Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e and Tables\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e show the results of implementing the proposed model on the IMDB, Twitter US Airline, and Sentiment140 datasets. We have used the confusion matrix to analyze the proposed model. The confusion matrix is a tool that can be used to evaluate the performance of ML classification models. The confusion matrix is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. As the confusion matrices show, the proposed model has been able to get acceptable results on the test data. It shows how well the model correctly classified the positive and negative samples. The negative, positive, and neutral polarities are represented in the confusion matrix with 0, 1, and 2 respectively.\u003c/p\u003e \u003cp\u003eAccuracy (ACC) = (TP\u0026thinsp;+\u0026thinsp;TN) / (TP\u0026thinsp;+\u0026thinsp;TN\u0026thinsp;+\u0026thinsp;FP\u0026thinsp;+\u0026thinsp;FN) (4)\u003c/p\u003e \u003cp\u003eWhere:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eTP\u0026thinsp;=\u0026thinsp;True Positives\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTN\u0026thinsp;=\u0026thinsp;True Negatives\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFP\u0026thinsp;=\u0026thinsp;False Positives\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFN\u0026thinsp;=\u0026thinsp;False Negatives\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eSensitivity\u0026thinsp;=\u0026thinsp;TP / (TP\u0026thinsp;+\u0026thinsp;FN) (5)\u003c/p\u003e \u003cp\u003eSpecificity\u0026thinsp;=\u0026thinsp;TN / (TN\u0026thinsp;+\u0026thinsp;FP) (6)\u003c/p\u003e \u003cp\u003ePrecision (P)\u0026thinsp;=\u0026thinsp;TP / (TP\u0026thinsp;+\u0026thinsp;FP) (7)\u003c/p\u003e \u003cp\u003eRecall (R)\u0026thinsp;=\u0026thinsp;TP / (TP\u0026thinsp;+\u0026thinsp;FN) (8)\u003c/p\u003e \u003cp\u003eF1-Score\u0026thinsp;=\u0026thinsp;2 \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\times\\)\u003c/span\u003e\u003c/span\u003e (P \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\times\\)\u003c/span\u003e\u003c/span\u003e R) / (P\u0026thinsp;+\u0026thinsp;R) (9)\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eHyperparameters for the proposed model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHyperparameters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eValue\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLearning rate:\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1e-5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeight decay:\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEpochs:\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBatch size:\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparing the results of other studies on the IMDB dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eACC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eYear\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLexicon-based (Giatsoglou et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2017\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.830\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.906\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2017\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCombination of CNN and LSTM (Yenter \u0026amp; Verma, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2017\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.895\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2017\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN-BiLSTM-Attention(Wang et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2019\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.903\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2019\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMultiple CNN and LSTM layers (Qaisar, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2020\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.899\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2020\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHybrid lexicon and neural networks (Shaukat et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.919\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2020\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN_GRU(Zouzou \u0026amp; El Azami, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2021\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.863\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2021\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDoc2vec\u0026thinsp;+\u0026thinsp;CNN\u0026thinsp;+\u0026thinsp;BiLSTM\u0026thinsp;+\u0026thinsp;Attention(Zhou et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2022\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.913\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnsemble hybrid deep learning model (Tan, Lee, Lim, et al., 2022)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.949\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN-LSTM(Mishra \u0026amp; Patil, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2023\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.850\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT-based (Arora et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2023\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.924\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.921\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.892\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.964\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMBi-GRUMCONV (Başarslan \u0026amp; Kayaalp, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2023\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.953\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOur model\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.958\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.980\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.970\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.990\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e2023\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparing the results of other studies on the Twitter US Airline dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eACC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eYear\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGRU (Hossen et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2021\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.785\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.720\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.730\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.710\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2021\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLSTM (Hossen et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2021\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.775\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.690\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.710\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.690\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2021\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN-LSTM (Jain et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.760\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.690\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.680\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.690\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2021\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRoBERTa-LSTM (Tan, Lee, Anbananthen, et al., 2022)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.913\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.910\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.910\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.910\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnsemble hybrid deep learning model (Tan, Lee, Lim, et al., 2022)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.917\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOur model\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.946\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.936\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.932\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.941\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e2023\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparing the results of other studies on the Sentiment140 dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eACC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eYear\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNaive Bayes (Unigram) (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.813\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2009\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMaxEnt (Unigram) (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.805\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2009\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVM (Unigram) (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.822\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2009\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNaive Bayes (Unigram\u0026thinsp;+\u0026thinsp;Bigram) (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.827\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2009\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMaxEnt (Unigram\u0026thinsp;+\u0026thinsp;Bigram) (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.830\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2009\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVM (Unigram\u0026thinsp;+\u0026thinsp;Bigram) (Go et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2009\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.816\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2009\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRule-basedclassifier (Siddiqua et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2016\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.877\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.884\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.843\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.923\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2016\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnsemble method (Saleena, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2018\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.733\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2018\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN-LSTM (Tyagi et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2020\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.812\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2020\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnsemble hybrid deep learning model (Tan, Lee, Lim, et al., 2022)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.898\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOur model\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.925\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.937\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.940\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.935\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e2023\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Discussion of results and implications\u003c/h2\u003e \u003cp\u003eIn this section, we delve into the theoretical and practical implications of our research, shedding light on its distinctive features in comparison to existing work.\u003c/p\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e4.1.1 Theoretical Implications\u003c/h2\u003e \u003cp\u003eOur study advances the field of SA by introducing a novel approach that combines BERT and GloVe word embedding models. By leveraging the contextual understanding of BERT and the global statistical patterns captured by GloVe, we offer a more comprehensive perspective on SA. This integration of diverse word embedding techniques provides a deeper understanding of text sentiment, contributing to the theoretical foundation of SA methodologies.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003e4.1.2 Practical implications\u003c/h2\u003e \u003cp\u003eFrom a practical standpoint, our research offers tangible benefits. The combination of BERT and GloVe significantly enhances SA accuracy, as demonstrated by our experimental results. This improved accuracy holds practical significance in various applications, including social media monitoring, customer feedback analysis, and product review sentiment assessment. Businesses and organizations can leverage our approach to gain more accurate insights into public sentiment, thereby making informed decisions and improving user experiences.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003e4.1.3 Distinguishing factors\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eWhile existing work often relies on single-word embedding models, our research innovatively integrates two prominent models. This distinctive approach yields improved sentiment analysis performance, as evidenced by the percentage improvement in accuracy on our sample dataset. The use of Bayesian Conditional for result integration further sets our work apart, offering a unique perspective on combining sentiment analysis outcomes. The proposed fusion method is not only effective, but it is simpler than other methods such as Dempster Shafer evidence theory, and it is also suitable for incomplete data and systems with uncertainty.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section3\"\u003e \u003ch2\u003e4.1.4 Implementation environment and hardware\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eWe have implemented our proposed model using the Python 3.x programming language and the TensorFlow library. The specifications of the computer system used are given in Table\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e9\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eInformation about the implementation environment\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSystem\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eConfiguration\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOperating system\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWindows 10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCPU\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIntel\u0026reg; Core\u0026trade; i7-7700K @ 4.2GHz\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRAM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e32 GB\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eH.D.D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1 TB\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eWe have covered the challenge of incomplete data, uncertainty, and increasing accuracy and reliability by presenting a hybrid method along with an ensemble learning technique. First, we presented a hybrid model of word embedding to apply the advantages of BERT and GloVe. Combining GloVe and BERT, two popular NLP techniques, could offer several advantages for various NLP tasks. which consists of semantic understanding, feature combination, efficiency, reduced Data requirements, and interpretable semantic features. It's important to note that the effectiveness of combining GloVe and BERT depends on the specific task and dataset. We used 3 classifiers based on ensemble CNN networks for classification. CNNs in NLP are particularly effective for tasks like text classification, sentiment analysis, topic modeling, and even sequence tagging. Ensemble techniques provide numerous benefits compared to individual models. These advantages encompass enhanced precision and accuracy, particularly when tackling intricate and chaotic issues. Furthermore, they mitigate the chances of falling into the traps of overfitting and underfitting by striking a balance between bias and variance. We fused the results obtained from the classifications at the decision level. Fusion at this level allows us to obtain the final result even if part of the data and results were insufficient. We did the fusion of results using Bayesian conditional and based on previous decisions and new decisions. Bayesian conditional fusion offers a flexible and adaptive approach to decision-making that can lead to more informed, accurate, and reliable outcomes, particularly in scenarios where both historical and new data are relevant. Our proposed model was evaluated on 3 datasets and obtained good results.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset used in this study can be made available with a reasonable request to the corresponding author.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAuthor declare that they have no conflict of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding Declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors did not receive support from any organization for the submitted work. No funding was received to assist with the preparation of this manuscript. No funding was received for conducting this study.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAlyoubi, K. H., \u0026amp; Sharma, A. (2023). A Deep CRNN-Based Sentiment Analysis System with Hybrid BERT Embedding. \u003cem\u003eInternational Journal of Pattern Recognition and Artificial Intelligence\u003c/em\u003e,\u003cem\u003e 37\u003c/em\u003e(05), 2352006. \u003c/li\u003e\n\u003cli\u003eArora, K., Gupta, N., \u0026amp; Pathak, S. (2023). Sentimental Analysis on IMDb Movies Review using BERT. 2023 4th International Conference on Electronics and Sustainable Communication Systems (ICESC), \u003c/li\u003e\n\u003cli\u003eBaşarslan, M. S., \u0026amp; Kayaalp, F. (2023). MBi-GRUMCONV: A novel Multi Bi-GRU and Multi CNN-Based deep learning model for social media sentiment analysis. \u003cem\u003eJournal of Cloud Computing\u003c/em\u003e,\u003cem\u003e 12\u003c/em\u003e(1), 5. \u003c/li\u003e\n\u003cli\u003eFersini, E., Messina, E., \u0026amp; Pozzi, F. A. (2014). Sentiment analysis: Bayesian ensemble learning. \u003cem\u003eDecision Support Systems\u003c/em\u003e,\u003cem\u003e 68\u003c/em\u003e, 26-38. \u003c/li\u003e\n\u003cli\u003eGhorbanali, A., \u0026amp; Sohrabi, M. K. (2023a). A comprehensive survey on deep learning-based approaches for multimodal sentiment analysis. \u003cem\u003eArtificial Intelligence Review\u003c/em\u003e, 1-34. \u003c/li\u003e\n\u003cli\u003eGhorbanali, A., \u0026amp; Sohrabi, M. K. (2023b). Exploiting bi-directional deep neural networks for multi-domain sentiment analysis using capsule network. \u003cem\u003eMultimedia Tools and Applications\u003c/em\u003e, 1-18. \u003c/li\u003e\n\u003cli\u003eGhorbanali, A., Sohrabi, M. K., \u0026amp; Yaghmaee, F. (2022). Ensemble transfer learning-based multimodal sentiment analysis using weighted convolutional neural networks. \u003cem\u003eInformation Processing \u0026amp; Management\u003c/em\u003e,\u003cem\u003e 59\u003c/em\u003e(3), 102929. \u003c/li\u003e\n\u003cli\u003eGiatsoglou, M., Vozalis, M. G., Diamantaras, K., Vakali, A., Sarigiannidis, G., \u0026amp; Chatzisavvas, K. C. (2017). Sentiment analysis leveraging emotions and word embeddings. \u003cem\u003eExpert Systems with Applications\u003c/em\u003e,\u003cem\u003e 69\u003c/em\u003e, 214-224. \u003c/li\u003e\n\u003cli\u003eGo, A., Bhayani, R., \u0026amp; Huang, L. (2009). Twitter sentiment classification using distant supervision. \u003cem\u003eCS224N project report, Stanford\u003c/em\u003e,\u003cem\u003e 1\u003c/em\u003e(12), 2009. \u003c/li\u003e\n\u003cli\u003eHossen, M. S., Jony, A. H., Tabassum, T., Islam, M. T., Rahman, M. M., \u0026amp; Khatun, T. (2021). Hotel review analysis for the prediction of business using deep learning approach. 2021 International Conference on Artificial Intelligence and Smart Systems (ICAIS), \u003c/li\u003e\n\u003cli\u003eHuang, F., Zhang, X., Zhao, Z., Xu, J., \u0026amp; Li, Z. (2019). Image\u0026ndash;text sentiment analysis via deep multimodal attentive fusion. \u003cem\u003eKnowledge-based systems\u003c/em\u003e,\u003cem\u003e 167\u003c/em\u003e, 26-37. \u003c/li\u003e\n\u003cli\u003eJain, P. K., Saravanan, V., \u0026amp; Pamula, R. (2021). A hybrid CNN-LSTM: A deep learning approach for consumer sentiment analysis using qualitative user-generated contents. \u003cem\u003eTransactions on Asian and Low-Resource Language Information Processing\u003c/em\u003e,\u003cem\u003e 20\u003c/em\u003e(5), 1-15. \u003c/li\u003e\n\u003cli\u003eJiang, W., Zhou, K., Xiong, C., Du, G., Ou, C., \u0026amp; Zhang, J. (2023). KSCB: A novel unsupervised method for text sentiment analysis. \u003cem\u003eApplied Intelligence\u003c/em\u003e,\u003cem\u003e 53\u003c/em\u003e(1), 301-311. \u003c/li\u003e\n\u003cli\u003eKim, Y. (2014). Convolutional neural networks for sentence classification. \u003cem\u003earXiv preprint arXiv:1408.5882\u003c/em\u003e. \u003c/li\u003e\n\u003cli\u003eMahto, D., \u0026amp; Yadav, S. C. (2023). Emotion prediction for textual data using GloVe based HeBi-CuDNNLSTM model. \u003cem\u003eMultimedia Tools and Applications\u003c/em\u003e, 1-26. \u003c/li\u003e\n\u003cli\u003eMishra, M., \u0026amp; Patil, A. (2023). Sentiment Prediction of IMDb Movie Reviews Using CNN-LSTM Approach. 2023 International Conference on Control, Communication and Computing (ICCC), \u003c/li\u003e\n\u003cli\u003eNaseem, U., Khan, S. K., Razzak, I., \u0026amp; Hameed, I. A. (2019). Hybrid words representation for airlines sentiment analysis. AI 2019: Advances in Artificial Intelligence: 32nd Australasian Joint Conference, Adelaide, SA, Australia, December 2\u0026ndash;5, 2019, Proceedings 32, \u003c/li\u003e\n\u003cli\u003eNi, R., \u0026amp; Cao, H. (2020). Sentiment Analysis based on GloVe and LSTM-GRU. 2020 39th Chinese control conference (CCC), \u003c/li\u003e\n\u003cli\u003eQaisar, S. M. (2020). Sentiment analysis of IMDb movie reviews using long short-term memory. 2020 2nd International Conference on Computer and Information Sciences (ICCIS), \u003c/li\u003e\n\u003cli\u003eSaleena, N. (2018). An ensemble classification system for twitter sentiment analysis. \u003cem\u003eProcedia Computer Science\u003c/em\u003e,\u003cem\u003e 132\u003c/em\u003e, 937-946. \u003c/li\u003e\n\u003cli\u003eShaukat, Z., Zulfiqar, A. A., Xiao, C., Azeem, M., \u0026amp; Mahmood, T. (2020). Sentiment analysis on IMDB using lexicon and neural networks. \u003cem\u003eSN Applied Sciences\u003c/em\u003e,\u003cem\u003e 2\u003c/em\u003e, 1-10. \u003c/li\u003e\n\u003cli\u003eSiddiqua, U. A., Ahsan, T., \u0026amp; Chy, A. N. (2016). Combining a rule-based classifier with ensemble of feature sets and machine learning techniques for sentiment analysis on microblog. 2016 19th international conference on computer and information technology (ICCIT), \u003c/li\u003e\n\u003cli\u003eTan, K. L., Lee, C. P., Anbananthen, K. S. M., \u0026amp; Lim, K. M. (2022). RoBERTa-LSTM: a hybrid model for sentiment analysis with transformer and recurrent neural network. \u003cem\u003eIEEE Access\u003c/em\u003e,\u003cem\u003e 10\u003c/em\u003e, 21517-21525. \u003c/li\u003e\n\u003cli\u003eTan, K. L., Lee, C. P., Lim, K. M., \u0026amp; Anbananthen, K. S. M. (2022). Sentiment analysis with ensemble hybrid deep learning model. \u003cem\u003eIEEE Access\u003c/em\u003e,\u003cem\u003e 10\u003c/em\u003e, 103694-103704. \u003c/li\u003e\n\u003cli\u003eTruşcǎ, M. M., Wassenberg, D., Frasincar, F., \u0026amp; Dekker, R. (2020). A hybrid approach for aspect-based sentiment analysis using deep contextual word embeddings and hierarchical attention. Web Engineering: 20th International Conference, ICWE 2020, Helsinki, Finland, June 9\u0026ndash;12, 2020, Proceedings 20, \u003c/li\u003e\n\u003cli\u003eTyagi, V., Kumar, A., \u0026amp; Das, S. (2020). Sentiment analysis on twitter data using deep learning approach. 2020 2nd international conference on advances in computing, communication control and networking (ICACCCN), \u003c/li\u003e\n\u003cli\u003eWang, L., Liu, C. H., Cai, D., Zhao, T., \u0026amp; Wang, M. (2019). Text sentiment analysis based on CNN-BiLSTM network and attention model. \u003cem\u003eJournal of Wuhan Institute of Technology\u003c/em\u003e,\u003cem\u003e 41\u003c/em\u003e(4), 386-391. \u003c/li\u003e\n\u003cli\u003eWang, Z., Yao, L., Shao, X., \u0026amp; Wang, H. (2023). A combination of TEXTCNN model and Bayesian classifier for microblog sentiment analysis. \u003cem\u003eJournal of Combinatorial Optimization\u003c/em\u003e,\u003cem\u003e 45\u003c/em\u003e(4), 109. \u003c/li\u003e\n\u003cli\u003eYenter, A., \u0026amp; Verma, A. (2017). Deep CNN-LSTM with combined kernels from multiple branches for IMDb review sentiment analysis. 2017 IEEE 8th annual ubiquitous computing, electronics and mobile communication conference (UEMCON), \u003c/li\u003e\n\u003cli\u003eYue, W., \u0026amp; Li, L. (2020). Sentiment analysis using Word2vec-CNN-BiLSTM classification. 2020 seventh international conference on social networks analysis, management and security (SNAMS), \u003c/li\u003e\n\u003cli\u003eZhang, Y., \u0026amp; Wallace, B. (2015). A sensitivity analysis of (and practitioners\u0026apos; guide to) convolutional neural networks for sentence classification. \u003cem\u003earXiv preprint arXiv:1510.03820\u003c/em\u003e. \u003c/li\u003e\n\u003cli\u003eZhao, H., Liu, Z., Yao, X., \u0026amp; Yang, Q. (2021). A machine learning-based sentiment analysis of online product reviews with a novel term weighting and feature selection approach. \u003cem\u003eInformation Processing \u0026amp; Management\u003c/em\u003e,\u003cem\u003e 58\u003c/em\u003e(5), 102656. \u003c/li\u003e\n\u003cli\u003eZhou, Y., Zhang, Q., Wang, D., \u0026amp; Gu, X. (2022). Text Sentiment Analysis Based on a New Hybrid Network Model. \u003cem\u003eComputational Intelligence and Neuroscience\u003c/em\u003e,\u003cem\u003e 2022\u003c/em\u003e. \u003c/li\u003e\n\u003cli\u003eZouzou, A., \u0026amp; El Azami, I. (2021). Text sentiment analysis with CNN \u0026amp; GRU model using GloVe. 2021 Fifth International Conference On Intelligent Computing in Data Sciences (ICDS), \u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Footnotes","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003e Feedforward Neural Networks\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e Convolutional Neural Networks\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e Recurrent Neural Networks\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e Long Short-Term Memory\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e Gated Recurrent Units\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e Generative Adversarial Networks\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e Bi-directional Long Short Term Memory\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ai.stanford.edu/~amaas/data/sentiment/\u003c/span\u003e\u003cspan address=\"https://ai.stanford.edu/~amaas/data/sentiment/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kaggle.com/datasets/crowdflower/twitter-airline-sentiment\u003c/span\u003e\u003cspan address=\"https://www.kaggle.com/datasets/crowdflower/twitter-airline-sentiment\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.crowdflower.com/data-for-everyone\u003c/span\u003e\u003cspan address=\"http://www.crowdflower.com/data-for-everyone\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Sentiment analysis, Opinion mining, BERT, GloVe, Ensemble, Bayesian","lastPublishedDoi":"10.21203/rs.3.rs-3961336/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3961336/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSentiment analysis (SA) of text holds a pivotal role in today's digital age, particularly within the realm of social media networks. The analysis of textual sentiments emerges as a critical facet of NLP. In social media, individuals extensively engage with a multitude of texts and opinions. SA empowers us to delve into and discover these opinions, sentiments, and viewpoints, thereby extracting valuable insights on a wide array of subjects. The significance of word embeddings for processing textual data lies in their ability to represent words as dense vectors, enabling machines to capture semantic relationships and contextual nuances, thereby enhancing various natural language processing tasks. There are two popular and famous models, BERT and GloVe, for embedding words. Currently, GloVe is considered one of the most precise approaches. However, this method does not take into account the sentiment information present in texts. Consequently, we opted to utilize pre-trained BERT models, which have been trained on extensive text corpora, in combination with the GloVe model to address this limitation. This study leverages a hybrid word embedding model combining BERT and GloVe. Several classifiers are employed to analyze text sentiment. At the decision level, we employ Bayesian Conditional to integrate current results with prior decisions. When combining previous decisions with new ones, the model achieves higher accuracy by refining or adjusting decisions in light of new evidence. Our approach demonstrates notable results, showcasing its practical significance. The results of the experiments on IMDB, Sentiment140, and Twitter US Airline datasets demonstrate that the proposed approach has achieved favorable results, with accuracies of 0.958, 0.925, and 0.946 respectively. These results are considered acceptable when compared to those of other similar studies.\u003c/p\u003e","manuscriptTitle":"Social network textual data classification through a hybrid word embedding approach and Bayesian conditional-based multiple classifiers","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-28 20:27:25","doi":"10.21203/rs.3.rs-3961336/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2bf62e78-b66a-4057-8283-89f01c41d99e","owner":[],"postedDate":"February 28th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-11-11T19:38:24+00:00","versionOfRecord":[],"versionCreatedAt":"2024-02-28 20:27:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3961336","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3961336","identity":"rs-3961336","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0