A Survey Of Face Emotion Recognition Using Deep Learning Methods | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article A Survey Of Face Emotion Recognition Using Deep Learning Methods Prof. Ketan Sarvakar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6794812/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The Deep learning techniques have significantly improved face emotion identification, a crucial component of human-computer interaction. This study examines facial emotion categorization with a variety of deep learning techniques, emphasizing well-known models such as VGG-19, ResNet-50, Inception-V3, and MobileNet. We investigate the effectiveness and constraints of these neural network systems through an analysis of different approaches for recognizing facial expressions. The pre-processing techniques, indicators of performance, and datasets used to assess these frameworks are all covered in the inquiry. This evaluation demonstrates the way The VGG model-19, ResNet-50, Inception-V3, and Mobile Network perform facial recognition of emotions tasks regarding precision, computational effectiveness, and real-time applications. The purpose of this study is to provide an in-depth review of the state-of-the-art and propose future lines of study for deep learning-driven face emotion detection development. Face Expressions Face Emotion Recognition Deep Learning VGG-19 ResNet-50 Inception-V3 MobileNet Figures Figure 1 Figure 2 1 Introduction A significant indirect communication tool is facial expression, which can convey a person's intentions and feelings. The goal of Face Emotion Recognition (FER) is to recognize these expressions, and it is relevant to fields such as human-computer interaction, security, and healthcare. The intricacy and diversity of human expressions provide difficulties for traditional FER approaches. Deep Learning, on the other hand, has revolutionized FER by offering more precise and effective solutions. Prominent deep learning architectures that have demonstrated exceptional performance in image classification tasks, including FER, are VGG-19, ResNet-50, Inception-V3, and MobileNet. ResNet-50 uses residual information to solve the vanishing gradient problem, while VGG-19 uses modest convolution filters in its deep network for great performance. MobileNet provides a lightweight, accurate model appropriate for embedded and mobile applications, while Inception-V3 uses a modular architecture to capture multi-scale information. This survey provides a thorough overview of current trends and future prospects in FER research by examining these cutting-edge deep learning techniques and assessing their architectures, strengths, and limits. 1.1 Deep Learning This research study explores the field of deep learning methods for the important problem of classifying emotions on faces. It examines several cutting-edge architectures, including VGG-19, ResNet-50, Inception-V3, and MobileNet, each with special benefits and challenges in precisely identifying facial emotions. In discussing how deep learning is transforming facial emotion recognition, the study emphasizes how this technology can learn intricate patterns from massive datasets and outperform conventional techniques in terms of precision. Also, it tackles issues like overfitting, computational complexity, and model optimization, opening the door for further developments in areas like sophisticated regularization strategies, simplified models, and real-time application optimization. All things considered, the study offers a thorough assessment of the state of deep learning techniques for facial emotion classification and sketches out future directions for development in this rapidly evolving area. Table 1 Comparative Analysis of Deep Learning, Machine Learning, and Artificial Intelligence ,Techniques in Face Emotion Classification Comparative Analysis of Deep Learning, Machine Learning, and Artificial Intelligence Techniques in Face Emotion Classification Aspect Deep Learning Machine Learning Artificial Intelligence Description Uses neural networks to learn hierarchically from facial expressions and automatically extract features. classifies facial expressions using statistical models and algorithms based on features that are extracted. Builds complete emotion identification systems with real-time processing capability by integrating ML and DL algorithms. Techniques - CNNs are used to automatically extract features. - RNNs for analysis of time. - Conventional classifiers such as k-NN and SVM. - Techniques for feature engineering. - Hybrid ML-DL techniques to increase precision. - Practical uses in sentiment analysis, HCI, and surveillance. Role - Improves face expression recognition speed through automatic extraction of features. - Classifies data more accurately. - Uses features that have been retrieved to classify emotions. - Offers a performance baseline for comparison. - Combines DL and ML for comprehensive emotion identification. - Makes emotion analysis in real time easier. Challenges - Excessive processing needs. - Large labeled datasets are required. - Poor performance with nuanced feelings. - Manual feature extraction takes a lot of time. - Ethical issues with data utilization and privacy. - Ensured the ability to process data in real-time. Future Directions - Effective architectures and learning transfer. - Making use of unsupervised learning to identify tiny clues. - Integration for enhanced performance with DL. - Engineered features automatically. - Emphasize moral AI with comprehensible results. - Developments in processing in real time. The three approaches—deep learning, machine learning, and artificial intelligence—for determining emotions from facial expressions are compared in Table 1 . Deep learning is the most accurate and fastest approach, but it requires a large amount of data and processing power. Machine learning is slower and less accurate than deep learning, despite being easier to use and requiring less data. Machine learning and deep learning are combined in artificial intelligence to achieve high accuracy and real-time processing. But this raises ethical concerns about the privacy of data. Generally speaking, the best technique for a specific application will depend on how well accuracy, speed, and ease of implementation are traded off. 2 Theory Of Emotion This study explores facial emotion identification, including face detection techniques such as Viola-Jones and Haar cascades, as well as dataset kinds, labeling, and size impact. [ 1 ] In an effort to improve accuracy, it covers feature improvement methods including augmentation and makes considerable use of CNNs and SVMs for emotion categorization. The study recommends additional investigation into the ways in which these components interact to improve understanding and efficacy in facial emotion recognition systems. This research presents a transfer learning-based facial emotion recognition system. For emotion recognition, the study uses pre-trained convolutional neural networks (CNNs) from the ImageNet database, such as VGG19, ResNet50, InceptionV3, and MobileNet. [ 2 ]The CK + database was used for the experiments, and the results showed that the following models had outstanding accuracy rates: ResNet50 (97.7%), InceptionV3 (98.5%), MobileNet (94.2%), and VGG19 (96%). Notably, out of all the pre-trained networks, MobileNet showed the highest accuracy.[ 2 ]The study suggests using these networks in the future to identify emotions in speech and EEG signals, thus expanding the system's potential. This study tackles the issues raised by deepfakes, emphasizing how they could harm people's perceptions of the media, their mental health, hate speech, misinformation, and political stability.[ 3 ] It addresses the difficulties brought about by the quick dissemination of false information and the availability of tools for creating deepfakes. It also covers new developments and methods for producing and identifying deepfakes.[ 3 ] The study intends to make a substantial contribution to AI research by offering workable strategies to lessen the negative consequences of deepfakes. This work uses the FER2013 dataset and a variety of models to classify facial expressions; using the Adam optimizer, it achieves an accuracy of 0.60. [ 4 ]It draws attention to the difficulties in predicting specific emotions because of the scarcity of training data and makes recommendations for data augmentation and study of frequently anticipated emotions to improve accuracy. [ 4 ]In addition, the study offers insights for future research areas by discussing the model's real-time capabilities and possible applications in other domains. This paper delves into the complexities and limitations of emotion recognition solely through facial expressions, highlighting the challenges posed by subtle facial movements and variations in expression across different datasets and populations.[ 5 ] It stresses the necessity for a more adaptable framework in emotion recognition and proposes a cross-dataset evaluation method to address biases and assess model generalization.[ 5 ] Despite achieving notable results, the study emphasizes the ongoing difficulties in accurately interpreting internal emotional states from facial expressions alone, advocating for a multidisciplinary approach and careful exploration of emotion recognition systems. The Multi-Branch Deep RBF Network, which incorporates local information into the recognition process through several RBF units coupled to a VGG-Face backbone, is introduced in this research as a way to improve CNNs.[ 6 ] Results from six ER datasets show that CNN performs better, particularly under difficult circumstances like significant class overlap and short sample sizes.[ 6 ] The model outperforms state-of-the-art techniques and provides opportunities for further study into the efficient integration of local data, the use of RBF centers for interpretability and explainability, and the mitigation of biases in facial emotion recognition systems. In this paper, a comprehensive methodology for realistically testing and assessing Facial Emotion Recognition (FER) models is presented.[ 7 ] To show off the framework's usability and make it easier to create new FER datasets, a web application was created. Using a lightweight CNN trained on the AffectNet dataset, remarkable accuracy that was almost human-level was attained.[ 7 ] It does, however, recognize the need for more study in order to optimize deep learning methods for FER, particularly when compared to more accurate models such as VGGNet versions. Informed permission and regulatory procedures in their web application are highlighted in the paper, which also highlights the significance of privacy and data management in intelligent behavior analysis. The article discusses the usefulness of metrics in identifying biases within both real and manipulated datasets, offering easily interpretable values for analysis. [ 8 ]Through a case study on the popular FER dataset Affectnet, the study uncovers heavy racial representational bias and stereotypical gender biases. Results indicate that the model is less affected by racial bias removal but significantly impacted by gender bias, highlighting the complexity of bias analysis in machine learning.[ 8 ] Future work includes expanding the analysis to more datasets, models, and training setups, developing new demographic datasets and models, exploring mitigation techniques, and refining metrics for broader applications in multi-class multi-demographic AI systems and medical diagnosis. The paper introduces a CNN-based approach with deep learning algorithms for emotional recognition through facial features. It enhances accuracy and classification methodology, especially in complex facial expression identification, by employing dimensional space reduction and kernel filters in preprocessing.[ 9 ] Data augmentation techniques like cropping and introducing noise, along with picture synthesis methods, were utilized to address overfitting and data scarcity.[ 9 ] Future plans include extending feature sets, recognizing additional emotions, and exploring automatic facial emotion recognition. This paper presents EMOCA, a self-supervised approach that uses emotion-rich image data to train single photographs to generate 3D facial expressions.[ 10 ] By integrating a distinct emotion similarity loss derived from deep features, it achieves better outcomes in 3D face shape reconstruction, especially when it comes to accurately expressing emotions. EMOCA outperforms existing methods for recognizing emotions in the wild and has potential uses in the gaming, cinema, AR/VR, and communication sectors.[ 10 ] In addition to addressing issues like deepfakes, the article highlights how crucial it is for human communication and avatar interactions to successfully portray emotions. This paper presents a deep learning framework that uses both conventional and novel CNN architectures to identify face emotions.[ 11 ] Accuracy was assessed on lab-controlled datasets, RAFD and KDEF, with results of 68.84% and 99.63%, respectively. For the KDEF dataset, a custom model outperformed the most advanced findings. In comparison to lab-controlled datasets, the accuracy of the non-lab-controlled datasets RAF-DB, SFEW, and AMFED + was 75.26%, 40.78%, and 54.15%, respectively.[ 11 ] Future developments, according to the study, should focus on enhancing accuracy by unsupervised pre-training using transfer learning, pre-processing, and feature extraction techniques. Additionally, advanced models like DBN, GANs, and Facial Action Units (AUs) should be investigated. This paper offers a thorough introduction to DeepFake technology, addressing its foundations, benefits, and related hazards with a special emphasis on GAN-based applications.[ 12 ] It talks about the difficulties in identifying DeepFakes, emphasizes the necessity for extra security measures to guarantee data integrity, and forecasts possible outcomes of AI employing DeepFakes to counter AI propaganda. This study explores the potential of smartphones combined with AI for aiding in the identification and management of depressive disorders and mental health. [ 13 ] It envisions personalized treatment recommendations, detection capabilities, and insights through digitally recognized expressions, with applications ranging from acute treatment classification to remote monitoring and preventative care.[ 13 ] Although it doesn't delve into natural language processing for depression diagnosis, the study acknowledges the potential of AI-based chatbots for responsive and contextual interactions, aiming for accessible mental health benefits at a low cost and quick deployment. The study presents an Efficient-SwishNet model for facial emotion recognition, highlighting its robustness, space efficiency, and affordability. [ 14 ] It achieves 100% recognition rate on FERG and CK datasets, outperforming previous algorithms on five separate datasets. The model does a great job managing photos with multiple orientations in addition to accurately detecting emotions from frontal face images. [ 14 ]The model's generalizability is demonstrated through cross-corpora examination. It has trouble with covered faces and occlusions, though, and the scientists intend to work on these issues in later research. Their objectives are to establish a specific FER dataset for real-time performance testing and refine the FER model for cross-corpora evaluation. This work presents a transfer learning and ResNet-based facial emotion detection system that achieves high classification rates in several facial expression categories. [ 15 ] The ResNet-18 architecture performed optimally, demonstrating the method's efficacy in classifying compound emotions.[ 15 ] Future plans call for expanding evaluations to other datasets such as Affectnet and iCV-MEFED, as well as doing cross-database tests. In order to combine deep learning with relational reasoning,[ 16 ] the paper presents a unified architecture that makes use of SqueezeNet for emotion recognition and Temporal Relation Networks (TRNs). [ 16 ] With its temporal relational reasoning, the TRN performs better in training and testing than other models, particularly the multi-scale TRN, when compared to single-scale TRN and multi-layer perceptron models. In order to achieve accurate face detection, the paper presents a unique deep learning-based approach that combines prediction with region-offering networks (RON).[ 17 ] It delivers good accuracy with a small model size and minimal computing load after being trained on the WIDER FACE dataset.[ 17 ] Future research will focus on developing real-time models for facial emotion identification using sophisticated architectures like 3D CNN, 3D U-Net, and YOLOv, as well as enhancing performance in difficult situations like low light and foggy pictures. This research outlines the advantages and disadvantages of the conventional ML-based and DL-based techniques to facial expression recognition (FER).[ 18 ]It highlights that 3D facial expression datasets are necessary for better results and talks about how DL approaches can be integrated with IoT sensors to boost FER capabilities for different sectors. This paper presents a multimodal emotion recognition model that uses CNN model for feature extraction and attention mechanism to merge facial expressions and EEG inputs.[ 19 ] More study is needed to improve face feature pre-training models and incorporate different modalities for better model performance, as demonstrated by the experimental results, [ 19 ]which show improved emotion identification accuracy when compared to utilizing either modality alone. This study shows that the VGG16 model (model B) achieves 100% and 73% success rates, respectively, with the fewest average losses of 0.011 and 0.1875, outperforming models C and A in face and emotion recognition. [ 20 ] However, because VGG16 has more layers, the training process takes longer.[ 20 ] Despite having a little lower accuracy (87% for face and 67% for emotion detection), Model C contains fewer parameters, which allows for a quicker training procedure. Real-time face and emotion detection systems for humanoid robots have effectively integrated both models. The study emphasizes the necessity of additional system enhancements as well as the significance of dataset size and illumination concerns for future advancements. This research presents a deep learning model-based real-time engagement detection method for online learners by analyzing facial emotions. [ 21 ] The system creates an engagement index (EI) that indicates whether a user is "engaged" or "disengaged" based on the monitoring of facial expressions through built-in web cams. It uses MFACXTOR for face point extraction and Faster R-CNN for face detection. It is trained on FER-2013, CK+, and custom datasets. Out of all the models that were assessed, ResNet-50, VGG-19, and Inception-V3 had the best accuracy (92.32%). [ 21 ]The method successfully determined involvement levels after testing on 20 students. Future developments could expand to larger datasets and students with specific needs, as well as incorporate data from more sensors. This overview highlights the difficulties in identifying emotions in real-world situations and covers robotic facial expression production and human facial expression detection. [ 22 ] To promote successful emotion transmission, future research should concentrate on increasing detection under diverse settings and expanding robots' capacity to express a greater variety of emotions, including mixed expressions and varying intensities. The study emphasizes that although face masks are useful in halting the spread of viruses, they have a negative impact on facial identity and social perceptions about emotion, age, and gender.[ 23 ] These results should not deter people from wearing masks when essential for medical reasons, such as during the COVID-19 epidemic. Rather, they highlight the necessity of cultural modification to reduce the psychological effects of mask wear.[ 23 ] Subsequent investigations ought to devise approaches to enhance mask-wearing communication, as it is vital for handling present and potential pandemics. To achieve better results, a proprietary expression detection model and the Viola-Jones method for face extraction across various datasets are used in the study's facial expression recognition framework, which makes use of the Jetson Nanodevice for law enforcement. [ 24 ]In the future, the framework will be expanded to include gender categorization, age prediction, and more deep-learning models. [ 24 ]Data augmentation will be taken into consideration for better performance on devices with limited resources. The research created a facial emotion recognition (FER) model by integrating a CNN with a Haar-Cascade classifier.[ 25 ] The model demonstrated 90% accuracy in image-based testing, but it encountered difficulties in real-time emotion identification, especially when dealing with masked faces.[ 25 ] In the future, 3D CNN, 3D U-Net, and YOLO settings for landmark-based emotion identification will be used to address biases and noise in facial expressions, with the goal of boosting accuracy in real-life scenarios. The paper highlights temporal dynamics and spatial asymmetry in the introduction of TSception, a multi-scale convolutional neural network intended for EEG emotion recognition.[ 26 ] TSception uses parallel multi-scale temporal kernels and hemisphere-specific kernels to improve emotional asymmetry pattern learning. Using fewer trainable parameters, evaluation on benchmark datasets showed improved classification performance.[ 26 ] Saliency maps facilitated further research into the impact of segment length and cross-individual generalization by identifying important informative locations and lateralization patterns. With the use of the AffectNet dataset and deep CNN models,[ 27 ] this research is able to identify emotions from masked faces with an accuracy of 69.3% and an average confusion matrix of 71.4%. In the future, the model will be improved via semi-supervised learning, attention processes without landmarks will be improved, [ 27 ]and a prototype device to help visually impaired people recognize their environment will be developed. Key issues in emotion identification across multiple modalities are highlighted in this survey, including the necessity for accurate emotional classification in physiological data, individual differences affecting facial emotion detection, and variability in speech emotion recognition (SER).[ 28 ] It highlights the superiority of deep learning on larger datasets and the significance of combining sensors for dependable multi-modal recognition. [ 28 ] Prospective avenues for advancement encompass enhancing resilience, precision, confidentiality, and user acceptance via multi-modal approaches and sophisticated learning methodologies such as unsupervised and reinforcement learning. This research presents a novel approach to remote facial video analysis employing RGB, NIR, and infrared cameras for emotion classification.[ 29 ] With a single feature from RGB films and a DL classifier, the study obtains encouraging results with 47.36% average accuracy by focusing on extracting features from pulsatile heartbeats in facial video frames.[ 29 ] The study highlights RGB cameras' potential for efficient emotion classification, but it also notes drawbacks including subject mobility and brief video inputs. For greater practicality, future work may incorporate facial tracking and registration. Using computer vision and deep learning techniques, this work conducts a systematic literature review (SLR) on emotion recognition, examining 77 academic papers.[ 30 ] It highlights trends in emotion recognition across face and body positions and highlights possible uses in law enforcement and healthcare. Convolutional Neural Networks (CNNs) are robust and accurate, yet they still face issues with limited technology and scarce data. [ 30 ] Promising techniques such as vision transformers are particularly useful for handling micro- and macro-expressions. Despite the constraints of the dataset, future study may include hand and body expressions. Continuous efforts are being made to improve accuracy and practical utility in real-world circumstances. 3 Material Methods 3.1 Face Emotion Recognition Method The Utilizing computer vision and deep learning approaches, the Face Emotion Recognition (FER) technology systematically analyzes face emotions. Consistency-checking face picture preprocessing, feature extraction using pre-trained CNNs like VGG 19, ResNet 50, MobileNet, and Inception V3, and emotion classification utilizing the derived features are all included. The technique seeks to improve emotion identification technology by precisely identifying and classifying the emotions portrayed in faces. In the Fig. 1 is represent that The flowchart displays a computer vision method for facial emotion recognition that begins with pre-processing photos in a dataset to guarantee consistency. As a feature extractor, a pre-trained convolutional neural network (CNN) such as VGG 19, ResNet 50, MobileNet, or Inception V3 examines the facial expressions in the pictures. A classifier uses these extracted features to classify the emotions that are represented in the faces. Ultimately, the system's efficacy is assessed to see how well it can identify emotions. Table 2 A Discussion Of The Typical Conventional FER Techniques In A Summary A Discussion Of The Typical Conventional FER Techniques In A Summary Citations Collections of Data Methods of Decision Making Characteristics Analysis of Emotions [ 31 ] FECR - HMM and the SVM classifier - PCA and LDA − 6 Emotions [ 32 ] CK Plus - Classifier using support vector machines - Gabor wavelets, nonlinear NLPCA, and the Haar wavelet transform − 6 Emotions [ 33 ] MMI , Japanese Female Facial Expression Database , CK + Data set - large margin classifier - ORB ,SIFT ,SURF − 7 Emotions [ 34 ] CK+ , MMI Facial Expression Corpus - SMO, MLP, and KNN for categorization - DCT, HOG − 7 Emotions [ 35 ] CK Plus - Denoising Sparse Autoencoder - Gabor Function, LBP, SIFT, and HOG − 7 Emotions [ 36 ] Dataset from a visual depth camera in real time - The DBN - Features of MLDP-GDA − 6 Emotions [ 37 ] BOSPHORUS - Euclidean distance - Geometric descriptor − 4 Emotions [ 38 ] CK+ , MMI Dataset , MUG - The AdaBoost-ELM, KLT, and EBGM - Salient geometric features − 7 Emotions [ 39 ] BVTKFER BCurtin Faces - The Random forest classifier - LBP − 6 Emotions [ 40 ] UVANEMO, SPOS, MMI, and BBC - Support Vector Machine - Four standard features—raw pixels, Gabor, HOG, and LBP—as well as RSTD − 2 Emotions Smile (genuine and fake) [ 41 ] CK + Facial Expression Dataset - Conditional Random Field and KNN - The Geometric descriptor − 7 Emotions [ 42 ] CK+, Japanese Female Facial Expression Database - EHMM - ASM and 2D DCT − 7 Emotions [ 43 ] Cohn-Kanade Plus , JAFFE Dataset, MUG - SVM, PCA, LDA, and K-NN - Gabor wavelets and geometric features − 7 Emotions [ 44 ] CK, JAFFE ,USTCNVIE, Yale, FEI - Markov Hidden Models - Chan-Vese energy function, Bhattacharyya distance function wavelet decomposition and SWLDA − 6 Emotions [ 45 ] Extended Cohn-Kanade Dataset - The SVR - Gabor wavelets, AAMS, and feature descriptors − 7 Emotions Table 2 provides an overview of the many conventional facial emotion recognition (FER) techniques that have been used in different research. The sources include a list of the datasets and methods utilized in the emotion categorization decision-making process. Commonly used datasets with classifiers such as HMM, SVM, KNN, and deep belief networks are CK+, MMI, and JAFFE. Feature extraction methods include Gabor wavelets, PCA and LDA, HOG, and LBP. These techniques often aim to categorize six or seven primary emotions, however some studies focus on a specific subset of emotions. The table illustrates the range of techniques and datasets used to identify emotions from facial expressions. 3.2 Comparing With The Modern Techniques The comparison with state-of-the-art methods highlights the efficacy of three CNN architectures in facial emotion recognition: VGG-19, ResNet-50, and Inception V3. The structural components of these models that are evaluated include convolutional layers for feature extraction, pooling layers for dimensionality reduction, batch normalization for stability, activation layers for non-linearity, and fully connected layers for classification. Because performance varies depending on a number of factors, including processing resources, dataset size, and accuracy requirements, there is no one ideal architecture. Rather, the effectiveness of each building is influenced by its own special arrangement and number of layers. 3.2.1 Emotion Recognition System Effectiveness Assessment on CFEE and RAF Databases Using the posed CFEE and spontaneous RAF databases, this part assesses the system's efficacy in identifying basic and compound emotions by contrasting the outcomes with sophisticated current approaches. The comparison of the CFEE database's basic emotion recognition (7 classes) with two cutting-edge investigations is shown in Table 3 Notably, a study that used the AlexNet CNN architecture was able to attain 74.799% test accuracy. Seventy percent of the photos in this study were used for training, and test and validation sets each had 159 images. Table 3 Evaluation Of The CFEE Database's Fundamental And Compound Emotions With Modern Methods Evaluation Of The CFEE Database's Fundamental And Compound Emotions With Modern Methods Ref. Approach The procedure Examples Classes Precision [ 46 ] Shape and appear + Nearest mean 10-fold cross-validation 1610 7-class 96.96% [ 47 ] The AlexNet CNN 70%-15%-15% (train/validation/test) 1127 245 238 74.79% [ 46 ] Shape and appearance + Nearest mean 10-fold cross-validation 5060 22-class 76.91% [ 48 ] The Highway-CNN 52.14% [ 15 ] The Resnet-18 Deep Features + SVM 10-fold cross-validation 1610 7-class 98.02%* 99.19%+ 70% train-15% test 1365 81.93% 10-fold cross-validation 5060 22-class 80.69% Using the RAF database, Table 4 offers a comparative comparison of the system's performance in identifying basic and compound emotions. The comparison incorporates findings from multiple cutting-edge methods. The top-performing tested methods are emphasized, showcasing their recognition rates against the spontaneous emotional expressions in the RAF database. These comparisons highlight the system's efficacy and offer a baseline against the highest reported recognition rates found in the literature review. Table 4 Comparison Of The RAF Database Applying Modern Methods Comparison Of The RAF Database Applying Modern Methods Ref The process The Samples Classes Recall The precision [ 49 ] Gabor + mSVM 15339 7-class 65% - Deep Locality-Preserving CNN + mSVM 74% 84.13% [ 50 ] Augmented data and cluster loss 76% - [ 51 ] Multi-Region Ensemble -CNN (VGG-16) 77% - Multi-Region Ensemble CNN (AlexNet) 75% - [ 52 ] Capsule Network 77% - [ 53 ] Double Completed-LBP (Double Cd-LBP) 78% - [ 54 ] Transfer learning Resnet-18 (AffectNet database) 80% - [ 55 ] Covariance pooling after final convolutional layers 79% 87.0% [ 56 ] Patch-Gated CNN (PG-CNN) - 83% [ 57 ] Conditional generative adversarial network-based EAU-Net Network (CGAN based EAU-Net) 81.83% - [ 58 ] Region Attention Network (RAN) - 86.90% [ 59 ] Pyramid With Super-Resolution (PSR) Network (VGG-16) 80.78% 88.98% [ 60 ] Gabor + mSVM 3954 11-class 33.76% - [ 49 ] Deep Locality-Preserving CNN + mSVM 44.55% 57.95% [ 61 ] LBP + NCMML 3171 16-class 30.10% 36.70% HOG + NCMML 36.90% 44.10% [ 15 ] Resnet-18 Deep Features + SVM 15339 7-class 86% 93.29% 3954 11-class 63.27% 76.52% 19059 16-class - 60.76% 4 The Recommended Models' Training Process 4.1 VGG-19 The Visual Graphics Group at Oxford introduced VGG-19, a convolutional neural network (CNN) that is well-known for being both straightforward and efficient in image identification applications, in 2014. With its modest receptive fields (3x3) and max-pooling layers, VGG-19 has a simple architecture made up of 19 layers (16 convolutional and 3 fully linked). The network performs admirably, particularly in picture classification tasks like the ImageNet challenge, where it obtained remarkable accuracy thanks to its depth and homogeneous structure. But the model's high number of parameters and deep water can result in computing complexity, which limits its usefulness for real-time applications on devices with limited resources. In spite of this, VGG-19 continues to be a cornerstone model in the field of deep learning research, acting as a standard and influencing later network architectures. 4.2 Resnet50 ResNet50, a groundbreaking deep learning model developed in 2015 by Kaiming He et al., changed computer vision thanks to its creative use of residual learning. With 50 convolutional layers split up into five phases, each containing bottleneck blocks, ResNet50 improved training efficiency by introducing identity shortcut connections to resolve gradient problems and allow for previously unattainable network depths. It gained greater accuracy on benchmarks such as ImageNet by using a combination of 1x1, 3x3, and 1x1 convolutions along with batch normalization. As a result, it became indispensable for tasks like object recognition (e.g., Faster R-CNN) and picture segmentation (e.g., Mask R-CNN). ResNet50 is a mainstay of contemporary computer vision systems because of its capacity to optimize deep networks. 4.3 Inception V3 Amongst the Inception family of CNNs, Inception V3, created by Google in 2015, stands out for its emphasis on object recognition and image classification. Its Inception module, which uses a variety of filter sizes to capture features, factorized convolutions for efficiency, auxiliary classifiers to help with deep network training, and the regular application of batch normalization and ReLU for quicker convergence are some of its salient features. Stem layers, Inception modules, auxiliary classifiers, and final layers for classification make up its architecture. Exceptional advantages include cutting-edge performance in picture tasks, less overfitting, and effective feature extraction. As a flexible and efficient model in contemporary computer vision, Inception V3 has applications in picture categorization, object recognition, and transfer learning. In the Fig. 2 is represent that Three CNN architectures used for facial emotion detection are compared in the graphic named VGG-19, ResNet-50, and Inception V3. Convolutional layers for feature extraction, pooling layers for dimensionality reduction, batch normalization for training stabilization, activation layers for adding non-linearity, and fully connected layers for classification are the structural elements shared by these models. The order and number of layers in each architecture vary: Inception V3 mixes convolutional and inception modules, ResNet-50 has fifty, while VGG-19 has sixteen convolutional layers. The size of the dataset, the amount of computing power needed, and the level of accuracy needed determine the optimal architecture for a given job; the image does not reveal the top-performing model for facial emotion identification. 4.4 MobileNet The google presented MobileNet in 2017 as a powerful deep-learning solution designed specifically for embedded and mobile devices. Its main innovation is depthwise separable convolutions, which significantly lower processing demands by splitting ordinary convolutions into depthwise and pointwise layers. Furthermore, MobileNet adds parameters such as width and resolution multipliers, which provide simple scaling and customization of model efficiency and size. With inverted residuals and linear bottlenecks, the transition to MobileNetV2 improves feature representation even more, preserving high accuracy at cheap computing costs. Because of its adaptability and suitability for a range of computer vision tasks, including object identification, segmentation, and image classification, MobileNet has become the standard for real-time applications on platforms with limited resources. 5 Implementation 5.1 Evaluation Metrics We computed the accuracy, mean precision, and mean recall—standard measures generally employed by modern techniques for face emotion recognition—in order to assess the performance of the suggested model. 5.1.1 Accuracy Accuracy is the most crucial statistic in multi-class classification. It is computed by dividing the total number of instances in the emotion class by the sum of the true negative (TN) and true positive (TP) instances. The classifier's accuracy is calculated as : $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:Overall\:Accuracy=\:\:\:\:\:\frac{{\sum\:}_{i=1}^{n}{TP}_{i}+{TN}_{i\:\:\:}}{TP+TN+FP+FN}\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\left(1\right)\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:$$ 5.1.2 Mean Precision For each emotion class, precision is the genuine positive prediction that the suggested model makes. In the event of multi-class classification, the model's mean precision is computed as follows: $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:Mean\:Precision=\:\:\:\:\:\frac{{\sum\:}_{i=1}^{n}{\:(TP}_{i})}{{\:\:{\sum\:}_{i=1}^{n}\:({TP}_{i\:\:\:}+\:FP}_{i})}\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\left(2\right)\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:$$ 5.1.3 Recall Identify the specific number of positive class predictions that were made from all of the dataset's positive instances of the emotion class. Here is how we calculated the mean recall: $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:Mean\:Precision=\:\:\:\:\:\frac{{\sum\:}_{i=1}^{n}{\:(TP}_{i})}{{\:\:{\sum\:}_{i=1}^{n}\:({TP}_{i\:\:\:}+\:FN}_{i})}\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\left(3\right)\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:$$ 5.1.4 Specificity It is defined as a portion of negative emotion classes that the classifier correctly categorizes as negative. Specificity is another name for true negative rate, or TNR. The mean specificity was calculated as follows. $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:Mean\:Specificity=\:\:\:\:\:\frac{{\sum\:}_{i=1}^{n}{\:(TN}_{i})}{{\:\:{\sum\:}_{i=1}^{n}\:({TN}_{i\:\:\:}+\:FP}_{i})}\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\left(4\right)\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:$$ 5.1.5 MACRO F-1 SCORE The computation of the macro F-1 score involves obtaining the unweighted arithmetic mean of the F-1 score for every emotion class. In terms of TP, false positives (FP), and false negatives (FN), the macro F-1 score is represented as : $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:Macro\:F-1\:Score=\:\:\:\:\:\frac{{\sum\:}_{i=1}^{n}{\:(TP}_{i})}{{\:\:{\sum\:}_{i=1}^{n}\:({TP}_{i\:\:}+0.5(\:FP}_{i}+{FN}_{i}\:\left)\right)}\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\left(5\right)\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:$$ In each of the equations (1) through (5), n is the total number of classes; \(\:{\:TP}_{i}\) and \(\:{\:TN}_{i}\) stand for true positives and true negatives, respectively, while \(\:{\:FP}_{i}\) and \(\:{\:FN}_{i}\) represent the number of false positives and false negatives for emotion class i. 6 Conclusion At some point, the study highlights the noteworthy progress made in facial emotion recognition (FER) using deep learning methods, especially with models such as VGG-19, ResNet-50, Inception-V3, and MobileNet. These models' exceptional accuracy in facial expression classification highlights their potential for practical uses in security systems, mental health monitoring, human-computer interaction, and other fields.While acknowledging the difficulties associated with overfitting, computational complexity, and model optimization, the paper also makes recommendations for future directions in the field, including feature augmentation techniques, transfer learning, and integration with other modalities including voice and EEG inputs. Through a comprehensive assessment of deep learning models, technical discussion, and application area identification, the research opens the door for future advancements and improvements in FER systems, leading to more accurate and efficient human-emotion recognition technologies. Declarations Funding Not applicable. Author information Authors and Affiliations Ketan Sarvakar Research Scholar , Gujarat Technological University, Ahmedabad, India Dr. Kaushik Rana Associate Professor,Gujarat Technological University, Ahmedabad, India 2 Contributions Conceptualization: Ketan Sarvakar; Methodology: Ketan Sarvakar; Formal anaysis and investigation: Ketan Sarvakar; Writingoriginal draft preparation: Ketan Sarvakar; Writing, review and editing: Ketan Sarvakar, Dr. Kaushik Rana; Resources: Ketan Sarvakar; Supervision: Dr. Kaushik Rana. Corresponding author Corresponding author: Ketan Sarvakar( [email protected] ) Ethics declarations Research involving human and /or animals Not applicable. Competing interests The authors declare no competing interests. References Naga P, Marri SD, Borreo R (Jan. 2023) Facial emotion recognition methods, datasets and technologies: A literature survey. Mater Today Proc 80:2824–2828. 10.1016/j.matpr.2021.07.046 Chowdary MK, Nguyen TN, Hemanth DJ (2023) Deep learning-based facial emotion recognition for human–computer interaction applications, Neural Comput Appl, vol. 35, no. 32, pp. 23311–23328, Nov. 10.1007/s00521-021-06012-8 Nguyen TT et al (2019) Deep Learning for Deepfakes Creation and Detection: A Survey. Sep. 10.1016/j.cviu.2022.103525 Saravanan A, Perichetla G, Gayathri DKS (2019) Facial Emotion Recognition using Convolutional Neural Networks, Oct. [Online]. Available: http://arxiv.org/abs/1910.05602 Dias W et al (2023) Cross-dataset emotion recognition from facial expressions through convolutional neural networks A R T I C L E I N F O. J Vis Commun Image Represent. 10.5281/zenodo.4696032 Hernández-Luquin F, Escalante HJ (2023) Multi-branch deep radial basis function networks for facial emotion recognition, Neural Comput Appl, vol. 35, no. 25, pp. 18131–18145, Sep. 10.1007/s00521-021-06420-w Siddiqui N, Dave R, Bauer T, Reither T, Black D, Hanson M A Robust Framework for Deep Learning Approaches to Facial Emotion Recognition and Evaluation, IEEE Dominguez-Catena I, Paternain D, Galar M Assessing Demographic Bias Transfer from Dataset to Model: A Case Study in Facial Expression Recognition, May 2022, [Online]. Available: http://arxiv.org/abs/2205.10049 Kumar T, Arora et al (2022) Optimal Facial Feature Based Emotional Recognition Using Deep Learning Algorithm, Comput Intell Neurosci, vol. 2022. 10.1155/2022/8379202 Daněček R, Black M EMOCA: Emotion Driven Monocular Face Capture and Animation. [Online]. Available: https://emoca.is.tue.mpg.de Appasaheb Borgalli R, Surve S (2022) Deep Learning Framework for Facial Emotion Recognition using CNN Architectures, in Proceedings of the International Conference on Electronics and Renewable Systems, ICEARS 2022, Institute of Electrical and Electronics Engineers Inc., pp. 1777–1784. 10.1109/ICEARS53579.2022.9751735 Malik A, Kuribayashi M, Abdullahi SM, Khan AN (2022) DeepFake Detection for Human Face Images and Videos: A Survey. IEEE Access 10:18757–18775. Institute of Electrical and Electronics Engineers Inc. 10.1109/ACCESS.2022.3151186 Lee YS, Park WH (Feb. 2022) Diagnosis of Depressive Disorder Model on Facial Expression Based on Fast R-CNN. Diagnostics 12(2). 10.3390/diagnostics12020317 Dar T, Javed A, Bourouis S, Hussein HS, Alshazly H (2022) Efficient-SwishNet Based System for Facial Emotion Recognition. IEEE Access 10:71311–71328. 10.1109/ACCESS.2022.3188730 K., R. Y., & M. R. (2022) C. facial emotional expression recognition using cnn deep features. E. L. 30(4), 1402–1416. Slimani, Compound facial emotional expression recognition using cnn deep features, 2022 Pise A, Vadapalli H, Sanders I (Aug. 2022) Facial emotion recognition using temporal relational network: an application to E-learning. Multimed Tools Appl 81:26633–26653. 10.1007/s11042-020-10133-y Mamieva D, Abdusalomov AB, Mukhiddinov M, Whangbo TK (Jan. 2023) Improved Face Detection Method via Learning Small Faces on Hard Images Based on a Deep Learning Approach. Sensors 23(1). 10.3390/s23010502 Khan AR (2022) Facial Emotion Recognition Using Conventional Machine Learning and Deep Learning Methods: Current Achievements, Analysis and Remaining Challenges. Inform (Switzerland). 13, 6. MDPI, Jun. 01 10.3390/info13060268 Wang S, Qu J, Zhang Y, Zhang Y (2023) Multimodal Emotion Recognition From EEG Signals and Facial Expressions. IEEE Access 11:33061–33068. 10.1109/ACCESS.2023.3263670 Dwijayanti S, Iqbal M, Suprapto BY (2022) Real-time Implementation of Face Recognition and Emotion Recognition in a Humanoid Robot Using a Convolutional Neural Network. IEEE Access. 10.1109/ACCESS.2022.3200762 Gupta S, Kumar P, Tekchandani RK (2023) Facial emotion recognition based real-time learner engagement detection system in online learning context using deep learning models, Multimed Tools Appl, vol. 82, no. 8, pp. 11365–11394, Mar. 10.1007/s11042-022-13558-9 Rawal N, Stock-Homburg RM (2022) Facial Emotion Expressions in Human–Robot Interaction: A Survey, Int J Soc Robot, vol. 14, no. 7, pp. 1583–1604, Sep. 10.1007/s12369-022-00867-0 Wong HK, Estudillo AJ (Dec. 2022) Face masks affect emotion categorisation, age estimation, recognition, and gender classification from faces. Cogn Res Princ Implic 7(1). 10.1186/s41235-022-00438-x Alsharekh MF (Aug. 2022) Facial Emotion Recognition in Verbal Communication Based on Deep Learning. Sensors 22(16). 10.3390/s22166105 Farkhod A, Abdusalomov AB, Mukhiddinov M, Cho YI (2022) Development of Real-Time Landmark-Based Emotion Recognition CNN for Masked Faces, Sensors, vol. 22, no. 22, Nov. 10.3390/s22228704 Ding Y, Robinson N, Zhang S, Zeng Q, Guan C (2023) TSception: Capturing Temporal Dynamics and Spatial Asymmetry From EEG for Emotion Recognition, IEEE Trans Affect Comput, vol. 14, no. 3, pp. 2238–2250, Jul. 10.1109/TAFFC.2022.3169001 Mukhiddinov M, Djuraev O, Akhmedov F, Mukhamadiyev A, Cho J (Feb. 2023) Masked Face Emotion Recognition Based on Facial Landmarks and Deep Learning Approaches for Visually Impaired People. Sensors 23(3). 10.3390/s23031080 Cai Y, Li X, Li J (2023) Emotion Recognition Using Different Sensors, Emotion Models, Methods and Datasets: A Comprehensive Review, Sensors, vol. 23, no. 5. MDPI, Mar. 01. 10.3390/s23052455 Talala S, Shvimmer S, Simhon R, Gilead M, Yitzhaky Y (2024) Emotion Classification Based on Pulsatile Images Extracted from Short Facial Videos via Deep Learning, Sensors, vol. 24, no. 8, Apr. 10.3390/s24082620 Pereira R et al (May 2024) Systematic Review of Emotion Detection with Computer Vision and Deep Learning. Sensors 24(11):3484. 10.3390/s24113484 Techno-Societal (2018) Springer International Publishing, 2020. 10.1007/978-3-030-16848-3 Reddy CVR, Reddy US, Kishore KVK (2019) Facial emotion recognition using NLPCA and SVM, Traitement du Signal, vol. 36, no. 1, pp. 13–22, Feb. 10.18280/ts.360102 Sajjad M, Nasir M, Ullah FUM, Muhammad K, Sangaiah AK, Baik SW (2019) Raspberry Pi assisted facial expression recognition framework for smart security in law-enforcement services, Inf Sci (N Y), vol. 479, pp. 416–431, Apr. 10.1016/j.ins.2018.07.027 Nazir M, Jan Z, Sajjad M (2017) Facial expression recognition using weber discrete wavelet transform. J Intell Fuzzy Syst 33(1):479–489. 10.3233/JIFS-161787 Zeng N, Zhang H, Song B, Liu W, Li Y, Dobaie AM (Jan. 2018) Facial expression recognition via learning deep sparse autoencoders. Neurocomputing 273:643–649. 10.1016/j.neucom.2017.08.043 Uddin MZ, Hassan MM, Almogren A, Zuair M, Fortino G, Torresen J (2017) A facial expression recognition system using robust face features from depth videos and deep learning, Computers and Electrical Engineering, vol. 63, pp. 114–125, Oct. 10.1016/j.compeleceng.2017.04.019 Al-agha SA, Saleh HH, Ghani RF (2017) Geometric-based Feature Extraction and Classification for Emotion Expressions of 3D Video Film. J Adv Inform Technol 74–79. 10.12720/jait.8.2.74-79 Ghimire D, Lee J, Li ZN, Jeong S (2017) Recognition of facial expressions based on salient geometric features and support vector machines, Multimed Tools Appl, vol. 76, no. 6, pp. 7921–7946, Mar. 10.1007/s11042-016-3428-9 Wang J, Yang H Face detection based on template matching and 2DPCA algorithm, in Proceedings – 1st International Congress on Image and Signal Processing, CISP 2008, 2008, pp. 575–579. 10.1109/CISP.2008.270 Wu Pping, Liu H, Zhang Xwu, Gao Y (2017) Spontaneous versus posed smile recognition via region-specific texture descriptor and geometric facial dynamics, Frontiers of Information Technology and Electronic Engineering, vol. 18, no. 7, pp. 955–967, Jul. 10.1631/FITEE.1600041 IEEE Computer Society and Institute of Electrical and Electronics Engineers 12th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition: FG 2017 : proceedings : 30 May – 3 June 2017, Washington, D.C Kim DJ (Jul. 2016) Facial expression recognition using ASM-based post-processing technique. Pattern Recognit Image Anal 26(3):576–581. 10.1134/S105466181603010X Cornejo JYR, Pedrini H, Flórez-Revuelta F Facial Expression Recognition with Occlusions based on Geometric Representation. Siddiqi MH, Ali R, Khan AM, Kim ES, Kim GJ, Lee S (2015) Facial expression recognition using active contour-based face detection, facial movement-based feature extraction, and non-linear feature selection, Multimed Syst, vol. 21, no. 6, pp. 541–555, Nov. 10.1007/s00530-014-0400-2 Chang KY, Chen CS, Hung YP (2013) Intensity rank estimation of facial expressions based on a single image, in Proceedings – 2013 IEEE International Conference on Systems, Man, and Cybernetics, SMC 2013, pp. 3157–3162. 10.1109/SMC.2013.538 Du S, Tao Y, Martinez AM (Apr. 2014) Compound facial expressions of emotion. Proc Natl Acad Sci U S A 111(15). 10.1073/pnas.1322355111 Expressions of Emotion Dataset [5] We trained the network on CFEE dataset and tested the trained model on RaFD. 2. Related Work. Slimani K, Lekdioui K, Messoussi R, Touahni R (2019) Compound facial expression recognition based on highway CNN, in ACM International Conference Proceeding Series, Association for Computing Machinery, Mar. 10.1145/3314074.3314075 Li S, Deng W, Du J Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. [Online]. Available: http://whdeng.cn/RAF/model1.html Institute of Electrical and Electronics Engineers and, Signal Processing IEEE, Society (2018) IEEE International Conference on Image Processing: proceedings : October 7–10, 2018, Megaron Athens International Conference Centre, Athens, Greece Fan Y, Lam JCK, Li VOK Multi-Region Ensemble Convolutional Neural Network for Facial Expression Recognition. Ghosh S, Dhall A, Sebe N, Automatic Group Affect Analysis in Images via Visual Attribute and Feature Networks, in Proceedings - International Conference on Image Processing, ICIP, Computer Society IEEE (2018) Aug. pp. 1967–1971. 10.1109/ICIP.2018.8451242 Institute of Electrical and Electronics Engineers and, Signal Processing IEEE, Society (2018) IEEE International Conference on Image Processing: proceedings : October 7–10, 2018, Megaron Athens International Conference Centre, Athens, Greece Vielzeuf V, Kervadec C, Pateux S, Lechervy A, Jurie F (2018) An Occam’s Razor View on Learning Audiovisual Emotion Recognition with Small Training Sets, Aug. [Online]. Available: http://arxiv.org/abs/1808.02668 Acharya D, Huang Z, Paudel DP, Gool V Covariance Pooling for Facial Expression Recognition. Li Y, Zeng J, Shan S, Chen X Patch-Gated CNN for Occlusion-aware Facial Expression Recognition. Deng J, Pang G, Zhang Z, Pang Z, Yang H, Yang G (2019) CGAN Based Facial Expression Recognition for Human-Robot Interaction. IEEE Access 7:9848–9859. 10.1109/ACCESS.2019.2891668 Wang K, Peng X, Yang J, Meng D, Qiao Y Region Attention Networks for Pose and Occlusion Robust Facial Expression Recognition, May 2019, [Online]. Available: http://arxiv.org/abs/1905.04075 Vo TH, Lee GS, Yang HJ, Kim SH (2020) Pyramid with Super Resolution for In-the-Wild Facial Expression Recognition. IEEE Access 8:131988–132001. 10.1109/ACCESS.2020.3010018 Li S, Deng W, Du J Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. [Online]. Available: http://whdeng.cn/RAF/model1.html You Z, Recognition B et al (eds) (2016) vol. 9967. in Lecture Notes in Computer Science, vol. 9967. Cham: Springer International Publishing, 10.1007/978-3-319-46654-5 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6794812","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":464744511,"identity":"e9516aad-23c8-40cc-b458-ae56f9c67558","order_by":0,"name":"Prof. Ketan Sarvakar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABYElEQVRIie3RsWrCQBjA8ZODuBy4XlCSVzgJ2ILWvMqFg3SxRSgUpyAUdFFcD1roQ3TtcOGgWYJdA7qUgFMHJYulHXqmEhML7Vpo/pBwH7kfdxAAysr+YDi3hurpZBNK33Q/iTyBBeJmO/aE/krk8Q4KjtPHMz/pA8884Swmb4/P3dnt1E82o2WDBNN4/bKVoDYWFdnPSB1JWOdANu8i13ImqwXjyznD/miFSBi0MKUS4JACyTNiYKYYEBWOhCWQWDAQ9YgiEunc1cCOgEhdGB2IGcN3BDyboyDxP8ScmVHP2qbkfgXXO2IWSR1DTZ0CHV6dWAwJ0SVRr5WeUsMaSC9GikSfsFYbEckUubIagtFmdOmehnNFkKth6p6jZugMcwQHfrxAA++Mw+qD/iq6thFdyGhwLW2t+gQ3207bMAIpkwP5ihyWzhAUv96osTL89ndy2eCIeD/tLisrK/sXfQLTDoOnILxs0QAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-4486-0224","institution":"","correspondingAuthor":true,"prefix":"","firstName":"Prof.","middleName":"Ketan","lastName":"Sarvakar","suffix":""}],"badges":[],"createdAt":"2025-06-01 09:11:15","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-6794812/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6794812/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":83832544,"identity":"624c7dfd-cc29-4b28-a798-d6a9933c65ad","added_by":"auto","created_at":"2025-06-03 12:16:15","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":135651,"visible":true,"origin":"","legend":"\u003cp\u003eFace Emotion Recognition Method\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6794812/v1/fef0088e6fee44229e9581b1.png"},{"id":83832546,"identity":"57ac14e4-674a-4c19-b0e4-bba84721f2a0","added_by":"auto","created_at":"2025-06-03 12:16:16","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":178489,"visible":true,"origin":"","legend":"\u003cp\u003eThe Deep Learning Models Utilized In Facial Emotion Recognition\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6794812/v1/b345ffc9c2a5cd8e27a6fc71.png"},{"id":83833356,"identity":"b00c2459-32b8-4dbc-b193-aecdabdb05dd","added_by":"auto","created_at":"2025-06-03 12:24:19","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2959878,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6794812/v1/27043ac4-1c79-4d9d-9930-a2658620ef01.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eA Survey Of Face Emotion Recognition Using Deep Learning Methods\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eA significant indirect communication tool is facial expression, which can convey a person's intentions and feelings. The goal of Face Emotion Recognition (FER) is to recognize these expressions, and it is relevant to fields such as human-computer interaction, security, and healthcare. The intricacy and diversity of human expressions provide difficulties for traditional FER approaches. Deep Learning, on the other hand, has revolutionized FER by offering more precise and effective solutions. Prominent deep learning architectures that have demonstrated exceptional performance in image classification tasks, including FER, are VGG-19, ResNet-50, Inception-V3, and MobileNet. ResNet-50 uses residual information to solve the vanishing gradient problem, while VGG-19 uses modest convolution filters in its deep network for great performance. MobileNet provides a lightweight, accurate model appropriate for embedded and mobile applications, while Inception-V3 uses a modular architecture to capture multi-scale information. This survey provides a thorough overview of current trends and future prospects in FER research by examining these cutting-edge deep learning techniques and assessing their architectures, strengths, and limits.\u003c/p\u003e \u003cdiv id=\"Sec2\" class=\"Section2\"\u003e \u003ch2\u003e1.1 Deep Learning\u003c/h2\u003e \u003cp\u003eThis research study explores the field of deep learning methods for the important problem of classifying emotions on faces. It examines several cutting-edge architectures, including VGG-19, ResNet-50, Inception-V3, and MobileNet, each with special benefits and challenges in precisely identifying facial emotions. In discussing how deep learning is transforming facial emotion recognition, the study emphasizes how this technology can learn intricate patterns from massive datasets and outperform conventional techniques in terms of precision. Also, it tackles issues like overfitting, computational complexity, and model optimization, opening the door for further developments in areas like sophisticated regularization strategies, simplified models, and real-time application optimization. All things considered, the study offers a thorough assessment of the state of deep learning techniques for facial emotion classification and sketches out future directions for development in this rapidly evolving area.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparative Analysis of Deep Learning, Machine Learning, and Artificial Intelligence ,Techniques in Face Emotion Classification\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eComparative Analysis of Deep Learning, Machine Learning, and Artificial Intelligence\u003c/p\u003e \u003cp\u003eTechniques in Face Emotion Classification\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAspect\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDeep Learning\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMachine Learning\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eArtificial Intelligence\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUses neural networks to learn\u003c/p\u003e \u003cp\u003ehierarchically from facial\u003c/p\u003e \u003cp\u003eexpressions and automatically\u003c/p\u003e \u003cp\u003eextract features.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eclassifies facial expressions using statistical models and \u003c/p\u003e \u003cp\u003ealgorithms based on features that are extracted.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBuilds complete emotion identification systems with real-time processing\u003c/p\u003e \u003cp\u003ecapability by integrating ML and DL algorithms.\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTechniques\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e- CNNs are used to automatically\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eextract features. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- RNNs for analysis of time.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Conventional classifiers such as k-NN and SVM. \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e- Techniques for feature\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eengineering.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Hybrid ML-DL techniques to increase\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eprecision. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- Practical uses in sentiment analysis, HCI, and surveillance.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eRole\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e- Improves face expression\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003erecognition speed through \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eautomatic extraction of features. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- Classifies data more accurately.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Uses features that have been\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eretrieved to classify emotions. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- Offers a performance baseline for comparison.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Combines DL and ML for comprehensive\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eemotion identification. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- Makes emotion analysis in real time easier.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eChallenges\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e- Excessive processing needs. \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e- Large labeled datasets are\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003erequired.\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Poor performance with nuanced feelings. \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e- Manual feature extraction takes a lot of time.\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Ethical issues with data utilization and privacy. \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e- Ensured the ability to process data in real-time.\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eFuture\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eDirections\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e- Effective architectures and\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003elearning transfer. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- Making use of unsupervised\u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003elearning to identify tiny clues.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Integration for enhanced performance with DL. \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e- Engineered features\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eautomatically.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Emphasize moral AI with comprehensible\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eresults. \u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003e- Developments in processing in real time.\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe three approaches\u0026mdash;deep learning, machine learning, and artificial intelligence\u0026mdash;for determining emotions from facial expressions are compared in Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Deep learning is the most accurate and fastest approach, but it requires a large amount of data and processing power. Machine learning is slower and less accurate than deep learning, despite being easier to use and requiring less data. Machine learning and deep learning are combined in artificial intelligence to achieve high accuracy and real-time processing. But this raises ethical concerns about the privacy of data. Generally speaking, the best technique for a specific application will depend on how well accuracy, speed, and ease of implementation are traded off.\u003c/p\u003e \u003c/div\u003e"},{"header":"2 Theory Of Emotion","content":"\u003cp\u003eThis study explores facial emotion identification, including face detection techniques such as Viola-Jones and Haar cascades, as well as dataset kinds, labeling, and size impact. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] In an effort to improve accuracy, it covers feature improvement methods including augmentation and makes considerable use of CNNs and SVMs for emotion categorization. The study recommends additional investigation into the ways in which these components interact to improve understanding and efficacy in facial emotion recognition systems.\u003c/p\u003e \u003cp\u003eThis research presents a transfer learning-based facial emotion recognition system. For emotion recognition, the study uses pre-trained convolutional neural networks (CNNs) from the ImageNet database, such as VGG19, ResNet50, InceptionV3, and MobileNet. [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]The CK\u0026thinsp;+\u0026thinsp;database was used for the experiments, and the results showed that the following models had outstanding accuracy rates: ResNet50 (97.7%), InceptionV3 (98.5%), MobileNet (94.2%), and VGG19 (96%). Notably, out\u003c/p\u003e \u003cp\u003eof all the pre-trained networks, MobileNet showed the highest accuracy.[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]The study suggests using these networks in the future to identify emotions in speech and EEG signals, thus expanding the system's potential.\u003c/p\u003e \u003cp\u003eThis study tackles the issues raised by deepfakes, emphasizing how they could harm people's perceptions of the media, their mental health, hate speech, misinformation, and political stability.[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] It addresses the difficulties brought about by the quick dissemination of false information and the availability of tools for creating deepfakes. It also covers new developments and methods for producing and identifying deepfakes.[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] The study intends to make a substantial contribution to AI research by offering workable strategies to lessen the negative consequences of deepfakes.\u003c/p\u003e \u003cp\u003eThis work uses the FER2013 dataset and a variety of models to classify facial expressions; using the Adam optimizer, it achieves an accuracy of 0.60. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]It draws attention to the difficulties in predicting specific emotions because of the scarcity of training data and makes recommendations for data augmentation and study of frequently anticipated emotions to improve accuracy. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]In addition, the study offers insights for future research areas by discussing the model's real-time capabilities and possible applications in other domains.\u003c/p\u003e \u003cp\u003eThis paper delves into the complexities and limitations of emotion recognition solely through facial expressions, highlighting the challenges posed by subtle facial movements and variations in expression across different datasets and populations.[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] It stresses the necessity for a more adaptable framework in emotion recognition and proposes a cross-dataset evaluation method to address biases and assess model generalization.[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] Despite achieving notable results, the study emphasizes the ongoing difficulties in accurately interpreting internal emotional states from facial expressions alone, advocating for a multidisciplinary approach and careful exploration of emotion recognition systems.\u003c/p\u003e \u003cp\u003eThe Multi-Branch Deep RBF Network, which incorporates local information into the recognition process through several RBF units coupled to a VGG-Face backbone, is introduced in this research as a way to improve CNNs.[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] Results from six ER datasets show that CNN performs better, particularly under difficult circumstances like significant class overlap and short sample sizes.[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] The model outperforms state-of-the-art techniques and provides opportunities for further study into the efficient integration of local data, the use of RBF centers for interpretability and explainability, and the mitigation of biases in facial emotion recognition systems.\u003c/p\u003e \u003cp\u003eIn this paper, a comprehensive methodology for realistically testing and assessing Facial Emotion Recognition (FER) models is presented.[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] To show off the framework's usability and make it easier to create new FER datasets, a web application was created. Using a lightweight CNN trained on the AffectNet dataset, remarkable accuracy that was almost human-level was attained.[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] It does, however, recognize the need for more study in order to optimize deep learning methods for FER, particularly when compared to more accurate models such as VGGNet versions. Informed permission and regulatory procedures in their web application are highlighted in the paper, which also highlights the significance of privacy and data management in intelligent behavior analysis.\u003c/p\u003e \u003cp\u003eThe article discusses the usefulness of metrics in identifying biases within both real and manipulated datasets, offering easily interpretable values for analysis. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]Through a case study on the popular FER dataset Affectnet, the study uncovers heavy racial representational bias and stereotypical gender biases. Results indicate that the model is less affected by racial bias removal but significantly impacted by gender bias, highlighting the complexity of bias analysis in machine learning.[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] Future work includes expanding the analysis to more datasets, models, and training setups, developing new demographic datasets and models, exploring mitigation techniques, and refining metrics for broader applications in multi-class multi-demographic AI systems and medical diagnosis.\u003c/p\u003e \u003cp\u003eThe paper introduces a CNN-based approach with deep learning algorithms for emotional recognition through facial features. It enhances accuracy and classification methodology, especially in complex facial expression identification, by employing dimensional space reduction and kernel filters in preprocessing.[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] Data augmentation techniques like cropping and introducing noise, along with picture synthesis methods, were utilized to address overfitting and data scarcity.[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] Future plans include extending feature sets, recognizing additional emotions, and exploring automatic facial emotion recognition.\u003c/p\u003e \u003cp\u003eThis paper presents EMOCA, a self-supervised approach that uses emotion-rich image data to train single photographs to generate 3D facial expressions.[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] By integrating a distinct emotion similarity loss derived from deep features, it achieves better outcomes in 3D face shape reconstruction, especially when it comes to accurately expressing emotions. EMOCA outperforms existing methods for recognizing emotions in the wild and has potential uses in the gaming, cinema, AR/VR, and communication sectors.[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] In addition to addressing issues like deepfakes, the article highlights how crucial it is for human communication and avatar interactions to successfully portray emotions.\u003c/p\u003e \u003cp\u003eThis paper presents a deep learning framework that uses both conventional and novel CNN architectures to identify face emotions.[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] Accuracy was assessed on lab-controlled datasets, RAFD and KDEF, with results of 68.84% and 99.63%, respectively. For the KDEF dataset, a custom model outperformed the most advanced findings. In comparison to lab-controlled datasets, the accuracy of the non-lab-controlled datasets RAF-DB, SFEW, and AMFED\u0026thinsp;+\u0026thinsp;was 75.26%, 40.78%, and 54.15%, respectively.[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] Future developments, according to the study, should focus on enhancing accuracy by unsupervised pre-training using transfer learning, pre-processing, and feature extraction techniques. Additionally, advanced models like DBN, GANs, and Facial Action Units (AUs) should be investigated.\u003c/p\u003e \u003cp\u003eThis paper offers a thorough introduction to DeepFake technology, addressing its foundations, benefits, and related hazards with a special emphasis on GAN-based applications.[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] It talks about the difficulties in identifying DeepFakes, emphasizes the necessity for extra security measures to guarantee data integrity, and forecasts possible outcomes of AI employing DeepFakes to counter AI propaganda.\u003c/p\u003e \u003cp\u003eThis study explores the potential of smartphones combined with AI for aiding in the identification and management of depressive disorders and mental health. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] It envisions personalized treatment recommendations, detection capabilities, and insights through digitally recognized expressions, with applications ranging from acute treatment classification to remote monitoring and preventative care.[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] Although it doesn't delve into natural language processing for depression diagnosis, the study acknowledges the potential of AI-based chatbots for responsive and contextual interactions, aiming for accessible mental health benefits at a low cost and quick deployment.\u003c/p\u003e \u003cp\u003eThe study presents an Efficient-SwishNet model for facial emotion recognition, highlighting its robustness, space efficiency, and affordability. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] It achieves 100% recognition rate on FERG and CK datasets, outperforming previous algorithms on five separate datasets. The model does a great job managing photos with multiple orientations in addition to accurately detecting emotions from frontal face images. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]The model's generalizability is demonstrated through cross-corpora examination. It has trouble with covered faces and occlusions, though, and the scientists intend to work on these issues in later research. Their objectives are to establish a specific FER dataset for real-time performance testing and refine the FER model for cross-corpora evaluation.\u003c/p\u003e \u003cp\u003eThis work presents a transfer learning and ResNet-based facial emotion detection system that achieves high classification rates in several facial expression categories. [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] The ResNet-18 architecture performed optimally, demonstrating the method's efficacy in classifying compound emotions.[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] Future plans call for expanding evaluations to other datasets such as Affectnet and iCV-MEFED, as well as doing cross-database tests.\u003c/p\u003e \u003cp\u003eIn order to combine deep learning with relational reasoning,[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] the paper presents a unified architecture that makes use of SqueezeNet for emotion recognition and Temporal Relation Networks (TRNs). [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] With its temporal relational reasoning, the TRN performs better in training and testing than other models, particularly the multi-scale TRN, when compared to single-scale TRN and multi-layer perceptron models.\u003c/p\u003e \u003cp\u003eIn order to achieve accurate face detection, the paper presents a unique deep learning-based approach that combines prediction with region-offering networks (RON).[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] It delivers good accuracy with a small model size and minimal computing load after being trained on the WIDER FACE dataset.[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] Future research will focus on developing real-time models for facial emotion identification using sophisticated architectures like 3D CNN, 3D U-Net, and YOLOv, as well as enhancing performance in difficult situations like low light and foggy pictures.\u003c/p\u003e \u003cp\u003eThis research outlines the advantages and disadvantages of the conventional ML-based and DL-based techniques to facial expression recognition (FER).[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]It highlights that 3D facial expression datasets are necessary for better results and talks about how DL approaches can be integrated with IoT sensors to boost FER capabilities for different sectors.\u003c/p\u003e \u003cp\u003eThis paper presents a multimodal emotion recognition model that uses CNN model for feature extraction and attention mechanism to merge facial expressions and EEG inputs.[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] More study is needed to improve face feature pre-training models and incorporate different modalities for better model performance, as demonstrated by the experimental results, [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]which show improved emotion identification accuracy when compared to utilizing either modality alone.\u003c/p\u003e \u003cp\u003eThis study shows that the VGG16 model (model B) achieves 100% and 73% success rates, respectively, with the fewest average losses of 0.011 and 0.1875, outperforming models C and A in face and emotion recognition. [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] However, because VGG16 has more layers, the training process takes longer.[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] Despite having a little lower accuracy (87% for face and 67% for emotion detection), Model C contains fewer parameters, which allows for a quicker training procedure. Real-time face and emotion detection systems for humanoid robots have effectively integrated both models. The study emphasizes the necessity of additional system enhancements as well as the significance of dataset size and illumination concerns for future advancements.\u003c/p\u003e \u003cp\u003eThis research presents a deep learning model-based real-time engagement detection method for online learners by analyzing facial emotions. [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] The system creates an engagement index (EI) that indicates whether a user is \"engaged\" or \"disengaged\" based on the monitoring of facial expressions through built-in web cams. It uses MFACXTOR for face point extraction and Faster R-CNN for face detection. It is trained on FER-2013, CK+, and custom datasets. Out of all the models that were assessed, ResNet-50, VGG-19, and Inception-V3 had the best accuracy (92.32%). [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]The method successfully determined involvement levels after testing on 20 students. Future developments could expand to larger datasets and students with specific needs, as well as incorporate data from more sensors.\u003c/p\u003e \u003cp\u003eThis overview highlights the difficulties in identifying emotions in real-world situations and covers robotic facial expression production and human facial expression detection. [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] To promote successful emotion transmission, future research should concentrate on increasing detection under diverse settings and expanding robots' capacity to express a greater variety of emotions, including mixed expressions and varying intensities.\u003c/p\u003e \u003cp\u003eThe study emphasizes that although face masks are useful in halting the spread of viruses, they have a negative impact on facial identity and social perceptions about emotion, age, and gender.[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] These results should not deter people from wearing masks when essential for medical reasons, such as during the COVID-19 epidemic. Rather, they highlight the necessity of cultural modification to reduce the psychological effects of mask wear.[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] Subsequent investigations ought to devise approaches to enhance mask-wearing communication, as it is vital for handling present and potential pandemics.\u003c/p\u003e \u003cp\u003eTo achieve better results, a proprietary expression detection model and the Viola-Jones method for face extraction across various datasets are used in the study's facial expression recognition framework, which makes use of the Jetson Nanodevice for law enforcement. [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]In the future, the framework will be expanded to include gender categorization, age prediction, and more deep-learning models. [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]Data augmentation will be taken into consideration for better performance on devices with limited resources.\u003c/p\u003e \u003cp\u003eThe research created a facial emotion recognition (FER) model by integrating a CNN with a Haar-Cascade classifier.[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] The model demonstrated 90% accuracy in image-based testing, but it encountered difficulties in real-time emotion identification, especially when dealing with masked faces.[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] In the future, 3D CNN, 3D U-Net, and YOLO settings for landmark-based emotion identification will be used to address biases and noise in facial expressions, with the goal of boosting accuracy in real-life scenarios.\u003c/p\u003e \u003cp\u003eThe paper highlights temporal dynamics and spatial asymmetry in the introduction of TSception, a multi-scale convolutional neural network intended for EEG emotion recognition.[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] TSception uses parallel multi-scale temporal kernels and hemisphere-specific kernels to improve emotional asymmetry pattern learning. Using fewer trainable parameters, evaluation on benchmark datasets showed improved classification performance.[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] Saliency maps facilitated further research into the impact of segment length and cross-individual generalization by identifying important informative locations and lateralization patterns.\u003c/p\u003e \u003cp\u003eWith the use of the AffectNet dataset and deep CNN models,[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] this research is able to identify emotions from masked faces with an accuracy of 69.3% and an average confusion matrix of 71.4%. In the future, the model will be improved via semi-supervised learning, attention processes without landmarks will be improved, [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]and a prototype device to help visually impaired people recognize their environment will be developed.\u003c/p\u003e \u003cp\u003eKey issues in emotion identification across multiple modalities are highlighted in this survey, including the necessity for accurate emotional classification in physiological data, individual differences affecting facial emotion detection, and variability in speech emotion recognition (SER).[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] It highlights the superiority of deep learning on larger datasets and the significance of combining sensors for dependable multi-modal recognition. [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] Prospective avenues for advancement encompass enhancing resilience, precision, confidentiality, and user acceptance via multi-modal approaches and sophisticated learning methodologies such as unsupervised and reinforcement learning.\u003c/p\u003e \u003cp\u003eThis research presents a novel approach to remote facial video analysis employing RGB, NIR, and infrared cameras for emotion classification.[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] With a single feature from RGB films and a DL classifier, the study obtains encouraging results with 47.36% average accuracy by focusing on extracting features from pulsatile heartbeats in facial video frames.[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] The study highlights RGB cameras' potential for efficient emotion classification, but it also notes drawbacks including subject mobility and brief video inputs. For greater practicality, future work may incorporate facial tracking and registration.\u003c/p\u003e \u003cp\u003eUsing computer vision and deep learning techniques, this work conducts a systematic literature review (SLR) on emotion recognition, examining 77 academic papers.[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] It highlights trends in emotion recognition across face and body positions and highlights possible uses in law enforcement and healthcare. Convolutional Neural Networks (CNNs) are robust and accurate, yet they still face issues with limited technology and scarce data. [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] Promising techniques such as vision transformers are particularly useful for handling micro- and macro-expressions. Despite the constraints of the dataset, future study may include hand and body expressions. Continuous efforts are being made to improve accuracy and practical utility in real-world circumstances.\u003c/p\u003e"},{"header":"3 Material Methods","content":"\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Face Emotion Recognition Method\u003c/h2\u003e \u003cp\u003eThe Utilizing computer vision and deep learning approaches, the Face Emotion Recognition (FER) technology systematically analyzes face emotions. Consistency-checking face picture preprocessing, feature extraction using pre-trained CNNs like VGG 19, ResNet 50, MobileNet, and Inception V3, and emotion classification utilizing the derived features are all included. The technique seeks to improve emotion identification technology by precisely identifying and classifying the emotions portrayed in faces.\u003c/p\u003e \u003cp\u003eIn the Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e is represent that The flowchart displays a computer vision method for facial emotion recognition that begins with pre-processing photos in a dataset to guarantee consistency. As a feature extractor, a pre-trained convolutional neural network (CNN) such as VGG 19, ResNet 50, MobileNet, or Inception V3 examines the facial expressions in the pictures. A classifier uses these extracted features to classify the emotions that are represented in the faces. Ultimately, the system's efficacy is assessed to see how well it can identify emotions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eA Discussion Of The Typical Conventional FER Techniques In A Summary\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003eA Discussion Of The Typical Conventional FER Techniques In A Summary\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCitations\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCollections of Data\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMethods of Decision Making\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCharacteristics\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAnalysis of Emotions\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eFECR\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- HMM and the SVM classifier\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- PCA and LDA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;6 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK Plus\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Classifier using support vector machines\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Gabor wavelets, nonlinear NLPCA, and the Haar wavelet transform\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;6 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMMI ,\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eJapanese Female Facial Expression Database ,\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eCK\u0026thinsp;+\u0026thinsp;Data set\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- large margin classifier\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- ORB ,SIFT ,SURF\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK+\u003c/b\u003e,\u003c/p\u003e \u003cp\u003e\u003cb\u003eMMI Facial Expression Corpus\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- SMO, MLP, and KNN for categorization\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- DCT, HOG\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK Plus\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Denoising Sparse Autoencoder\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Gabor Function, LBP, SIFT, and HOG\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eDataset from a visual depth camera in real time\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- The DBN\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Features of MLDP-GDA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;6 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eBOSPHORUS\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Euclidean distance\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Geometric descriptor\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;4 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK+ ,\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eMMI Dataset ,\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eMUG\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- The AdaBoost-ELM, KLT, and EBGM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Salient geometric features\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eBVTKFER \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eBCurtin Faces\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- The Random forest classifier\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- LBP\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;6 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eUVANEMO, SPOS, MMI, and BBC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Support Vector Machine\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Four standard features\u0026mdash;raw pixels, Gabor, HOG, and LBP\u0026mdash;as well as RSTD\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;2 Emotions Smile \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e(genuine and fake)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK\u0026thinsp;+\u0026thinsp;Facial Expression Dataset\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Conditional Random Field and KNN\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- The Geometric descriptor\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK+, Japanese Female Facial Expression Database\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- EHMM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- ASM and 2D DCT\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCohn-Kanade Plus\u003c/b\u003e,\u003c/p\u003e \u003cp\u003e\u003cb\u003eJAFFE Dataset, MUG\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- SVM, PCA, LDA, and K-NN\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Gabor wavelets and geometric features\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCK, JAFFE ,USTCNVIE, Yale, FEI\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- Markov Hidden Models\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Chan-Vese energy function, Bhattacharyya distance function \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003ewavelet decomposition and SWLDA\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;6 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eExtended Cohn-Kanade Dataset\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e- The SVR\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e- Gabor wavelets, AAMS, and feature descriptors\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e\u0026minus;\u0026thinsp;7 Emotions\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e provides an overview of the many conventional facial emotion recognition (FER) techniques that have been used in different research. The sources include a list of the datasets and methods utilized in the emotion categorization decision-making process. Commonly used datasets with classifiers such as HMM, SVM, KNN, and deep belief networks are CK+, MMI, and JAFFE. Feature extraction methods include Gabor wavelets, PCA and LDA, HOG, and LBP. These techniques often aim to categorize six or seven primary emotions, however some studies focus on a specific subset of emotions. The table illustrates the range of techniques and datasets used to identify emotions from facial expressions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Comparing With The Modern Techniques\u003c/h2\u003e \u003cp\u003eThe comparison with state-of-the-art methods highlights the efficacy of three CNN architectures in facial emotion recognition: VGG-19, ResNet-50, and Inception V3. The structural components of these models that are evaluated include convolutional layers for feature extraction, pooling layers for dimensionality reduction, batch normalization for stability, activation layers for non-linearity, and fully connected layers for classification. Because performance varies depending on a number of factors, including processing resources, dataset size, and accuracy requirements, there is no one ideal architecture. Rather, the effectiveness of each building is influenced by its own special arrangement and number of layers.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1 Emotion Recognition System Effectiveness Assessment on CFEE and RAF Databases\u003c/h2\u003e \u003cp\u003eUsing the posed CFEE and spontaneous RAF databases, this part assesses the system's efficacy in identifying basic and compound emotions by contrasting the outcomes with sophisticated current approaches. The comparison of the CFEE database's basic emotion recognition (7 classes) with two cutting-edge investigations is shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e Notably, a study that used the AlexNet CNN architecture was able to attain 74.799% test accuracy. Seventy percent of the photos in this study were used for training, and test and validation sets each had 159 images.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eEvaluation Of The CFEE Database's Fundamental And Compound Emotions With Modern Methods\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"6\" nameend=\"c6\" namest=\"c1\"\u003e \u003cp\u003e Evaluation Of The CFEE Database's Fundamental And Compound Emotions With Modern Methods\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRef.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eApproach\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eThe procedure\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eExamples\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eClasses\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eShape and appear\u0026thinsp;+\u0026thinsp;Nearest mean\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e10-fold cross-validation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1610\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e7-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e96.96%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eThe AlexNet CNN\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e70%-15%-15% \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e(train/validation/test)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1127\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e245\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e238\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e74.79%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eShape and appearance\u0026thinsp;+\u0026thinsp;Nearest mean\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e10-fold cross-validation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e5060\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e22-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e76.91%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eThe Highway-CNN\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e52.14%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003eThe Resnet-18 Deep Features\u0026thinsp;+\u0026thinsp;SVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e10-fold cross-validation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e1610\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e\u003cb\u003e7-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e98.02%*\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e99.19%+\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e70% train-15% test\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e1365\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e81.93%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e10-fold cross-validation\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e5060\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e22-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e80.69%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eUsing the RAF database, Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e offers a comparative comparison of the system's performance in identifying basic and compound emotions. The comparison incorporates findings from multiple cutting-edge methods. The top-performing tested methods are emphasized, showcasing their recognition rates against the spontaneous emotional expressions in the RAF database. These comparisons highlight the system's efficacy and offer a baseline against the highest reported recognition rates found in the literature review.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison Of The RAF Database Applying Modern Methods\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"6\" nameend=\"c6\" namest=\"c1\"\u003e \u003cp\u003eComparison Of The RAF Database Applying Modern Methods\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRef\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eThe process\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eThe Samples\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eClasses\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eThe precision\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eGabor\u0026thinsp;+\u0026thinsp;mSVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\" morerows=\"12\" rowspan=\"13\"\u003e \u003cp\u003e\u003cb\u003e15339\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"12\" rowspan=\"13\"\u003e \u003cp\u003e\u003cb\u003e7-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e65%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eDeep Locality-Preserving CNN\u0026thinsp;+\u0026thinsp;mSVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e74%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e84.13%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eAugmented data and cluster loss\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e76%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMulti-Region Ensemble -CNN (VGG-16)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e77%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eMulti-Region Ensemble CNN (AlexNet)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e75%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCapsule Network\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e77%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eDouble Completed-LBP (Double Cd-LBP)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e78%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eTransfer learning Resnet-18 (AffectNet database)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e80%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eCovariance pooling after final convolutional layers\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e79%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e87.0%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003ePatch-Gated CNN (PG-CNN)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e83%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eConditional generative adversarial network-based \u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eEAU-Net Network (CGAN based EAU-Net)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e81.83%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eRegion Attention Network (RAN)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e86.90%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003ePyramid With Super-Resolution (PSR) Network (VGG-16)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e80.78%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e88.98%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eGabor\u0026thinsp;+\u0026thinsp;mSVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e3954\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e11-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e33.76%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eDeep Locality-Preserving CNN\u0026thinsp;+\u0026thinsp;mSVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e44.55%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e57.95%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eLBP\u0026thinsp;+\u0026thinsp;NCMML\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e3171\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003e16-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e30.10%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e36.70%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eHOG\u0026thinsp;+\u0026thinsp;NCMML\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e36.90%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e44.10%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e\u003cb\u003eResnet-18 Deep Features\u0026thinsp;+\u0026thinsp;SVM\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e15339\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e7-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e86%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e93.29%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e3954\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e11-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e63.27%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e76.52%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e19059\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e16-class\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e60.76%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4 The Recommended Models' Training Process","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e4.1 VGG-19\u003c/h2\u003e \u003cp\u003eThe Visual Graphics Group at Oxford introduced VGG-19, a convolutional neural network (CNN) that is well-known for being both straightforward and efficient in image identification applications, in 2014. With its modest receptive fields (3x3) and max-pooling layers, VGG-19 has a simple architecture made up of 19 layers (16 convolutional and 3 fully linked). The network performs admirably, particularly in picture classification tasks like the ImageNet challenge, where it obtained remarkable accuracy thanks to its depth and homogeneous structure. But the model's high number of parameters and deep water can result in computing complexity, which limits its usefulness for real-time applications on devices with limited resources. In spite of this, VGG-19 continues to be a cornerstone model in the field of deep learning research, acting as a standard and influencing later network architectures.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Resnet50\u003c/h2\u003e \u003cp\u003eResNet50, a groundbreaking deep learning model developed in 2015 by Kaiming He et al., changed computer vision thanks to its creative use of residual learning. With 50 convolutional layers split up into five phases, each containing bottleneck blocks, ResNet50 improved training efficiency by introducing identity shortcut connections to resolve gradient problems and allow for previously unattainable network depths. It gained greater accuracy on benchmarks such as ImageNet by using a combination of 1x1, 3x3, and 1x1 convolutions along with batch normalization. As a result, it became indispensable for tasks like object recognition (e.g., Faster R-CNN) and picture segmentation (e.g., Mask R-CNN). ResNet50 is a mainstay of contemporary computer vision systems because of its capacity to optimize deep networks.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Inception V3\u003c/h2\u003e \u003cp\u003eAmongst the Inception family of CNNs, Inception V3, created by Google in 2015, stands out for its emphasis on object recognition and image classification. Its Inception module, which uses a variety of filter sizes to capture features, factorized convolutions for efficiency, auxiliary classifiers to help with deep network training, and the regular application of batch normalization and ReLU for quicker convergence are some of its salient features. Stem layers, Inception modules, auxiliary classifiers, and final layers for classification make up its architecture. Exceptional advantages include cutting-edge performance in picture tasks, less overfitting, and effective feature extraction. As a flexible and efficient model in contemporary computer vision, Inception V3 has applications in picture categorization, object recognition, and transfer learning.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e is represent that Three CNN architectures used for facial emotion detection are compared in the graphic named VGG-19, ResNet-50, and Inception V3. Convolutional layers for feature extraction, pooling layers for dimensionality reduction, batch normalization for training stabilization, activation layers for adding non-linearity, and fully connected layers for classification are the structural elements shared by these models. The order and number of layers in each architecture vary: Inception V3 mixes convolutional and inception modules, ResNet-50 has fifty, while VGG-19 has sixteen convolutional layers. The size of the dataset, the amount of computing power needed, and the level of accuracy needed determine the optimal architecture for a given job; the image does not reveal the top-performing model for facial emotion identification.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e4.4 MobileNet\u003c/h2\u003e \u003cp\u003eThe google presented MobileNet in 2017 as a powerful deep-learning solution designed specifically for embedded and mobile devices. Its main innovation is depthwise separable convolutions, which significantly lower processing demands by splitting ordinary convolutions into depthwise and pointwise layers. Furthermore, MobileNet adds parameters such as width and resolution multipliers, which provide simple scaling and customization of model efficiency and size. With inverted residuals and linear bottlenecks, the transition to MobileNetV2 improves feature representation even more, preserving high accuracy at cheap computing costs. Because of its adaptability and suitability for a range of computer vision tasks, including object identification, segmentation, and image classification, MobileNet has become the standard for real-time applications on platforms with limited resources.\u003c/p\u003e \u003c/div\u003e"},{"header":"5 Implementation","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e5.1 Evaluation Metrics\u003c/h2\u003e \u003cp\u003eWe computed the accuracy, mean precision, and mean recall\u0026mdash;standard measures generally employed by modern techniques for face emotion recognition\u0026mdash;in order to assess the performance of the suggested model.\u003c/p\u003e \u003cdiv id=\"Sec15\" class=\"Section3\"\u003e \u003ch2\u003e5.1.1 Accuracy\u003c/h2\u003e \u003cp\u003eAccuracy is the most crucial statistic in multi-class classification. It is computed by dividing the total number of instances in the emotion class by the sum of the true negative (TN) and true positive (TP) instances. The classifier's accuracy is calculated as :\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:Overall\\:Accuracy=\\:\\:\\:\\:\\:\\frac{{\\sum\\:}_{i=1}^{n}{TP}_{i}+{TN}_{i\\:\\:\\:}}{TP+TN+FP+FN}\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\left(1\\right)\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e5.1.2 Mean Precision\u003c/h2\u003e \u003cp\u003eFor each emotion class, precision is the genuine positive prediction that the suggested model makes. In the event of multi-class classification, the model's mean precision is computed as follows:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:Mean\\:Precision=\\:\\:\\:\\:\\:\\frac{{\\sum\\:}_{i=1}^{n}{\\:(TP}_{i})}{{\\:\\:{\\sum\\:}_{i=1}^{n}\\:({TP}_{i\\:\\:\\:}+\\:FP}_{i})}\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\left(2\\right)\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003e5.1.3 Recall\u003c/h2\u003e \u003cp\u003eIdentify the specific number of positive class predictions that were made from all of the dataset's positive instances of the emotion class. Here is how we calculated the mean recall:\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:Mean\\:Precision=\\:\\:\\:\\:\\:\\frac{{\\sum\\:}_{i=1}^{n}{\\:(TP}_{i})}{{\\:\\:{\\sum\\:}_{i=1}^{n}\\:({TP}_{i\\:\\:\\:}+\\:FN}_{i})}\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\left(3\\right)\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003e5.1.4 Specificity\u003c/h2\u003e \u003cp\u003eIt is defined as a portion of negative emotion classes that the classifier correctly categorizes as negative. Specificity is another name for true negative rate, or TNR. The mean specificity was calculated as follows.\u003cdiv id=\"Equd\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equd\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:Mean\\:Specificity=\\:\\:\\:\\:\\:\\frac{{\\sum\\:}_{i=1}^{n}{\\:(TN}_{i})}{{\\:\\:{\\sum\\:}_{i=1}^{n}\\:({TN}_{i\\:\\:\\:}+\\:FP}_{i})}\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\left(4\\right)\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section3\"\u003e \u003ch2\u003e5.1.5 MACRO F-1 SCORE\u003c/h2\u003e \u003cp\u003eThe computation of the macro F-1 score involves obtaining the unweighted arithmetic mean of the F-1 score for every emotion class. In terms of TP, false positives (FP), and false negatives (FN), the macro F-1 score is represented as :\u003cdiv id=\"Eque\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Eque\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:Macro\\:F-1\\:Score=\\:\\:\\:\\:\\:\\frac{{\\sum\\:}_{i=1}^{n}{\\:(TP}_{i})}{{\\:\\:{\\sum\\:}_{i=1}^{n}\\:({TP}_{i\\:\\:}+0.5(\\:FP}_{i}+{FN}_{i}\\:\\left)\\right)}\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\left(5\\right)\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn each of the equations (1) through (5), n is the total number of classes; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\:TP}_{i}\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\:TN}_{i}\\)\u003c/span\u003e\u003c/span\u003estand for true positives and true negatives, respectively, while\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\:FP}_{i}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\:FN}_{i}\\)\u003c/span\u003e\u003c/span\u003e represent the number of false positives and false negatives for emotion class i.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"6 Conclusion","content":"\u003cp\u003eAt some point, the study highlights the noteworthy progress made in facial emotion recognition (FER) using deep learning methods, especially with models such as VGG-19, ResNet-50, Inception-V3, and MobileNet. These models' exceptional accuracy in facial expression classification highlights their potential for practical uses in security systems, mental health monitoring, human-computer interaction, and other fields.While acknowledging the difficulties associated with overfitting, computational complexity, and model optimization, the paper also makes recommendations for future directions in the field, including feature augmentation techniques, transfer learning, and integration with other modalities including voice and EEG inputs. Through a comprehensive assessment of deep learning models, technical discussion, and application area identification, the research opens the door for future advancements and improvements in FER systems, leading to more accurate and efficient human-emotion recognition technologies.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors and Affiliations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eKetan Sarvakar\u003c/p\u003e\n\u003cp\u003eResearch Scholar , Gujarat Technological University, Ahmedabad, India\u003c/p\u003e\n\u003cp\u003eDr. Kaushik Rana\u003c/p\u003e\n\u003cp\u003eAssociate Professor,Gujarat Technological University, Ahmedabad, India \u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eContributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization: Ketan Sarvakar; Methodology: Ketan Sarvakar; Formal anaysis and investigation: Ketan Sarvakar; Writingoriginal draft preparation: Ketan Sarvakar; Writing, review and editing: Ketan Sarvakar, Dr. Kaushik Rana; Resources: Ketan Sarvakar; Supervision: Dr. Kaushik Rana.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCorresponding author\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCorresponding author: Ketan Sarvakar(
[email protected])\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics declarations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResearch involving human and /or animals\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eNaga P, Marri SD, Borreo R (Jan. 2023) Facial emotion recognition methods, datasets and technologies: A literature survey. Mater Today Proc 80:2824\u0026ndash;2828. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.matpr.2021.07.046\u003c/span\u003e\u003cspan address=\"10.1016/j.matpr.2021.07.046\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChowdary MK, Nguyen TN, Hemanth DJ (2023) Deep learning-based facial emotion recognition for human\u0026ndash;computer interaction applications, Neural Comput Appl, vol. 35, no. 32, pp. 23311\u0026ndash;23328, Nov. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00521-021-06012-8\u003c/span\u003e\u003cspan address=\"10.1007/s00521-021-06012-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen TT et al (2019) Deep Learning for Deepfakes Creation and Detection: A Survey. Sep. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cviu.2022.103525\u003c/span\u003e\u003cspan address=\"10.1016/j.cviu.2022.103525\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaravanan A, Perichetla G, Gayathri DKS (2019) Facial Emotion Recognition using Convolutional Neural Networks, Oct. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1910.05602\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1910.05602\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDias W et al (2023) Cross-dataset emotion recognition from facial expressions through convolutional neural networks A R T I C L E I N F O. J Vis Commun Image Represent. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5281/zenodo.4696032\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.4696032\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHern\u0026aacute;ndez-Luquin F, Escalante HJ (2023) Multi-branch deep radial basis function networks for facial emotion recognition, Neural Comput Appl, vol. 35, no. 25, pp. 18131\u0026ndash;18145, Sep. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00521-021-06420-w\u003c/span\u003e\u003cspan address=\"10.1007/s00521-021-06420-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiddiqui N, Dave R, Bauer T, Reither T, Black D, Hanson M A Robust Framework for Deep Learning Approaches to Facial Emotion Recognition and Evaluation, IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDominguez-Catena I, Paternain D, Galar M Assessing Demographic Bias Transfer from Dataset to Model: A Case Study in Facial Expression Recognition, May 2022, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/2205.10049\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/2205.10049\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKumar T, Arora et al (2022) Optimal Facial Feature Based Emotional Recognition Using Deep Learning Algorithm, Comput Intell Neurosci, vol. 2022. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1155/2022/8379202\u003c/span\u003e\u003cspan address=\"10.1155/2022/8379202\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDaněček R, Black M EMOCA: Emotion Driven Monocular Face Capture and Animation. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://emoca.is.tue.mpg.de\u003c/span\u003e\u003cspan address=\"https://emoca.is.tue.mpg.de\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAppasaheb Borgalli R, Surve S (2022) Deep Learning Framework for Facial Emotion Recognition using CNN Architectures, in Proceedings of the International Conference on Electronics and Renewable Systems, ICEARS 2022, Institute of Electrical and Electronics Engineers Inc., pp. 1777\u0026ndash;1784. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ICEARS53579.2022.9751735\u003c/span\u003e\u003cspan address=\"10.1109/ICEARS53579.2022.9751735\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMalik A, Kuribayashi M, Abdullahi SM, Khan AN (2022) DeepFake Detection for Human Face Images and Videos: A Survey. IEEE Access 10:18757\u0026ndash;18775. Institute of Electrical and Electronics Engineers Inc.\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2022.3151186\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2022.3151186\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee YS, Park WH (Feb. 2022) Diagnosis of Depressive Disorder Model on Facial Expression Based on Fast R-CNN. Diagnostics 12(2). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/diagnostics12020317\u003c/span\u003e\u003cspan address=\"10.3390/diagnostics12020317\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDar T, Javed A, Bourouis S, Hussein HS, Alshazly H (2022) Efficient-SwishNet Based System for Facial Emotion Recognition. IEEE Access 10:71311\u0026ndash;71328. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2022.3188730\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2022.3188730\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK., R. Y., \u0026amp; M. R. (2022) C. facial emotional expression recognition using cnn deep features. E. L. 30(4), 1402\u0026ndash;1416. Slimani, Compound facial emotional expression recognition using cnn deep features, 2022\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePise A, Vadapalli H, Sanders I (Aug. 2022) Facial emotion recognition using temporal relational network: an application to E-learning. Multimed Tools Appl 81:26633\u0026ndash;26653. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s11042-020-10133-y\u003c/span\u003e\u003cspan address=\"10.1007/s11042-020-10133-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMamieva D, Abdusalomov AB, Mukhiddinov M, Whangbo TK (Jan. 2023) Improved Face Detection Method via Learning Small Faces on Hard Images Based on a Deep Learning Approach. Sensors 23(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s23010502\u003c/span\u003e\u003cspan address=\"10.3390/s23010502\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan AR (2022) Facial Emotion Recognition Using Conventional Machine Learning and Deep Learning Methods: Current Achievements, Analysis and Remaining Challenges. Inform (Switzerland). 13, 6. MDPI, Jun. 01 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/info13060268\u003c/span\u003e\u003cspan address=\"10.3390/info13060268\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang S, Qu J, Zhang Y, Zhang Y (2023) Multimodal Emotion Recognition From EEG Signals and Facial Expressions. IEEE Access 11:33061\u0026ndash;33068. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2023.3263670\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2023.3263670\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDwijayanti S, Iqbal M, Suprapto BY (2022) Real-time Implementation of Face Recognition and Emotion Recognition in a Humanoid Robot Using a Convolutional Neural Network. IEEE Access. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2022.3200762\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2022.3200762\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGupta S, Kumar P, Tekchandani RK (2023) Facial emotion recognition based real-time learner engagement detection system in online learning context using deep learning models, Multimed Tools Appl, vol. 82, no. 8, pp. 11365\u0026ndash;11394, Mar. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s11042-022-13558-9\u003c/span\u003e\u003cspan address=\"10.1007/s11042-022-13558-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRawal N, Stock-Homburg RM (2022) Facial Emotion Expressions in Human\u0026ndash;Robot Interaction: A Survey, Int J Soc Robot, vol. 14, no. 7, pp. 1583\u0026ndash;1604, Sep. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s12369-022-00867-0\u003c/span\u003e\u003cspan address=\"10.1007/s12369-022-00867-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWong HK, Estudillo AJ (Dec. 2022) Face masks affect emotion categorisation, age estimation, recognition, and gender classification from faces. Cogn Res Princ Implic 7(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s41235-022-00438-x\u003c/span\u003e\u003cspan address=\"10.1186/s41235-022-00438-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlsharekh MF (Aug. 2022) Facial Emotion Recognition in Verbal Communication Based on Deep Learning. Sensors 22(16). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s22166105\u003c/span\u003e\u003cspan address=\"10.3390/s22166105\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFarkhod A, Abdusalomov AB, Mukhiddinov M, Cho YI (2022) Development of Real-Time Landmark-Based Emotion Recognition CNN for Masked Faces, Sensors, vol. 22, no. 22, Nov. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s22228704\u003c/span\u003e\u003cspan address=\"10.3390/s22228704\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing Y, Robinson N, Zhang S, Zeng Q, Guan C (2023) TSception: Capturing Temporal Dynamics and Spatial Asymmetry From EEG for Emotion Recognition, IEEE Trans Affect Comput, vol. 14, no. 3, pp. 2238\u0026ndash;2250, Jul. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TAFFC.2022.3169001\u003c/span\u003e\u003cspan address=\"10.1109/TAFFC.2022.3169001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMukhiddinov M, Djuraev O, Akhmedov F, Mukhamadiyev A, Cho J (Feb. 2023) Masked Face Emotion Recognition Based on Facial Landmarks and Deep Learning Approaches for Visually Impaired People. Sensors 23(3). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s23031080\u003c/span\u003e\u003cspan address=\"10.3390/s23031080\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCai Y, Li X, Li J (2023) Emotion Recognition Using Different Sensors, Emotion Models, Methods and Datasets: A Comprehensive Review, Sensors, vol. 23, no. 5. MDPI, Mar. 01. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s23052455\u003c/span\u003e\u003cspan address=\"10.3390/s23052455\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTalala S, Shvimmer S, Simhon R, Gilead M, Yitzhaky Y (2024) Emotion Classification Based on Pulsatile Images Extracted from Short Facial Videos via Deep Learning, Sensors, vol. 24, no. 8, Apr. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s24082620\u003c/span\u003e\u003cspan address=\"10.3390/s24082620\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePereira R et al (May 2024) Systematic Review of Emotion Detection with Computer Vision and Deep Learning. Sensors 24(11):3484. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/s24113484\u003c/span\u003e\u003cspan address=\"10.3390/s24113484\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTechno-Societal (2018) Springer International Publishing, 2020. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-030-16848-3\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-16848-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReddy CVR, Reddy US, Kishore KVK (2019) Facial emotion recognition using NLPCA and SVM, Traitement du Signal, vol. 36, no. 1, pp. 13\u0026ndash;22, Feb. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.18280/ts.360102\u003c/span\u003e\u003cspan address=\"10.18280/ts.360102\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSajjad M, Nasir M, Ullah FUM, Muhammad K, Sangaiah AK, Baik SW (2019) Raspberry Pi assisted facial expression recognition framework for smart security in law-enforcement services, Inf Sci (N Y), vol. 479, pp. 416\u0026ndash;431, Apr. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ins.2018.07.027\u003c/span\u003e\u003cspan address=\"10.1016/j.ins.2018.07.027\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNazir M, Jan Z, Sajjad M (2017) Facial expression recognition using weber discrete wavelet transform. J Intell Fuzzy Syst 33(1):479\u0026ndash;489. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3233/JIFS-161787\u003c/span\u003e\u003cspan address=\"10.3233/JIFS-161787\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeng N, Zhang H, Song B, Liu W, Li Y, Dobaie AM (Jan. 2018) Facial expression recognition via learning deep sparse autoencoders. Neurocomputing 273:643\u0026ndash;649. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.neucom.2017.08.043\u003c/span\u003e\u003cspan address=\"10.1016/j.neucom.2017.08.043\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUddin MZ, Hassan MM, Almogren A, Zuair M, Fortino G, Torresen J (2017) A facial expression recognition system using robust face features from depth videos and deep learning, Computers and Electrical Engineering, vol. 63, pp. 114\u0026ndash;125, Oct. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.compeleceng.2017.04.019\u003c/span\u003e\u003cspan address=\"10.1016/j.compeleceng.2017.04.019\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAl-agha SA, Saleh HH, Ghani RF (2017) Geometric-based Feature Extraction and Classification for Emotion Expressions of 3D Video Film. J Adv Inform Technol 74\u0026ndash;79. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.12720/jait.8.2.74-79\u003c/span\u003e\u003cspan address=\"10.12720/jait.8.2.74-79\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGhimire D, Lee J, Li ZN, Jeong S (2017) Recognition of facial expressions based on salient geometric features and support vector machines, Multimed Tools Appl, vol. 76, no. 6, pp. 7921\u0026ndash;7946, Mar. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s11042-016-3428-9\u003c/span\u003e\u003cspan address=\"10.1007/s11042-016-3428-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang J, Yang H Face detection based on template matching and 2DPCA algorithm, in Proceedings \u0026ndash;\u0026thinsp;1st International Congress on Image and Signal Processing, CISP 2008, 2008, pp. 575\u0026ndash;579. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/CISP.2008.270\u003c/span\u003e\u003cspan address=\"10.1109/CISP.2008.270\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu Pping, Liu H, Zhang Xwu, Gao Y (2017) Spontaneous versus posed smile recognition via region-specific texture descriptor and geometric facial dynamics, Frontiers of Information Technology and Electronic Engineering, vol. 18, no. 7, pp. 955\u0026ndash;967, Jul. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1631/FITEE.1600041\u003c/span\u003e\u003cspan address=\"10.1631/FITEE.1600041\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIEEE Computer Society and Institute of Electrical and Electronics Engineers 12th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition: FG 2017 : proceedings : 30 May \u0026ndash;\u0026thinsp;3 June 2017, Washington, D.C\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim DJ (Jul. 2016) Facial expression recognition using ASM-based post-processing technique. Pattern Recognit Image Anal 26(3):576\u0026ndash;581. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1134/S105466181603010X\u003c/span\u003e\u003cspan address=\"10.1134/S105466181603010X\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCornejo JYR, Pedrini H, Fl\u0026oacute;rez-Revuelta F Facial Expression Recognition with Occlusions based on Geometric Representation.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiddiqi MH, Ali R, Khan AM, Kim ES, Kim GJ, Lee S (2015) Facial expression recognition using active contour-based face detection, facial movement-based feature extraction, and non-linear feature selection, Multimed Syst, vol. 21, no. 6, pp. 541\u0026ndash;555, Nov. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s00530-014-0400-2\u003c/span\u003e\u003cspan address=\"10.1007/s00530-014-0400-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang KY, Chen CS, Hung YP (2013) Intensity rank estimation of facial expressions based on a single image, in Proceedings \u0026ndash;\u0026thinsp;2013 IEEE International Conference on Systems, Man, and Cybernetics, SMC 2013, pp. 3157\u0026ndash;3162. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/SMC.2013.538\u003c/span\u003e\u003cspan address=\"10.1109/SMC.2013.538\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu S, Tao Y, Martinez AM (Apr. 2014) Compound facial expressions of emotion. Proc Natl Acad Sci U S A 111(15). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1073/pnas.1322355111\u003c/span\u003e\u003cspan address=\"10.1073/pnas.1322355111\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eExpressions of Emotion Dataset [5] We trained the network on CFEE dataset and tested the trained model on RaFD. 2. Related Work.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSlimani K, Lekdioui K, Messoussi R, Touahni R (2019) Compound facial expression recognition based on highway CNN, in ACM International Conference Proceeding Series, Association for Computing Machinery, Mar. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1145/3314074.3314075\u003c/span\u003e\u003cspan address=\"10.1145/3314074.3314075\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi S, Deng W, Du J Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://whdeng.cn/RAF/model1.html\u003c/span\u003e\u003cspan address=\"http://whdeng.cn/RAF/model1.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInstitute of Electrical and Electronics Engineers and, Signal Processing IEEE, Society (2018) IEEE International Conference on Image Processing: proceedings : October 7\u0026ndash;10, 2018, Megaron Athens International Conference Centre, Athens, Greece\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan Y, Lam JCK, Li VOK Multi-Region Ensemble Convolutional Neural Network for Facial Expression Recognition.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGhosh S, Dhall A, Sebe N, Automatic Group Affect Analysis in Images via Visual Attribute and Feature Networks, in Proceedings - International Conference on Image Processing, ICIP, Computer Society IEEE (2018) Aug. pp. 1967\u0026ndash;1971. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ICIP.2018.8451242\u003c/span\u003e\u003cspan address=\"10.1109/ICIP.2018.8451242\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInstitute of Electrical and Electronics Engineers and, Signal Processing IEEE, Society (2018) IEEE International Conference on Image Processing: proceedings : October 7\u0026ndash;10, 2018, Megaron Athens International Conference Centre, Athens, Greece\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVielzeuf V, Kervadec C, Pateux S, Lechervy A, Jurie F (2018) An Occam\u0026rsquo;s Razor View on Learning Audiovisual Emotion Recognition with Small Training Sets, Aug. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1808.02668\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1808.02668\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAcharya D, Huang Z, Paudel DP, Gool V Covariance Pooling for Facial Expression Recognition.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Zeng J, Shan S, Chen X Patch-Gated CNN for Occlusion-aware Facial Expression Recognition.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeng J, Pang G, Zhang Z, Pang Z, Yang H, Yang G (2019) CGAN Based Facial Expression Recognition for Human-Robot Interaction. IEEE Access 7:9848\u0026ndash;9859. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2019.2891668\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2019.2891668\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang K, Peng X, Yang J, Meng D, Qiao Y Region Attention Networks for Pose and Occlusion Robust Facial Expression Recognition, May 2019, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://arxiv.org/abs/1905.04075\u003c/span\u003e\u003cspan address=\"http://arxiv.org/abs/1905.04075\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVo TH, Lee GS, Yang HJ, Kim SH (2020) Pyramid with Super Resolution for In-the-Wild Facial Expression Recognition. IEEE Access 8:131988\u0026ndash;132001. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2020.3010018\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2020.3010018\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi S, Deng W, Du J Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://whdeng.cn/RAF/model1.html\u003c/span\u003e\u003cspan address=\"http://whdeng.cn/RAF/model1.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYou Z, Recognition B et al (eds) (2016) vol. 9967. in Lecture Notes in Computer Science, vol. 9967. Cham: Springer International Publishing, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-319-46654-5\u003c/span\u003e\u003cspan address=\"10.1007/978-3-319-46654-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Face Expressions, Face Emotion Recognition, Deep Learning, VGG-19, ResNet-50, Inception-V3, MobileNet","lastPublishedDoi":"10.21203/rs.3.rs-6794812/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6794812/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe Deep learning techniques have significantly improved face emotion identification, a crucial component of human-computer interaction. This study examines facial emotion categorization with a variety of deep learning techniques, emphasizing well-known models such as VGG-19, ResNet-50, Inception-V3, and MobileNet. We investigate the effectiveness and constraints of these neural network systems through an analysis of different approaches for recognizing facial expressions. The pre-processing techniques, indicators of performance, and datasets used to assess these frameworks are all covered in the inquiry. This evaluation demonstrates the way The VGG model-19, ResNet-50, Inception-V3, and Mobile Network perform facial recognition of emotions tasks regarding precision, computational effectiveness, and real-time applications. The purpose of this study is to provide an in-depth review of the state-of-the-art and propose future lines of study for deep learning-driven face emotion detection development.\u003c/p\u003e","manuscriptTitle":"A Survey Of Face Emotion Recognition Using Deep Learning Methods","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-03 12:16:11","doi":"10.21203/rs.3.rs-6794812/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"566e8828-404c-47d9-9156-a37ebd48dc2e","owner":[],"postedDate":"June 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-06-03T12:16:11+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-03 12:16:11","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6794812","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6794812","identity":"rs-6794812","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.