Speaker-Aware Emotion Recognition in Dialogues via SemGloVe- BERT and Graph Attention Networks

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Emotion recognition in dialogues play a critical role in understanding human communications, with impactful applications in domains such as customer support, virtual assistants, and mental health monitoring. Traditional deep learning (DL) approaches like LSTM, Bi-GRU, and BiLSTM have shown promise in identifying emotions from text. However, these models often fail to capture long-range contextual dependencies and speaker-level interactions inherent in multi-turn conversations, leading to suboptimal performance. To address these limitations, this study proposes a novel framework that integrates DialogueGCN (Graph Convolutional Networks for Dialogue) with Graph Attention Networks (GAT) for improved emotion recognition. DialogueGCN is specifically designed to model both temporal dynamics and speaker-specific dependencies within a dialogue using graph structures. GAT enhances this representation by assigning varying levels of attention to different nodes, thereby emphasizing more relevant speaker interactions. The proposed model was implemented using Python and evaluated on the Daily Dialog dataset. The architecture outperforms conventional models significantly, achieving 93% accuracy. Compared to existing methods—LSTM (85%), Bi-GRU (87%), BiLSTM (89%), and BERT (75%)—the proposed DialogueGCN+GAT model also demonstrated superior accuracy(93%), precision (94%), recall (99%), and F1-score (97%). These findings validate the strength of graph-based approaches in emotion recognition tasks, particularly in handling complex dialogue structures. The results suggest that the proposed model offers a more context-aware and speaker-sensitive solution, making it highly effective for real-world dialogue systems. Future work aims to extend this model to other multimodal datasets to further evaluate its generalizability and performance.
Full text 12,182 characters · extracted from preprint-html · click to expand
Speaker-Aware Emotion Recognition in Dialogues via SemGloVe- BERT and Graph Attention Networks | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Speaker-Aware Emotion Recognition in Dialogues via SemGloVe- BERT and Graph Attention Networks Sakunthala Prabha K S, Suguna Marappan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6876811/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Emotion recognition in dialogues play a critical role in understanding human communications, with impactful applications in domains such as customer support, virtual assistants, and mental health monitoring. Traditional deep learning (DL) approaches like LSTM, Bi-GRU, and BiLSTM have shown promise in identifying emotions from text. However, these models often fail to capture long-range contextual dependencies and speaker-level interactions inherent in multi-turn conversations, leading to suboptimal performance. To address these limitations, this study proposes a novel framework that integrates DialogueGCN (Graph Convolutional Networks for Dialogue) with Graph Attention Networks (GAT) for improved emotion recognition. DialogueGCN is specifically designed to model both temporal dynamics and speaker-specific dependencies within a dialogue using graph structures. GAT enhances this representation by assigning varying levels of attention to different nodes, thereby emphasizing more relevant speaker interactions. The proposed model was implemented using Python and evaluated on the Daily Dialog dataset. The architecture outperforms conventional models significantly, achieving 93% accuracy. Compared to existing methods—LSTM (85%), Bi-GRU (87%), BiLSTM (89%), and BERT (75%)—the proposed DialogueGCN+GAT model also demonstrated superior accuracy(93%), precision (94%), recall (99%), and F1-score (97%). These findings validate the strength of graph-based approaches in emotion recognition tasks, particularly in handling complex dialogue structures. The results suggest that the proposed model offers a more context-aware and speaker-sensitive solution, making it highly effective for real-world dialogue systems. Future work aims to extend this model to other multimodal datasets to further evaluate its generalizability and performance. Physical sciences/Energy science and technology Physical sciences/Engineering DialogueGCN Graph Attention Networks Emotion Recognition Speaker-Level Variations Dialogue Dependencies Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6876811","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":517397040,"identity":"664efe20-effa-445a-8fc3-f646476378ac","order_by":0,"name":"Sakunthala Prabha K S","email":"","orcid":"","institution":"Vellore Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"Sakunthala","middleName":"Prabha K","lastName":"S","suffix":""},{"id":517397041,"identity":"95fc2f3c-d0fa-4817-a750-e6d301091df9","order_by":1,"name":"Suguna Marappan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAz0lEQVRIiWNgGAWjYDADfgbGBjCDjWgtkg2MjQ2kaTE4ALOGEJCfkX5NuqCGIdr4RnL7A4YaOwY+aQI6DW7klEnPOMaQu+1GItBhx5IZ2GQOENAikZMmzcMG1HLmIFAL2wEGNokEQg4DafnHkLu5B6TlHxFaGG6kH5PmbWPI3cDe2NjA2EaEFoMzb5itefskcmccb2yckdiXzEPYYe3pD2/zfLPJ7W9mf/Dhwzc7OfkZhBzGwGMAJCQgbKBiHkLqgYD9ARGKRsEoGAWjYEQDAHMIPixA7KfUAAAAAElFTkSuQmCC","orcid":"","institution":"Vellore Institute of Technology","correspondingAuthor":true,"prefix":"","firstName":"Suguna","middleName":"","lastName":"Marappan","suffix":""}],"badges":[],"createdAt":"2025-06-12 05:53:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6876811/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6876811/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":91949098,"identity":"ed693010-4f57-4bde-b25f-911d599e25b1","added_by":"auto","created_at":"2025-09-23 06:13:25","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1453406,"visible":true,"origin":"","legend":"","description":"","filename":"SpeakerAwareEmotionRecognitioninDialoguesfinalversion1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6876811/v1/3d43f4be44b7466afd914641.pdf"},{"id":91949097,"identity":"337466f6-0c8a-459f-a7db-0c1f98b5dc1a","added_by":"auto","created_at":"2025-09-23 06:13:25","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5500,"visible":true,"origin":"","legend":"","description":"","filename":"4e7dda87931042799ca54f93eae5dfdd.json","url":"https://assets-eu.researchsquare.com/files/rs-6876811/v1/309e6ec39da9bd68e602f74a.json"},{"id":98214687,"identity":"ae5449fe-6d18-4d2e-a857-c595b0a71e18","added_by":"auto","created_at":"2025-12-15 10:10:38","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1061437,"visible":true,"origin":"","legend":"","description":"","filename":"SpeakerAwareEmotionRecognitioninDialoguesfinalversion1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6876811/v1_covered_b2849970-7a5a-4aa9-b278-f16e0cdc5cfd.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Speaker-Aware Emotion Recognition in Dialogues via SemGloVe- BERT and Graph Attention Networks","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"DialogueGCN, Graph Attention Networks, Emotion Recognition, Speaker-Level Variations, Dialogue Dependencies","lastPublishedDoi":"10.21203/rs.3.rs-6876811/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6876811/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Emotion recognition in dialogues play a critical role in understanding human communications, with impactful applications in domains such as customer support, virtual assistants, and mental health monitoring. Traditional deep learning (DL) approaches like LSTM, Bi-GRU, and BiLSTM have shown promise in identifying emotions from text. However, these models often fail to capture long-range contextual dependencies and speaker-level interactions inherent in multi-turn conversations, leading to suboptimal performance. To address these limitations, this study proposes a novel framework that integrates DialogueGCN (Graph Convolutional Networks for Dialogue) with Graph Attention Networks (GAT) for improved emotion recognition. DialogueGCN is specifically designed to model both temporal dynamics and speaker-specific dependencies within a dialogue using graph structures. GAT enhances this representation by assigning varying levels of attention to different nodes, thereby emphasizing more relevant speaker interactions. The proposed model was implemented using Python and evaluated on the Daily Dialog dataset. The architecture outperforms conventional models significantly, achieving 93% accuracy. Compared to existing methods—LSTM (85%), Bi-GRU (87%), BiLSTM (89%), and BERT (75%)—the proposed DialogueGCN+GAT model also demonstrated superior accuracy(93%), precision (94%), recall (99%), and F1-score (97%). These findings validate the strength of graph-based approaches in emotion recognition tasks, particularly in handling complex dialogue structures. The results suggest that the proposed model offers a more context-aware and speaker-sensitive solution, making it highly effective for real-world dialogue systems. Future work aims to extend this model to other multimodal datasets to further evaluate its generalizability and performance.","manuscriptTitle":"Speaker-Aware Emotion Recognition in Dialogues via SemGloVe- BERT and Graph Attention Networks","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-23 06:13:21","doi":"10.21203/rs.3.rs-6876811/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"650cd23b-022e-447d-b3ad-639812ee568a","owner":[],"postedDate":"September 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":54983180,"name":"Physical sciences/Energy science and technology"},{"id":54983181,"name":"Physical sciences/Engineering"}],"tags":[],"updatedAt":"2025-12-15T10:10:14+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-23 06:13:21","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6876811","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6876811","identity":"rs-6876811","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00