Integration of Audio video Speech Recognition using LSTM and Feed Forward Convolutional Neural Network

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract In the current scenario, audio visual speech recognition is one of the emerging fields of research, but there is still deficiency of appropriate visual features for recognition of visual speech. Human lip-readers are increasingly being presented as useful in the gathering of forensic evidence but, like all human, suffer from unreliability in analyzing the lip movement. Here we used a custom dataset and design the system in such a way that it predicts the output for the lip reading. The problem of speaker independent lip-reading is very demanding due to unpredictable variations between people. Also due to recent developments and advances in the fields of signal processing and computer vision. The task of automating the lip reading is becoming a field of great interest. Here we use MFCC techniques for audio processing and LSTM method for visual speech recognition and finally integrate the audio and video using feed forward neural network (FFNN) and also got good accuracy. That is why the AVSR technique attract a great attention as a reliable solution for the speech detection problem. The final model was capable of taking more appropriate decision while predicting the spoken word. We were able to get a good accuracy of about 92.38% for the final model.
Full text 16,251 characters · extracted from preprint-html · click to expand
Integration of Audio video Speech Recognition using LSTM and Feed Forward Convolutional Neural Network | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Integration of Audio video Speech Recognition using LSTM and Feed Forward Convolutional Neural Network Shashidhar R, Sudarshan Patil Kulkarni This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-173380/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract In the current scenario, audio visual speech recognition is one of the emerging fields of research, but there is still deficiency of appropriate visual features for recognition of visual speech. Human lip-readers are increasingly being presented as useful in the gathering of forensic evidence but, like all human, suffer from unreliability in analyzing the lip movement. Here we used a custom dataset and design the system in such a way that it predicts the output for the lip reading. The problem of speaker independent lip-reading is very demanding due to unpredictable variations between people. Also due to recent developments and advances in the fields of signal processing and computer vision. The task of automating the lip reading is becoming a field of great interest. Here we use MFCC techniques for audio processing and LSTM method for visual speech recognition and finally integrate the audio and video using feed forward neural network (FFNN) and also got good accuracy. That is why the AVSR technique attract a great attention as a reliable solution for the speech detection problem. The final model was capable of taking more appropriate decision while predicting the spoken word. We were able to get a good accuracy of about 92.38% for the final model. Electrical Engineering Audio visual speech recognition lip reading FFNN LSTM Deep neural network Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Full Text Due to technical limitations, full-text HTML conversion of this manuscript could not be completed. However, the latest manuscript can be downloaded and accessed as a PDF. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-173380","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":17925327,"identity":"4101dac1-2419-4a56-9955-cd6844922a13","order_by":0,"name":"Shashidhar R","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+ElEQVRIiWNgGAWjYDACCQYGZgYD5gQw5wMDG1TYgEgtjDOI18IA0cLMQ4y75Gc3P/5cUGCdxz/7ANtnmz988vwMzA8/MBTcwanF4M4xM+kZBunFEucSmGfntrEZzmxgM5ZgMHiGW4tEghkzj8HhxIYzDMzMuQ1sCQYHGMyA4odxO2xG+ufPIC3zQVos/rAl2B9g/4ZXC8ONHANpkJYNIC0MbEBbGHjw22JwI6cMqCU9ceMZxmbGXqBfZhzmKZZIwO+wzZ95/lgnzjvDfJjhx59j8vzt7Rs/fPiDx2EIwNgAJI6Bo4khgRgNUFBDgtpRMApGwSgYKQAAhd9LAlTTd7oAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-3737-7819","institution":"JSS Science and Technology University,Sri Jayachamarajendra College of Engineering","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Shashidhar","middleName":"","lastName":"R","suffix":""},{"id":17925328,"identity":"3d54c9cc-f617-43df-a42b-c16e2585c423","order_by":1,"name":"Sudarshan Patil Kulkarni","email":"","orcid":"","institution":"JSS Science and Technology University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sudarshan","middleName":"Patil","lastName":"Kulkarni","suffix":""}],"badges":[],"createdAt":"2021-01-28 11:11:07","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-173380/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-173380/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":7364729,"identity":"52cedbe6-ee63-4546-b4e6-f0dc832b6db4","added_by":"auto","created_at":"2021-03-25 20:11:05","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":150415,"visible":true,"origin":"","legend":"Preprocessing Step in single Frame.","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/f8c7c7bc1ed2861d0977a33d.png"},{"id":7364431,"identity":"bf52690e-99ce-420e-84ac-047f7b3cca97","added_by":"auto","created_at":"2021-03-25 20:08:05","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":62051,"visible":true,"origin":"","legend":"Block Diagram of AVSR","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/7c16906a5489e1d199f7de97.png"},{"id":7364432,"identity":"67e39629-6dd8-464b-b693-354f7381b712","added_by":"auto","created_at":"2021-03-25 20:08:05","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":74586,"visible":true,"origin":"","legend":"Complete Diagram of the Simple Recurrent Network unit (a) and a Long Short-Term Memory block (b) as used in the hidden layers of a recurrent neural network","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/ca32c25c7d3a824e900b0a1e.png"},{"id":7364425,"identity":"5e3e3e39-6d9c-4d33-87bc-1cf565690274","added_by":"auto","created_at":"2021-03-25 20:08:05","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":325654,"visible":true,"origin":"","legend":" Feed Forward Neural Architecture of Audio Model\n\n","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/6ab4da64e79678886c7e5488.png"},{"id":7364728,"identity":"abd1bc52-9ccf-4783-b326-b8c78045cf55","added_by":"auto","created_at":"2021-03-25 20:11:05","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":252101,"visible":true,"origin":"","legend":"Confusion matrixes for audio model","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/4ea109ddb99da6b0b01c7c1d.png"},{"id":7364915,"identity":"4eeff8e4-f355-4289-9386-07bb4c2f6023","added_by":"auto","created_at":"2021-03-25 20:17:05","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":128355,"visible":true,"origin":"","legend":"Confusion matrixes for video model\n\n","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/f08215ec1be65cb3f10b6d74.png"},{"id":7364428,"identity":"06999ebb-aa0a-4879-92fa-ddd358848d00","added_by":"auto","created_at":"2021-03-25 20:08:05","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":55421,"visible":true,"origin":"","legend":"Accuracy curve of combined Model","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/c5edf32d79474c5783c29a09.png"},{"id":7364734,"identity":"4fd97a8c-36ee-4905-9517-b89897f7cd31","added_by":"auto","created_at":"2021-03-25 20:11:05","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":48986,"visible":true,"origin":"","legend":"Loss curve of Combined Model","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/e707817878e32df610262342.png"},{"id":7364731,"identity":"98dcc66a-4fbe-4495-a4dc-3da7771ba41c","added_by":"auto","created_at":"2021-03-25 20:11:05","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":97638,"visible":true,"origin":"","legend":"Accuracy curve for LSTM model","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/7c4bdba36fa70f8be3467ab5.png"},{"id":7364842,"identity":"68e1a6a1-3239-448b-8bbc-96653068d918","added_by":"auto","created_at":"2021-03-25 20:14:05","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":109224,"visible":true,"origin":"","legend":"Loss curve for LSTM model","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/cfbc268f8cc458144090dc37.png"},{"id":7364840,"identity":"7853091b-7cc1-4021-8333-652c54a88cd5","added_by":"auto","created_at":"2021-03-25 20:14:05","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":197190,"visible":true,"origin":"","legend":"Confusion matrix of combined model","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1/2e71c7a3c094bdabaafce6a9.png"},{"id":13613247,"identity":"36bf5415-b4cd-441b-acd0-a9527a89f55c","added_by":"auto","created_at":"2021-09-17 06:36:22","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":856152,"visible":true,"origin":"","legend":"","description":"","filename":"ManscriptFinalWPC.pdf","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1_covered.pdf"},{"id":11904089,"identity":"d12d46d3-8f4a-4302-8264-59ec6d328971","added_by":"auto","created_at":"2021-07-29 03:12:52","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":852523,"visible":true,"origin":"","legend":"","description":"","filename":"ManscriptFinalWPC.pdf","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1_covered.pdf"},{"id":7364959,"identity":"0d86fa22-1e81-4b58-9f79-3d9ef73b55b7","added_by":"auto","created_at":"2021-03-25 20:20:09","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":990305,"visible":true,"origin":"","legend":"","description":"","filename":"ManscriptFinalWPC.pdf","url":"https://assets-eu.researchsquare.com/files/rs-173380/v1_stamped.pdf"}],"financialInterests":"","formattedTitle":"Integration of Audio video Speech Recognition using LSTM and Feed Forward Convolutional Neural Network","fulltext":[{"header":"Full Text","content":"Due to technical limitations, full-text HTML conversion of this manuscript could not be completed. However, the latest manuscript can be downloaded and \u003ca href='/article/rs-173380/latest.pdf' target='_blank'\u003e accessed as a PDF.\u003c/a\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Audio visual speech recognition, lip reading, FFNN, LSTM, Deep neural network","lastPublishedDoi":"10.21203/rs.3.rs-173380/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-173380/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn the current scenario, audio visual speech recognition is one of the emerging fields of research, but there is still deficiency of appropriate visual features for recognition of visual speech. Human lip-readers are increasingly being presented as useful in the gathering of forensic evidence but, like all human, suffer from unreliability in analyzing the lip movement. Here we used a custom dataset and design the system in such a way that it predicts the output for the lip reading. The problem of speaker independent lip-reading is very demanding due to unpredictable variations between people. Also due to recent developments and advances in the fields of signal processing and computer vision. The task of automating the lip reading is becoming a field of great interest. Here we use MFCC techniques for audio processing and LSTM method for visual speech recognition and finally integrate the audio and video using feed forward neural network (FFNN) and also got good accuracy. That is why the AVSR technique attract a great attention as a reliable solution for the speech detection problem. The final model was capable of taking more appropriate decision while predicting the spoken word. We were able to get a good accuracy of about 92.38% for the final model.\u003c/p\u003e","manuscriptTitle":"Integration of Audio video Speech Recognition using LSTM and Feed Forward Convolutional Neural Network","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-03-25 20:08:02","doi":"10.21203/rs.3.rs-173380/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7bf7889b-784b-442c-8d3a-03b3141bac06","owner":[],"postedDate":"March 25th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":3171272,"name":"Electrical Engineering"}],"tags":[],"updatedAt":"2021-07-29T03:12:38+00:00","versionOfRecord":[],"versionCreatedAt":"2021-03-25 20:08:02","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-173380","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-173380","identity":"rs-173380","version":["v1"]},"buildId":"ApUGefWb6u5IBVtyqm6d5","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0