Analysis: Serving Individuals with Language Impairments using Automatic Speech Recognition Models and Large Language Models: Challenges and Opportunities

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Large language models (LLMs) have attracted much attention for healthcare applications, demonstrating strong potential in automating conversational interactions. However, cloud-hosted LLMs pose major data privacy concerns when processing Protected Health Information. Moreover, current LLM-based systems rely on text input/output, creating substantial barriers for users, such as children and older adults, who may have difficulty typing. To mitigate these challenges, there has been growing interest in developing edge device-based, voice-enabled LLM systems. Running LLMs on edge devices minimizes the risks of PHI leaking to the cloud, while automatic speech recognition (ASR) eliminates the need for text-based inputs. Despite these advantages, existing ASRs convert speech into word-by-word text, which often contains disfluencies and fillers (e.g., ”um”, ”hum”) and grammatical errors, especially for individuals with language impairments. This noisy input can significantly degrade the performance of LLMs, yet this chained issue remains under-explored in healthcare applications. To address this critical gap, we conducted a systematic analysis through comparison studies and ablation experiments to identify key factors affecting the performance of edge-based ASR-LLM systems when used by individuals with language impairments. Furthermore, we proposed an evaluation framework for speech-enabled AI healthcare to emphasize both interpretability and robustness, paving the way for more inclusive and secure conversational healthcare solutions.
Full text 14,177 characters · extracted from preprint-html · click to expand
Analysis: Serving Individuals with Language Impairments using Automatic Speech Recognition Models and Large Language Models: Challenges and Opportunities | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Analysis: Serving Individuals with Language Impairments using Automatic Speech Recognition Models and Large Language Models: Challenges and Opportunities Yiyu Shi, Ruiyang Qin, Haoxinran Yu, Lixuan Wei, Yuxuan Liu, Dancheng Liu, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7069967/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Large language models (LLMs) have attracted much attention for healthcare applications, demonstrating strong potential in automating conversational interactions. However, cloud-hosted LLMs pose major data privacy concerns when processing Protected Health Information. Moreover, current LLM-based systems rely on text input/output, creating substantial barriers for users, such as children and older adults, who may have difficulty typing. To mitigate these challenges, there has been growing interest in developing edge device-based, voice-enabled LLM systems. Running LLMs on edge devices minimizes the risks of PHI leaking to the cloud, while automatic speech recognition (ASR) eliminates the need for text-based inputs. Despite these advantages, existing ASRs convert speech into word-by-word text, which often contains disfluencies and fillers (e.g., ”um”, ”hum”) and grammatical errors, especially for individuals with language impairments. This noisy input can significantly degrade the performance of LLMs, yet this chained issue remains under-explored in healthcare applications. To address this critical gap, we conducted a systematic analysis through comparison studies and ablation experiments to identify key factors affecting the performance of edge-based ASR-LLM systems when used by individuals with language impairments. Furthermore, we proposed an evaluation framework for speech-enabled AI healthcare to emphasize both interpretability and robustness, paving the way for more inclusive and secure conversational healthcare solutions. Health sciences/Health care/Medical ethics Physical sciences/Engineering Full Text Additional Declarations There is NO Competing Interest. This is determined to be not a human subject study by our IRB as it is based on public domain dataset with no identifiable information. Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7069967","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":482212192,"identity":"045f1068-f150-41fd-a98d-ae21ea482af6","order_by":0,"name":"Yiyu Shi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAApUlEQVRIiWNgGAWjYFACHoYDEhUMjA0MDAYkaLE4Q6oWhso2UrQYnD978MDNeXdkG9ibt0kQp+VGXsLBmdueGTfwHCsjVguPwWHJbYcTGyRyzIjUcv6MweG/c4Ba5N8Qq+VAjsEByQaQLTxEapG8AdQiceywcRtPWrEFUVr4zp8x/iBRc1i2n/3wxhtEaVE4AGWwEaUcBOQbiFY6CkbBKBgFIxYAAB45NMUJFFFHAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-6788-9823","institution":"University of Notre Dame","correspondingAuthor":true,"prefix":"","firstName":"Yiyu","middleName":"","lastName":"Shi","suffix":""},{"id":482212193,"identity":"11c9695c-fee7-40a7-b1e1-486b658569b0","order_by":1,"name":"Ruiyang Qin","email":"","orcid":"","institution":"Villanova University","correspondingAuthor":false,"prefix":"","firstName":"Ruiyang","middleName":"","lastName":"Qin","suffix":""},{"id":482212194,"identity":"28566407-c1d0-4077-bcc0-c5479f9418ad","order_by":2,"name":"Haoxinran Yu","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Haoxinran","middleName":"","lastName":"Yu","suffix":""},{"id":482212195,"identity":"8f1da50f-27cc-4ab1-87a3-ca35e93f22f4","order_by":3,"name":"Lixuan Wei","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Lixuan","middleName":"","lastName":"Wei","suffix":""},{"id":482212196,"identity":"a6d6ef1b-dd32-45c6-b52b-041a751fbc38","order_by":4,"name":"Yuxuan Liu","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Yuxuan","middleName":"","lastName":"Liu","suffix":""},{"id":482212197,"identity":"ee7c18a0-45c4-4ba5-bc1e-cf0e14ba3b5d","order_by":5,"name":"Dancheng Liu","email":"","orcid":"https://orcid.org/0009-0003-6107-5612","institution":"State University of New York at Buffalo","correspondingAuthor":false,"prefix":"","firstName":"Dancheng","middleName":"","lastName":"Liu","suffix":""},{"id":482212198,"identity":"3c1b831c-0a3d-468e-abeb-1accce3513de","order_by":6,"name":"Chenhui Xu","email":"","orcid":"","institution":"State University of New York at Buffalo","correspondingAuthor":false,"prefix":"","firstName":"Chenhui","middleName":"","lastName":"Xu","suffix":""},{"id":482212199,"identity":"3a15523c-538c-4130-a419-f15cc7bd1dde","order_by":7,"name":"Jiajie Li","email":"","orcid":"","institution":"State University of New York at Buffalo","correspondingAuthor":false,"prefix":"","firstName":"Jiajie","middleName":"","lastName":"Li","suffix":""},{"id":482212200,"identity":"fd1fa5a0-a748-400b-9580-2aaf42fd73f3","order_by":8,"name":"Gelei Xu","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Gelei","middleName":"","lastName":"Xu","suffix":""},{"id":482212201,"identity":"74b15ff6-c79f-409d-a845-df16ccc05ca1","order_by":9,"name":"Ahmed Abbasi","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Ahmed","middleName":"","lastName":"Abbasi","suffix":""},{"id":482212202,"identity":"8ab13e5f-21f0-4867-a1e2-94e6d8e36bb2","order_by":10,"name":"Jinjun Xiong","email":"","orcid":"","institution":"State University of New York at Buffalo","correspondingAuthor":false,"prefix":"","firstName":"Jinjun","middleName":"","lastName":"Xiong","suffix":""},{"id":482212203,"identity":"a59c0efe-ae45-42ce-9bcd-519b0b186dc0","order_by":11,"name":"Xiufan Yu","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Xiufan","middleName":"","lastName":"Yu","suffix":""},{"id":482212204,"identity":"309ebe27-ef94-46e6-acd3-0ff70d97f290","order_by":12,"name":"Zhi Zheng","email":"","orcid":"","institution":"University of Notre Dame","correspondingAuthor":false,"prefix":"","firstName":"Zhi","middleName":"","lastName":"Zheng","suffix":""}],"badges":[],"createdAt":"2025-07-08 03:10:07","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7069967/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7069967/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":87453083,"identity":"76d90af6-a9d3-438b-bdc8-a918f9fcd7dc","added_by":"auto","created_at":"2025-07-24 03:26:14","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":8680280,"visible":true,"origin":"","legend":"Article File","description":"","filename":"NMEDASRLLMFinal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7069967/v1_covered_f863504b-0c7a-4e0a-93ba-8ff278bbe12b.pdf"}],"financialInterests":"\u003cp\u003eThere is \u003cstrong\u003eNO\u003c/strong\u003e Competing Interest.\u003c/p\u003e\n\u003cp\u003eThis is determined to be not a human subject study by our IRB as it is based on public domain dataset with no identifiable information.\u003c/p\u003e","formattedTitle":"Analysis: Serving Individuals with Language Impairments using Automatic Speech Recognition Models and\nLarge Language Models: Challenges and Opportunities","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-7069967/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7069967/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Large language models (LLMs) have attracted much attention for healthcare applications, demonstrating strong potential in automating conversational interactions. However, cloud-hosted LLMs pose major data privacy concerns when processing Protected Health Information. Moreover, current LLM-based systems rely on text input/output, creating substantial barriers for users, such as children and older adults, who may have difficulty typing. To mitigate these challenges, there has been growing interest in developing edge device-based, voice-enabled LLM systems. Running LLMs on edge devices minimizes the risks of PHI leaking to the cloud, while automatic speech recognition (ASR) eliminates the need for text-based inputs. Despite these advantages, existing ASRs convert speech into word-by-word text, which often contains disfluencies and fillers (e.g., ”um”, ”hum”) and grammatical errors, especially for individuals with language impairments. This noisy input can significantly degrade the performance of LLMs, yet this chained issue remains under-explored in healthcare applications. To address this critical gap, we conducted a systematic analysis through comparison studies and ablation experiments to identify key factors affecting the performance of edge-based ASR-LLM systems when used by individuals with language impairments. Furthermore, we proposed an evaluation framework for speech-enabled AI healthcare to emphasize both interpretability and robustness, paving the way for\r\nmore inclusive and secure conversational healthcare solutions.","manuscriptTitle":"Analysis: Serving Individuals with Language Impairments using Automatic Speech Recognition Models and\nLarge Language Models: Challenges and Opportunities","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-07-24 03:18:05","doi":"10.21203/rs.3.rs-7069967/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"0373ee96-a2c6-4113-971a-f31b5024c386","owner":[],"postedDate":"July 24th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":51199705,"name":"Health sciences/Health care/Medical ethics"},{"id":51199706,"name":"Physical sciences/Engineering"}],"tags":[],"updatedAt":"2025-07-24T03:18:05+00:00","versionOfRecord":[],"versionCreatedAt":"2025-07-24 03:18:05","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7069967","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7069967","identity":"rs-7069967","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-04T02:00:05.705006+00:00
License: CC-BY-4.0