Normalization-Aware Multi-Input Neural Architecture for Token-Level Hinglish Language Identification

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract A popular form of informal digital communication in South Asia is code-mixed text, or a mixture of languages within the same utterance. Hinglish (Hindi-English in Roman script) is particularly challenging for token-level language identification (LID) due to spelling variation, lack of standard orthography, and phonetic inconsistencies. The current methods make use of either fixed vocabularies, with prohibited generalization, or massive transformer models, which perform poorly in the noisy and low-resource environment. To overcome this, we present a normalization-aware hybrid neural model which minimizes lexical variability prior to coding. Graph-based fuzzy matching algorithm groups between variants of spelling and maps them to canonical forms. The model combines three streams of features, namely, the original token, normalized form, and a character-level BiLSTM representation. These are processed using a sequence BiLSTM with additive attention. Our model, tested on the L3Cube-HingLID dataset (approximately 1 Million tokens) has an accuracy of 97% and an English F1-score of 0.95, which is 29% lower in classification error than a token-only BiLSTM baseline. The results demonstrate the effectiveness of normalization and multi-level feature representations for code-mixed LID.
Full text 10,042 characters · extracted from preprint-html · click to expand
Normalization-Aware Multi-Input Neural Architecture for Token-Level Hinglish Language Identification | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Normalization-Aware Multi-Input Neural Architecture for Token-Level Hinglish Language Identification Satyaki Pal, Ira Nath This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9476439/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract A popular form of informal digital communication in South Asia is code-mixed text, or a mixture of languages within the same utterance. Hinglish (Hindi-English in Roman script) is particularly challenging for token-level language identification (LID) due to spelling variation, lack of standard orthography, and phonetic inconsistencies. The current methods make use of either fixed vocabularies, with prohibited generalization, or massive transformer models, which perform poorly in the noisy and low-resource environment. To overcome this, we present a normalization-aware hybrid neural model which minimizes lexical variability prior to coding. Graph-based fuzzy matching algorithm groups between variants of spelling and maps them to canonical forms. The model combines three streams of features, namely, the original token, normalized form, and a character-level BiLSTM representation. These are processed using a sequence BiLSTM with additive attention. Our model, tested on the L3Cube-HingLID dataset (approximately 1 Million tokens) has an accuracy of 97% and an English F1-score of 0.95, which is 29% lower in classification error than a token-only BiLSTM baseline. The results demonstrate the effectiveness of normalization and multi-level feature representations for code-mixed LID. Code-mixing Hinglish Language identification BiLSTM normalization Attention mechanism Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9476439","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":629161340,"identity":"e0f84573-3c5c-4d9b-a57d-e521022324d7","order_by":0,"name":"Satyaki Pal","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHElEQVRIiWNgGAWjYNCCAxJAgrn58R8eNjkw/wFe5cwwLYxtBjwyfMZgLQmEtYAYjA0SPDZyiQ0gNj4t8u3nj0n8OGORxz8jscFAIscsfX7Y4YdAW+zkdBuwazE4k8wm2XNDoljiRmLDA4Mzabkbb6cZALUkG5sdwKGFIZlNgueDRGIDUItBYs+x3I2zE0BaDiRuw6FFvv8xm+QfoJb5QC0SB//9Tzecnf4BrxaGG8ls0jw3JBI3ALVINvCwJchL5+C3xeDGY2NrmTMSxYZnHrYZM/CwGW6Qzik4kGCA2y/y/YkPb745Vpcndzz58GOgFnn52embP3yosJPDpQUIWEDxmICw9wAkWPAB5g8oWuQb8KoeBaNgFIyCEQgA5D5nesAAyC8AAAAASUVORK5CYII=","orcid":"","institution":"JIS College of Engineering","correspondingAuthor":true,"prefix":"","firstName":"Satyaki","middleName":"","lastName":"Pal","suffix":""},{"id":629161343,"identity":"b18c9d12-913e-40a1-a13a-cd322b11886e","order_by":1,"name":"Ira Nath","email":"","orcid":"","institution":"JIS College of Engineering","correspondingAuthor":false,"prefix":"","firstName":"Ira","middleName":"","lastName":"Nath","suffix":""}],"badges":[],"createdAt":"2026-04-20 21:08:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9476439/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9476439/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109426731,"identity":"6128cb8f-99b9-4814-9c5d-9c7de844f0d1","added_by":"auto","created_at":"2026-05-18 03:06:24","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":287872,"visible":true,"origin":"","legend":"","description":"","filename":"NormalizationAwareTokenLevelLIDFullLength.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9476439/v1_covered_44950f74-21df-485a-9425-58795e506b20.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Normalization-Aware Multi-Input Neural Architecture for Token-Level Hinglish Language Identification","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Code-mixing, Hinglish, Language identification, BiLSTM, normalization, Attention mechanism","lastPublishedDoi":"10.21203/rs.3.rs-9476439/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9476439/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eA popular form of informal digital communication in South Asia is code-mixed text, or a mixture of languages within the same utterance. Hinglish (Hindi-English in Roman script) is particularly challenging for token-level language identification (LID) due to spelling variation, lack of standard orthography, and phonetic inconsistencies. The current methods make use of either fixed vocabularies, with prohibited generalization, or massive transformer models, which perform poorly in the noisy and low-resource environment. To overcome this, we present a normalization-aware hybrid neural model which minimizes lexical variability prior to coding. Graph-based fuzzy matching algorithm groups between variants of spelling and maps them to canonical forms. The model combines three streams of features, namely, the original token, normalized form, and a character-level BiLSTM representation. These are processed using a sequence BiLSTM with additive attention. Our model, tested on the L3Cube-HingLID dataset (approximately 1 Million tokens) has an accuracy of 97% and an English F1-score of 0.95, which is 29% lower in classification error than a token-only BiLSTM baseline. The results demonstrate the effectiveness of normalization and multi-level feature representations for code-mixed LID.\u003c/p\u003e","manuscriptTitle":"Normalization-Aware Multi-Input Neural Architecture for Token-Level Hinglish Language Identification","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-18 03:06:17","doi":"10.21203/rs.3.rs-9476439/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b098cbdb-2f71-4888-823d-f2cc242a3e20","owner":[],"postedDate":"May 18th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-05-18T03:06:17+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-18 03:06:17","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9476439","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9476439","identity":"rs-9476439","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00