Towards Secure Social Platforms: Hate Speech Detection and Classification in Indian Languages Using Hybrid Soft Computing Techniques | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Towards Secure Social Platforms: Hate Speech Detection and Classification in Indian Languages Using Hybrid Soft Computing Techniques Purbani Kar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7085357/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The widespread adoption of high-speed internet has fueled a surge in social media usage. However, the absence of robust regulations has allowed abusive and offensive content to proliferate on these platforms. Existing research predominantly focuses on English, overlooking the rich linguistic diversity of India. The difficulties of multilingualism and code-mixing have made it more difficult to identify hate speech in Indian languages, which has led to a lack of resources. For the purpose of detecting hate speech in Indian languages, traditional and deep learning techniques have been utilized despite these obstacles. For the purpose of identifying and classifying hate speech in Indian languages, we propose a novel strategy that makes use of hybrid soft computing methods to address these difficulties. Our model comprises three key processes: gathering meaningful information, feature extraction, and prediction. Initially, we leverage BERT for code conversion and the modified Meerkat optimization (MMO) algorithm for similarity checks to discern the nature of tweets. Subsequently, we employ UNet with the multi-color shark optimization (MCSO) algorithm for feature learning, facilitating the extraction and selection of optimal features from the gathered information. Additionally, we introduce the Bayesian tensorized neural network (BTNN) for classifying hate speech in Indian languages, including Tamil, Malayalam, Kannada, Hindi, Bengali, and Marathi. To evaluate the effectiveness of our method, we utilize publicly available datasets, DravidianCodeMix, Gold-standard, L3Cube, and HASOC 2020. The simulation results shows that the UNet + BTNN model consistently outperforms other models, achieving average accuracies of 98.452%, 97.856%, 98.154%, 97.579%, 96.898% and 98.565% for Tamil, Malayalam, Kannada, Hindi, Bengali, and Marathi, respectively. hate speech prediction model Indian language code conversion similarity checking Full Text Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7085357","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":489942472,"identity":"c9f826af-22d8-4cdd-8b60-f41dea2fd7f5","order_by":0,"name":"Purbani Kar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8ElEQVRIiWNgGAWjYHACNiCWMAAS6R8+gLjsJGh5xjgDxGUmTgsDUAvjM2YeEJOQFnOJ3GcPfu6wMOaffTjtsc2vbfJ8zAyMHz7m4NZiOSPd3LD3jISZxLm0dOPcvtuGbcwMzJIzt+HWYnAjjU2Ct03ChuEMT4J0bs9tRqAWNmZeAlok/wK1yJ/h/yBt2XPbnigt0kBbzAzOMKRJM/y4nUhQi2XPM3Zj2TYJY8MzDMmGvQ23k9uYGZvx+sWcPY3t4du2OsN5ZxgSH/z4c9t2fnvzwQ8f8TkMhcfYBiYbcKvH0MLwB6/iUTAKRsEoGKEAAEPsS34f4eBHAAAAAElFTkSuQmCC","orcid":"","institution":"Techno College of Engineering Agartala","correspondingAuthor":true,"prefix":"","firstName":"Purbani","middleName":"","lastName":"Kar","suffix":""}],"badges":[],"createdAt":"2025-07-09 15:29:51","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7085357/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7085357/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":90295576,"identity":"f23d4584-efeb-4252-9967-fcba7e6d9732","added_by":"auto","created_at":"2025-09-01 08:12:00","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":547739,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7085357/v1_covered_e68addf1-b9cd-40b7-a8e6-7f3efea2b7a4.pdf"}],"financialInterests":"","formattedTitle":"Towards Secure Social Platforms: Hate Speech Detection and Classification in Indian Languages Using Hybrid Soft Computing Techniques","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"hate speech, prediction model, Indian language, code conversion, similarity checking","lastPublishedDoi":"10.21203/rs.3.rs-7085357/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7085357/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe widespread adoption of high-speed internet has fueled a surge in social media usage. However, the absence of robust regulations has allowed abusive and offensive content to proliferate on these platforms. Existing research predominantly focuses on English, overlooking the rich linguistic diversity of India. The difficulties of multilingualism and code-mixing have made it more difficult to identify hate speech in Indian languages, which has led to a lack of resources. For the purpose of detecting hate speech in Indian languages, traditional and deep learning techniques have been utilized despite these obstacles. For the purpose of identifying and classifying hate speech in Indian languages, we propose a novel strategy that makes use of hybrid soft computing methods to address these difficulties. Our model comprises three key processes: gathering meaningful information, feature extraction, and prediction. Initially, we leverage BERT for code conversion and the modified Meerkat optimization (MMO) algorithm for similarity checks to discern the nature of tweets. Subsequently, we employ UNet with the multi-color shark optimization (MCSO) algorithm for feature learning, facilitating the extraction and selection of optimal features from the gathered information. Additionally, we introduce the Bayesian tensorized neural network (BTNN) for classifying hate speech in Indian languages, including Tamil, Malayalam, Kannada, Hindi, Bengali, and Marathi. To evaluate the effectiveness of our method, we utilize publicly available datasets, DravidianCodeMix, Gold-standard, L3Cube, and HASOC 2020. The simulation results shows that the UNet\u0026thinsp;+\u0026thinsp;BTNN model consistently outperforms other models, achieving average accuracies of 98.452%, 97.856%, 98.154%, 97.579%, 96.898% and 98.565% for Tamil, Malayalam, Kannada, Hindi, Bengali, and Marathi, respectively.\u003c/p\u003e","manuscriptTitle":"Towards Secure Social Platforms: Hate Speech Detection and Classification in Indian Languages Using Hybrid Soft Computing Techniques","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-07-25 14:53:45","doi":"10.21203/rs.3.rs-7085357/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0e683892-849c-4223-8fd5-a4b4a68a7e19","owner":[],"postedDate":"July 25th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-09-01T08:03:52+00:00","versionOfRecord":[],"versionCreatedAt":"2025-07-25 14:53:45","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7085357","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7085357","identity":"rs-7085357","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.