Leveraging Pāṇinian Grammar and Neural Models for Morphologically Rich Sanskrit NLP | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Leveraging Pāṇinian Grammar and Neural Models for Morphologically Rich Sanskrit NLP Yashawant Pathak, Jagdish Makhijani This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7947104/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Sanskrit’s rule-based grammatical precision and morphological richness make it a compelling foundation for developing linguistically informed Natural Language Processing (NLP) systems. Rooted in Pāṇini’s Aṣṭādhyāyī, the language encodes syntactic and semantic relations directly within its word forms, offering structural advantages over modern, statistically modeled languages. This study introduces a hybrid framework that integrates symbolic Pāṇinian grammar with neural architectures to enhance preprocessing and downstream language understanding. Specifically, sandhi splitting (euphonic decomposition) is re-engineered as an alternative to conventional stopword removal, preserving semantic integrity while improving feature granularity. The framework combines rule-based segmentation (sanskrit_parser) with data-driven sequence models (CharSS and ByT5), evaluated on the SandhiKosh and Digital Corpus of Sanskrit datasets for tasks such as word segmentation, sentiment classification, and syntactic parsing. Experimental results demonstrate that sandhi-aware preprocessing yields up to 8 percent higher F1 scores compared with conventional pipelines, confirming the synergistic potential of grammatical formalism and deep learning. Beyond advancing Sanskrit NLP, the proposed approach contributes a transferable methodology for processing other morphologically rich, low-resource languages, bridging ancient linguistic theory with modern computational intelligence. Sanskrit NLP Pāṇinian Grammar Sandhi Splitting Hybrid Rule-based and Neural Models Morphologically Rich Languages Low-Resource Language Processing Linguistically Informed Preprocessing Computational Linguistics Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7947104","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":535352055,"identity":"20eb3eb7-1d72-41c5-8f8f-d5962b31ebae","order_by":0,"name":"Yashawant Pathak","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/0lEQVRIiWNgGAWjYHACNhCRwMDDfODABxCXnTgtBkAtbIkHZ4C4zMRr4TE+zAPiE9JicPz4s8cFNX/y+HsOGBy2+bVNno+ZgfHDxxw8Ws7kmBvPOGZQLHG2IeFwbt9twzZmBmbJmdvwaDmQwybNw2aQ2HCe4cDh3J7bjEAtbMy8+LScf/5MmuefQeL884wNhy17btsT1nIjwUyat80gccPZZobDDD9uJxLUInnjjZn0zD7jxI1njjEc7G24ndzGzNiM1y9859OfSRd8k0ucdyb/84cff27bzm9vPvjhIx4tCgeQI4KxDUw24FYPBPINKHH3B6/iUTAKRsEoGKEAAGGcWNLT+Gi9AAAAAElFTkSuQmCC","orcid":"","institution":"Rustamji Institute of Technology","correspondingAuthor":true,"prefix":"","firstName":"Yashawant","middleName":"","lastName":"Pathak","suffix":""},{"id":535352056,"identity":"ff62473c-520a-4f58-aaf2-64b19825a6f6","order_by":1,"name":"Jagdish Makhijani","email":"","orcid":"","institution":"Rustamji Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"Jagdish","middleName":"","lastName":"Makhijani","suffix":""}],"badges":[],"createdAt":"2025-10-27 10:48:45","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7947104/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7947104/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":94766469,"identity":"93fb19c6-406b-454d-ba14-a6d56c7199e4","added_by":"auto","created_at":"2025-10-30 12:55:55","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2817111,"visible":true,"origin":"","legend":"","description":"","filename":"SanskritPaperwithauthornames.docx","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/17485cad71f498aeedd10043.docx"},{"id":94766457,"identity":"58c9c908-4dc9-45d9-b928-6b7d8f4963a4","added_by":"auto","created_at":"2025-10-30 12:55:54","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4467,"visible":true,"origin":"","legend":"","description":"","filename":"a9b5833ae314475ba32272ecab716298.json","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/a01b26e69ed53f1113039c9a.json"},{"id":94766468,"identity":"481fe7cf-4546-4572-99a3-cb2e93b94a6e","added_by":"auto","created_at":"2025-10-30 12:55:55","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":163365,"visible":true,"origin":"","legend":"","description":"","filename":"a9b5833ae314475ba32272ecab7162981enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/f7a066fca275f63d9bc4024d.xml"},{"id":94766459,"identity":"03a95c1c-2de1-4095-a196-bb1c688d1af3","added_by":"auto","created_at":"2025-10-30 12:55:54","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1126459,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/dd355b70eaf02d3892860a12.png"},{"id":94766458,"identity":"7a71c311-a7b7-40a0-a1df-e152112930c8","added_by":"auto","created_at":"2025-10-30 12:55:54","extension":"jpeg","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1074,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/687fa6ee84ba8447c391c7eb.jpeg"},{"id":94766466,"identity":"5c4019b0-9509-4f93-9f75-151539d8efd1","added_by":"auto","created_at":"2025-10-30 12:55:55","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1421117,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/d400071a667fc241a6ff2890.png"},{"id":94824429,"identity":"0c6f8079-ac71-49e9-861c-64f8c2af71ad","added_by":"auto","created_at":"2025-10-31 06:48:59","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":108234,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/036b1f09532e472e358f9193.png"},{"id":94766461,"identity":"812af12d-1824-42b9-a328-ac1baae4d23b","added_by":"auto","created_at":"2025-10-30 12:55:54","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":90803,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/af19b6661ae847192cc82b2b.png"},{"id":94823795,"identity":"36cf0cbf-36b9-4887-94bc-b440b52b1edb","added_by":"auto","created_at":"2025-10-31 06:47:59","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":61875,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/d97ee84256a2a9c646afd6b2.png"},{"id":94823943,"identity":"c6282867-39c4-456f-a293-d8f49b986a20","added_by":"auto","created_at":"2025-10-31 06:48:18","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":935,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/4c2eeaf6fd77b11f572e5d43.png"},{"id":94823597,"identity":"7f040461-4dc9-4b7b-a088-01321ca2c746","added_by":"auto","created_at":"2025-10-31 06:47:38","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":50712,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/e7d89a57c2fba6c29a215391.png"},{"id":94823846,"identity":"50273f40-713e-43b2-8092-0d2a8856fe69","added_by":"auto","created_at":"2025-10-31 06:48:06","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":30292,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/88923e4cca9a84cbb44fcd44.png"},{"id":94766464,"identity":"5dd1b88a-a536-4cee-840c-78ee891ac334","added_by":"auto","created_at":"2025-10-30 12:55:55","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":27619,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/7f79df5add2b29300a8b3599.png"},{"id":94766471,"identity":"4a9ea56c-c760-45d4-a70d-657c013f16c4","added_by":"auto","created_at":"2025-10-30 12:55:55","extension":"xml","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":161981,"visible":true,"origin":"","legend":"","description":"","filename":"a9b5833ae314475ba32272ecab7162981structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/ae1ba35351bb14dea07ebf58.xml"},{"id":94766470,"identity":"6c0c84de-0e70-461b-a701-c204e9ecc81d","added_by":"auto","created_at":"2025-10-30 12:55:55","extension":"html","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":191010,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1/2f95a4a3c96849e8828b42e2.html"},{"id":94827173,"identity":"2d417e97-c955-4f30-8eef-8347fa688709","added_by":"auto","created_at":"2025-10-31 06:55:22","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":956756,"visible":true,"origin":"","legend":"","description":"","filename":"SanskritPaperwithauthornames.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7947104/v1_covered_d16bacaa-c676-466a-9f5a-8e5dd71e0b83.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Leveraging Pāṇinian Grammar and Neural Models for Morphologically Rich Sanskrit NLP","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Sanskrit NLP, Pāṇinian Grammar, Sandhi Splitting, Hybrid Rule-based and Neural Models, Morphologically Rich Languages, Low-Resource Language Processing, Linguistically Informed Preprocessing, Computational Linguistics","lastPublishedDoi":"10.21203/rs.3.rs-7947104/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7947104/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSanskrit\u0026rsquo;s rule-based grammatical precision and morphological richness make it a compelling foundation for developing linguistically informed Natural Language Processing (NLP) systems. Rooted in Pāṇini\u0026rsquo;s Aṣṭādhyāyī, the language encodes syntactic and semantic relations directly within its word forms, offering structural advantages over modern, statistically modeled languages. This study introduces a hybrid framework that integrates symbolic Pāṇinian grammar with neural architectures to enhance preprocessing and downstream language understanding. Specifically, sandhi splitting (euphonic decomposition) is re-engineered as an alternative to conventional stopword removal, preserving semantic integrity while improving feature granularity. The framework combines rule-based segmentation (sanskrit_parser) with data-driven sequence models (CharSS and ByT5), evaluated on the SandhiKosh and Digital Corpus of Sanskrit datasets for tasks such as word segmentation, sentiment classification, and syntactic parsing. Experimental results demonstrate that sandhi-aware preprocessing yields up to 8 percent higher F1 scores compared with conventional pipelines, confirming the synergistic potential of grammatical formalism and deep learning. Beyond advancing Sanskrit NLP, the proposed approach contributes a transferable methodology for processing other morphologically rich, low-resource languages, bridging ancient linguistic theory with modern computational intelligence.\u003c/p\u003e","manuscriptTitle":"Leveraging Pāṇinian Grammar and Neural Models for Morphologically Rich Sanskrit NLP","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-30 12:55:50","doi":"10.21203/rs.3.rs-7947104/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2c26bbfd-9638-4279-b23e-8ee8fb76512a","owner":[],"postedDate":"October 30th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-10-30T12:55:50+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-30 12:55:50","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7947104","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7947104","identity":"rs-7947104","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.