A Unified Intelligent Information System for Clause Extraction, Risk Identification, and Consistency Analysis in Legal and Policy Documents Using Multi Model LLM Integration and Structured Knowledge Representation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Unified Intelligent Information System for Clause Extraction, Risk Identification, and Consistency Analysis in Legal and Policy Documents Using Multi Model LLM Integration and Structured Knowledge Representation Veerababu Reddy, Pravallika Bhosale, Sreeja Alle, Anees Abdul, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9278472/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Modern organizations increasingly rely on large scale legal and policy documents for compliance, governance, and regulatory decision making. However, manual analysis is time consuming, error prone, and insufficient for capturing complex semantic relationships and cross clause dependencies. While advances in large language models (LLMs) have improved automated text understanding, existing approaches treat clause extraction, risk identification, and consistency verification as separate tasks, limiting document level reasoning. Additionally, cloud based processing raises privacy concerns, and limited interpretability reduces user trust. This paper proposes a unified intelligent information system that integrates multiple LLMs with structured data storage for context aware legal document analysis. The framework introduces a multi model architecture using shared semantic representations to enable joint clause interpretation, contextual risk assessment, and cross clause consistency analysis via natural language inference. A structured storage layer supports efficient indexing and retrieval of clause level knowledge, while an explainability module provides evidence grounded reasoning to enhance transparency. The system is designed for privacy preserving offline deployment, ensuring secure processing of sensitive data. Experimental results on benchmark datasets show that the approach achieves 92.4% F1-score in clause extraction, 89.7% F1-score in risk classification, and 91.2% precision in consistency detection, outperforming traditional machine learning, transformer based models, and standalone LLM pipelines. These findings demonstrate improved document level understanding, scalability, and interpretability. DOI for code and datasets: https://doi.org/10.5281/zenodo.19239518 , GitHub repository link: https://github.com/annuu005/cogdoc.git . Intelligent Information Systems Large Language Models Legal and Policy Document Analysis Structured Data Storage Explainable Artificial Intelligence Full Text Additional Declarations No competing interests reported. Supplementary Files Demovideo.mp4 Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9278472","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":615196772,"identity":"0fe1c483-fce9-42b1-8fa6-c23bdb5e142e","order_by":0,"name":"Veerababu Reddy","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHklEQVRIie3RMWuDQBTA8SeCLgZXhWK+womQDKX4VU6Ezo4dMpwIdpHMfoxMLdkMD5JF6OrgYBFculgCKYUQehYCQkylW6H3H94dx/tNByAS/d0yIN1R8amDxHrvY4QSAib7HeEH+WHtu3nqb+qHRTmdG1QxaHB0ncKP3oJFCfpjJmFwSW6Ke9/Ot429TjtCiPdUePFtum3AyClgekkMLZ+ZTEFpVVCZcEJn+SZ2NAUBCgDUhsjLxyc7oXsmrpN05IQwvUbURJHCGD1OpIoTaaWGUT2JEchVEjtmuGz8dVJ1xPHSXRjLkyVqdu6xQSLLr+/sUN49qxSy9mi5eqTWe+2AlrVD3A+QXnp7vikGH3y596djye34jkgkEv2jvgAni21Z0IDXAQAAAABJRU5ErkJggg==","orcid":"","institution":"Vignan's Lara Institute of Technology and Science","correspondingAuthor":true,"prefix":"","firstName":"Veerababu","middleName":"","lastName":"Reddy","suffix":""},{"id":615196776,"identity":"640f27ab-9776-4d7a-a792-cb0cfa24a929","order_by":1,"name":"Pravallika Bhosale","email":"","orcid":"","institution":"Vignan's Lara Institute of Technology and Science","correspondingAuthor":false,"prefix":"","firstName":"Pravallika","middleName":"","lastName":"Bhosale","suffix":""},{"id":615196780,"identity":"7e483ce4-b2a4-4754-b2de-0a01338c5759","order_by":2,"name":"Sreeja Alle","email":"","orcid":"","institution":"Vignan's Lara Institute of Technology and Science","correspondingAuthor":false,"prefix":"","firstName":"Sreeja","middleName":"","lastName":"Alle","suffix":""},{"id":615196783,"identity":"b99624ab-0bcb-4c33-a4e2-c950b1a22964","order_by":3,"name":"Anees Abdul","email":"","orcid":"","institution":"Vignan's Lara Institute of Technology and Science","correspondingAuthor":false,"prefix":"","firstName":"Anees","middleName":"","lastName":"Abdul","suffix":""},{"id":615196785,"identity":"d4c049e5-851b-4649-aff1-ca558b514e74","order_by":4,"name":"Purna Chandu Repana","email":"","orcid":"","institution":"Vignan's Lara Institute of Technology and Science","correspondingAuthor":false,"prefix":"","firstName":"Purna","middleName":"Chandu","lastName":"Repana","suffix":""},{"id":615196787,"identity":"1f31d42d-fdc3-4e2c-800e-43493acc261f","order_by":5,"name":"Sri Vardhini Nidamanuri","email":"","orcid":"","institution":"Vignan's Lara Institute of Technology and Science","correspondingAuthor":false,"prefix":"","firstName":"Sri","middleName":"Vardhini","lastName":"Nidamanuri","suffix":""}],"badges":[],"createdAt":"2026-03-31 10:38:07","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9278472/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9278472/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106430347,"identity":"0e96bc9e-5902-49f6-a297-84600fd5949b","added_by":"auto","created_at":"2026-04-08 12:58:32","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1022616,"visible":true,"origin":"","legend":"","description":"","filename":"UnifiedIntelligentInformationSystem.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9278472/v1_covered_84f4e74e-fc28-4daa-8efa-acb9f69ebc27.pdf"},{"id":106002859,"identity":"e1327dc6-023f-42e4-bf98-f848c3e75188","added_by":"auto","created_at":"2026-04-02 10:20:57","extension":"mp4","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":17174180,"visible":true,"origin":"","legend":"","description":"","filename":"Demovideo.mp4","url":"https://assets-eu.researchsquare.com/files/rs-9278472/v1/fd4f84a9d82f3b728f74aed7.mp4"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Unified Intelligent Information System for Clause Extraction, Risk Identification, and Consistency Analysis in Legal and Policy Documents Using Multi Model LLM Integration and Structured Knowledge Representation","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Intelligent Information Systems, Large Language Models, Legal and Policy Document Analysis, Structured Data Storage, Explainable Artificial Intelligence","lastPublishedDoi":"10.21203/rs.3.rs-9278472/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9278472/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Modern organizations increasingly rely on large scale legal and policy documents for compliance, governance, and regulatory decision making. However, manual analysis is time consuming, error prone, and insufficient for capturing complex semantic relationships and cross clause dependencies. While advances in large language models (LLMs) have improved automated text understanding, existing approaches treat clause extraction, risk identification, and consistency verification as separate tasks, limiting document level reasoning. Additionally, cloud based processing raises privacy concerns, and limited interpretability reduces user trust. This paper proposes a unified intelligent information system that integrates multiple LLMs with structured data storage for context aware legal document analysis. The framework introduces a multi model architecture using shared semantic representations to enable joint clause interpretation, contextual risk assessment, and cross clause consistency analysis via natural language inference. A structured storage layer supports efficient indexing and retrieval of clause level knowledge, while an explainability module provides evidence grounded reasoning to enhance transparency. The system is designed for privacy preserving offline deployment, ensuring secure processing of sensitive data. Experimental results on benchmark datasets show that the approach achieves 92.4% F1-score in clause extraction, 89.7% F1-score in risk classification, and 91.2% precision in consistency detection, outperforming traditional machine learning, transformer based models, and standalone LLM pipelines. These findings demonstrate improved document level understanding, scalability, and interpretability. DOI for code and datasets: https://doi.org/10.5281/zenodo.19239518, GitHub repository link: https://github.com/annuu005/cogdoc.git.","manuscriptTitle":"A Unified Intelligent Information System for Clause Extraction, Risk Identification, and Consistency Analysis in Legal and Policy Documents Using Multi Model LLM Integration and Structured Knowledge Representation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-02 10:20:53","doi":"10.21203/rs.3.rs-9278472/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b3988f64-c7aa-4c86-bd18-1685f20b5173","owner":[],"postedDate":"April 2nd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-08T12:57:40+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-02 10:20:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9278472","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9278472","identity":"rs-9278472","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.