AutoClusRE: An automatic clustering-based method for relation extraction and knowledge graph construction | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article AutoClusRE: An automatic clustering-based method for relation extraction and knowledge graph construction TongYao Wang, Donghui Shi, Jose Aguilar, Jozef Zurada This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7705090/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The knowledge graph plays a vital role in integrating and representing information. As a core component in the construction of a knowledge graph, relation extraction is responsible for identifying and establishing semantic relationships between entities. However, existing methods rely heavily on annotated data. When dealing with unlabeled text, they fail to automatically perform relation extraction and construct a knowledge graph. This poses a significant challenge for knowledge graph construction. To address the above challenges, this paper proposes a domain-adaptive entity-relation extraction and knowledge graph construction framework, named AutoClusRE (Automatic Clustering-based Relation Extraction), which is based on a large language model and designed for extracting relational triples from unlabeled text. This method requires no additional training and can directly perform adaptive entity–relation extraction and knowledge graph construction on unlabeled text. Specifically, AutoClusRE constructs an entity type table and a relation type table, which are embedded into the prompt templates of the large language model and dynamically updates them during the extraction process to serve as contextual guidance for the large language model. In addition, a semantics-based clustering algorithm is introduced to effectively merge duplicates and control excessive expansion of the type tables. Experiments conducted on the public dataset FOBIE and the custom dataset HSA (Huizhou-style architecture) demonstrated that AutoClusRE achieved competitive extraction performance without relying on any annotated data. On the public FOBIE dataset, AutoClusRE outperformed traditional deep learning models by 14.8% and DeepSeek by 9.5% in terms of F1 score. On the custom dataset HSA, AutoClusRE’s F1 score improved by 17.2% over traditional deep learning models and by 2% over DeepSeek. This study not only introduces a novel relation extraction method for information management in unknown domains but also constructs a knowledge graph in the field of Huizhou-style architecture based on AutoClusRE, contributing to the digital research and application of domain knowledge in this field. Relation Extraction Semantic Clustering Knowledge Graph Large Language Model Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7705090","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":538304893,"identity":"322a8779-0d78-41d0-a5f0-c20c5941ac4d","order_by":0,"name":"TongYao Wang","email":"","orcid":"","institution":"Anhui Jianzhu University","correspondingAuthor":false,"prefix":"","firstName":"TongYao","middleName":"","lastName":"Wang","suffix":""},{"id":538304895,"identity":"d8820da0-6d65-4bb0-ac76-c7597c8112e3","order_by":1,"name":"Donghui Shi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/UlEQVRIiWNgGAWjYDACCcYGCSCVwMDMfICZgQcsZkCsFrYEYrWAEVALA48BM1QMvxb52c2NNz7uqM0zOM7z8XOBTF1iA3vzNgmGmjs4tRjcOdhsOfPM8WKDw7ybpWfwsCU28Bwrk2A49gy3FonENmnetmOJGw7zbpDm4eFJbJDIMQN68DBuh82Aa+F5/JuHRyKxQf4Nfi0MN8BaakBa2IC2GABt4cGvxeBGItAvbQcSZx5mM7Pm4UkwbuNJK7ZIOIbPYekPb3xsq0vsO3/48W3enjrZfvbDG298qMHjMAiAKmDsYWBgAzESCGlgYKiD0j8IKx0Fo2AUjIKRBwBqmFSkB3nKOQAAAABJRU5ErkJggg==","orcid":"","institution":"Anhui Jianzhu University","correspondingAuthor":true,"prefix":"","firstName":"Donghui","middleName":"","lastName":"Shi","suffix":""},{"id":538304896,"identity":"8459c2dc-761e-468a-9193-63cae01a334f","order_by":2,"name":"Jose Aguilar","email":"","orcid":"","institution":"IMDEA Networks","correspondingAuthor":false,"prefix":"","firstName":"Jose","middleName":"","lastName":"Aguilar","suffix":""},{"id":538304897,"identity":"4e965084-7aaa-4f60-9bfc-b633bb69544b","order_by":3,"name":"Jozef Zurada","email":"","orcid":"","institution":"University of Louisville","correspondingAuthor":false,"prefix":"","firstName":"Jozef","middleName":"","lastName":"Zurada","suffix":""}],"badges":[],"createdAt":"2025-09-24 15:08:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7705090/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7705090/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":95107724,"identity":"e5427ba8-4512-4c90-9e6e-e98e22c598b2","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1853083,"visible":true,"origin":"","legend":"","description":"","filename":"AutoClusREAutomaticClusteringMethodforEntityRelationshipExtractionandKnowledgeGraphConstruction.docx","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/0c471df54cb43f7b20666a22.docx"},{"id":95107720,"identity":"0144db85-bdfa-4378-8d0f-68fbfb6cbdc0","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6822,"visible":true,"origin":"","legend":"","description":"","filename":"64d8520ff4d142df95fab38d2ee5441b.json","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/1b741ab82802e5e0373a0689.json"},{"id":95224796,"identity":"babdd0b7-870e-4150-8f07-c31abb8ba48d","added_by":"auto","created_at":"2025-11-05 16:24:18","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":162152,"visible":true,"origin":"","legend":"","description":"","filename":"64d8520ff4d142df95fab38d2ee5441b1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/085364ed910be5dbaad2757b.xml"},{"id":95107721,"identity":"9a0f3ecc-a4fd-4ecf-aa75-883717a6f98b","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"jpeg","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":335934,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/2abfcf67b17165bcf45e839b.jpeg"},{"id":95107719,"identity":"06addc96-29e2-48f5-b79e-8ec70f15c84a","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"jpeg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":450214,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/0149e2c69e7de5e15af381ba.jpeg"},{"id":95107730,"identity":"1c671d3b-9ce6-46b0-94a4-6e2b8c923c6a","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":71641,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/4d58af2e73bcfb5391416845.png"},{"id":95107726,"identity":"879b8dca-c757-4300-8467-3af6993b4eaf","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":101907,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/8c46175c7256b34549abeb6f.png"},{"id":95224424,"identity":"d30a77d3-7c79-4f0d-a68f-f3329967d651","added_by":"auto","created_at":"2025-11-05 16:23:44","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":34620,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/674186117b855e77fc894772.png"},{"id":95224296,"identity":"f78f6a06-2a15-4222-924c-6b518724f4e2","added_by":"auto","created_at":"2025-11-05 16:23:34","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":602183,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/c21ab6757225ab2c38b05f20.png"},{"id":95107723,"identity":"07bef802-91e7-4e30-a0d9-cca481c0dc71","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":55596,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/b67598f378cce9c738bb6068.png"},{"id":95224929,"identity":"45309f67-92d1-47e1-bb59-b9de41d15e3f","added_by":"auto","created_at":"2025-11-05 16:24:28","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":76425,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/ae966a04e2c16c71c741bed6.png"},{"id":95107727,"identity":"4c415ee4-0571-4860-a0e6-02a16838bf4c","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":23268,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/f87500c0c08a2ef388401df0.png"},{"id":95107728,"identity":"119bc41f-8ec6-49db-8c79-cbebf05df2af","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":43421,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/3ed79e4e7ecc97a3bdda19c7.png"},{"id":95107729,"identity":"3496afa2-150e-4e08-9035-26f8775c7369","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":15599,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/ed23c3155aa11675e749b915.png"},{"id":95107733,"identity":"d479dd0e-8055-419f-b602-4544f8e70c41","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":143062,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/f155cd3fc96b3dfff7724a57.png"},{"id":95107734,"identity":"b9efb2ee-1c34-493e-919d-b3c2fd95fd84","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"xml","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":161397,"visible":true,"origin":"","legend":"","description":"","filename":"64d8520ff4d142df95fab38d2ee5441b1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/238414a5d9023fee932b6de0.xml"},{"id":95107735,"identity":"8809ba09-305c-4fdc-a53a-b1d3e4686e2b","added_by":"auto","created_at":"2025-11-04 11:17:14","extension":"html","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":180954,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1/da5ef3c2b95d2e3c81546adc.html"},{"id":95653964,"identity":"e330a4ea-1776-4b82-897f-8f5a5d17a5d9","added_by":"auto","created_at":"2025-11-11 16:06:54","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1138094,"visible":true,"origin":"","legend":"","description":"","filename":"AutoClusREAutomaticClusteringMethodforEntityRelationshipExtractionandKnowledgeGraphConstruction.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7705090/v1_covered_398085ad-a534-4b53-8bd8-6339ae38af74.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"AutoClusRE: An automatic clustering-based method for relation extraction and knowledge graph construction","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Relation Extraction, Semantic Clustering, Knowledge Graph, Large Language Model","lastPublishedDoi":"10.21203/rs.3.rs-7705090/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7705090/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe knowledge graph plays a vital role in integrating and representing information. As a core component in the construction of a knowledge graph, relation extraction is responsible for identifying and establishing semantic relationships between entities. However, existing methods rely heavily on annotated data. When dealing with unlabeled text, they fail to automatically perform relation extraction and construct a knowledge graph. This poses a significant challenge for knowledge graph construction. To address the above challenges, this paper proposes a domain-adaptive entity-relation extraction and knowledge graph construction framework, named AutoClusRE (Automatic Clustering-based Relation Extraction), which is based on a large language model and designed for extracting relational triples from unlabeled text. This method requires no additional training and can directly perform adaptive entity\u0026ndash;relation extraction and knowledge graph construction on unlabeled text. Specifically, AutoClusRE constructs an entity type table and a relation type table, which are embedded into the prompt templates of the large language model and dynamically updates them during the extraction process to serve as contextual guidance for the large language model. In addition, a semantics-based clustering algorithm is introduced to effectively merge duplicates and control excessive expansion of the type tables. Experiments conducted on the public dataset FOBIE and the custom dataset HSA (Huizhou-style architecture) demonstrated that AutoClusRE achieved competitive extraction performance without relying on any annotated data. On the public FOBIE dataset, AutoClusRE outperformed traditional deep learning models by 14.8% and DeepSeek by 9.5% in terms of F1 score. On the custom dataset HSA, AutoClusRE\u0026rsquo;s F1 score improved by 17.2% over traditional deep learning models and by 2% over DeepSeek. This study not only introduces a novel relation extraction method for information management in unknown domains but also constructs a knowledge graph in the field of Huizhou-style architecture based on AutoClusRE, contributing to the digital research and application of domain knowledge in this field.\u003c/p\u003e","manuscriptTitle":"AutoClusRE: An automatic clustering-based method for relation extraction and knowledge graph construction","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-04 11:17:09","doi":"10.21203/rs.3.rs-7705090/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c20ce331-2e53-4634-8d81-1cc0a0e2593c","owner":[],"postedDate":"November 4th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-11-10T10:27:37+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-04 11:17:09","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7705090","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7705090","identity":"rs-7705090","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.