Ontology Extension with NLP-based Concept Extraction for Domain Experts in Catalytic Sciences

preprint OA: closed
Full text JSON View at publisher

Abstract

Ontologies store semantic knowledge in a machine-readable way and represent domain knowledge in controlled vocabulary. In this work, a workflow is set up to derive classes from a text dataset using natural language processing (NLP) methods. Furthermore, ontologies and thesauri are browsed for those classes and corresponding existing textual definitions are extracted. A base ontology is selected to be extended with knowledge from catalysis science, while word similarity is used to introduce new classes to the ontology based on the class candidates. Relations are introduced to automatically reference them to already existing classes in the selected ontology. The workflow is conducted for a text dataset related to catalysis research on methanation of CO 2 and seven semantic artifacts assisting ontology extension by domain experts. Undefined concepts and unstructured relations can be more easily introduced automatically into existing ontologies. Domain experts can then revise the resulting extended ontology by choosing the best fitting definition of a class and specifying suggested relations between concepts of catalyst research. A structured extension of ontologies supported by NLP methods is made possible to facilitate a Findable, Accessible, Interoperable, Reusable (FAIR) data management workflow.
Full text 13,232 characters · extracted from preprint-html · click to expand
Ontology Extension with NLP-based Concept Extraction for Domain Experts in Catalytic Sciences | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Ontology Extension with NLP-based Concept Extraction for Domain Experts in Catalytic Sciences Alexander S. Behr, Marc Völkenrath, Norbert Kockmann This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2457909/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Jul, 2023 Read the published version in Knowledge and Information Systems → Version 1 posted 10 You are reading this latest preprint version Abstract Ontologies store semantic knowledge in a machine-readable way and represent domain knowledge in controlled vocabulary. In this work, a workflow is set up to derive classes from a text dataset using natural language processing (NLP) methods. Furthermore, ontologies and thesauri are browsed for those classes and corresponding existing textual definitions are extracted. A base ontology is selected to be extended with knowledge from catalysis science, while word similarity is used to introduce new classes to the ontology based on the class candidates. Relations are introduced to automatically reference them to already existing classes in the selected ontology. The workflow is conducted for a text dataset related to catalysis research on methanation of CO 2 and seven semantic artifacts assisting ontology extension by domain experts. Undefined concepts and unstructured relations can be more easily introduced automatically into existing ontologies. Domain experts can then revise the resulting extended ontology by choosing the best fitting definition of a class and specifying suggested relations between concepts of catalyst research. A structured extension of ontologies supported by NLP methods is made possible to facilitate a Findable, Accessible, Interoperable, Reusable (FAIR) data management workflow. Ontology Natural Language Processing Automated Ontology Annotation Information Extraction CO2 Methanation Catalytic Conversion Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 14 Jul, 2023 Read the published version in Knowledge and Information Systems → Version 1 posted Editorial decision: Major revision 08 May, 2023 Reviews received at journal 04 May, 2023 Reviewers agreed at journal 25 Apr, 2023 Reviews received at journal 24 Apr, 2023 Reviewers agreed at journal 05 Feb, 2023 Reviewers agreed at journal 02 Feb, 2023 Reviewers invited by journal 30 Jan, 2023 Editor assigned by journal 25 Jan, 2023 Submission checks completed at journal 09 Jan, 2023 First submitted to journal 09 Jan, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2457909","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":166231091,"identity":"314154db-db15-4996-9647-e695597df9e2","order_by":0,"name":"Alexander S. Behr","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABIUlEQVRIiWNgGAWjYDACdjDJDMQ8INIGxDOAyiUwSGDTwoyqJU0CoiWBeC2HCWvhb2Y+JvFzh7WceQPvwceFe87XGRxg3vi48se2xO3sCYw3PmBqkTjMlmzYeybdWOYAX7LxjGe3JQwOsBUbnkm4nbiz5wGz5Qws1hzmMXzA23Y4cQYDj5k0zwGglvtvzCQbgFo23Ehgk+bB1CF/mP/Dwb9th+uBWsx/8xw4B7SFB0nLH0wtBod5GB8DbUmQANrCzHPgAJoWLO4yPMxmbCzblm44g5kvGeiwZMmZIL80pN023nDmYbNlD6YWuePNzyTftlnLS7D3HvzMc8COnw8YYg8bbG7LbjiefPDGDyzWwAEzphBjAz4No2AUjIJRMApwAwAcTme+jP2L5wAAAABJRU5ErkJggg==","orcid":"","institution":"TU Dortmund University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Alexander","middleName":"S.","lastName":"Behr","suffix":""},{"id":166231092,"identity":"1157512b-5145-4421-9a5b-0d7274aed5e0","order_by":1,"name":"Marc Völkenrath","email":"","orcid":"","institution":"TU Dortmund University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Marc","middleName":"","lastName":"Völkenrath","suffix":""},{"id":166231093,"identity":"1a65a3ce-8e0b-4f07-b40e-b7e918ddad82","order_by":2,"name":"Norbert Kockmann","email":"","orcid":"","institution":"TU Dortmund University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Norbert","middleName":"","lastName":"Kockmann","suffix":""}],"badges":[],"createdAt":"2023-01-09 08:14:40","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2457909/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2457909/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s10115-023-01919-1","type":"published","date":"2023-07-15T01:08:14+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":41739920,"identity":"ef84e346-8c5e-4264-9cf3-9f8bcf7003fa","added_by":"auto","created_at":"2023-08-18 04:20:36","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":513128,"visible":true,"origin":"","legend":"","description":"","filename":"NLPPaper2022SpringerNature.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2457909/v1_covered_46a5bf00-1392-421b-be6d-1c8c7fecb366.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Ontology Extension with NLP-based Concept Extraction for Domain Experts in Catalytic Sciences","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"knowledge-and-information-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"kais","sideBox":"Learn more about [Knowledge and Information Systems](http://link.springer.com/journal/10115)","snPcode":"10115","submissionUrl":"https://submission.nature.com/new-submission/10115/3","title":"Knowledge and Information Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Ontology, Natural Language Processing, Automated Ontology Annotation, Information Extraction, CO2 Methanation, Catalytic Conversion","lastPublishedDoi":"10.21203/rs.3.rs-2457909/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2457909/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eOntologies store semantic knowledge in a machine-readable way and represent domain knowledge in controlled vocabulary. In this work, a workflow is set up to derive classes from a text dataset using natural language processing (NLP) methods. Furthermore, ontologies and thesauri are browsed for those classes and corresponding existing textual definitions are extracted. A base ontology is selected to be extended with knowledge from catalysis science, while word similarity is used to introduce new classes to the ontology based on the class candidates. Relations are introduced to automatically reference them to already existing classes in the selected ontology. The workflow is conducted for a text dataset related to catalysis research on methanation of CO\u003csub\u003e2\u003c/sub\u003e and seven semantic artifacts assisting ontology extension by domain experts. Undefined concepts and unstructured relations can be more easily introduced automatically into existing ontologies. Domain experts can then revise the resulting extended ontology by choosing the best fitting definition of a class and specifying suggested relations between concepts of catalyst research. A structured extension of ontologies supported by NLP methods is made possible to facilitate a Findable, Accessible, Interoperable, Reusable (FAIR) data management workflow.\u003c/p\u003e","manuscriptTitle":"Ontology Extension with NLP-based Concept Extraction for Domain Experts in Catalytic Sciences","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-01-17 15:32:18","doi":"10.21203/rs.3.rs-2457909/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-05-08T07:01:56+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-05-04T11:11:40+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"10872008-61ae-4eaa-8f87-5a6338fd4fbc","date":"2023-04-25T20:34:00+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-04-24T10:21:48+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"01fe4059-3224-41d3-8362-ed78ab7f11ec","date":"2023-02-05T08:48:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"8b72b5e4-5cca-4ba4-ae08-0eada9c8512d","date":"2023-02-02T08:43:42+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-01-30T08:40:14+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-01-25T12:33:51+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-01-09T14:02:14+00:00","index":"","fulltext":""},{"type":"submitted","content":"Knowledge and Information Systems","date":"2023-01-09T08:14:32+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"knowledge-and-information-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"kais","sideBox":"Learn more about [Knowledge and Information Systems](http://link.springer.com/journal/10115)","snPcode":"10115","submissionUrl":"https://submission.nature.com/new-submission/10115/3","title":"Knowledge and Information Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"6ff45969-8b98-45ab-8427-5d8383a6b447","owner":[],"postedDate":"January 17th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-08-18T04:12:51+00:00","versionOfRecord":{"articleIdentity":"rs-2457909","link":"https://doi.org/10.1007/s10115-023-01919-1","journal":{"identity":"knowledge-and-information-systems","isVorOnly":false,"title":"Knowledge and Information Systems"},"publishedOn":"2023-07-15 01:08:14","publishedOnDateReadable":"July 15th, 2023"},"versionCreatedAt":"2023-01-17 15:32:18","video":"","vorDoi":"10.1007/s10115-023-01919-1","vorDoiUrl":"https://doi.org/10.1007/s10115-023-01919-1","workflowStages":[]},"version":"v1","identity":"rs-2457909","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2457909","identity":"rs-2457909","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00