Extracting Non-Taxonomic and Ternary Relations from Patient-Generated Texts for Semantic Interoperability

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Patient-generated texts usually contain useful information, but lack rigorous internal cell structure making them unfit for structured databases. Current research work has focused on identifying hierarchical/taxonomic relations, consequently ignoring non-hierarchical and ternary relations, which are equally crucial for comprehensive semantic understanding. This study addresses this through semantic alignment based on non-taxonomic and ternary relational components. The work adopts a Design Science Research (DSR) approach, with a pragmatic research philosophy. We develop and evaluate a knowledge-infused neural framework for cross-domain ontology integration that supports capturing and representing non-taxonomic and ternary relationships beyond general hierarchical relations. The framework adopts a four-layered architecture. A key contribution is the implementation of the delayed fusion strategy that balances the need for contextual neural learning, the interpretable rule-based dictionary knowledge, and incorporating BioBERT as a relations validator to ensure domain factual grounding in the integration. The framework was evaluated on 38,115 documents from anxiety and depression datasets, of which 27,183 were key phrases. The hybrid per class adaptive strategy extracted 113 unique concepts, prioritizing a more conservative prediction. The hybrid union extracted 222 unique concepts prioritizing wider coverage of domain concepts for the construction of the knowledge graph. The framework achieved an accuracy of 98.91% and an F1 score of 77.6%, a 10% improvement compared to the BiLSTM F1 score (67.7%). The framework also validated 384 semantic relations with a validation rate of 92.7%. Of the semantic relations validated, 240 were ternary relations that captured multiple contexts of interactions between the concepts. Non-taxonomic relations were 144 in total, organized into different semantic categories including associative, functional, causal, risk factors, and statistical. The framework transforms unstructured patient-generated texts into structured interoperable knowledge, while preserving different clinical contexts. This helped advance semantic interoperability of health data by improving clinical decision making and biomedical knowledge reuse.
Full text 149,181 characters · extracted from preprint-html · click to expand
Extracting Non-Taxonomic and Ternary Relations from Patient-Generated Texts for Semantic Interoperability | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Extracting Non-Taxonomic and Ternary Relations from Patient-Generated Texts for Semantic Interoperability Jael Gudu, Joseph Balikuddembe, Johnson Mwebaze, Daniel Opiyo This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9087634/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Patient-generated texts usually contain useful information, but lack rigorous internal cell structure making them unfit for structured databases. Current research work has focused on identifying hierarchical/taxonomic relations, consequently ignoring non-hierarchical and ternary relations, which are equally crucial for comprehensive semantic understanding. This study addresses this through semantic alignment based on non-taxonomic and ternary relational components. The work adopts a Design Science Research (DSR) approach, with a pragmatic research philosophy. We develop and evaluate a knowledge-infused neural framework for cross-domain ontology integration that supports capturing and representing non-taxonomic and ternary relationships beyond general hierarchical relations. The framework adopts a four-layered architecture. A key contribution is the implementation of the delayed fusion strategy that balances the need for contextual neural learning, the interpretable rule-based dictionary knowledge, and incorporating BioBERT as a relations validator to ensure domain factual grounding in the integration. The framework was evaluated on 38,115 documents from anxiety and depression datasets, of which 27,183 were key phrases. The hybrid per class adaptive strategy extracted 113 unique concepts, prioritizing a more conservative prediction. The hybrid union extracted 222 unique concepts prioritizing wider coverage of domain concepts for the construction of the knowledge graph. The framework achieved an accuracy of 98.91% and an F1 score of 77.6%, a 10% improvement compared to the BiLSTM F1 score (67.7%). The framework also validated 384 semantic relations with a validation rate of 92.7%. Of the semantic relations validated, 240 were ternary relations that captured multiple contexts of interactions between the concepts. Non-taxonomic relations were 144 in total, organized into different semantic categories including associative, functional, causal, risk factors, and statistical. The framework transforms unstructured patient-generated texts into structured interoperable knowledge, while preserving different clinical contexts. This helped advance semantic interoperability of health data by improving clinical decision making and biomedical knowledge reuse. Knowledge-Infused neural framework ontology integration non-taxonomic ternary semantic interoperability Figures Figure 1 Introduction Natural language texts usually contain useful information, but in an unstructured way, lacking a rigorous internal cell structure, making them unfit for structured databases. However, they have the potential to provide unique value yet they remain untapped, unexplored [1]. Within the architecture of data representation and storage systems, it is crucial to ensure the data in natural language texts is well-organized, easily discoverable, and effectively utilized in its entirety [2]. This creates a need for interaction and collaboration with other data storage systems [3]. As a result, systems are increasingly relying on the integration of vast amounts of external data to support comprehensive and accurate decision-making [4]. However, challenges hindering meeting this analytical need include: 1) combining cross-domain data to multiply its value (integration), 2) an awareness of the associations and relations between the different pieces of data [5], [6], [7]. Additionally, most of the research work has focused on identifying taxonomic relations, such as broader/narrower, consequently ignoring other relations such as non-hierarchical and structural-level relations that are equally crucial for comprehensive semantic understanding [8], [9]. Mechanisms and approaches have been adopted that address these challenges including deep learning, transformer based approaches (BioBERT), Large Language Frameworks (LLMs), graphical representation of data (knowledge graphs). However, even with the advancements, ternary (N-ary) and non-taxonomic relations and community structural dependencies and interactions remains underexplored, leading to incomplete mappings. These relationships between texts and graph structures are crucial for understanding the complexity of clinical events. The need for scalable, relation-aware ontology learning and integration that can manage intricate relationships across several ontologies is what has motivated this study. The inability of current approaches to capture rich, complex semantic relationships results in uneven and disjointed knowledge integration. Our goal is to improve ontology networks' quality (completeness, expressiveness, and scalability) by creating a comprehensive, relation-aware approach that will make them more resilient and dependable for semantic integration. The function of the proposed extraction framework is to extract relevant clinical concepts for each stream or input, classify the relations between them into binary, ternary and non-taxonomic, within the relevant domain local network. The framework demonstrated high effectiveness, achieving a test accuracy of 98.91% and a strong recall of 92.8% and F1 Score of 77.6%. This highlights the framework’s reliability and robustness in identifying semantically relevant medical concepts existing in real-world noisy patient-generate texts. The framework as it is serves a practical tool for data engineering and semantic data design that can be used to transform real-world unstructured patient-generated text into structured and interoperable semantic knowledge. This addresses the core problem of semantic interoperability of mental health (anxiety and depression) data. Additionally, the framework is best suited for high recall information retrieval and screening of clinical concepts in the real-world context. Background and Related Works In patient-generated text, semantic relationships are defined based on context. Context, in this case, is defined as the details surrounding a concept that gives details of a particular situation, at a particular instance of time and place [ 10 ]. Within such texts, context is not confined to a single sentence, in some instances spans multiple sentences or an entire document [ 11 ]. For example, in these sentences An individual has suffered from anxiety for over a year and a half. It has mainly affected their sleep. The concept ‘ anxiety’ is very clear within the first sentence, but in the subsequent sentence the ‘it’ needs more surrounding concepts (context) for meaning to be understood. This characteristic of patient-generated texts is one of the primary causes of semantic ambiguity and mapping inconsistencies mentioned in Chap. 1, thus a barrier to seamless semantic interoperability. Without proper tools and mechanisms that reason and learn context, cross-dependency and co-reference in patient-generated texts fail. This leads to missed concepts, misclassification of concepts, and mapping inaccuracies. 2.1 Types of Semantic Relationships Within such texts, semantic relations is not confined to explicit expressions, in some instances some are latent in meaning and in words (implicit). They are often difficult to realize as they are hidden in between different sentences spread across a document [ 12 ], [ 13 ]. This study adopts an arity dimensional categorization approach to semantic relationships. 2.1.1 Arity or Degree of Relationship This dimension defines the instances of concepts participating in the relationship in the text [ 14 ]. It encompasses Binary and N-ary relations. Binary Relation are the primary focus of most relation extraction works. N-ary Relations describing ternary and higher-ary relations where n is greater than two [ 15 ]. For Example in this sentence, A patient has suffered anxiety for a long time now, it has affected their sleep. The person previously tried cognitive behavioural therapy to help but it failed after a few month. Instead the situation became worse. The person then consulted a doctor who then prescribed mirtazapine. The medication worked effectively and the person was able to sleep. However, the effect of the medication after a few days has been dreadful causing loss of energy. The person is aware that these are some of the symptoms. In the first two sentences, three concepts “anxiety, affected sleep, and cognitive behavioural therapy” are participating in a ternary relationship. However, as the sentences increase it moves to an N-ary relationship with 6 concepts involved. These complexities underline the necessity for advanced frameworks capable of handling n-ary relations, as they would also facilitate the creation of more comprehensive and accurate biomedical knowledge graphs, thereby enhancing the level of expressiveness (scope) and effectiveness of integration [ 15 ]. 2.2 Relation Extraction Techniques for Semantic Relations In various domains such as healthcare, security, software, and education, relations extraction from unstructured texts has largely been driven by deep learning (CNNs and RNNs) and pre-trained language frameworks (BERTs and LLMs) for contextual embedding. They have become the core technology for Named Entity Recognition and knowledge graph construction [ 15 ], [ 16 ]. 2.2.1 Deep Learning Approaches There are various deep learning techniques including Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), Long-Short Term Memory (LSTMs), Gated Recurrent Unit (GRU) [ 17 ], [ 18 ]. These techniques often require more training to match the precision and domain-specific adaptability of rule-based systems. Additionally, they still require substantial computational resources [ 19 ]. Frameworks such as CNN and Long Short-Term Memory (LSTM) shows significant performance over rule-based and dictionary-based approaches in incorporating contextual information. They can be utilized independently or in tandem with explicit semantic domain knowledge. They also have the ability to inference for unspecified relationships, due to their ability to generalize a sentence or across sentences without requiring any patient- or situation-specific training data, thus offering a scalable approach [ 17 ], [ 20 ]. While CNNs and RNNs are effective at capturing local features within a sentence, they struggle with long-distance dependencies. Despite handling longer sequences better than RNN, LSTMs and BiLSTM face issues with very long-range dependencies. This contributes to the persistent semantic interoperability gap in unstructured texts that have cross-sentences dependencies leading to missed concepts and non-taxonomic relations. For this task a BiLSTM leverages on its sequential processing and labeling richness is well suited for the linear format with clear, concise, logical flow of information found in patient expressions [ 20 ]. Additionally, the BiLSTM leverages on its ability to capture local-feature and mid-range dependencies within efficient timelines based on its low computing resource consumption architecture. A Comparative study results by [ 21 ] reveals that in one training, the train epoch of BiLSTM is 178.41 seconds, while a Transformer training epoch is 1657.75 seconds. This makes a BiLSTM suitable for devices and scenes that require low or less computing resource consumption. This creates the need for a more efficient fusion scheme(s) that balances the performance and compute cost. 2.2.2 Transformer based Approaches The emergence of transformer based approaches including Generative Pre-trained Transformers frameworks such as Bidirectional Encoder Representation from Transformers (BERT) and Large Language Frameworks (LLMs) has revolutionized the concept and relation extraction landscape, with their exceptional state-of-the-art performance in various applications [ 22 ]. Biomedical adaptations of BERT like BioBERT [ 23 ] and ClinicalBERT [ 24 ] has enabled more accurate interpretation of clinical narratives thus advancing clinical decision support, and disease prediction. These frameworks leverage on the innate capabilities of their architecture consisting of an encoder and a decoder to understand contextual relationships within text [ 25 ]. 2.2.3 Hybrid Approaches The recent and ongoing integration of deep learning and transformer based frameworks into the relation extraction process represents a flexible and effective architectural combination for tasks requiring both global context and fine-grained sequence processing [ 26 ], [ 27 ]. This is because of their unique combination of deep learning and semantic web technologies. By autonomously learning hierarchical and non-hierarchical representations from large datasets, deep learning and transformer-based framework have drastically improved the quality, robustness, and automation of the knowledge fusion process. [ 28 ] used a adopted a transformer and BiLSTM approach for text mining and visualization of underground engineering text reports. [ 29 ] combined common NLP techniques with traditional latent thematic analysis (text mining approach) to classify UGC drawn from an interactive HIV mHealth environment. In the instances given the integration of different approaches in the relation extraction process improved the precision and also made full use of the potential value information hidden behind massive text. Methodology The study derives its corpus from online medical forums including: patient.info health forum [ 30 ], Health Unlocked [ 31 ], and Mental Health Forum [ 32 ], Depression Dataset on Kaggle [ 33 ]. It was a combination of anxiety disorder (AnxD) from different online forums and open dataset called Depression Dataset. Simple random sampling was used to avoid selection bias on the texts. The dataset was a collection of 38,115 sentences with 27,183 key phrases. From the key phrases 21,746 were used as the training set, and 5,437 as the test set. The division of the dataset is shown in Table 1 . Table 1 Size of datasets Dataset Total Documents Relevant Set Training Set Testing Set Proof of concept 5,822 3,474 2,779 694 AnxD and Depression 38,115 27,183 21,746 5,437 3.1 Ontology Dictionary Construction In the experiment significant words from the corpus in the different variations and synonyms are bootstrapped in a local concept dictionary (a specialized clinical vocabulary), as seed words to act as sematic triggers and anchor terms for each of the categories. This is operationalized by compiling data from the training dataset (sentence and document-level) and was also linked to external concept dictionaries such as SNOMED CT. The local concept dictionaries were then expanded and enriched using 5 WordNet synonyms for each word variant to prevent the explosion of the synonyms. The study opted for WordNet due to the many variants of concepts found in text, especially when identifying the different symptoms variants found in the texts. 3.2 Relation Extraction Framework Architecture The BiLSTM’s architecture was deliberately designed with relatively low complexity with 2 input layers receiving total parameters 7,019,457 (26.78 MB). The 26.78MB comprised of 4,019,101(15.33MB) trainable parameters and 3,000,356 (11.45MB) non-trainable parameters. The practicability of this is that the total parameter size of 26.78MBs represents a compact neural network architecture. The framework implements a dual-stream feature-fusion architecture with a delayed fusion strategy. This allows the freezing of a larger portion (non-trainable) of the parameters (15.33MB) and focus on the training and learning capacity on the attention fusion. The framework implements a dual-stream feature-fusion architecture with a delayed fusion strategy. This balances the need for contextual neural learning and dictionary and rule-based knowledge (knowledge-infusion method) integration. As seen in Fig. 6, the framework receives two inputs seq_input (sequential text input) and dict_feat that contains engineered features obtained from medical dictionaries, WordNet which are processed independently. The seq_input undergoes linguistic processing of the text sequence, while dict_feat is prepared for attention weighting. BioBERT acted as a biomedical relations precision sieve serving as a cleaner and refiner of the proposed relations, filtering out what was not semantically and factually sound. Adopting the zero-shot validation strategy allowed the model to perform the validation by understanding relational information and semantic similarity based on BioBERT’s innate pre-trained knowledge, without the need for additional training [ 34 ]. To assess the performance of the proposed zero-shot inferencing and validation strategy, the model was exposed to relations, which contained several multi-class predicates initially extracted by BiLSTM. In addition, the study simplified the task to binary non-taxonomic and ternary relations validation, and not multi-class relation classification task. Adopting the zero-shot validation strategy allowed the model to perform the validation by understanding relational information and semantic similarity based on BioBERT’s innate pre-trained knowledge, without the need for additional training [ 34 ]. Selecting the confidence threshold was an important step in determining and controlling the grounds for accepting a valid relation or rejecting a relation by BioBERT. This is because the choice of threshold ensured the trade-off between having the model capture many relations as possible (recall) and ensuring quality and accurate biomedical relations (precision). In addition, the thresholding mechanism allowed the model to make practical decisions while balancing false positive errors [ 35 ]. A one-stage threshold setting approach was adopted in this study. For this work, the threshold was set to 0.4, a low threshold that was focused on giving priority to recall and ensuring as many candidate relations as possible were detected and accepted for further validation processing. This threshold was set purely as a cutoff parameter. The goal was to ensure potential candidate relations were not discarded prematurely. At this stage, the model compared each confidence score against the predefined threshold and only relations that were equal to or above this threshold proceeded to the validation process, while those below were left out. The input given to the zero-shot validator were 414 candidate relations of which 384 met the 0.4 threshold and were retained as the final validation output. Of the 414 candidate relations, 30 relations failed to meet the threshold and were discarded. Results Given an input sentence, the framework employed a hybrid knowledge-infused neural approach to predict a set of candidate relation types meant to guide the subsequent joint entity-relation extractor. The initial extractor, powered by the BiLSTM identified 5089 of the relations comprising ternary relations and non-taxonomic relations. The extraction was based on real-world patient contexts that guided the customized predicate labels used by the BiLSTM. The 5089 relations were then subjected to normalization and deduplication to ensure only clean relationships remained. This led to a significant drop in the total number or relations to 414 non-taxonomic and ternary relations. Table 2 shows the extraction results of the different relations extracted. Table 2 Relations Extracted from Anxiety and Depression Dataset Type of relationship Relationship Count Sample relations Ternary 55 'treatedWith', 'managesConditionContext', 'relievedInContext' , Non-Taxonomic 359 interacts_with, similar_to, co_symptom_of, participates_in Total 414 The 55 ternary relations revealed valuable context-aware associations from the texts. This captured higher levels of relations, especially the ternary, that provided in-depth understanding to different patient experiences, medical conditions and treatments. Relation Extraction Rules The extraction process was governed by the following rules: First, an entity–relationship triple ‘’ was only considered valid and sound if its relation and the concept names of the head entity and tail entity were accurate. Secondly, when the candidate relations identified did not fit the concepts in the sentence, the extractor did not force a relation to be selected from the incorrect candidate relations, but instead assumed that no relation exists between those entities. For example, given a sentence below with no relations between anxiety , anxiety disorder and the words sadly, and treatment in the biomedical context the framework should assume there is no relation between them. According to a national survey in 2010, 32 percent of adolescents in the United States have an anxiety disorder. [ 2 ] Sadly, majority of children with anxiety never receive treatment , [ 36 ], [ 37 ]. This was meant to ensure that only logical and meaningful relations between concepts were included in the ontology. It also aimed at formalizing domain knowledge and establishing logical constraints within the ontology. To illustrate extracted binary non-taxonomic relations the following example is used: (fatigue → anxiety, suggests ) expressed in triple as ('fatigue', 'suggests', 'anxiety'). Consider this relevant sentence subset, The brain gets deprived of glucose; you develop fatigue, mind-fog, dizziness, and as the process progresses you can develop restlessness as the brain searches for natural alternate energy sources such as adrenaline, which can then cause anxiety and even panic attacks. In this example, framework performed a sentence-level extraction and found an association suggests between a concept fatigue and domain concepts anxiety and panic attacks. This means the existence of fatigue in a sentence suggests a medical condition panic attacks or anxiety. (anxiety → sadness, hasSymptom ) expressed in triple as ('anxiety', 'sadness', 'hasSymptom') or ('anxiety', 'hasSymptom', 'sadness') (depression → fluoxetine, treatedWith ) 4.1 Relation Validation Results Each candidate relation was presented to the BioBERT for semantic and factual validation, specifically within the patient expressions on anxiety and depression disorders. As a result, from the candidate relations (ternary and non-taxonomic), the BioBERT powered validation identified 384 relations, with confidence values ranging from 75–98%. The validation results are shown in Table 3 . Table 3 Summary of Biomedical validated Relations Type of relation BiLSTM Relations BioBERT Validated Relations Sample validated relations Ternary 55 240 {'symptom': 'low energy', 'disease': 'anxiety', 'treatment': 'negative thought challenge', 'relation': 'managesConditionContext'} Non-taxonomic 359 144 Depression, treatedWith, fluoxetine , fatigue, suggests,anxiety Total 414 384 Validation rate:92.7% The validation results reveal distinctions between the two categories of relations. The BioBERT validated all the ternary relations (100% validation rate) initially extracted by BiLSTM. However, due to the pre-training on biomedical literature BioBERT was able to capture more valid ternary relations contexts, therefore identifying 240 ternary relations. This was 185 more entries than what the BiLSTM had extracted. This does not indicate a failure by the BiLSTM, but instead confirms that BiLSTM was recall focused, and was limited in precision (section 5.3.5). The 100% validation rate shows that the BiLSTM was performed significantly well in extracting complex relations from the unstructured text; that align to existing biomedical knowledge. In contrast, BioBERT validated 144 of the 359 non-taxonomic relations, with a 40.1% validation rate. 4.2 Ternary Relation Extraction Unlike binary non-taxonomic associations that seek general connections between concepts, in ternary relation structure the relation covers a multi-concept clinical context. The multi-concept clinical context here refers to clinical integral truth conditions (with all the concepts from all the three classes, disease, symptom and treatment participating). This context justified the specific situation why, where and when the relation holds true, without which the relation is invalid. The validated ternary relations had a consistent pattern and structure that comprised of a subject concept, relation (explaining the nature of the link), object concept with the context attached to it. This was expressed as: concept, associated_with, concept | {clinical context} Table 10 explicitly demonstrates this ternary structure from the text given. A person has suffered anxiety for a long time now, it has affected their sleep. They had. previously tried cognitive behavioural therapy to help but it failed after a short while. Instead, the situation became worse. The person then consulted a doctor who prescribed mirtazapine the person’s first time on such a medication. The medication worked effectively improving their sleep. However, the effect of the medication after a few days has been dreadful causing loss of energy. The person is aware that these are some of the symptoms. Table 4 shows the sample ternary relations. Table 4 Sample Results of Ternary Relations with clinical context. The table shows the ternary relations based on subject, objects and the relations between them. It also shows the types and classes where the subjects belong and the contextual qualifiers that make the ternary relation hold between the subject and the object. Subject Concept Subject type Relation Object Concept Object type Anxiety Disease affects sleep Symptom | {for a long time > 1 year, treated using CBT (treatment)} loss of energy Symptom caused_by mirtazapine Treatment | {Treated anxiety (disease) temporarily, dreadful effect after few days} Anxiety Disease treated_by Cognitive Behavioural Therapy (CBT) Treatment | {unable to sleep (symptom), for a long time} Anxiety Disease treated_by mirtazapine Treatment | {improved sleep, temporarily effective} From the text, a document-level extraction reveals more than one ternary relations as seen in table 5.4.3.2. These are based on the structure form: Anxiety (Disease), treated_by, mirtazapine (treatment) | {long time with anxiety, prior CBT failure (treatment), affected sleep (symptom)} Mirtazapine (treatment), causes, loss of energy (symptom) | {used for a short while, offered temporary relief for anxiety (disease), delayed dreadful effect} In the two ternary relations, the framework resolves a contradiction that Mirtazapine (treatment) can bring relief and can cause loss of energy that would arise if they were expressed purely as binary non-taxonomic relations such as Mirtazapine (treatment), causes, loss of energy Mirtazapine (treatment), relieves, anxiety Any interaction with the two binary non-taxonomic associations would not know how to distinguish these two relations if context was not provided. The context here include past medical history (using CBT to treat anxiety), temporal evidence of the treatment offering relief for a short while, and also the symptom evolves after a short while. These extracted ternary relations show that richer information and contexts can be discovered by with the aim of supporting clinical inferencing and integration with existing domain ontologies. The 240 ternary relations preserve clinical contextual structures during semantic integration address a semantic interoperability challenge: the ambiguity of isolated clinical facts. Here, the study defines ambiguity of clinical fact as a situation that a concept is found in different conditions [ 38 ]. For example, information such as Mirtazapine (treatment), causes, loss of energy may lead to several interpretations without this other piece of information Mirtazapine (treatment), relieves, anxiety and the reverse is also true. Therefore the ternary structure loss of energy, causedBy, mirtazapine | {Treated anxiety (disease) temporarily, dreadful effect after few days} resolves these ambiguity. Clinical decisions occur within complex interrelated information including diagnosis results, patient histories, pharmacological interactions, and social and psychological factors [ 39 ], [ 40 ]. Any system supporting a clinical or patient monitoring decision is equipped with the context (all relevant and valid information) to guide its reasoning towards a semantically consistent interpretation across different health systems. 4.3 Confidence Score Analysis This section analyses the model’s level of certainty in the relations it has validated. Table 5 gives the confidence range for the relations together with average confidence for the different categories. Table 5 Summary of the Confidence Scores per Relation Category Relation Type Confidence Range Average Confidence (%) Associative 0.917–0.984 96 Functional 0.697–0.965 88 Causal 0.790–0.962 92 Risk Factors 0.704–0.947 86 Statistical 0.608–0.851 75 Overall Average Confidence 87.4 The results reveal that all the confidence scores were above the 0.4 threshold that was set. The lowest score was 0.60, which was still higher than the threshold. This shows that the 0.4 threshold cast a wider net, prioritizing recall without compromising the quality of the relations. An analysis of the results reveals that over 90.5% of the validated relations were above 90%, which indicates that BioBERT zero-shot validator captured high quality relations that existed within the pre-existing biomedical knowledge, confirming its role in the framework as a precision focused optimizer. The highest confidence score range was 0.917–0.984, assigned to ‘ associatedWith’ . This indicates that BioBERT was highly confident in validating this type of non-taxonomic relations between different disease, symptom and treatment concepts. This high confidence score can be attributed to co-occurrence relationships, which accelerate the recognition of term associations, has been well documented in biomedical literature [ 41 ]. For example, the co-occurrence of anxiety and panic, with respiratory challenges (chronic breathlessness) have been well documented [ 42 ]. The consistent high confidence score within the 54 associative relation category made them a more reliable semantic link within the knowledge graph. The relatively low confidence scores largely for statistical relations (0.60) and risk factor relations (0.70) suggests relations that are indicative and explain factors that influence the occurrence of disorders, but might not be the final or controlling factor to the existence of the disorder [ 43 ]. The moderate confidence scores (0.79–0.85) predominantly for causal and some statistical relations, indicated that these relations were reliable for inclusion in the knowledge graph, but might require more evidence and verification to increase the level of certainty. 4.4 Performance Evaluation When compared to the BiLSTM results (recall = 75.6%, precision = 61.3%, and F1 score = 67.7%), the hybrid framework achieved a 10% improvement on the F1 score, a 17.2% improvement on the recall and a 5.4% improvement on precision. Table 6 shows a summary of the performance evaluation. The improvement on the precision and recall indicates the significance of infusing domain knowledge and BioBERT to the framework. The framework focused on optimizing the precision and ensuring the relations were semantically relevant and valid to the biomedical domain. Table 6 Summary of the overall performance of the framework Dataset Performance Hybrid Completeness and Consistency Check Precision (%) Recall (%) F1-Score (%) Hamming loss Jaccard AnxD and Depression Dataset 66.7 92.8 77.6 0.029 0.667 The Jaccard in this case was used context of relations validation to confirm what was currently existing in biomedical literature and what BiLSTM had extracted. The Jaccard of 0.667, a value closer to 1 means the predicates and true relation labels existing in the domain were nearly identical [ 44 ], resulting in reliable relations. This mean the results produced by the model are reliable Discussions The hybrid knowledge infused neural approach, integrating domain knowledge with a BiLSTM framework fused with an attention mechanism demonstrated high levels of efficacy in the extraction and classification of complex medical relations found in free texts. In experimental evaluation, the hybrid framework achieved an accuracy of 98.91% on the test set, a significant improvement over traditional RNNs and standalone BiLSTM algorithms [ 45 ]. The results of the experimental evaluation indicate that the developed framework achieved, a significant improvement in accuracy over traditional RNNs and standalone BiLSTM algorithms on the test dataset. The framework produces the best results in terms of recall and F1 score highlighting its reliability and robustness in handling unstructured test dataset. A knowledge graph representing the data was created using the hybrid union strategy. An analysis of the graph reveals highly interconnected concepts within the graph with a high density and average degree. This consequently provides a qualitative foundation to evaluate of the knowledge contained in the graph and the reality the knowledge represents. Limitations and Future Works An inherent challenge encountered during the study was the issue of class imbalance in medical datasets, which complicated accurate concept prediction and classification capabilities of the framework. Class imbalance occurs when the training dataset contains classes that are unequally distributed [ 46 ], [ 47 ]. This arose due to the nature of the patient-generated texts used. This type of data naturally represents the patient’s subjective experience that captured the voice of the patient when describing medical situations, including figures of speech, idioms, or lay terms [ 48 ]. This bias is confirmed by the resulting class distribution where symptoms (53 concepts) was the majority (prevalent) class, followed by treatment (40 concepts), while disease (18) was the minority (rare) class. This mirrored how often the balance is unreached in real-world datasets. With this type of distribution, it was expected that the framework would have failed to make meaningful predictions for the minority classes [ 49 ]. The F1 score results of 0.789 for treatment, 0.768 for symptoms, and 0.657 for disease, confirmed the expectation and showed the difficulty the framework faced in accomplishing its objective of accurately predicting the minority class of interest. Another significant challenge experienced in this study was semantic overlap and ambiguity. A situation that arises when certain classes tend to share the same vocabulary, resulting in misclassification of domain concepts [ 50 ]. This limitation contributed to the low precision especially in disease category (47%), an indication misclassification, despite correctly identifying the broader context. For example, Dataset 1: When we experience an involuntary high degree stress response, the symptoms can be so profound that we think we are having a medical emergency , which anxious personalities react to with more fear . And when we become more afraid, the body is going to produce another stress response, which causes more changes and symptoms , which we can react to with more fear , experience more symptoms, and so on. From the excerpt “ which anxious personalities react to with more fear” , the word “anxious” describing a temporary feeling of fear and worry about something [ 51 ] was reasoned out and misclassified as a possible disease class “anxiety”. The disease concept anxiety is linked to being fearful and anxious in different circumstances over time [ 52 ]. The framework’s error was occasioned by the lack temporal pattern within the specific excerpt to give the contextual guidance for an appropriate classification decision. Because of this semantic ambiguity, the framework extracted the concept “anxious” and classified it within the disease category. This was because the concept has a strong lexical association to anxiety (existence of fear) while overriding the context that indicated a symptom in this case. Such cases of semantic overlap are common in the medical domain, a problem that this framework did not resolve fully. Potential solutions include adding clear descriptive text and framework refinement [ 53 ]. However, due to the broad nature of the patient expressions within the texts exploring these complex solutions can be considered for further work. Refining relation directionality will be a key area of focus in future studies building upon this foundational work. This is to ensure ontological consistency to prevent semantic and inferencing errors, conflicts, lower quality, and knowledge inconsistencies that arise from poorly integrated information. Extending this study by developing a model to improve the BioBERT fine-tuning classification and prediction performance and to compare the new results to the present ones. Based on the promising results, the study was inspired and aim to adapt the proposed framework to patient contexts that go beyond pure clinical associations. The key question to be answered is: would the experiment improve the scope of BiOBERT validated relations beyond the 393 if pure patient contexts are to be factored in? Conclusion The results shown make the framework useful for knowledge extraction tasks requiring comprehensive semantic coverage. This is critical in ensuring information needed for: knowledge representation, informed decision making, semantic dependency building is available. In such applications and tasks, overlooking key concepts can be more costly than retrieving occasional false positives making the trade-off a justifiable choice for this study. Declarations Author Contribution All authors reviewed the manuscript Funding This work was supported by METEGA. Ethics declarations Competing Interests The authors declare no competing interests. Ethics Approval and Consent to Participate Not applicable. Consent for Publication Not applicable. References Khine PP, Wang Z, Shun (2018) Data lake: a new ideology in big data era. ITM Web Conf 17:03025. 10.1051/itmconf/20181703025 Naseem O Metadata management in data lakes, Data Science Central. Accessed: Oct. 29, 2024. [Online]. Available: https://www.datasciencecentral.com/metadata-management-in-data-lakes/ Chen Z, Wan Y, Liu Y, Valera-Medina A (Jan. 2024) A knowledge graph-supported information fusion approach for multi-faceted conceptual modelling. Inf Fusion 101:101985. 10.1016/j.inffus.2023.101985 Rousanb W Leveraging Big Data for Enhanced Government Security: A Deep Dive into Use Cases, Datahub Analytics. [Online]. Available: https://datahubanalytics.com/leveraging-big-data-for-enhanced-government-security/#:~:text=Big%20Data%20enables%20the%20integration%20and%20an alysis,resource%20allocation%2 C%20risk%20assessment%2C%2 0and%20decision%2Dmaking%20p rocesses Koukaras P (2025) Data Integration and Storage Strategies in Heterogeneous Analytical Systems: Architectures, Methods, and Interoperability Challenges. Information 16(11). 10.3390/info16110932 Moslemi MH, Mousavi A, Behkamal B, Milani M Heterogeneity in Entity Matching: A Survey and Experimental Analysis. 2025. [Online]. Available: https://arxiv.org/abs/2508.08076 Subramaniam P, Ma Y, Li C, Mohanty I, Fernandez RC Comprehensive and Comprehensible Data Catalogs: The What, Who, Where, When, Why, and How of Metadata Management, ArXiv , vol. abs/2103.07532, 2021, [Online]. Available: https://api.semanticscholar.org/CorpusID:232233695 Kolli M (Dec. 2023) Formal modeling of multi-viewpoint ontology alignment by mappings composition. Acta Univ Sapientiae Inf 15:181–204. 10.2478/ausi-2023-0013 Osman I, Pileggi SF, Ben Yahia S, Diallo G (2021) An Alignment-Based Implementation of a Holistic Ontology Integration Method. MethodsX 8:101460. 10.1016/j.mex.2021.101460 Kowalski P, Jousselme A-L (2023) Context-awareness for information correction and reasoning in evidence theory. Int J Approx Reason 153:29–48. https://doi.org/10.1016/j.ijar.2022.11.009 Xie T et al (2024) ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity, in., IEEE International Conference on Data Mining Workshops (ICDMW) , Los Alamitos, CA, USA: IEEE Computer Society, Dec. 2024, pp. 907–911. 10.1109/ICDMW65004.2024.00126 Alam R Understanding Semantic Relations: Types, Examples, and Applications, Linkedin. Accessed: Aug. 02, 2025. [Online]. Available: https://www.linkedin.com/pulse/understanding-semantic-relations-types-examples-md-rabby-alam-ndewf/ Zemla JC (2022) Knowledge representations derived from semantic fluency data. Front Psychol 13:815860 Lentschat M, Buche P, Dibie-Barthelemy J, Roche M (Dec. 2022) A new method to extract n-Ary relation instances from scientific documents. Expert Syst Appl 209:118332. 10.1016/j.eswa.2022.118332 Fernandes JLM (2025) Extracting n-ary Relations from Biomedical Literature using Deep-Learning Techniques, PhD Thesis, UNIVERSIDADE DE LISBOA Choi S, Jung Y (2025) Knowledge Graph Construction: Extraction, Learning, and Evaluation. Appl Sci 15(7). 10.3390/app15073727 Li Y, Luan Z, Liu Y, Liu H, Qi J, Han D (2024) Automated information extraction model enhancing traditional Chinese medicine RCT evidence extraction (Evi-BERT): algorithm development and validation. Front Artif Intell 7:1454945. 10.3389/frai.2024.1454945 Zengeya T, Fonou Dombeu JV (Jan. 2024) A Review of State of the Art Deep Learning Models for Ontology Construction. IEEE Access 1–1. 10.1109/ACCESS.2024.3406426 Qian XX, Chau PH, Fong DYT, Ho M, Woo J (2025) Development and Validation of a Rule-Based Natural Language Processing Algorithm to Identify Falls in Inpatient Records of Older Adults: Retrospective Analysis, JMIR Aging , vol. 8, p. e65195, Jul. 10.2196/65195 Ji J, Chen B, Jiang H (Sep. 2020) Fully-connected LSTM–CRF on medical concept extraction. Int J Mach Learn Cybern 11. 10.1007/s13042-020-01087-6 Lei H (2025) The Difference between BiLSTM and Transformer-The Discussion about Suitable Scenarios for BiLSTM and Transformer. Sci Technol Eng Chem Environ Prot, 1, 1 Cho HN et al (2024) Task-specific transformer-based language models in health care: scoping review. JMIR Med Inf 12:e49724 Lee J et al (2020) BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4):1234–1240 Wang G et al (2023) Optimized glycemic control of type 2 diabetes with reinforcement learning: a proof-of-concept trial. Nat Med 29(10):2633–2642 Visweswaraiah V, Everyday AI (Sep. 2025) Real-World Applications of Transformer- Based Language Models. Int J Comput Trends Technol 73:19–27. 10.14445/22312803/IJCTT-V73I9P103 Mienye ID, Swart TG (2024) A Comprehensive Review of Deep Learning: Architectures, Recent Advances, and Applications. Information 15(12). 10.3390/info15120755 Zubair M, Hussain M, Albashrawi MA, Bendechache M, Owais M (2025) A comprehensive review of techniques, algorithms, advancements, challenges, and clinical applications of multi-modal medical image fusion for improved diagnosis, Comput. Methods Programs Biomed. , vol. 272, p. 109014, Dec. 10.1016/j.cmpb.2025.109014 Shao R, Lin P, Xu Z (Oct. 2024) Integrated natural language processing method for text mining and visualization of underground engineering text reports. Autom Constr 166:105636. 10.1016/j.autcon.2024.105636 Skeen SJ, Jones SS, Cruse CM, Horvath KJ (2022) Integrating Natural Language Processing and Interpretive Thematic Analyses to Gain Human-Centered Design Insights on HIV Mobile Health: Proof-of-Concept Analysis., JMIR Hum. Factors , vol. 9, no. 3, p. e37350, Jul. 10.2196/37350 Health N, Navigate Health Anxiety disorders,. Accessed: Jul. 26, 2025. [Online]. Available: https://community.patient.info/tag/anxiety%20disorders/ HealthUnlocked Anxiety and Depression Support, HealthUnlocked. Accessed: Jul. 26, 2025. [Online]. Available: https://healthunlocked.com/anxiety-depression-support Together 4 Change Limited Mental Health Forum. [Online]. Available: https://www.mentalhealthforum.net/forum/threads/do-you-experience-anxiety-approved-research.394720/ França DS (2024) Depression Dataset. Kaggle, [Online]. Available: https://www.kaggle.com/datasets/diegosilvadefrana/depression-dataset Dong X, Zhao D, Meng J, Guo B, Lin H (2025) SyRACT: zero-shot biomedical document-level relation extraction with synergistic RAG and CoT. Bioinformatics 41(7):btaf356 White H Computer Vision Basics: Confidence & Accuracy, Leverage. [Online]. Available: https://www.leverege.com/blogpost/computer-vision-basics-how-confidence-accuracy-and-thresholds-impact-performance#:~:text=Confidence%20Threshold%20Trade%2Doffs,threshold%20depends%20on%20the%20application Merikangas KR et al (2010) Lifetime prevalence of mental disorders in U.S. adolescents: results from the National Comorbidity Survey Replication–Adolescent Supplement (NCS-A)., J. Am. Acad. Child Adolesc. Psychiatry , vol. 49, no. 10, pp. 980–989, Oct. 10.1016/j.jaac.2010.05.017 Weir K, American Psychological Association Brighter futures for anxious kids,. Accessed: Feb. 08, 2026. [Online]. Available: https://www.apa.org/monitor/2017/03/anxious-kids Libertin CR et al (Jun. 2023) Data structuring may prevent ambiguity and improve personalized medical prognosis. Pers Med 91:101142. 10.1016/j.mam.2022.101142 Weiner SJ (Mar. 2022) Contextualizing care: An essential and measurable clinical competency. Patient Educ Couns 105(3):594–598. 10.1016/j.pec.2021.06.016 Shankar R (2025) Context is King: From Prompt Engineering to Context Engineering in Healthcare AI. Available SSRN 5365971 Yang H et al (Nov. 2017) A framework for exploring associations between biomedical terms in PubMed. Oncotarget 8(61):103100–103107. 10.18632/oncotarget.21532 Schloesser K et al (2023) Interaction of panic and episodic breathlessness among patients with life-limiting diseases: a cross-sectional study, Ann. Palliat. Med. , vol. 12, no. 5, [Online]. Available: https://apm.amegroups.org/article/view/116813 Albuquerque F, Monteiro E, Rodrigues MAB (2023) The Explanatory Factors of Risk Disclosure in the Integrated Reports of Listed Entities in Brazil. Risks 11(6). 10.3390/risks11060108 Jadeja M Jaccard Similarity Made Simple: A Beginner’s Guide to Data Comparison, Medium. [Online]. Available: https://mayurdhvajsinhjadeja.medium.com/jaccard-similarity-34e2c15fb524 Ruiye Z (2024) Volleyball training video classification description using the BiLSTM fusion attention mechanism, Heliyon , vol. 10, no. 15, p. e34735, Aug. 10.1016/j.heliyon.2024.e34735 Salmi M, Atif D, Oliva D, Abraham A, Ventura S (Sep. 2024) Handling imbalanced medical datasets: review of a decade of research. Artif Intell Rev 57(10):273. 10.1007/s10462-024-10884-2 Yan S (2020) Context awareness and embedding for biomedical event extraction. Bioinformatics 36(2):637–643. 10.1093/bioinformatics/btz607 Forbush TB et al (2013) ‘Sitting on pins and needles’: characterization of symptom descriptions in clinical notes., AMIA Jt. Summits Transl. Sci. Proc. AMIA Jt. Summits Transl. Sci. , vol. pp. 67–71, 2013 Alkhawaldeh RS, AlShaqsi J, Al-Ahmad B, Alkhawaldeh SM, Drogham O (Sep. 2025) An attention-based multi-residual and BiLSTM architecture for early diagnosis of autism spectrum disorder. Sci Rep 15(1):33608. 10.1038/s41598-025-19006-6 Wahba Y, Madhavji N, Steinbacher J (2022) Reducing Misclassification Due to Overlapping Classes in Text Classification via Stacking Classifiers on Different Feature Subsets. 406–419. 10.1007/978-3-030-98015-3_28 Herndon J Having Anxiety vs. Feeling Anxious: What’s the Difference? Healthline. [Online]. Available: https://www.healthline.com/health/anxiety/anxiety-vs-anxious Suma P, Chand, Marwaha R, Anxiety (2023) StatPearls Internet , [Online]. Available: https://www.ncbi.nlm.nih.gov/books/NBK470361/ Dolin RH, Alschuler L (Feb. 2011) Approaching semantic interoperability in Health Level Seven. J Am Med Inf Assoc JAMIA 18(1):99–103. 10.1136/jamia.2010.007864 He B, Yang Y, Wang L, Zhou J (2024) The text classification method based on BiLSTM and multi-scale CNN. Comput Life 12(2):43–49 Leo J, Luhanga E, Michael K (2019) Machine Learning Model for Imbalanced Cholera Dataset in Tanzania, Sci. World J. , vol. no. 1, p. 9397578, Jan. 2019. 10.1155/2019/9397578 Johnson JM, Khoshgoftaar TM (Mar. 2019) Survey on deep learning with class imbalance. J Big Data 6(1):27. 10.1186/s40537-019-0192-5 Figeys M, Koubasi F, Hwang D, Hunder A, Miguel-Cruz A, Ríos Rincón A (2023) Challenges and promises of mixed-reality interventions in acquired brain injury rehabilitation: A scoping review, Int. J. Med. Inf. , vol. 179, p. 105235, Nov. 10.1016/j.ijmedinf.2023.105235 Rajput D, Wang W-J, Chen C-C (Feb. 2023) Evaluation of a decided sample size in machine learning applications. BMC Bioinformatics 24(1):48. 10.1186/s12859-023-05156-9 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9087634","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":607195834,"identity":"787ac728-e096-4e52-b562-6d196a5c24bf","order_by":0,"name":"Jael Gudu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA10lEQVRIiWNgGAWjYHACAwYJBgYZfhAzoYAELTySDSAtBsRqAQIegwNwNgEgH5G8TcKixobH+PzqxA8PDBjk+cUO4NdieCOtTELiWBqP2Y23myWADjOcOTuBgJYZOWYSEmyHgVrObgBpSTC4TZSWf/95jGec3fyDKC3yEkAtkm0HeAz4e7cRZ4sBz7NiC8m+ZB6JG7zbLBIMJAj7Rb49eeNtiW92cvz9Zzff/FFhI88vTciWAwwMzBIglgRYpQR+5WBbGhgYGD+AWPwHCKseBaNgFIyCkQkA7CY/vPrxx7cAAAAASUVORK5CYII=","orcid":"","institution":"Makerere University","correspondingAuthor":true,"prefix":"","firstName":"Jael","middleName":"","lastName":"Gudu","suffix":""},{"id":607195840,"identity":"ebb5d6c2-0bf1-4295-b503-6d07970863dc","order_by":1,"name":"Joseph Balikuddembe","email":"","orcid":"","institution":"Makerere University","correspondingAuthor":false,"prefix":"","firstName":"Joseph","middleName":"","lastName":"Balikuddembe","suffix":""},{"id":607195841,"identity":"4d7223cb-e4f1-4dd6-86eb-2d3695aa2e86","order_by":2,"name":"Johnson Mwebaze","email":"","orcid":"","institution":"Makerere University","correspondingAuthor":false,"prefix":"","firstName":"Johnson","middleName":"","lastName":"Mwebaze","suffix":""},{"id":607195842,"identity":"958f62d9-0f6e-40cb-a731-d134784b91ad","order_by":3,"name":"Daniel Opiyo","email":"","orcid":"","institution":"Rongo University","correspondingAuthor":false,"prefix":"","firstName":"Daniel","middleName":"","lastName":"Opiyo","suffix":""}],"badges":[],"createdAt":"2026-03-10 20:38:44","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9087634/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9087634/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104877002,"identity":"c4f43024-c752-4d14-88dd-d89b4062b944","added_by":"auto","created_at":"2026-03-18 08:44:26","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":333739,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe architectural design of the relation extractor that was used to identify and extract the non-taxonomic and ternary relations from the patient-generated texts\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9087634/v1/4e5ddd44e855a8eb077d3d51.png"},{"id":105750012,"identity":"a821dcf9-25ed-4dc9-81d2-0972106ce1cc","added_by":"auto","created_at":"2026-03-30 15:10:47","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1338282,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9087634/v1/050f8732-e2f6-4906-876c-e9d29ac7387b.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Extracting Non-Taxonomic and Ternary Relations from Patient-Generated Texts for Semantic Interoperability","fulltext":[{"header":"Introduction","content":"\u003cp\u003eNatural language texts usually contain useful information, but in an unstructured way, lacking a rigorous internal cell structure, making them unfit for structured databases. However, they have the potential to provide unique value yet they remain untapped, unexplored [1]. Within the architecture of data representation and storage systems, it is crucial to ensure the data in natural language texts is well-organized, easily discoverable, and effectively utilized in its entirety [2].\u0026nbsp;This creates a need for interaction and collaboration with other data storage systems\u0026nbsp;[3]. As a result, systems are increasingly relying on the integration of vast amounts of external data to support comprehensive and accurate decision-making\u0026nbsp;[4]. However, challenges hindering meeting this analytical need include: 1) combining cross-domain data to multiply its value (integration), 2) an awareness of the associations and relations between the different pieces of data\u0026nbsp;[5], [6], [7].\u0026nbsp;Additionally, most of the research work has focused on identifying taxonomic relations, such as broader/narrower, consequently ignoring other relations such as non-hierarchical and structural-level relations that are equally crucial for comprehensive semantic understanding\u0026nbsp;[8], [9].\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMechanisms and approaches have been adopted that address these challenges including deep learning, transformer based approaches (BioBERT), Large Language Frameworks (LLMs), graphical representation of data (knowledge graphs). However, even with the advancements, ternary\u0026nbsp;(N-ary) and\u0026nbsp;non-taxonomic\u0026nbsp;relations\u0026nbsp;and community structural\u0026nbsp;dependencies and interactions remains underexplored, leading to incomplete mappings. These relationships between texts and graph structures\u0026nbsp;are crucial for understanding the complexity of\u0026nbsp;clinical events. The need for scalable, relation-aware ontology learning and integration that can manage intricate relationships across several ontologies\u0026nbsp;is what has motivated this study. The inability of current approaches to capture rich, complex semantic relationships results in uneven and disjointed knowledge integration. Our goal is to improve ontology networks\u0026apos; quality (completeness, expressiveness, and scalability) by creating a comprehensive, relation-aware approach that will make them more resilient and dependable for semantic integration.\u003c/p\u003e\n\u003cp\u003eThe function of the proposed extraction framework is to extract relevant clinical concepts for each stream or input, classify the relations between them into binary, ternary and non-taxonomic, within the relevant domain local network. The framework demonstrated high effectiveness, achieving a test accuracy of 98.91% and a strong recall of 92.8% and F1 Score of 77.6%. This highlights the framework\u0026rsquo;s reliability and robustness in identifying semantically relevant medical concepts existing in real-world noisy patient-generate texts. The framework as it is serves a practical tool for data engineering and semantic data design that can be used to transform real-world unstructured patient-generated text into structured and interoperable semantic knowledge. This addresses the core problem of semantic interoperability of mental health (anxiety and depression) data. Additionally, the framework is best suited for high recall information retrieval and screening of clinical concepts in the real-world context.\u0026nbsp;\u003c/p\u003e"},{"header":"Background and Related Works","content":"\u003cp\u003eIn patient-generated text, semantic relationships are defined based on context. Context, in this case, is defined as the details surrounding a concept that gives details of a particular situation, at a particular instance of time and place [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Within such texts, context is not confined to a single sentence, in some instances spans multiple sentences or an entire document [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. For example, in these sentences\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eAn individual has suffered from\u003c/em\u003e \u003cb\u003eanxiety\u003c/b\u003e \u003cem\u003efor over a year and a half.\u003c/em\u003e \u003cb\u003eIt\u003c/b\u003e \u003cem\u003ehas mainly affected their sleep.\u003c/em\u003e\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe concept \u0026lsquo;\u003cem\u003eanxiety\u0026rsquo;\u003c/em\u003e is very clear within the first sentence, but in the subsequent sentence the \u003cem\u003e\u0026lsquo;it\u0026rsquo;\u003c/em\u003e needs more surrounding concepts (context) for meaning to be understood. This characteristic of patient-generated texts is one of the primary causes of semantic ambiguity and mapping inconsistencies mentioned in Chap.\u0026nbsp;1, thus a barrier to seamless semantic interoperability. Without proper tools and mechanisms that reason and learn context, cross-dependency and co-reference in patient-generated texts fail. This leads to missed concepts, misclassification of concepts, and mapping inaccuracies.\u003c/p\u003e \u003cdiv id=\"Sec2\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Types of Semantic Relationships\u003c/h2\u003e \u003cp\u003eWithin such texts, semantic relations is not confined to explicit expressions, in some instances some are latent in meaning and in words (implicit). They are often difficult to realize as they are hidden in between different sentences spread across a document [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. This study adopts an arity dimensional categorization approach to semantic relationships.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section3\"\u003e \u003ch2\u003e2.1.1 Arity or Degree of Relationship\u003c/h2\u003e \u003cp\u003eThis dimension defines the instances of concepts participating in the relationship in the text [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. It encompasses \u003cem\u003eBinary and N-ary\u003c/em\u003e relations. \u003cem\u003eBinary Relation\u003c/em\u003e are the primary focus of most relation extraction works. \u003cem\u003eN-ary Relations\u003c/em\u003e describing ternary and higher-ary relations where n is greater than two [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. For Example in this sentence,\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eA patient has suffered anxiety for a long time now, it has affected their sleep. The person previously tried cognitive behavioural therapy to help but it failed after a few month. Instead the situation became worse. The person then consulted a doctor who then prescribed mirtazapine. The medication worked effectively and the person was able to sleep. However, the effect of the medication after a few days has been dreadful causing loss of energy. The person is aware that these are some of the symptoms.\u003c/em\u003e \u003c/p\u003e\u003cp\u003eIn the first two sentences, three concepts \u003cem\u003e\u0026ldquo;anxiety, affected sleep, and cognitive behavioural therapy\u0026rdquo;\u003c/em\u003e are participating in a ternary relationship. However, as the sentences increase it moves to an N-ary relationship with 6 concepts involved. These complexities underline the necessity for advanced frameworks capable of handling n-ary relations, as they would also facilitate the creation of more comprehensive and accurate biomedical knowledge graphs, thereby enhancing the level of expressiveness (scope) and effectiveness of integration [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Relation Extraction Techniques for Semantic Relations\u003c/h2\u003e \u003cp\u003eIn various domains such as healthcare, security, software, and education, relations extraction from unstructured texts has largely been driven by deep learning (CNNs and RNNs) and pre-trained language frameworks (BERTs and LLMs) for contextual embedding. They have become the core technology for Named Entity Recognition and knowledge graph construction [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1 Deep Learning Approaches\u003c/h2\u003e \u003cp\u003eThere are various deep learning techniques including Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), Long-Short Term Memory (LSTMs), Gated Recurrent Unit (GRU) [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. These techniques often require more training to match the precision and domain-specific adaptability of rule-based systems. Additionally, they still require substantial computational resources [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Frameworks such as CNN and Long Short-Term Memory (LSTM) shows significant performance over rule-based and dictionary-based approaches in incorporating contextual information. They can be utilized independently or in tandem with explicit semantic domain knowledge. They also have the ability to inference for unspecified relationships, due to their ability to generalize a sentence or across sentences without requiring any patient- or situation-specific training data, thus offering a scalable approach [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWhile CNNs and RNNs are effective at capturing local features within a sentence, they struggle with long-distance dependencies. Despite handling longer sequences better than RNN, LSTMs and BiLSTM face issues with very long-range dependencies. This contributes to the persistent semantic interoperability gap in unstructured texts that have cross-sentences dependencies leading to missed concepts and non-taxonomic relations. For this task a BiLSTM leverages on its sequential processing and labeling richness is well suited for the linear format with clear, concise, logical flow of information found in patient expressions [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Additionally, the BiLSTM leverages on its ability to capture local-feature and mid-range dependencies within efficient timelines based on its low computing resource consumption architecture. A Comparative study results by [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] reveals that in one training, the train epoch of BiLSTM is 178.41 seconds, while a Transformer training epoch is 1657.75 seconds. This makes a BiLSTM suitable for devices and scenes that require low or less computing resource consumption. This creates the need for a more efficient fusion scheme(s) that balances the performance and compute cost.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e2.2.2 Transformer based Approaches\u003c/h2\u003e \u003cp\u003eThe emergence of transformer based approaches including Generative Pre-trained Transformers frameworks such as Bidirectional Encoder Representation from Transformers (BERT) and Large Language Frameworks (LLMs) has revolutionized the concept and relation extraction landscape, with their exceptional state-of-the-art performance in various applications [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Biomedical adaptations of BERT like BioBERT [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] and ClinicalBERT [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] has enabled more accurate interpretation of clinical narratives thus advancing clinical decision support, and disease prediction. These frameworks leverage on the innate capabilities of their architecture consisting of an encoder and a decoder to understand contextual relationships within text [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.2.3 Hybrid Approaches\u003c/h2\u003e \u003cp\u003eThe recent and ongoing integration of deep learning and transformer based frameworks into the relation extraction process represents a flexible and effective architectural combination for tasks requiring both global context and fine-grained sequence processing [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. This is because of their unique combination of deep learning and semantic web technologies. By autonomously learning hierarchical and non-hierarchical representations from large datasets, deep learning and transformer-based framework have drastically improved the quality, robustness, and automation of the knowledge fusion process. [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] used a adopted a transformer and BiLSTM approach for text mining and visualization of underground engineering text reports. [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] combined common NLP techniques with traditional latent thematic analysis (text mining approach) to classify UGC drawn from an interactive HIV mHealth environment. In the instances given the integration of different approaches in the relation extraction process improved the precision and also made full use of the potential value information hidden behind massive text.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Methodology","content":"\u003cp\u003eThe study derives its corpus from online medical forums including: patient.info health forum [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], Health Unlocked [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e], and Mental Health Forum [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], Depression Dataset on Kaggle [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. It was a combination of anxiety disorder (AnxD) from different online forums and open dataset called Depression Dataset. Simple random sampling was used to avoid selection bias on the texts. The dataset was a collection of 38,115 sentences with 27,183 key phrases. From the key phrases 21,746 were used as the training set, and 5,437 as the test set. The division of the dataset is shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSize of datasets\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal Documents\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRelevant Set\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTraining Set\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTesting Set\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProof of concept\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5,822\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3,474\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2,779\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e694\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAnxD and Depression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e38,115\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e27,183\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e21,746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5,437\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Ontology Dictionary Construction\u003c/h2\u003e \u003cp\u003e In the experiment significant words from the corpus in the different variations and synonyms are bootstrapped in a local concept dictionary (a specialized clinical vocabulary), as seed words to act as sematic triggers and anchor terms for each of the categories. This is operationalized by compiling data from the training dataset (sentence and document-level) and was also linked to external concept dictionaries such as SNOMED CT. The local concept dictionaries were then expanded and enriched using 5 WordNet synonyms for each word variant to prevent the explosion of the synonyms. The study opted for WordNet due to the many variants of concepts found in text, especially when identifying the different symptoms variants found in the texts.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Relation Extraction Framework Architecture\u003c/h2\u003e \u003cp\u003eThe BiLSTM\u0026rsquo;s architecture was deliberately designed with relatively low complexity with 2 input layers receiving total parameters 7,019,457 (26.78 MB). The 26.78MB comprised of 4,019,101(15.33MB) trainable parameters and 3,000,356 (11.45MB) non-trainable parameters. The practicability of this is that the total parameter size of 26.78MBs represents a compact neural network architecture. The framework implements a dual-stream feature-fusion architecture with a delayed fusion strategy. This allows the freezing of a larger portion (non-trainable) of the parameters (15.33MB) and focus on the training and learning capacity on the attention fusion. The framework implements a dual-stream feature-fusion architecture with a delayed fusion strategy. This balances the need for contextual neural learning and dictionary and rule-based knowledge (knowledge-infusion method) integration. As seen in Fig.\u0026nbsp;6, the framework receives two inputs \u003cem\u003eseq_input\u003c/em\u003e (sequential text input) and \u003cem\u003edict_feat\u003c/em\u003e that contains engineered features obtained from medical dictionaries, WordNet which are processed independently. The \u003cem\u003eseq_input\u003c/em\u003e undergoes linguistic processing of the text sequence, while \u003cem\u003edict_feat\u003c/em\u003e is prepared for attention weighting.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eBioBERT acted as a biomedical relations precision sieve serving as a cleaner and refiner of the proposed relations, filtering out what was not semantically and factually sound. Adopting the zero-shot validation strategy allowed the model to perform the validation by understanding relational information and semantic similarity based on BioBERT\u0026rsquo;s innate pre-trained knowledge, without the need for additional training [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo assess the performance of the proposed zero-shot inferencing and validation strategy, the model was exposed to relations, which contained several multi-class predicates initially extracted by BiLSTM. In addition, the study simplified the task to binary non-taxonomic and ternary relations validation, and not multi-class relation classification task. Adopting the zero-shot validation strategy allowed the model to perform the validation by understanding relational information and semantic similarity based on BioBERT\u0026rsquo;s innate pre-trained knowledge, without the need for additional training [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSelecting the confidence threshold was an important step in determining and controlling the grounds for accepting a valid relation or rejecting a relation by BioBERT. This is because the choice of threshold ensured the trade-off between having the model capture many relations as possible (recall) and ensuring quality and accurate biomedical relations (precision). In addition, the thresholding mechanism allowed the model to make practical decisions while balancing false positive errors [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. A one-stage threshold setting approach was adopted in this study. For this work, the threshold was set to 0.4, a low threshold that was focused on giving priority to recall and ensuring as many candidate relations as possible were detected and accepted for further validation processing. This threshold was set purely as a cutoff parameter. The goal was to ensure potential candidate relations were not discarded prematurely. At this stage, the model compared each confidence score against the predefined threshold and only relations that were equal to or above this threshold proceeded to the validation process, while those below were left out. The input given to the zero-shot validator were 414 candidate relations of which 384 met the 0.4 threshold and were retained as the final validation output. Of the 414 candidate relations, 30 relations failed to meet the threshold and were discarded.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eGiven an input sentence, the framework employed a hybrid knowledge-infused neural approach to predict a set of candidate relation types meant to guide the subsequent joint entity-relation extractor. The initial extractor, powered by the BiLSTM identified 5089 of the relations comprising ternary relations and non-taxonomic relations. The extraction was based on real-world patient contexts that guided the customized predicate labels used by the BiLSTM. The 5089 relations were then subjected to normalization and deduplication to ensure only clean relationships remained. This led to a significant drop in the total number or relations to 414 non-taxonomic and ternary relations. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the extraction results of the different relations extracted.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eRelations Extracted from Anxiety and Depression Dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eType of relationship\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRelationship Count\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSample relations\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTernary\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003e'treatedWith', 'managesConditionContext', 'relievedInContext'\u003c/em\u003e,\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNon-Taxonomic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e359\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003einteracts_with, similar_to, co_symptom_of, participates_in\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e414\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe 55 ternary relations revealed valuable context-aware associations from the texts. This captured higher levels of relations, especially the ternary, that provided in-depth understanding to different patient experiences, medical conditions and treatments.\u003c/p\u003e \u003cp\u003e \u003cb\u003eRelation Extraction Rules\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe extraction process was governed by the following rules:\u003c/p\u003e \u003cp\u003eFirst, an entity\u0026ndash;relationship triple \u0026lsquo;\u0026lt;h, r, t\u0026gt;\u0026rsquo; was only considered valid and sound if its relation and the concept names of the head entity and tail entity were accurate. Secondly, when the candidate relations identified did not fit the concepts in the sentence, the extractor did not force a relation to be selected from the incorrect candidate relations, but instead assumed that no relation exists between those entities. For example, given a sentence below with no relations between \u003cem\u003eanxiety\u003c/em\u003e, \u003cem\u003eanxiety disorder\u003c/em\u003e and the words \u003cem\u003esadly, and treatment\u003c/em\u003e in the biomedical context the framework should assume there is no relation between them.\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eAccording to a national survey in 2010, 32 percent of adolescents in the United States have an anxiety disorder.\u003c/em\u003e[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] \u003cem\u003eSadly, majority of children with anxiety never receive treatment\u003c/em\u003e, [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThis was meant to ensure that only logical and meaningful relations between concepts were included in the ontology. It also aimed at formalizing domain knowledge and establishing logical constraints within the ontology.\u003c/p\u003e \u003cp\u003eTo illustrate extracted binary non-taxonomic relations the following example is used:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e(fatigue \u0026rarr; anxiety, \u003cem\u003esuggests\u003c/em\u003e) expressed in triple as ('fatigue', 'suggests', 'anxiety').\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eConsider this relevant sentence subset,\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eThe brain gets deprived of glucose; you develop fatigue, mind-fog, dizziness, and as the process progresses you can develop restlessness as the brain searches for natural alternate energy sources such as adrenaline, which can then cause anxiety and even panic attacks.\u003c/em\u003e \u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn this example, framework performed a sentence-level extraction and found an association \u003cem\u003esuggests\u003c/em\u003e between a concept fatigue and domain concepts anxiety and panic attacks. This means the existence of fatigue in a sentence suggests a medical condition panic attacks or anxiety.\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e(anxiety \u0026rarr; sadness, \u003cem\u003ehasSymptom\u003c/em\u003e) expressed in triple as ('anxiety', 'sadness', 'hasSymptom') or ('anxiety', 'hasSymptom', 'sadness')\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e(depression \u0026rarr; fluoxetine, \u003cem\u003etreatedWith\u003c/em\u003e)\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Relation Validation Results\u003c/h2\u003e \u003cp\u003eEach candidate relation was presented to the BioBERT for semantic and factual validation, specifically within the patient expressions on anxiety and depression disorders. As a result, from the candidate relations (ternary and non-taxonomic), the BioBERT powered validation identified 384 relations, with confidence values ranging from 75\u0026ndash;98%. The validation results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSummary of Biomedical validated Relations\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eType of relation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBiLSTM Relations\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBioBERT Validated Relations\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSample validated relations\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTernary\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e240\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003e{'symptom': 'low energy', 'disease': 'anxiety', 'treatment': 'negative thought challenge', 'relation': 'managesConditionContext'}\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNon-taxonomic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e359\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e144\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eDepression, treatedWith, fluoxetine\u003c/em\u003e,\u003c/p\u003e \u003cp\u003e\u003cem\u003efatigue, suggests,anxiety\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eTotal\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e414\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e384\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eValidation rate:92.7%\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe validation results reveal distinctions between the two categories of relations. The BioBERT validated all the ternary relations (100% validation rate) initially extracted by BiLSTM. However, due to the pre-training on biomedical literature BioBERT was able to capture more valid ternary relations contexts, therefore identifying 240 ternary relations. This was 185 more entries than what the BiLSTM had extracted. This does not indicate a failure by the BiLSTM, but instead confirms that BiLSTM was recall focused, and was limited in precision (section 5.3.5). The 100% validation rate shows that the BiLSTM was performed significantly well in extracting complex relations from the unstructured text; that align to existing biomedical knowledge. In contrast, BioBERT validated 144 of the 359 non-taxonomic relations, with a 40.1% validation rate.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Ternary Relation Extraction\u003c/h2\u003e \u003cp\u003eUnlike binary non-taxonomic associations that seek general connections between concepts, in ternary relation structure the relation covers a multi-concept clinical context. The multi-concept clinical context here refers to clinical integral truth conditions (with all the concepts from all the three classes, disease, symptom and treatment participating). This context justified the specific situation why, where and when the relation holds true, without which the relation is invalid.\u003c/p\u003e \u003cp\u003eThe validated ternary relations had a consistent pattern and structure that comprised of a subject concept, relation (explaining the nature of the link), object concept with the context attached to it. This was expressed as:\u003c/p\u003e \u003cp\u003e \u003cem\u003econcept, associated_with, concept\u003c/em\u003e | \u003cem\u003e{clinical context}\u003c/em\u003e\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;10 explicitly demonstrates this ternary structure from the text given.\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eA person has suffered anxiety for a long time now, it has affected their sleep. They had. previously tried cognitive behavioural therapy to help but it failed after a short while. Instead, the situation became worse. The person then consulted a doctor who prescribed mirtazapine the person\u0026rsquo;s first time on such a medication. The medication worked effectively improving their sleep. However, the effect of the medication after a few days has been dreadful causing loss of energy. The person is aware that these are some of the symptoms.\u003c/em\u003e \u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e shows the sample ternary relations.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSample Results of Ternary Relations with clinical context. The table shows the ternary relations based on subject, objects and the relations between them. It also shows the types and classes where the subjects belong and the contextual qualifiers that make the ternary relation hold between the subject and the object.\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubject Concept\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSubject type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRelation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eObject Concept\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eObject type\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eAnxiety\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eDisease\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003eaffects\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003esleep\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cem\u003eSymptom\u003c/em\u003e | \u003cem\u003e{for a long time\u0026thinsp;\u0026gt;\u0026thinsp;1 year, treated using CBT (treatment)}\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eloss of energy\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eSymptom\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003ecaused_by\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003emirtazapine\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cem\u003eTreatment\u003c/em\u003e | \u003cem\u003e{Treated anxiety (disease) temporarily, dreadful effect after few days}\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eAnxiety\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eDisease\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003etreated_by\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eCognitive Behavioural Therapy (CBT)\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cem\u003eTreatment\u003c/em\u003e | \u003cem\u003e{unable to sleep (symptom), for a long time}\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eAnxiety\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eDisease\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cem\u003etreated_by\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003emirtazapine\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cem\u003eTreatment\u003c/em\u003e | \u003cem\u003e{improved sleep, temporarily effective}\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e \u003c/p\u003e\u003cp\u003eFrom the text, a document-level extraction reveals more than one ternary relations as seen in table 5.4.3.2. These are based on the structure form:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eAnxiety (Disease), treated_by, mirtazapine (treatment)\u003c/em\u003e | \u003cem\u003e{long time with anxiety, prior CBT failure (treatment), affected sleep (symptom)}\u003c/em\u003e\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cem\u003eMirtazapine (treatment), causes, loss of energy (symptom)\u003c/em\u003e | \u003cem\u003e{used for a short while, offered temporary relief for anxiety (disease), delayed dreadful effect}\u003c/em\u003e\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eIn the two ternary relations, the framework resolves a contradiction that \u003cem\u003eMirtazapine (treatment)\u003c/em\u003e can bring relief and can cause loss of energy that would arise if they were expressed purely as binary non-taxonomic relations such as\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eMirtazapine (treatment), causes, loss of energy\u003c/em\u003e \u003c/p\u003e\u003cp\u003e \u003cem\u003eMirtazapine (treatment), relieves, anxiety\u003c/em\u003e \u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e Any interaction with the two binary non-taxonomic associations would not know how to distinguish these two relations if context was not provided. The context here include past medical history (using CBT to treat anxiety), temporal evidence of the treatment offering relief for a short while, and also the symptom evolves after a short while. These extracted ternary relations show that richer information and contexts can be discovered by with the aim of supporting clinical inferencing and integration with existing domain ontologies.\u003c/p\u003e \u003cp\u003eThe 240 ternary relations preserve clinical contextual structures during semantic integration address a semantic interoperability challenge: the ambiguity of isolated clinical facts. Here, the study defines ambiguity of clinical fact as a situation that a concept is found in different conditions [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. For example, information such as \u003cem\u003eMirtazapine (treatment), causes, loss of energy\u003c/em\u003e may lead to several interpretations without this other piece of information \u003cem\u003eMirtazapine (treatment), relieves, anxiety\u003c/em\u003e and the reverse is also true. Therefore the ternary structure \u003cem\u003eloss of energy, causedBy, mirtazapine\u003c/em\u003e | \u003cem\u003e{Treated anxiety (disease) temporarily, dreadful effect after few days}\u003c/em\u003e resolves these ambiguity. Clinical decisions occur within complex interrelated information including diagnosis results, patient histories, pharmacological interactions, and social and psychological factors [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. Any system supporting a clinical or patient monitoring decision is equipped with the context (all relevant and valid information) to guide its reasoning towards a semantically consistent interpretation across different health systems.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Confidence Score Analysis\u003c/h2\u003e \u003cp\u003eThis section analyses the model\u0026rsquo;s level of certainty in the relations it has validated. Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e gives the confidence range for the relations together with average confidence for the different categories.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSummary of the Confidence Scores per Relation Category\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRelation Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eConfidence Range\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAverage Confidence (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAssociative\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.917\u0026ndash;0.984\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFunctional\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.697\u0026ndash;0.965\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e88\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCausal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.790\u0026ndash;0.962\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e92\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRisk Factors\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.704\u0026ndash;0.947\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStatistical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.608\u0026ndash;0.851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eOverall Average Confidence\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e87.4\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe results reveal that all the confidence scores were above the 0.4 threshold that was set. The lowest score was 0.60, which was still higher than the threshold. This shows that the 0.4 threshold cast a wider net, prioritizing recall without compromising the quality of the relations. An analysis of the results reveals that over 90.5% of the validated relations were above 90%, which indicates that BioBERT zero-shot validator captured high quality relations that existed within the pre-existing biomedical knowledge, confirming its role in the framework as a precision focused optimizer.\u003c/p\u003e \u003cp\u003eThe highest confidence score range was 0.917\u0026ndash;0.984, assigned to \u0026lsquo;\u003cem\u003eassociatedWith\u0026rsquo;\u003c/em\u003e. This indicates that BioBERT was highly confident in validating this type of non-taxonomic relations between different disease, symptom and treatment concepts. This high confidence score can be attributed to co-occurrence relationships, which accelerate the recognition of term associations, has been well documented in biomedical literature [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. For example, the co-occurrence of anxiety and panic, with respiratory challenges (chronic breathlessness) have been well documented [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. The consistent high confidence score within the 54 associative relation category made them a more reliable semantic link within the knowledge graph.\u003c/p\u003e \u003cp\u003eThe relatively low confidence scores largely for statistical relations (0.60) and risk factor relations (0.70) suggests relations that are indicative and explain factors that influence the occurrence of disorders, but might not be the final or controlling factor to the existence of the disorder [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. The moderate confidence scores (0.79\u0026ndash;0.85) predominantly for causal and some statistical relations, indicated that these relations were reliable for inclusion in the knowledge graph, but might require more evidence and verification to increase the level of certainty.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e4.4 Performance Evaluation\u003c/h2\u003e \u003cp\u003eWhen compared to the BiLSTM results (recall\u0026thinsp;=\u0026thinsp;75.6%, precision\u0026thinsp;=\u0026thinsp;61.3%, and F1 score\u0026thinsp;=\u0026thinsp;67.7%), the hybrid framework achieved a 10% improvement on the F1 score, a 17.2% improvement on the recall and a 5.4% improvement on precision. Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows a summary of the performance evaluation. The improvement on the precision and recall indicates the significance of infusing domain knowledge and BioBERT to the framework. The framework focused on optimizing the precision and ensuring the relations were semantically relevant and valid to the biomedical domain.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSummary of the overall performance of the framework\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003eDataset\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c4\" namest=\"c2\"\u003e \u003cp\u003ePerformance\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"3\" nameend=\"c7\" namest=\"c5\"\u003e \u003cp\u003eHybrid Completeness and Consistency Check\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003ePrecision (%)\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eRecall (%)\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eF1-Score (%)\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eHamming loss\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003eJaccard\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAnxD and Depression Dataset\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e66.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e92.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e77.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.029\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.667\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe Jaccard in this case was used context of relations validation to confirm what was currently existing in biomedical literature and what BiLSTM had extracted. The Jaccard of 0.667, a value closer to 1 means the predicates and true relation labels existing in the domain were nearly identical [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e], resulting in reliable relations. This mean the results produced by the model are reliable\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussions","content":"\u003cp\u003eThe hybrid knowledge infused neural approach, integrating domain knowledge with a BiLSTM framework fused with an attention mechanism demonstrated high levels of efficacy in the extraction and classification of complex medical relations found in free texts. In experimental evaluation, the hybrid framework achieved an accuracy of 98.91% on the test set, a significant improvement over traditional RNNs and standalone BiLSTM algorithms [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe results of the experimental evaluation indicate that the developed framework achieved, a significant improvement in accuracy over traditional RNNs and standalone BiLSTM algorithms on the test dataset. The framework produces the best results in terms of recall and F1 score highlighting its reliability and robustness in handling unstructured test dataset. A knowledge graph representing the data was created using the hybrid union strategy. An analysis of the graph reveals highly interconnected concepts within the graph with a high density and average degree. This consequently provides a qualitative foundation to evaluate of the knowledge contained in the graph and the reality the knowledge represents.\u003c/p\u003e"},{"header":"Limitations and Future Works","content":"\u003cp\u003eAn inherent challenge encountered during the study was the issue of class imbalance in medical datasets, which complicated accurate concept prediction and classification capabilities of the framework. Class imbalance occurs when the training dataset contains classes that are unequally distributed [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e], [\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. This arose due to the nature of the patient-generated texts used. This type of data naturally represents the patient\u0026rsquo;s subjective experience that captured the voice of the patient when describing medical situations, including figures of speech, idioms, or lay terms [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. This bias is confirmed by the resulting class distribution where symptoms (53 concepts) was the majority (prevalent) class, followed by treatment (40 concepts), while disease (18) was the minority (rare) class. This mirrored how often the balance is unreached in real-world datasets. With this type of distribution, it was expected that the framework would have failed to make meaningful predictions for the minority classes [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]. The F1 score results of 0.789 for treatment, 0.768 for symptoms, and 0.657 for disease, confirmed the expectation and showed the difficulty the framework faced in accomplishing its objective of accurately predicting the minority class of interest.\u003c/p\u003e \u003cp\u003eAnother significant challenge experienced in this study was semantic overlap and ambiguity. A situation that arises when certain classes tend to share the same vocabulary, resulting in misclassification of domain concepts [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. This limitation contributed to the low precision especially in disease category (47%), an indication misclassification, despite correctly identifying the broader context.\u003c/p\u003e \u003cp\u003eFor example,\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cem\u003eDataset 1: When we experience an involuntary high degree stress response, the symptoms can be so profound that we think we are having a medical emergency\u003c/em\u003e, \u003cb\u003ewhich anxious personalities react to with more fear\u003c/b\u003e. \u003cem\u003eAnd when we become more afraid, the body is going to produce another stress response, which causes more changes and symptoms\u003c/em\u003e, \u003cb\u003ewhich we can react to with more fear\u003c/b\u003e, \u003cem\u003eexperience more symptoms, and so on.\u003c/em\u003e\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eFrom the excerpt \u0026ldquo;\u003cem\u003ewhich anxious personalities react to with more fear\u0026rdquo;\u003c/em\u003e, the word \u0026ldquo;anxious\u0026rdquo; describing a temporary feeling of fear and worry about something [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e] was reasoned out and misclassified as a possible disease class \u0026ldquo;anxiety\u0026rdquo;. The disease concept anxiety is linked to being fearful and anxious in different circumstances over time [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe framework\u0026rsquo;s error was occasioned by the lack temporal pattern within the specific excerpt to give the contextual guidance for an appropriate classification decision. Because of this semantic ambiguity, the framework extracted the concept \u0026ldquo;anxious\u0026rdquo; and classified it within the disease category. This was because the concept has a strong lexical association to anxiety (existence of fear) while overriding the context that indicated a symptom in this case. Such cases of semantic overlap are common in the medical domain, a problem that this framework did not resolve fully. Potential solutions include adding clear descriptive text and framework refinement [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. However, due to the broad nature of the patient expressions within the texts exploring these complex solutions can be considered for further work. Refining relation directionality will be a key area of focus in future studies building upon this foundational work. This is to ensure ontological consistency to prevent semantic and inferencing errors, conflicts, lower quality, and knowledge inconsistencies that arise from poorly integrated information. Extending this study by developing a model to improve the BioBERT fine-tuning classification and prediction performance and to compare the new results to the present ones. Based on the promising results, the study was inspired and aim to adapt the proposed framework to patient contexts that go beyond pure clinical associations. The key question to be answered is: would the experiment improve the scope of BiOBERT validated relations beyond the 393 if pure patient contexts are to be factored in?\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThe results shown make the framework useful for knowledge extraction tasks requiring comprehensive semantic coverage. This is critical in ensuring information needed for: knowledge representation, informed decision making, semantic dependency building is available. In such applications and tasks, overlooking key concepts can be more costly than retrieving occasional false positives making the trade-off a justifiable choice for this study.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors reviewed the manuscript\u003c/p\u003e\u003cp\u003eFunding\u003c/p\u003e\n\u003cp\u003eThis work was supported by METEGA.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eEthics declarations\u003c/p\u003e\n\u003cp\u003eCompeting Interests\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003eEthics Approval and Consent to Participate\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003eConsent for Publication\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eKhine PP, Wang Z, Shun (2018) Data lake: a new ideology in big data era. ITM Web Conf 17:03025. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1051/itmconf/20181703025\u003c/span\u003e\u003cspan address=\"10.1051/itmconf/20181703025\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaseem O Metadata management in data lakes, Data Science Central. Accessed: Oct. 29, 2024. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.datasciencecentral.com/metadata-management-in-data-lakes/\u003c/span\u003e\u003cspan address=\"https://www.datasciencecentral.com/metadata-management-in-data-lakes/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen Z, Wan Y, Liu Y, Valera-Medina A (Jan. 2024) A knowledge graph-supported information fusion approach for multi-faceted conceptual modelling. Inf Fusion 101:101985. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.inffus.2023.101985\u003c/span\u003e\u003cspan address=\"10.1016/j.inffus.2023.101985\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRousanb W Leveraging Big Data for Enhanced Government Security: A Deep Dive into Use Cases, Datahub Analytics. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://datahubanalytics.com/leveraging-big-data-for-enhanced-government-security/#:~:text=Big%20Data%20enables%20the%20integration%20and%20an\nalysis,resource%20allocation%2\nC%20risk%20assessment%2C%2\n0and%20decision%2Dmaking%20p\nrocesses\u003c/span\u003e\u003cspan address=\"https://datahubanalytics.com/l\neveraging-big-data-for-enhanced-gov\nernment-security/#:~:text=Big%20Data%\n20ena\nbles%20the%20integration%20and%2\n0analysis,resource%20allocation%2C%\n20risk%20assessment%2C%20and%20\ndecision%2Dmaking%20processes\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoukaras P (2025) Data Integration and Storage Strategies in Heterogeneous Analytical Systems: Architectures, Methods, and Interoperability Challenges. Information 16(11). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/info16110932\u003c/span\u003e\u003cspan address=\"10.3390/info16110932\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoslemi MH, Mousavi A, Behkamal B, Milani M Heterogeneity in Entity Matching: A Survey and Experimental Analysis. 2025. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://arxiv.org/abs/2508.08076\u003c/span\u003e\u003cspan address=\"https://arxiv.org/abs/2508.08076\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSubramaniam P, Ma Y, Li C, Mohanty I, Fernandez RC Comprehensive and Comprehensible Data Catalogs: The What, Who, Where, When, Why, and How of Metadata Management, \u003cem\u003eArXiv\u003c/em\u003e, vol. abs/2103.07532, 2021, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://api.semanticscholar.org/CorpusID:232233695\u003c/span\u003e\u003cspan address=\"https://api.semanticscholar.org/CorpusID:232233695\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKolli M (Dec. 2023) Formal modeling of multi-viewpoint ontology alignment by mappings composition. Acta Univ Sapientiae Inf 15:181\u0026ndash;204. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2478/ausi-2023-0013\u003c/span\u003e\u003cspan address=\"10.2478/ausi-2023-0013\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOsman I, Pileggi SF, Ben Yahia S, Diallo G (2021) An Alignment-Based Implementation of a Holistic Ontology Integration Method. MethodsX 8:101460. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.mex.2021.101460\u003c/span\u003e\u003cspan address=\"10.1016/j.mex.2021.101460\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKowalski P, Jousselme A-L (2023) Context-awareness for information correction and reasoning in evidence theory. Int J Approx Reason 153:29\u0026ndash;48. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ijar.2022.11.009\u003c/span\u003e\u003cspan address=\"10.1016/j.ijar.2022.11.009\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXie T et al (2024) ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity, in., \u003cem\u003eIEEE International Conference on Data Mining Workshops (ICDMW)\u003c/em\u003e, Los Alamitos, CA, USA: IEEE Computer Society, Dec. 2024, pp. 907\u0026ndash;911. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ICDMW65004.2024.00126\u003c/span\u003e\u003cspan address=\"10.1109/ICDMW65004.2024.00126\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlam R Understanding Semantic Relations: Types, Examples, and Applications, Linkedin. Accessed: Aug. 02, 2025. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.linkedin.com/pulse/understanding-semantic-relations-types-examples-md-rabby-alam-ndewf/\u003c/span\u003e\u003cspan address=\"https://www.linkedin.com/pulse/understanding-semantic-relations-types-examples-md-rabby-alam-ndewf/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZemla JC (2022) Knowledge representations derived from semantic fluency data. Front Psychol 13:815860\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLentschat M, Buche P, Dibie-Barthelemy J, Roche M (Dec. 2022) A new method to extract n-Ary relation instances from scientific documents. Expert Syst Appl 209:118332. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.eswa.2022.118332\u003c/span\u003e\u003cspan address=\"10.1016/j.eswa.2022.118332\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFernandes JLM (2025) Extracting n-ary Relations from Biomedical Literature using Deep-Learning Techniques, PhD Thesis, UNIVERSIDADE DE LISBOA\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChoi S, Jung Y (2025) Knowledge Graph Construction: Extraction, Learning, and Evaluation. Appl Sci 15(7). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/app15073727\u003c/span\u003e\u003cspan address=\"10.3390/app15073727\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Luan Z, Liu Y, Liu H, Qi J, Han D (2024) Automated information extraction model enhancing traditional Chinese medicine RCT evidence extraction (Evi-BERT): algorithm development and validation. Front Artif Intell 7:1454945. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/frai.2024.1454945\u003c/span\u003e\u003cspan address=\"10.3389/frai.2024.1454945\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZengeya T, Fonou Dombeu JV (Jan. 2024) A Review of State of the Art Deep Learning Models for Ontology Construction. IEEE Access 1\u0026ndash;1. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2024.3406426\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2024.3406426\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQian XX, Chau PH, Fong DYT, Ho M, Woo J (2025) Development and Validation of a Rule-Based Natural Language Processing Algorithm to Identify Falls in Inpatient Records of Older Adults: Retrospective Analysis, \u003cem\u003eJMIR Aging\u003c/em\u003e, vol. 8, p. e65195, Jul. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/65195\u003c/span\u003e\u003cspan address=\"10.2196/65195\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi J, Chen B, Jiang H (Sep. 2020) Fully-connected LSTM\u0026ndash;CRF on medical concept extraction. Int J Mach Learn Cybern 11. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s13042-020-01087-6\u003c/span\u003e\u003cspan address=\"10.1007/s13042-020-01087-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLei H (2025) The Difference between BiLSTM and Transformer-The Discussion about Suitable Scenarios for BiLSTM and Transformer. Sci Technol Eng Chem Environ Prot, 1, 1\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCho HN et al (2024) Task-specific transformer-based language models in health care: scoping review. JMIR Med Inf 12:e49724\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee J et al (2020) BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4):1234\u0026ndash;1240\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang G et al (2023) Optimized glycemic control of type 2 diabetes with reinforcement learning: a proof-of-concept trial. Nat Med 29(10):2633\u0026ndash;2642\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVisweswaraiah V, Everyday AI (Sep. 2025) Real-World Applications of Transformer- Based Language Models. Int J Comput Trends Technol 73:19\u0026ndash;27. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.14445/22312803/IJCTT-V73I9P103\u003c/span\u003e\u003cspan address=\"10.14445/22312803/IJCTT-V73I9P103\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMienye ID, Swart TG (2024) A Comprehensive Review of Deep Learning: Architectures, Recent Advances, and Applications. Information 15(12). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/info15120755\u003c/span\u003e\u003cspan address=\"10.3390/info15120755\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZubair M, Hussain M, Albashrawi MA, Bendechache M, Owais M (2025) A comprehensive review of techniques, algorithms, advancements, challenges, and clinical applications of multi-modal medical image fusion for improved diagnosis, \u003cem\u003eComput. Methods Programs Biomed.\u003c/em\u003e, vol. 272, p. 109014, Dec. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cmpb.2025.109014\u003c/span\u003e\u003cspan address=\"10.1016/j.cmpb.2025.109014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShao R, Lin P, Xu Z (Oct. 2024) Integrated natural language processing method for text mining and visualization of underground engineering text reports. Autom Constr 166:105636. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.autcon.2024.105636\u003c/span\u003e\u003cspan address=\"10.1016/j.autcon.2024.105636\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSkeen SJ, Jones SS, Cruse CM, Horvath KJ (2022) Integrating Natural Language Processing and Interpretive Thematic Analyses to Gain Human-Centered Design Insights on HIV Mobile Health: Proof-of-Concept Analysis., \u003cem\u003eJMIR Hum. Factors\u003c/em\u003e, vol. 9, no. 3, p. e37350, Jul. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/37350\u003c/span\u003e\u003cspan address=\"10.2196/37350\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHealth N, Navigate Health Anxiety disorders,. Accessed: Jul. 26, 2025. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://community.patient.info/tag/anxiety%20disorders/\u003c/span\u003e\u003cspan address=\"https://community.patient.info/tag/anxiety%20disorders/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHealthUnlocked Anxiety and Depression Support, HealthUnlocked. Accessed: Jul. 26, 2025. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://healthunlocked.com/anxiety-depression-support\u003c/span\u003e\u003cspan address=\"https://healthunlocked.com/anxiety-depression-support\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTogether 4 Change Limited Mental Health Forum. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.mentalhealthforum.net/forum/threads/do-you-experience-anxiety-approved-research.394720/\u003c/span\u003e\u003cspan address=\"https://www.mentalhealthforum.net/forum/threads/do-you-experience-anxiety-approved-research.394720/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFran\u0026ccedil;a DS (2024) Depression Dataset. Kaggle, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kaggle.com/datasets/diegosilvadefrana/depression-dataset\u003c/span\u003e\u003cspan address=\"https://www.kaggle.com/datasets/diegosilvadefrana/depression-dataset\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDong X, Zhao D, Meng J, Guo B, Lin H (2025) SyRACT: zero-shot biomedical document-level relation extraction with synergistic RAG and CoT. Bioinformatics 41(7):btaf356\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWhite H Computer Vision Basics: Confidence \u0026amp; Accuracy, Leverage. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.leverege.com/blogpost/computer-vision-basics-how-confidence-accuracy-and-thresholds-impact-performance#:~:text=Confidence%20Threshold%20Trade%2Doffs,threshold%20depends%20on%20the%20application\u003c/span\u003e\u003cspan address=\"https://www.leverege.com/blogpost/computer-vision-basics-how-confidence-accuracy-and-thresholds-impact-performance#:~:text=Confidence%20Threshold%20Trade%2Doffs,threshold%20depends%20on%20the%20application\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMerikangas KR et al (2010) Lifetime prevalence of mental disorders in U.S. adolescents: results from the National Comorbidity Survey Replication\u0026ndash;Adolescent Supplement (NCS-A)., \u003cem\u003eJ. Am. Acad. Child Adolesc. Psychiatry\u003c/em\u003e, vol. 49, no. 10, pp. 980\u0026ndash;989, Oct. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jaac.2010.05.017\u003c/span\u003e\u003cspan address=\"10.1016/j.jaac.2010.05.017\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeir K, American Psychological Association Brighter futures for anxious kids,. Accessed: Feb. 08, 2026. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.apa.org/monitor/2017/03/anxious-kids\u003c/span\u003e\u003cspan address=\"https://www.apa.org/monitor/2017/03/anxious-kids\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLibertin CR et al (Jun. 2023) Data structuring may prevent ambiguity and improve personalized medical prognosis. Pers Med 91:101142. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.mam.2022.101142\u003c/span\u003e\u003cspan address=\"10.1016/j.mam.2022.101142\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeiner SJ (Mar. 2022) Contextualizing care: An essential and measurable clinical competency. Patient Educ Couns 105(3):594\u0026ndash;598. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.pec.2021.06.016\u003c/span\u003e\u003cspan address=\"10.1016/j.pec.2021.06.016\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShankar R (2025) Context is King: From Prompt Engineering to Context Engineering in Healthcare AI. Available SSRN \u003cem\u003e5365971\u003c/em\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang H et al (Nov. 2017) A framework for exploring associations between biomedical terms in PubMed. Oncotarget 8(61):103100\u0026ndash;103107. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.18632/oncotarget.21532\u003c/span\u003e\u003cspan address=\"10.18632/oncotarget.21532\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchloesser K et al (2023) Interaction of panic and episodic breathlessness among patients with life-limiting diseases: a cross-sectional study, \u003cem\u003eAnn. Palliat. Med.\u003c/em\u003e, vol. 12, no. 5, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://apm.amegroups.org/article/view/116813\u003c/span\u003e\u003cspan address=\"https://apm.amegroups.org/article/view/116813\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlbuquerque F, Monteiro E, Rodrigues MAB (2023) The Explanatory Factors of Risk Disclosure in the Integrated Reports of Listed Entities in Brazil. Risks 11(6). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/risks11060108\u003c/span\u003e\u003cspan address=\"10.3390/risks11060108\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJadeja M Jaccard Similarity Made Simple: A Beginner\u0026rsquo;s Guide to Data Comparison, Medium. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://mayurdhvajsinhjadeja.medium.com/jaccard-similarity-34e2c15fb524\u003c/span\u003e\u003cspan address=\"https://mayurdhvajsinhjadeja.medium.com/jaccard-similarity-34e2c15fb524\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRuiye Z (2024) Volleyball training video classification description using the BiLSTM fusion attention mechanism, \u003cem\u003eHeliyon\u003c/em\u003e, vol. 10, no. 15, p. e34735, Aug. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.heliyon.2024.e34735\u003c/span\u003e\u003cspan address=\"10.1016/j.heliyon.2024.e34735\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalmi M, Atif D, Oliva D, Abraham A, Ventura S (Sep. 2024) Handling imbalanced medical datasets: review of a decade of research. Artif Intell Rev 57(10):273. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10462-024-10884-2\u003c/span\u003e\u003cspan address=\"10.1007/s10462-024-10884-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan S (2020) Context awareness and embedding for biomedical event extraction. Bioinformatics 36(2):637\u0026ndash;643. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btz607\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btz607\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eForbush TB et al (2013) \u0026lsquo;Sitting on pins and needles\u0026rsquo;: characterization of symptom descriptions in clinical notes., \u003cem\u003eAMIA Jt. Summits Transl. Sci. Proc. AMIA Jt. Summits Transl. Sci.\u003c/em\u003e, vol. pp. 67\u0026ndash;71, 2013\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlkhawaldeh RS, AlShaqsi J, Al-Ahmad B, Alkhawaldeh SM, Drogham O (Sep. 2025) An attention-based multi-residual and BiLSTM architecture for early diagnosis of autism spectrum disorder. Sci Rep 15(1):33608. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-025-19006-6\u003c/span\u003e\u003cspan address=\"10.1038/s41598-025-19006-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWahba Y, Madhavji N, Steinbacher J (2022) Reducing Misclassification Due to Overlapping Classes in Text Classification via Stacking Classifiers on Different Feature Subsets. 406\u0026ndash;419. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/978-3-030-98015-3_28\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-98015-3_28\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHerndon J Having Anxiety vs. Feeling Anxious: What\u0026rsquo;s the Difference? Healthline. [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.healthline.com/health/anxiety/anxiety-vs-anxious\u003c/span\u003e\u003cspan address=\"https://www.healthline.com/health/anxiety/anxiety-vs-anxious\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuma P, Chand, Marwaha R, Anxiety (2023) \u003cem\u003eStatPearls Internet\u003c/em\u003e, [Online]. Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/books/NBK470361/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/books/NBK470361/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDolin RH, Alschuler L (Feb. 2011) Approaching semantic interoperability in Health Level Seven. J Am Med Inf Assoc JAMIA 18(1):99\u0026ndash;103. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/jamia.2010.007864\u003c/span\u003e\u003cspan address=\"10.1136/jamia.2010.007864\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe B, Yang Y, Wang L, Zhou J (2024) The text classification method based on BiLSTM and multi-scale CNN. Comput Life 12(2):43\u0026ndash;49\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeo J, Luhanga E, Michael K (2019) Machine Learning Model for Imbalanced Cholera Dataset in Tanzania, \u003cem\u003eSci. World J.\u003c/em\u003e, vol. no. 1, p. 9397578, Jan. 2019. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1155/2019/9397578\u003c/span\u003e\u003cspan address=\"10.1155/2019/9397578\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJohnson JM, Khoshgoftaar TM (Mar. 2019) Survey on deep learning with class imbalance. J Big Data 6(1):27. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s40537-019-0192-5\u003c/span\u003e\u003cspan address=\"10.1186/s40537-019-0192-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFigeys M, Koubasi F, Hwang D, Hunder A, Miguel-Cruz A, R\u0026iacute;os Rinc\u0026oacute;n A (2023) Challenges and promises of mixed-reality interventions in acquired brain injury rehabilitation: A scoping review, \u003cem\u003eInt. J. Med. Inf.\u003c/em\u003e, vol. 179, p. 105235, Nov. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ijmedinf.2023.105235\u003c/span\u003e\u003cspan address=\"10.1016/j.ijmedinf.2023.105235\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRajput D, Wang W-J, Chen C-C (Feb. 2023) Evaluation of a decided sample size in machine learning applications. BMC Bioinformatics 24(1):48. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12859-023-05156-9\u003c/span\u003e\u003cspan address=\"10.1186/s12859-023-05156-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Knowledge-Infused neural framework, ontology integration, non-taxonomic, ternary, semantic interoperability","lastPublishedDoi":"10.21203/rs.3.rs-9087634/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9087634/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePatient-generated texts usually contain useful information, but lack rigorous internal cell structure making them unfit for structured databases. Current research work has focused on identifying hierarchical/taxonomic relations, consequently ignoring non-hierarchical and ternary relations, which are equally crucial for comprehensive semantic understanding. This study addresses this through semantic alignment based on non-taxonomic and ternary relational components. The work adopts a Design Science Research (DSR) approach, with a pragmatic research philosophy. We develop and evaluate a knowledge-infused neural framework for cross-domain ontology integration that supports capturing and representing non-taxonomic and ternary relationships beyond general hierarchical relations. The framework adopts a four-layered architecture. A key contribution is the implementation of the delayed fusion strategy that balances the need for contextual neural learning, the interpretable rule-based dictionary knowledge, and incorporating BioBERT as a relations validator to ensure domain factual grounding in the integration. The framework was evaluated on 38,115 documents from anxiety and depression datasets, of which 27,183 were key phrases. The hybrid per class adaptive strategy extracted 113 unique concepts, prioritizing a more conservative prediction. The hybrid union extracted 222 unique concepts prioritizing wider coverage of domain concepts for the construction of the knowledge graph. The framework achieved an accuracy of 98.91% and an F1 score of 77.6%, a 10% improvement compared to the BiLSTM F1 score (67.7%). The framework also validated 384 semantic relations with a validation rate of 92.7%. Of the semantic relations validated, 240 were ternary relations that captured multiple contexts of interactions between the concepts. Non-taxonomic relations were 144 in total, organized into different semantic categories including associative, functional, causal, risk factors, and statistical. The framework transforms unstructured patient-generated texts into structured interoperable knowledge, while preserving different clinical contexts. This helped advance semantic interoperability of health data by improving clinical decision making and biomedical knowledge reuse.\u003c/p\u003e","manuscriptTitle":"Extracting Non-Taxonomic and Ternary Relations from Patient-Generated Texts for Semantic Interoperability","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-18 08:44:04","doi":"10.21203/rs.3.rs-9087634/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e44c05a4-dab3-4006-bf8a-0e0f17745ff9","owner":[],"postedDate":"March 18th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-30T15:09:31+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-18 08:44:04","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9087634","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9087634","identity":"rs-9087634","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0