Research on multi-attribute attentional civil aviation entity alignment based on Siamese network

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

This paper proposes a Siamese network with a font pre-training model and multi-attribute attention alignment to improve the integrity and accuracy of civil aviation knowledge graphs for emergency management.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

The preprint studies how to construct a more reliable, information-rich knowledge graph for civil aviation emergencies, focusing specifically on entity alignment across multi-source, heterogeneous domain data. Using a Siamese-network based Multiple Attribute Attention Entity Alignment (SMAAEA) model, it incorporates a font pre-training model with a character feature layer and uses two sub-networks to learn complete entity semantics and the relative importance of different attributes to improve symmetry and reduce ambiguity. The authors report that, compared with other models, their approach achieves an F1 value of 96.97%, supporting feasibility. As a caveat, the work is presented as an under-review preprint and includes limited methodological detail in the provided text. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

With the growth of demand for air-space integrated transportation industry, the civil aviation transportation industry plays an increasingly important role in the national economy and transportation. The diversity and diversified development of civil aviation emergencies have greatly affected the decision-making efficiency for emergencies in civil aviation emergency management system. Constructing a knowledge graph of civil aviation and improving the reliability and richness of the knowledge graph have become an urgent problem to be solved. There are some problems during the construction of civil aviation domain knowledge graph such as long entity length, existing hybrid and composite entities, similarity of the font features of domain entity names, difference of information between entities, separated coding between entities, and error prone in the transmission process of coding. So we research the construction method of knowledge graph for civil aviation emergencies to resolve these problems. (1) We propose a font pre-training model and incorporate a character feature layer into the embedding layer. (2) We propose a multi-attribute attention alignment method based on Siamese network; entities with similar structure in civil aviation knowledge base are input into the font layer for pre-training; the complete semantic information of entities and the importance of different attributes to entities are learned through two sub-networks respectively to improve the integrity of civil aviation knowledge graph. The experimental results show that, compared with other models, the F1 value of the proposed model can reach 96.97%, which verifies the feasibility of our model.
Full text 183,539 characters · extracted from preprint-html · click to expand
Research on multi-attribute attentional civil aviation entity alignment based on Siamese network | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Research on multi-attribute attentional civil aviation entity alignment based on Siamese network Jintao Wang, Jiayi Qu, Zuyi Zhao, Xiao Dong This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3066877/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 9 You are reading this latest preprint version Abstract With the growth of demand for air-space integrated transportation industry, the civil aviation transportation industry plays an increasingly important role in the national economy and transportation. The diversity and diversified development of civil aviation emergencies have greatly affected the decision-making efficiency for emergencies in civil aviation emergency management system. Constructing a knowledge graph of civil aviation and improving the reliability and richness of the knowledge graph have become an urgent problem to be solved. There are some problems during the construction of civil aviation domain knowledge graph such as long entity length, existing hybrid and composite entities, similarity of the font features of domain entity names, difference of information between entities, separated coding between entities, and error prone in the transmission process of coding. So we research the construction method of knowledge graph for civil aviation emergencies to resolve these problems. ( 1 ) We propose a font pre-training model and incorporate a character feature layer into the embedding layer. ( 2 ) We propose a multi-attribute attention alignment method based on Siamese network; entities with similar structure in civil aviation knowledge base are input into the font layer for pre-training; the complete semantic information of entities and the importance of different attributes to entities are learned through two sub-networks respectively to improve the integrity of civil aviation knowledge graph. The experimental results show that, compared with other models, the F1 value of the proposed model can reach 96.97%, which verifies the feasibility of our model. Physical sciences/Mathematics and computing/Computer science Physical sciences/Mathematics and computing/Information technology Physical sciences/Engineering Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Figure 16 Introduction The civil transport industry has developed rapidly in recent years, especially in the aviation field. The total transport turnover increases continually. As a main part of transport, the scale of aviation keeps expanding. The construction of civil aviation emergency management system will face new challenges and development opportunities 1 . In 2020, due to the impact of the COVID-19, airline passenger volume dropped significantly. Meanwhile, due to the weak operation capacity of airlines, airport equipment failure, high-altitude traffic control, and bad weather, the safety situation of civil aviation is severe and complicated and faces many challenges: ( 1 ) Due to the particularity of the aviation field, serious consequences will occur in case of sudden emergency or timely and efficient measures are not taken; ( 2 ) The causes of civil aviation accidents are complex, such as various types of natural disasters, accidental disasters, human factors and problems with the aircraft itself, etc; ( 3 ) The imperfection of aviation management system for emergencies, and the emergency management ability needs to be further improved. Due to the particularity of civil aviation, it is of great significance to study the construction of knowledge map of civil aviation emergencies for the emergency management system to make rapid and effective decisions in the face of complex and diverse emergencies. The key components of civil aviation knowledge map include knowledge extraction and knowledge fusion. In the process of constructing domain knowledge graph, there are still some problems, such as lack of communication between information and has great difference between information. Moreover, most knowledge is usually crawled through web pages, or comes from some non-standard data sets, which is easy to cause some problems of text information knowledge, such as content cross, multi-source heterogeneous, etc. Therefore, further expanding knowledge fusion on the basis of entity recognition and relationship extraction is the key to establish and expand the knowledge map of civil aviation emergencies. By using the knowledge fusion technology, the description from different entities is unified. Then these entities are input into the target knowledge base to make the knowledge map more accurate and enrich the content of knowledge, which is of great significance to further improve the efficient decision-making when facing emergencies in the field of civil aviation and form a complete and scientific emergency management system. The construction technology of knowledge graph for civil aviation emergencies is researched in this paper. Aiming at the problems of multi-source and heterogeneous knowledge are difficult to organize and integrate effectively in civil aviation, further fusion is carried out on the basis of knowledge extraction. On the background of civil aviation emergencies, we focuses on the entity alignment method in knowledge fusion based on knowledge extraction to achieve information symmetry between entities and eliminate ambiguity. The main research content is as follows: For a large number of complex and mixed entities in civil aviation emergencies, and the entities cannot be extracted automatically and the names are long, it leads to low extraction accuracy. Thus, a Siamese-network Multiple Attribute Attention Entity Alignment (SMAAEA) model is proposed. The information between entities can be symmetrical, and the diversity of knowledge graph is improved, so as to overcome the problems of low recall rate and the contribution to entities of important attributes cannot be distinguished, which are caused by the inability to use entities to describe the details of text information. It improves the quality and confidence of the knowledge graph, and provides theoretical support for the construction of the knowledge graph in the civil aviation field. The remainder of this paper is organized as follows: In section 1 we summarize the research background and significance of this paper. Then, the process and key technology of entity alignment is introduced, and proposes the entity alignment method with SMAAEA model in section 2. In section 3, experiments are carried out to verify the proposed model. We summary our work briefly in section 4. Related work Research status of knowledge graph The historical evolution of knowledge graphs is broadly divided into three phases. In the first phase (late 1950s and early 1960s), the concept of Semantic Network was first proposed by M. Ross Quillian and Robert F. Simmons 2 , when the Semantic Network was used to achieve the storage of data information through the form of graphs, which were easier to understand and display, and the large amount of information using the Semantic Web is able to more concise and effective way to achieve information expression and storage by means of diagrams. In the second phase (mid-1980s and early 2000s), the Semantic Web has developed rapidly. The Knowledge Graph was developed from the Semantic Web to enable computers to access and use knowledge as humans do. The category of ontology, as "an explicit specification of a conceptual model 3 , describes knowledge at the semantic level, and computers can process and recognize ontologies, so that information processing, flow, and interaction between humans and computers can be achieved, making human-computer interaction simpler and clearer. The third phase (early 1920s-present) is a boom phase in the development of knowledge graphs, such as YAGO 4 , DBpedia 5 , Freebase 6 , etc. In the academic world, the first knowledge graph construction project that adopts the combination of Chinese and English with different language features is Xlore 7 and OpenKE platform, which was built by Li Juanzi's team at Tsinghua University and created a system system for media event analysis and mining, entity linking, etc. based on Xlore; Fudan University and Shanghai Jiaotong University rely on Wikipedia and Baidu Encyclopedia as the basis, considering that Based on Wikipedia and Baidu, Fudan University and Shanghai Jiao Tong University created zhishi.me 8 and CN-pedia 9 Chinese knowledge graph platforms, considering the complexity of obtaining information from encyclopedias; the Institute of Computer Technology of the Chinese Academy of Sciences applied Open KN to build a prototype system of "people cube, matter cube, and knowledge cube"; In 2017, the cnSchema.org project 10–11 , which is used to maintain the schema layer of open domain knowledge graphs, was initiated by several universities in China. The main advantages of these knowledge graphs above are that they can be applied in different domains with a wide range of applications, with search functions and question and answer service systems. However, there are relatively few open data links and databases, and there is a lack of rich structured Chinese data. Meanwhile, in the commercial world, the search engines established by Sogou and Baidu are widely used. The development of these search engines has made the search method more convenient and intelligent, satisfying the needs of users and bringing a good experience at the same time. In recent years, knowledge graph technology 12–19 has continued to develop and improve, and has been widely used in many fields such as finance, medicine, social media, fisheries, animal health, military industry, ocean, and agriculture. Research status of knowledge fusion Two main methods, entity alignment and entity linking, are used in the knowledge fusion process of multi-source knowledge bases. Early studies by researchers on entity alignment techniques20 mainly use the relationships between entities to determine, whether the entities are successfully aligned with each other, such as rule-based entity alignment techniques, which require manual intervention to customize the rules by themselves. The literature 21proposes that the degree of similarity of attributes can be determined by the sides of a triangle. The literature 22 proposes to find matching laws in a continuous iterative learning process when there is a part of entity seeds that have already been matched. For the text information that is not annotated, this matching law is used to annotate and get more candidate matching seeds. The literature 23 evaluates entities by labeling and scoring them. The methods are supervised learning, semi-supervised learning and unsupervised learning, respectively. The literature24 proposes the use of semi-supervised learning for joint training of entities to achieve entity alignment. The literature25 proposes an entity alignment approach in Conditional Random Field (CRF), in which the distance function of similarity between entities is computed so as to adjust the learning parameters and the characteristic function to maximize their product. The most representative models are Support Vector Machine (SVM)26, integrated learning27, and decision trees28. The literature29 proposes a support vector machine based entity alignment approach under supervised learning. When there is no way to practice and analyze a large amount of data, more practice data can be obtained by active learning with continuous information exchange with the annotators. The literature30 proposed the ALIAS system to deal with the redundancy problem between similar entities through information exchange between humans and machines. As the research progresses, the Deep Learning (DP) based entity alignment model represents the features and attribute values of entities as vectors, which in turn form dense vectors to determine whether the two are the same entity between them. The literature 31 proposes a TextCNN convolutional neural network that randomly initializes the text vectors in the text to derive a feature vector representation of each text, extracts the feature representation of the text information using the convolutional layer, and outputs the final vector representation in the fully connected layer. The literature32 introduces TextRNN, which can recognize text information containing time series, input the text information into the RNN in an orderly manner to obtain the sequence feature representation, and the number of neurons in the hidden layer can be used as an abstract representation of the text information. The literature33 mentions FastText, which performs text vectorization representation of text data in the word2vec34 pre-training model, and introduces an n-grams model based on this, which is used to count the probability of occurrence of text information, and finally sums all the input feature vector representations and then averages them as the final input text vector representation, and then finally uses the softmax layers. The current research approach suffers from the problems arising from over-reliance on machine learning-based and rule-based entity alignment tasks: first, there is no way to use a large number of entities for detailed portrayal of text information for manually limited features, which results in low recall rate, and second, there is no way to distinguish the contribution of important attributes to entities. In this paper, to address the problem of information asymmetry among civil aviation emergency entities and the lack of differentiation of the importance of different attributes to entities, SMAAEA is proposed to improve the symmetry of information among entities based on knowledge extraction. By mapping entities based on different knowledge bases in the civil aviation domain to the same entity in the target knowledge base, the problem of information asymmetry among entities is overcome and the importance of different attributes to the same entity can be distinguished. The general research idea of entity alignment Suppose the set of entities obtained by knowledge extraction is \(E=\left\{{e}_{1},{e}_{2},\dots {e}_{M}\right\}\) , and the set of entities in the civil aviation knowledge base is \(K=\left\{{k}_{1},{k}_{2},\dots {k}_{N}\right\}\) , the goal of entity alignment is to align the different entities \({e}_{i}\) in the set of entities \(E\) of the civil aviation database obtained by knowledge extraction through crawling corresponds to the different entities \({k}_{j}\) in the civil aviation knowledge base, as shown in Fig. 1 . Overall idea : Firstly, we pre-process the data in the civil aviation emergency text set and the civil aviation emergency text set obtained after knowledge extraction, and then we label the processed data and train the vectors; further, we construct the Siamese BERT network multi-attribute attention alignment model, and input the processed entity vectors into the alignment model to obtain the complete serialized entity with importance distinction vector with importance distinction. Data preprocessing The pre-processing process of the data is: firstly, the civil aviation emergency text data are organized and cleaned; secondly, the data are used to create a dataset by rule matching, and finally, the entities in the dataset are annotated by automatic annotation method. The rule-based entity alignment is shown in Fig. 2 . The specific process is: firstly, the new entities obtained after crawling and knowledge extraction. Secondly, the entity names of the new entities are matched separately and simultaneously by a combination of exact and fuzzy matching methods to find the entities in the civil aviation database, and the new entities are compared with the original candidate entities in turn, and finally the four attribute values such as flight number, departure location, arrival location, and flight time generated by the entities are matched separately by a more detailed multi-level matching method. The specific rules are as follows. Firstly, if the flight number of the paired entities can be matched, then they can be mapped to the same entity; secondly, if the flight number of the aircraft cannot be matched, then the flight departure location of the paired entities are matched, and if they can be matched, then they can be mapped to the same entity; finally, if both the flight number of the aircraft and the flight departure location cannot be Finally, if neither the flight number of the aircraft nor the departure location of the aircraft can be matched, the arrival location of the aircraft of the paired entity is matched, and if it can be matched, it means that they can be mapped as the same entity. And so on, if the flight number, flight departure location and flight arrival location cannot complete the matching, then the matching of flight time is completed, if the matching is completed, it indicates that the entities can be aligned, otherwise it is considered that the entities cannot be aligned. The number of alignable and non-alignable entities can be obtained in the above way, which are 11119 and 48350 in order. the alignable entities, non-alignable entities and entities in the civil aviation knowledge base are counted to generate a new civil aviation dataset, with successful alignment between entities represented by \(classify=1\) and failed alignment between entities represented by \(classify=0\) . And each set of candidate entities with \(classify=0\) is detected by matching technique for further alignment. The positive and negative examples in the dataset are shown in Table 1 . Table 1 Statistical table of emergency data sets in the field of civil aviation Training set Test set 9000 Positive example: 4500 Negative example༚4500 2000 Positive example: 1000 Negative example: 1000 10000 Positive example: 5000 Negative example༚5000 11000 Positive example: 5500 Negative example༚5500 At the same time, to better improve the quality of word separation, we match the airline name text base, flight number text base and starting place name text base in the domain dictionary, so that we can better construct a professional dictionary for the civil aviation domain. In order to align the entities more accurately, the data corpus is labeled by BIEO (Begin, Inside, End, Outside) method, which can effectively determine the boundary of the entities. Among them, B represents the first character of the entity participle of the successfully predicted category, I represents the entity participle character in the middle position of the successfully predicted category, E represents the end part character of the successfully predicted category, and O represents the predicted character that is not an entity. An example of the labeling is shown in Fig. 3 . Build Siamese network based multi-attribute attentional entity alignment model Build Siamese network based model The Siamese network is proposed to solve the problem of similarity metrics in face verification. The core idea is that if a set of data is input and mapped into a high-dimensional feature space, then all its samples of text information can be described as pairs of high-dimensional feature information, so that the high-dimensional feature information can be used to determine the similarity degree of sample pairs. For example, Google's BERT pre-training model can match two sentences based on their semantic similarity, and it has achieved good results so far. However, if N sentences are input at the same time, it takes O(N*N) time to match the semantic similarity of each sentence one by one and find out the most similar sentence among them, which greatly increases the query time. Since the SBEART network system can handle the problem of semantic similarity between texts, we put it through the Siamese BERT network model which can also capture the connection between sentences and perform the complete representation, significantly reducing the amount of operations. Since the entities used for entity alignment in the construction of the knowledge graph of civil aviation domain are all from the civil aviation knowledge base, and the structures of these entities are similar, the above-mentioned characteristics of Siamese network can be used to encode these structurally similar entities. As shown in Fig. 4 , entity 1 of the Siamese network is input to one of the sub-networks for encoding to get the entity representation vector u, and entity 2 of the Siamese network is input to the other sub-networks to get the entity representation vector v. The two high-dimensional feature information representation vectors u and v are spliced, and the sigmoid activation function is used in the fully connected layer to compute the high-dimensional feature vector for Classification. In this regard, the two-part network model structure of the Siamese network is shown in Fig. 5 , and the two different neural networks of the Siamese network encode the entities correspondingly. When encoding the Bi-LSTM network, the attribute name vector representation and the attribute value vector representation containing character features are spliced and input to the Bi-LSTM network, so that not only the text information can be described but also the information contained in the encoding can be learned completely. When encoding the Bi-GRU network, an attention mechanism is introduced to assign weights to the encoding, and the resulting high-dimensional feature entity vector representation outputs the entity feature vector representation with importance. Build Entity Encoding with Siamese Network In this paper, SMAAEA model is proposed based on Siamese network, and the two parts of entity vectors in SMAAEA model are based on different encoding methods. One is to input the mixed vector of words containing character features into Bi-LSTM for encoding, which can improve the integrity of encoding by using Bi-LSTM to get the timing information in both positive and negative directions; the other is to input the shorter attribute text into Bi-GRU + Attention mechanism for encoding, and get the final The final encoding is obtained by the weight parameter in the Attention mechanism. The encoding of the two molecular networks is summed to obtain a high-dimensional feature information vector representation. The calculation is shown in Eq. ( 1 ): $${e}_{i}=Add\left({e}_{i1},{e}_{i2}\right)$$ 1 \({e}_{i1}\) denotes the Govett vector of entities output by the Bi-LSTM network, and \({e}_{i2}\) denotes the high-dimensional feature vector obtained from the output of the Bi-GRU + Attention network, and the two parts of the encoding will be elaborated separately in the following. Build Bi-lstm-based Entity Encoding Based on the characteristics of the entity structure of civil aviation emergencies and considering the network characteristics of Bi-LSTM, it can not only distinguish the front-to-back order of words in, but also encode from back to front, so in the input part of the Bi-LSTM encoding layer, the attributes of the entities and different attributes are spliced for finer-grained classification as the starting input. Considering the long length of entities in civil aviation field and the existence of a large number of mixed and composite entities, firstly, a word embedding layer and a char feature layer are added to the initial coding layer, and the initial feature vector containing character information can be obtained in the embedding layer, and then the initial feature vector containing character information is used as the input of the Bi-LSTM layer. Finally, for each word vector \({w}_{ij}\) , the information in the context of the two temporal directions before and after the entity can be captured using the bi-directional LSTM, and the intensity relationship between different words can be captured. The feature vectors containing character information are stitched to obtain the new representation \(\overrightarrow{{\text{h}}_{\text{i}}}=\left[\overrightarrow{{\text{h}}_{\text{i}\text{j}}},\overleftarrow{{\text{h}}_{\text{i}\text{j}}}\right]\) , and this new representation contains bidirectional semantic relationships. The output of the model is a stitching of high-dimensional feature vectors of bidirectional semantics in textual information from the civil aviation domain, and the structure of Bi-LSTM is shown in Fig. 6 . The Bi-LSTM-based entity vector \({e}_{i}\) is calculated as shown in Eq. ( 2 ): $${e}_{i}=BiLSTM\left({w}_{i1},{w}_{i2},\dots .{w}_{iL}\right)$$ 2 Among them, the Embedding layer can reduce the dimensionality of the data by matrix multiplication without changing the information contained in the word vectors. Because of the similarity of some entities and the similarity of words contained in the names of two entities, the word feature vector representation is combined with the character feature vector representation in the Embedding layer, and the vectorized text information with character features is finally obtained by vectorizing the Char Embedding layer. Among them, the word embedding layer is used to extract the context vectors of the words in the text, and the char feature layer is used to extract the conformational feature vector and the glyph feature vector of each word, and the vectors and character feature vectors in the civil aviation text are obtained through the char feature layer and the word embedding layer, and then the char embedding layer to incorporate the character feature vectors and lexical information into the character representation, as shown in Fig. 7 . The specific training process is shown in Fig. 8 . Firstly, the civil text subwords are extracted, and then the feature vectors of different words and characters are extracted and pre-trained. Since in one-hot coding, each word corresponds to an n-dimensional feature vector representation, the sparse matrix obtained by such a representation will cause a waste of resources, therefore, the word feature vector representation is one-to-one corresponded with the character feature vector representation in the Embeding layer to turn the sparse matrix into a dense matrix, and by adding the word vectors containing character feature information with the subvector sequence, the The vector representation of the civil aviation textualization. Build Entity Encoding with Bi-GRU + Attention We use the Bi-LSTM entity encoding method, and the final obtained entity feature representation vector representation only contains the information of attributes and attribute values, but there is no precise differentiation of the specific impact of different attribute values on each attribute. In this paper, we change the input of Char Embedding layer, instead of just inputting words, we input a shorter sentence containing attribute features, and the encoding structure is shown in Fig. 9 . Since each entity not only consists of ontology, but also contains different attributes and attribute values, the feature vector representation output by Bi-GRU is input to the attribute encoder, and the entity information of the feature vector representation is encoded to output the feature vector text information representation with attribute features. Then, the obtained text information representations are interacted with the Attention mechanism to obtain feature vector text information representations with different attributes. Since GRU belongs to the derivative structure of LSTM, the number of gating units used is reduced by merging the input gate control unit and forgetting gate control unit in LSTM, and the parameters are reduced accordingly, so the training speed of textualized information can be improved without changing the learned temporal information. Based on the above analysis, the entity coding based on Bi-GRU + Attention is proposed, and the attribute information is added in the coding process. The reset gate in GRU can better capture the dependencies at shorter distances in time order, and the update gate in GRU can better capture the dependencies at longer times in time order. bi-GRU network can obtain the timing information in two opposite directions, and compared with Bi-LSTM, Bi-GRU has fewer network parameters, which can improve the training speed. In practice, the GRU performance does not differ much from the LSTM performance, and the training speed is improved. the working structure of GRU is shown in Fig. 10 . The calculation process of reset gate control unit and update gate control single gate is shown in Equations ( 3 ) and ( 4 ). $${r}_{t}=\sigma ({W}_{xr}\cdot {x}_{t}+{W}_{hr}\cdot {h}_{t-1}+{b}_{r})$$ 3 $${z}_{t}=\sigma \left({w}_{xz}{X}_{t}+{w}_{hc}\cdot {h}_{t-1}+{b}_{z}\right)$$ 4 The candidate hidden layer unit is calculated by first multiplying the output of the reset gate at the current time point by the hidden layer unit at the same time, then connecting the current time point sequence, and finally deriving the sequence unit of the candidate hidden layer by the operation on the tanh activation function, as shown in Eq. ( 5 ). $${h}_{t}^{{\prime }}={tan}h({W}_{xh}\cdot {x}_{t}+({r}_{t}\cdot {h}_{t-1})\cdot {W}_{hh}+{b}_{h})$$ 5 At moment \(t\) , the hidden layer state is \({h}_{t}\) , and the update gate control unit \({z}_{t}\) of the current time series is multiplied with the hidden state feature vector \({h}_{t-1}\) at the previous moment and multiplied with the hidden layer state feature vector \({h}_{t}^{{\prime }}\) at the current moment for the integrated operation as shown in Eq. ( 6 ). $${h}_{t}={z}_{t}\cdot {h}_{t-1}+(1-{z}_{t})\cdot {h}_{t}^{{\prime }}$$ 6 For the word text information vectors, after inputting them in the order they are in the sentence, for different words \({w}_{ij}\) , input to the Bi-GRU network, the text information feature vector representation containing two different directions before and after can also be obtained, and then these textualized feature vectors are stitched together to obtain the new feature vector representation \({h}_{i}=\left[\overrightarrow{{h}_{\text{i}j}},\Leftarrow {h}_{\text{i}j}\right]\) . The calculation procedure is shown in equations ( 7 ) and ( 8 ). $${h}_{ij}=BiGRU\left({w}_{i1},{w}_{i2},\dots .{w}_{iL}\right)$$ 7 $$att{r}_{i}=Add\left({h}_{i1},{h}_{i2},\dots {h}_{iL}\right)$$ 8 After the Bi-GRU network, the different attributes of each entity feature vector \({e}_{i}\) are passed through the attribute encoder to obtain \(N\) feature vector text representations containing attribute information, namely \(att{r}_{i1},att{r}_{i2},...,att{r}_{iN}\) . Finally, the final encoding is obtained by integrating the feature text vectors containing attributes through the Attention mechanism. The computation process is shown in Eq. ( 9 ). $${e}_{i}=Attention(att{r}_{i1},att{r}_{i2},att{r}_{iN})$$ 9 Attention can be specifically divided into three steps, and the specific process is shown in Fig. 11 . In the first stage, the relationship between the two is calculated based on the Query and Key values, as shown in Eq. ( 10 ): $${a}_{i}=softmax\left(si{m}_{i}\right)=\frac{{e}^{si{m}_{i}}}{{\varSigma }_{j=1}^{{L}_{x}}{e}^{si{m}_{j}}}$$ 10 In the second stage, the operator of \(Softmax\) is introduced, whose purpose is to do a numerical transformation of the results of the first paragraph, calculated as shown in Eq. ( 11 ): $$Attention(Query,Source)={\sum }_{i=1}^{{L}_{x}}{a}_{i}\cdot Vlau{e}_{i}$$ 11 In the third stage, the corresponding values are weighted by their weights to obtain the attention vector. Build Classification module For a specified set of candidate entities, such as Entity 1 and Entity 2, the first entity is first input into the Siamese network to obtain the feature vector representation \({e}_{11}\) with completeness by Bi-LSTM coding; then the same entity vector is input into the Bi-GRU + Attention network model to obtain the feature representation vector \({e}_{12}\) with importance; finally, the feature vector \({e}_{11}\) and the feature vector \({e}_{12}\) are added together to obtain the complete high-dimensional feature representation vector \({e}_{1}\) and the feature vector \({e}_{12}\) are added to obtain the complete high-dimensional feature representation vector \({e}_{1}\) with importance representation. and so on, the same operation is performed for another entity to obtain two high-dimensional feature vector representations. The two feature vector representations are stitched in the fully connected layer, and the stitched feature vector representation is predicted. The SMAAEA model training process is shown in Fig. 12 . The calculation steps are shown in Equations ( 12 ), ( 13 ), ( 14 ), and ( 15 ). $${e}_{1}=TN({w}_{11},{w}_{12},\dots {w}_{1L};att{r}_{11},att{r}_{12},\dots ,att{r}_{1M})$$ 12 $${e}_{2}=TN({w}_{21},{w}_{22},\dots {w}_{2L};att{r}_{21},att{r}_{22},\dots ,att{r}_{2M})$$ 13 $$e=Concate\left({e}_{1},{e}_{2}\right)$$ 14 $$py=sigmoid\left(e\right)$$ 15 In the above Equations, \(TN\) is represented as Siamese network, \(w\) is character feature vector representation, attr is attribute feature vector representation, and \(att{r}_{ij}=BiGRU({w}_{i1},{w}_{i2},\dots {w}_{is})\) , \(sigmod\left(y\right)=\frac{1}{1+{\text{e}}^{-y}}\) , and the final prediction result \(py\) by two classification method on the high-dimensional feature vector representation. In the model encoding process, a group of entities share the weights and parameters in the model. Build Entity alignment fusion module Due to the specificity of the civil aviation domain and the fact that the data used are jointly constructed by knowledge extraction and civil aviation knowledge base data, it results in certain differences in the types of entities and the names of the attributes corresponding to the entities in the data from different sources. If knowledge alignment fusion is not performed on these different representations of the same entity, it will lead to the confusion of entity links and reduce the quality of the knowledge graph, so it is necessary to further process the entities through knowledge fusion, and in this paper, the SMAAEA model is applied to build the alignment fusion module in knowledge fusion. Alignment fusion is the integration of information from different sources. When fusing the external civil aviation knowledge base with the original civil aviation knowledge base, two aspects need to be considered: one is the fusion of data quality, i.e., whether the entities can be aligned, their attributes and categories are considered, and the redundancy and conflicts between entities are resolved, and then the new entities obtained are fused into the civil aviation knowledge base through the Schema layer; the other is the merging of knowledge bases in the same domain, and the fusion of historical data with similar structure of historical data are fused and the data are converted into RDF triples. The process of alignment fusion is shown in Fig. 13 . First, the table of stored entities is checked to detect whether there is any entity data different from that in the original civil aviation knowledge base, and if a new entity representation is detected, then the entity is processed according to the table of stored entities. The new entities contain two parts, one part is detected from the external entity repository and the other part is from the local civil aviation entity repository, and the entities in the local civil aviation entity repository are used as candidate entities. The candidate entities are matched in multiple directions, and if the matching is successful, the candidate entities are input into the SMAAEA network model together with the new entities detected from the external knowledge base, from which it is judged whether the two describe the same object. If the entities are successfully aligned, the next aligned entities are unified in terms of attributes and input into the local civil aviation knowledge base together. Experiment and Analysis Experimental dataset The text information of civil aviation accidents used in the experiments of this paper is more than 15,000 data, which are jointly constructed by the crawled investigation tracking reports about safety accidents in aviation field released by the Civil Aviation Safety Information System of China 1 , the expanded dataset, and the dataset obtained by knowledge extraction. These text information are divided into 9000, 10000, and 11000 data for training, and the number of text data in the test set is 2000 data for testing. The SMAAEA model will divide the data according to 9:1 during training, and the entity labeling method is shown in Table 2 . Table 2 Methods of entity annotation Entity types Entity start Inside the entity Entity ending Time of incident B-DATE I-DATE E-DATE Airlines B-ARIL I-ARIL E-ARIL Type of aircraft B-FTYPE I-FTYPE E-FTYPE Flight number B-FNUM I-FNUM I-FNUM Registration number B-REG I-REG E-REG Departure airport B-ORIG I-ORIG E-ORIG Destination B-DEST I-DEST E-DEST Aircrew B-CREW I-CREW E-CREW Number of passengers B-PSG I-PSG E-PSG casualties B-CASU I-CASU E-CASU Cause of accident B-REA I-REA E-REA Accident result B-RES I-RES E-RES Nonentity type O O O Experimental environment Loss function The cross-entropy loss function is used to evaluate the SMAAEA network model since it can provide a measure of the variability of the probability distribution of two samples. For a set of entity pairs \((px,py)\) , the candidate sample prediction results are denoted by py, and the set of prediction results taking values \(\left\{\text{0,1}\right\}\) . A sample with actual label information quantity \(y\) , the probability of \(y=1\) is denoted by \(p\) , and the probability of \(y=0\) is denoted by \((1-p)\) , and the formula of the cross-entropy loss function is shown in Eq. ( 16 ). $${log}\left(\left.y\right|p\right)=-\left[y*{log}\left(p\right)+\left(1-y\right){log}\left(1-p\right)\right]$$ 16 For all text data, the smaller the cross-entropy loss is, the better the SMAAEA model proposed in this paper is proved to be, and the loss function of the whole model is expressed as the average of all sample loss functions, as shown in Eq. ( 17 ). $$loss=\frac{1}{m}\sum _{i}^{m}\text{log}\left(\left.{y}_{i}\right|{p}_{i}\right)$$ 17 Parameter setting The parameter settings for this experiment are shown in Table 3 : Table 3 Model parameter settings Parameters Setting values max_length 512 max_attribute 8 batch_size 64 learning_rate 1e-3 epoch 30 dropout 0.2 optimizer Adam Experimental environment The experimental environment for this experiment is shown in Table 4 . Table 4 Experimental Environment parameter settings Experimental environment Parameter setting System environment Windows 10 GPU parameters 24 cores、128G、1080Ti Python version 3.6.7 Keras version 2.2.5 TensorFlow version 1.14.0 Experimental evaluation index Performance metrics In this paper, the harmonic mean F1 is used to measure the precision and recall, and is used as the performance index of the SMAAEA model. The accuracy rate P represents the proportion of the real results in the sample that are correct to the predicted results that are positive, as shown in Eq. ( 18 ). $$P=\frac{TP}{TP+FP}=\frac{\begin{array}{c}The EquationNumber of instances where the true class is positive \\ and the model discrimination is also positive\end{array}}{\text{T}\text{h}\text{e} \text{n}\text{u}\text{m}\text{b}\text{e}\text{r} \text{o}\text{f} \text{i}\text{n}\text{s}\text{t}\text{a}\text{n}\text{c}\text{e}\text{s} \text{w}\text{h}\text{e}\text{r}\text{e} \text{t}\text{h}\text{e} \text{p}\text{r}\text{e}\text{d}\text{i}\text{c}\text{t}\text{i}\text{o}\text{n} \text{r}\text{e}\text{s}\text{u}\text{l}\text{t} \text{i}\text{s} \text{p}\text{o}\text{s}\text{i}\text{t}\text{i}\text{v}\text{e}}$$ 18 Recall rate R represents the proportion of predicted correct results to actual correct results in the sample, as shown in Eq. ( 19 ). $$R=\frac{TP}{TP+FN}=\frac{\begin{array}{c}The EquationNumber of instances where the true class is positive \\ and the model discrimination is positive\end{array}}{\text{T}\text{h}\text{e} \text{n}\text{u}\text{m}\text{b}\text{e}\text{r} \text{o}\text{f} \text{i}\text{n}\text{s}\text{t}\text{a}\text{n}\text{c}\text{e}\text{s} \text{t}\text{h}\text{a}\text{t} \text{a}\text{r}\text{e} \text{t}\text{r}\text{u}\text{l}\text{y} \text{p}\text{o}\text{s}\text{i}\text{t}\text{i}\text{v}\text{e}}$$ 19 The harmonic mean F1 is calculated as shown in Eq. ( 20 ). $$F1=\frac{2\text{*}P\text{*}R}{P+R}$$ 20 Comparison model Experiments are conducted on the constructed dataset to compare several typical neural network models with the SMAAEA network model proposed in this paper, and the models are described as follows. a. SVM: The support vector machine based model mentioned in reference 26. b. CNN: The initialized feature vector representation of the character and word mixed features in the civil aviation text was used as the input, the number of convolution kernels was set to 3, 4, and 5, respectively, and the textual vector features were extracted to obtain the high-dimensional feature representation vector and input to the fully connected layer. c. RNN: The input part is the same as the CNN network model, the initialized feature representation vector is obtained and input to the hidden layer to extract the feature information of the vector, and the output part is the same as the CNN model. d. LSTM: The input and output parts are the same as those of the RNN network model, but different from the RNN network when extracting information, LSTM can extract information in the front and back directions. e. NSMAAEA: The network model proposed in this paper is not used for training. ( 3 )Ablation experiment In order to verify the enhancement effect of the character feature word pre-training module, the Attention module and the Bi-LSTM module in the SMAAEA model on the whole model, ablation experiments were conducted, and the comparative models used in the experiments are explained as follows. a. SMAAEA-char embedding: no pre-training containing character feature word vectors; b. SMAAEA-attention: without Bi-GRU and Attention-based entity coding methods; c. SMAAEA-bilstm: no Bi-LSTM-based entity coding method is used. In order to verify the enhancement effects of character feature word pre-training module, Attention module and Bi-LSTM module in the SMAAEA model on the entire model, ablation experiments were conducted. The comparison model used in the experiment was explained as follows. a. SMAAEA-char embedding: Do not carry out the pre-training of the character feature word vector; b. SMAAEA-attention: The entity coding method based on Bi-GRU and Attention is not adopted. c. SMAAEA-bilstm: The entity coding method based on Bi-LSTM is not adopted. Simulation experiment and performance analysis First, the SMAAEA model was trained on the training set, and then 2000 text data were selected as the test set to test the training results, and the P, R, and F1 values of the entity types were counted, as shown in Table 5 . Table 5 Results of the civil aviation emergency entity alignment experiment Entity types Precision/% Recall/% F1/% Time of incident 97.46 99.98 98.67 Airlines 97.76 86.98 92.28 Type of aircraft 91.91 96.36 94.11 Flight number 87.63 91.58 88.48 Registration number 91.79 94.43 93.25 Departure airport 95.14 88.46 91.67 Destination 84.78 79.56 82.23 Aircrew 93.17 95.58 94.47 Number of passengers 95.69 95.69 95.69 Casualties 84.42 84.42 84.42 Cause of accident 88.58 85.62 87.03 Accident result 92.57 93.18 92.89 Overall indicator 96.12 97.89 96.97 From the results, it can be seen that the reconciled mean F1 for the time of the incident is the highest, and the reconciled mean F1 values for the number of crew and passengers are 95.69% and 94.47%, respectively, which are only 2.5% and 2.23% lower than the overall index. The main reason is that the entity types of the number of passengers and crew are specific values and the entity types are relatively simple. The entity type of incident time is date type is also specific number, and the expression form of these entity types is fixed and simpler compared with text-based data, which is not easy to generate ambiguity. The R-value of destination entity alignment is relatively low, with a difference of 18.33% compared to the overall index, which is a large difference. The main reason is that the entity names are relatively long, there is similarity in the glyph names, and the semantics are more complex and prone to ambiguity. Since the number of casualties on the civil aviation event dataset is prone to ambiguity, its F1 value is also relatively smaller. In order to verify the feasibility of SMAAEA network model, this paper will conduct a comprehensive analysis of SMAAEA network model from three aspects: reference experiments, comparison experiments and stability experiments. The first group is the reference experiment, in which the SMAAEA model is compared with machine learning models and other classical neural network models to prove the effectiveness of SMAAEA model in improving recall and accuracy; the second group is the ablation experiment, in which SMAAEA model is ablated with different parts to analyze the specific roles of different modules and prove the effectiveness of each part of the module; the third is the stability experiment. Five different sets of training data are selected respectively, and the SMAAEA model is compared with its simplified model to verify the reliability and stability of the model. Reference experiment Training different comparison model networks and recording the F1 values after every 5 rounds of iterations of the dataset, training found that the training effect grows faster in rounds 0–15, and the growth of the model training effect tends to be stable after the 20th round, training found that the F1 values are optimal after 25 rounds, and the effect of the 30th round is not significant with the 25th round. Comparing the F1 value performance of SMAAEA model with other models under different training rounds, as shown in Fig. 14 . Table 6 SMAAEA and comparative model experimental results Model Precision/% Recall/% F1/% SVM 94.56 82.34 87.08 CNN 96.79 93.20 95.10 RNN 92.38 93.20 92.91 LSTM 96.99 89.90 93.42 NSMAAEA 96.59 93.38 94.83 SMAAEA 96.12 97.89 96.97 a. The experimental results in Table 6 show that the SMAAEA model achieves an accuracy rate of 96.12% and a recall rate of 97.89% compared with machine learning models and other neural network models on the civil aviation emergencies dataset, both of which are higher than other comparison models. The SVM-based entity alignment method has the lowest F1 value compared with the F1 values of other neural network models, mainly because only the same attributes can exchange information, and there are restrictions on the exchange, and information cannot be exchanged between different attributes, resulting in an F1 value of only 87.08%. the SMAAEA model and other comparison models can enable the exchange of information between various attributes of the same entity, thus improving the The SMAAEA model and other comparative models are able to exchange information among the attributes of the same entity, thus improving the recall rate and achieving better experimental results. b. From the experimental results, it can be concluded that the F1 value of the proposed SMAAEA model is improved by 1.87%, 4.06%, and 3.55% compared with CNN, RNN, and LSTM, respectively. The effectiveness of the SMAAEA model proposed in this thesis can be proved. c. The number of parameters of NSMAAEA model is 3424337, and the number of parameters of SMAAEA model is 1796374. The difference between the two is whether the Siamese network is used or not. This can prove that the SMAAEA network model is beneficial to achieve the entity alignment task and realize the information exchange and connection between entities. d. From the comparison results of the whole experiment, the F1 value of SMAAEA, the model proposed in this paper, is higher compared with other models. The reason for this is mainly that the two sub-networks of the Siamese network distinguish the importance of the same entity according to the integrity of encoding and different attributes, respectively, and combine the high-dimensional feature vector representations obtained from the two sub-networks to learn entity encoding from multiple perspectives, thus improving the recall rate of the model. Ablation experiment Training different simplified model networks and recording the F1 values after every 5 iterations of the dataset, training found that the training effect grows faster in rounds 0–10, and the growth of the model training effect tends to be stable after round 15. The F1 values were found to be optimal after 25 rounds, and the effect of the 30th round was not significant with the 25th round. Comparing the F1 value performance of SMAAEA model with its simplified model under different training rounds, as shown in Fig. 15 . Table 7 SMAAEA and simplified model experimental results Model Precision Recall F1 SMAAEA-char embedding 95.89 94.90 95.97 SMAAEA-attention 92.82 95.40 94.01 SMAAEA-bilstm 97.34 88.30 92.59 SMAAEA 96.12 97.89 96.97 a. In Table 7 ,By comparing and analyzing the SMAAEA-embedding model with the SMAAEA model in this paper, we find that the F1 value of model SMAAEA is improved by 0.75% compared with model SAAEA-embedding, which proves that adding the Emedding layer to pre-train the word feature representation vector and character feature vector improves the experimental results very effective. b. The SMAAEA-attention model is compared with the SMAAEA model, i.e., the attention mechanism is removed from the SMAAEA model and then compared with the SMAAEA model. the F1 value of the SMAAEA model is improved by 2.71%. It can be seen that the introduction of the attention mechanism can make the model pay attention to relatively more important attributes in the learning process, so that the important attributes can play a greater role in the classification process, thus improving the effectiveness of the model. c. The SMAAEA-bilstm model and SMAAEA model are compared and analyzed, i.e., the Bi-LSTM network is removed from the SMAAEA model, and then it is compared with the SMAAEA model. The F1 value of model SMAAEA increased by 4.13% compared with model SMAAEA-bilstm. This is mainly due to the fact that Bi-LSTM can learn the information in both directions before and after the feature representation vector, which ensures that the integrity of the encoding can be learned, thus improving the learning quality of the model. d. By comparing and analyzing the above simplified model with the SMAAEA model, the introduction of the Bi-LSTM modular model has the best improvement of 4.13%, which proves the important role of the Bi-LSTM network model for coding integrity learning. Stability experiment In order to verify the stability of SMAAEA model, five groups of text data with different numbers of entries are selected for the experiment, and the comparison models are the three simplified models mentioned in the ablation experiment, and the number of text data set for training is 9000, 9500, 10000, 10500 and 11000, and the number of text data used for testing is 1000. As shown in Fig. 16 , the horizontal coordinates in the figure indicate the number of text data trained and the vertical coordinates indicate the F1 values corresponding to the different models. As can be seen from Fig. 16 , when the number of text data is 10000, the F1 values of SMAAEA-embedding model and SMAAEA model are close to each other, and SMAAEA-embedding is slightly higher than SMAAEA model, but the difference between their F1 values is only about 0.04%, and in other cases with different number of text data, the F1 values of SMAAEA model are all different. The F1 values of SMAAEA model are higher than the other models, which proves that SMAAEA model shows better stability and reliability for different training items of text information. Conclusion Knowledge fusion is an important foundation for the construction of the knowledge graph of civil aviation emergencies. In this paper, we propose an entity alignment method based on knowledge extraction to address the special characteristics of civil aviation domain and the problems of long and complex entities in civil aviation domain, the difficulty of fast entity extraction and the low extraction accuracy in the current knowledge fusion process. It is verified and analyzed by reference experiments, comparison experiments and stability experiments. This paper firstly analyzes the entity alignment task based on civil aviation emergencies and carries out the general design. Secondly, in constructing the dataset, the rule-based entity alignment method is combined with the annotation of entities through lexicon matching to improve the accuracy and recall rate of entity alignment. Then SMAAEA network model is proposed to investigate the entity alignment task. Due to the special characteristics of civil aviation domain, entities in the domain may have similar glyph structures or the same radicals, and entity alignment is relatively difficult. In the process of entity coding and improving the pre-training model, a char feature layer is introduced to obtain a word mixture feature vector containing character features. And the obtained character-word mixed feature vector representation is input to the Bi-LSTM layer and the shorter attribute sentences are input to Bi-GRU + Attention respectively, and finally the high-dimensional vector representation is obtained. After entity alignment, knowledge fusion is performed to fuse the new knowledge with the knowledge in the target civil aviation knowledge base to further improve the quality of the knowledge spectrum graph. Finally, the feasibility and reliability of the SMAAEA model proposed in this thesis are verified by designing experiments and analyzing the results. In the future work, we consider whether feature matching and partition indexing of entities can be performed in the process of entity alignment to make entity alignment more efficient and accurate, but this method requires a large amount of computational resources and is usually applied to the construction of general domain knowledge graphs, and the method requires high data quality and needs to be calculated by various types of similarity functions, and the next step can be to apply the The next step is to apply the structural relationship between entities in this method to the civil aviation knowledge base. Declarations Author contributions statement J.W. designed the study and writed the main manuscript text. J.Q. performed the experiment and analyzed the results, Z.Z collected the data and analyzed the results.X.D. prepared fgures and tables, and all authors reviewed the manuscript. Competing interests The authors declare no competing interests. Data availibility The datasets used and analysed during the current study will be available from the corresponding author on reasonable request References World Accident Investigation and Tracking[EB/OL]Aviation Safety Information System of CAAC.[2018-09]. http://safety.caac.gov.cn /index /initpage.act. Sowa J F. Principles of semantic networks: Explorations in the representation of knowledge[M]. Morgan Kaufmann, 2014. Gruber T R. Toward principles for the design of ontologies used for knowledge sharing[J]. International journal of human-computer studies, 1995, 43(5-6): 907-928. Suchanek F M, Kasneci G, Weikum G. Yago: a core of semantic knowledge[C]//Proceedings of the 16th international conference on World Wide Web. 2007: 697-706. Bizer C, Lehmann J, Kobilarov G, et al. Dbpedia-a crystallization point for the web of data[J]. Journal of web semantics, 2009, 7(3): 154-165. Bollacker K, Evans C, Paritosh P, et al. Freebase: a collaboratively created graph database for structuring human knowledge[C]//Proceedings of the 2008 ACM SIGMOD international conference on Management of data. 2008: 1247-1250. Wang Z, Li J, Wang Z, et al. XLore: A Large-scale English-Chinese Bilingual Knowledge Graph[C]//International semantic web conference (Posters & Demos). 2013, 1035: 121-124. Niu X, Sun X, Wang H, et al. Zhishi. me-weaving chinese linking open data[C]//International Semantic Web Conference. Springer, Berlin, Heidelberg, 2021: 205-220. Xu B, Xu Y, Liang J, et al. CN-DBpedia: A never-ending Chinese knowledge extraction system[C]//International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems. Springer, Cham, 2017: 428-438. LiDing. Schema[EB/OL][2019,3,25].https//github.com/cnschema/cnschema/wiki/Schema. Kejriwal M. Domain-specific knowledge graph construction[M]. Cham: Springer International Publishing, 2019. Plank B, Goldberg Y. Multilingual part-of-speech tagging with bidirectional long short-term memory models and auxiliary loss[J]. arXiv preprint arXiv:1604.05529, 2016. Bengio Y, Senécal J S. Adaptive importance sampling to accelerate training of a neural probabilistic language model[J]. IEEE Transactions on Neural Networks, 2008, 19(4): 713-722. Castro M J, Prat F. New directions in connectionist language modeling[C]//International Work-Conference on Artificial Neural Networks. Springer, Berlin, Heidelberg, 2021: 598-605. Mikolov T, Karafiát M, Burget L, et al. Recurrent neural network based language model[C]//Interspeech. 2010, 2(3): 1045-1048. Mikolov T, Kombrink S, Burget L, et al. Extensions of recurrent neural network language model[C]//2011 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2011: 5528-5531. Cho K, Van Merriënboer B, Gulcehre C, et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation[J]. arXiv preprint arXiv:1406.1078, 2014. Chopra S, Hadsell R, Le Cun Y. Learning a similarity metric discriminatively, with application to face verification[C]//2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05). IEEE, 2005, 1: 539-546. Niu X, Sun X, Wang H, et al. me-weaving chinese linking open data[C]//International Semantic Web Conference. Springer, Berlin, Heidelberg, 2021: 205-220. Lin H L, Wang Y Z, Jia Y T, et al. Network big data oriented knowledge fusion methods: A survey[J]. Chinese Journal of Computers, 2017, 40(1): 1-27. Ngomo A C N, Auer S. LIMES—a time-efficient approach for large-scale link discovery on the web of data[C]//Twenty-Second International Joint Conference on Artificial Intelligence. 2011: 905-912. Niu X, Rong S, Wang H, et al. An effective rule miner for instance matching in a web of data[C]//Proceedings of the 21st ACM international conference on Information and knowledge management. 2022: 1085-1094. Wang X, Liu K, He S, et al. Multi-source knowledge bases entity alignment by leveraging semantic tags[J]. Chinese Journal of Computers, 2017, 40(3): 701-711. Zhang Weili, Huang Tinglei, Liang Xiao Entity alignment of encyclopedia knowledge base based on semi supervised collaborative training [J] Computer and Modernization, 2017 (12): 88-93. [Yuan Man, Cao Yang, Chen Ping. Research on the Standard Vocabulary Reference Model in Construction of Educational Knowledge Graph [J]. E education Research , 2022, 41 (03): 76-84. Vapnik V. The nature of statistical learning theory[M]. Springer science & business media, 1999. Kantardzic M. Data mining: concepts, models, methods, and algorithms[M]. John Wiley & Sons, 2021. Han J, Pei J, Kamber M. Data mining: concepts and techniques[M]. Elsevier, 2021. Hu Fanghuai Research on the Construction Method of Chinese Knowledge Graph Based on Multiple Data Sources [D] Doctoral Dissertation Shanghai: Department of Computer Application Technology, East China University of Science and Technology, 2015. Sarawagi S, Bhamidipaty A. Interactive deduplication using active learning[C]//Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining. 2002: 269-278. Kim Y. Convolutional Neural Networks for Sentence Classification[C]//empirical methods in natural language processing. 2014: 1746-1751. Liu P, Qiu X, Huang X, et al. Recurrent neural network for text classification with multi-task learning[C]//international joint conference on artificial intelligence. 2016: 2873-2879. Joulin A, Grave E, Bojanowski P, et al. Bag of Tricks for Efficient Text Classification[C]//conference of the european chapter of the association for computational linguistics. 2017: 427-431. Mikolov T, Sutskever I, Chen K, et al. Distributed representations of words and phrases and their compositionality[C]//Advances in neural information processing systems. 2023: 3111-3119. Song, Yu, Gui, Mingyu, et al. BG-INT: An Entity Alignment Interaction Model Based on BERT and GCN[C]//Communications in Computer and Information Science, Volume 1772, 2023:67-81. Munne, Rumana Ferdous,Ichise, Ryutaro. Entity alignment via summary and attribute embeddings[C]Logic Journal of the IGPL, Volume 31, Issue 2, 2023:314-324. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major revision 08 Sep, 2023 Reviews received at journal 01 Sep, 2023 Reviewers agreed at journal 27 Aug, 2023 Reviewers agreed at journal 27 Aug, 2023 Reviewers invited by journal 24 Aug, 2023 Editor assigned by journal 19 Aug, 2023 Editor invited by journal 22 Jun, 2023 Submission checks completed at journal 22 Jun, 2023 First submitted to journal 15 Jun, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3066877","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":212137530,"identity":"0c15bd8d-4b8e-4f4c-9dc6-581210247c84","order_by":0,"name":"Jintao Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYLACCQMQyXwAwjtApBYJBga2BBK0gHQxMPAYEKfF4PjZwy8sCu7U8bf3fP7ws41Bju9GAuPnAnxazuSlWUgYPJOQOHN2m2RvG4Ox5I0EZukZeLSYHcgxM5AwOCzBcCN3GwNvG0PihhsJbMw8+LScfwPRIn8j5/HHv20M9YS13MgxfgDSYnAjh0EaaEuCASEt9jfemAED+bDkxjPHzKRlzkkYzjzzsFkanxbJ/hzjzxJ/DvPLHW9+/PFNmY083/Hkg5/xaQECNmkJBAfEZGzArwGYUD5+IKRkFIyCUTAKRjYAAKgXTU+8deCkAAAAAElFTkSuQmCC","orcid":"","institution":"Shenyang Aerospace University","correspondingAuthor":true,"prefix":"","firstName":"Jintao","middleName":"","lastName":"Wang","suffix":""},{"id":212137531,"identity":"5704e4b5-ea2b-41d1-abb9-d12343d249da","order_by":1,"name":"Jiayi Qu","email":"","orcid":"","institution":"Shenyang Aerospace University","correspondingAuthor":false,"prefix":"","firstName":"Jiayi","middleName":"","lastName":"Qu","suffix":""},{"id":212137532,"identity":"06e493eb-28ce-492c-b8c3-b01301d1df02","order_by":2,"name":"Zuyi Zhao","email":"","orcid":"","institution":"Shenyang Aerospace University","correspondingAuthor":false,"prefix":"","firstName":"Zuyi","middleName":"","lastName":"Zhao","suffix":""},{"id":212137533,"identity":"b9e68fcf-ba6e-4d3c-ad4f-2dade1486175","order_by":3,"name":"Xiao Dong","email":"","orcid":"","institution":"Northeastern University","correspondingAuthor":false,"prefix":"","firstName":"Xiao","middleName":"","lastName":"Dong","suffix":""}],"badges":[],"createdAt":"2023-06-15 09:14:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3066877/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3066877/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":39194290,"identity":"e8c721a6-5fcd-4b7e-a9c5-c4fa1d1cb99f","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":39183,"visible":true,"origin":"","legend":"\u003cp\u003eEntity alignment tasks\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/4f34bb6e9b02f947a13c452b.png"},{"id":39194292,"identity":"b81aa520-545b-4599-a5b8-a1a401431b1d","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":103361,"visible":true,"origin":"","legend":"\u003cp\u003eRules-based entity alignment diagram\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/a68c7277d911482253c9b319.png"},{"id":39195091,"identity":"657b4f4d-c139-40fd-bc3a-9e46d46e3135","added_by":"auto","created_at":"2023-06-27 20:58:12","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":7621,"visible":true,"origin":"","legend":"\u003cp\u003eLabeling examples\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/878af8bb9c998ee39e4a6938.png"},{"id":39194291,"identity":"98c6e91e-10ba-49bb-b806-1d3c3dff7c2a","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":31125,"visible":true,"origin":"","legend":"\u003cp\u003eThe overall structure is based on the SMAAEA model\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/cea4f5d59407d4d4f0345b9e.png"},{"id":39194298,"identity":"9680c1dc-a0be-4597-a4d1-765583c0de18","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":17923,"visible":true,"origin":"","legend":"\u003cp\u003eOverall model diagram of entity encoding\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/cee9aa086b16a48d9aea17cf.png"},{"id":39195094,"identity":"798640bf-75c2-433c-8b4b-c5a5169a77a4","added_by":"auto","created_at":"2023-06-27 20:58:12","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":37124,"visible":true,"origin":"","legend":"\u003cp\u003eBased on Bi-LSTM entity encoding\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/22622d0e9161cda0f87bb98e.png"},{"id":39194294,"identity":"8c100bb5-d000-4492-85d0-a4cb300e46b1","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":14744,"visible":true,"origin":"","legend":"\u003cp\u003eChar Embedding Layer Model Structure Diagram\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/89f315e853ea66e93fe445bc.png"},{"id":39195092,"identity":"ccf9c0a0-9351-484d-9961-e426480aa9d1","added_by":"auto","created_at":"2023-06-27 20:58:12","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":23514,"visible":true,"origin":"","legend":"\u003cp\u003eExample of an Embedding layer process\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/7a61c573de5ca7bc26ff166c.png"},{"id":39194301,"identity":"0966cba7-2f96-4658-be1f-597186d088dd","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":32960,"visible":true,"origin":"","legend":"\u003cp\u003eEntity coding structure diagram based on Bi-GRU and Attention\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/c32cea2ff5b982cbf902c790.png"},{"id":39194296,"identity":"412d8216-3acf-4e20-a0f2-8f9df1b2e2fb","added_by":"auto","created_at":"2023-06-27 20:50:12","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":23245,"visible":true,"origin":"","legend":"\u003cp\u003eGRU Network Structure Diagram\u003c/p\u003e","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/f4b4ffca9783eb9d9bd10080.png"},{"id":39195352,"identity":"06af186f-bbd4-4784-b4a3-e39ff8d90c12","added_by":"auto","created_at":"2023-06-27 21:06:12","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":27323,"visible":true,"origin":"","legend":"\u003cp\u003eThe calculation process of Attention\u003c/p\u003e","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/71b1a7092436dfcfc6edaef0.png"},{"id":39195093,"identity":"bb188d25-3853-4eb6-a0ce-07415ca49914","added_by":"auto","created_at":"2023-06-27 20:58:12","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":58554,"visible":true,"origin":"","legend":"\u003cp\u003eSMAAEA model training flowchart\u003c/p\u003e","description":"","filename":"12.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/00360c54e214dfcc99f70a21.png"},{"id":39194304,"identity":"c1fd7b9a-346a-4f27-b62b-d57717453c32","added_by":"auto","created_at":"2023-06-27 20:50:13","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":218082,"visible":true,"origin":"","legend":"\u003cp\u003eOverall flowchart of knowledge fusion\u003c/p\u003e","description":"","filename":"13.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/fab8fa76e2310c6782a51f47.png"},{"id":39195499,"identity":"7551f534-13a9-4f4a-bc6f-db4d35334a71","added_by":"auto","created_at":"2023-06-27 21:14:12","extension":"png","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":20647,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of F1 values of SMAAEA with other network models under different iterations\u003c/p\u003e","description":"","filename":"14.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/94ed1b6a5447fc230ca08b0e.png"},{"id":39195097,"identity":"7da4556c-5d18-4260-b9a2-e5fbeaa6304c","added_by":"auto","created_at":"2023-06-27 20:58:13","extension":"png","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":23251,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of F1 values of SMAAEA and its simplified model with different iterations\u003c/p\u003e","description":"","filename":"15.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/a2f7cefba9e9fabd1219f352.png"},{"id":39194303,"identity":"da9803e9-c7b3-4fe4-9e03-2733fe71cd2a","added_by":"auto","created_at":"2023-06-27 20:50:13","extension":"png","order_by":16,"title":"Figure 16","display":"","copyAsset":false,"role":"figure","size":35812,"visible":true,"origin":"","legend":"\u003cp\u003eVariation of F1 value in the civil aviation text test set\u003c/p\u003e","description":"","filename":"16.png","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/2b422f5f8d12c0de9f03f29c.png"},{"id":39195550,"identity":"4507c2f0-2a3a-4d8a-b055-8baf77c8ab04","added_by":"auto","created_at":"2023-06-27 21:14:21","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1138793,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3066877/v1/01c65c61-21a7-4c63-acbb-49970efe9df2.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Research on multi-attribute attentional civil aviation entity alignment based on Siamese network","fulltext":[{"header":"Introduction","content":"\u003cp\u003eThe civil transport industry has developed rapidly in recent years, especially in the aviation field. The total transport turnover increases continually. As a main part of transport, the scale of aviation keeps expanding. The construction of civil aviation emergency management system will face new challenges and development opportunities\u003csup\u003e1\u003c/sup\u003e. In 2020, due to the impact of the COVID-19, airline passenger volume dropped significantly. Meanwhile, due to the weak operation capacity of airlines, airport equipment failure, high-altitude traffic control, and bad weather, the safety situation of civil aviation is severe and complicated and faces many challenges: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) Due to the particularity of the aviation field, serious consequences will occur in case of sudden emergency or timely and efficient measures are not taken; (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) The causes of civil aviation accidents are complex, such as various types of natural disasters, accidental disasters, human factors and problems with the aircraft itself, etc; (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) The imperfection of aviation management system for emergencies, and the emergency management ability needs to be further improved. Due to the particularity of civil aviation, it is of great significance to study the construction of knowledge map of civil aviation emergencies for the emergency management system to make rapid and effective decisions in the face of complex and diverse emergencies. The key components of civil aviation knowledge map include knowledge extraction and knowledge fusion. In the process of constructing domain knowledge graph, there are still some problems, such as lack of communication between information and has great difference between information.\u003c/p\u003e \u003cp\u003eMoreover, most knowledge is usually crawled through web pages, or comes from some non-standard data sets, which is easy to cause some problems of text information knowledge, such as content cross, multi-source heterogeneous, etc. Therefore, further expanding knowledge fusion on the basis of entity recognition and relationship extraction is the key to establish and expand the knowledge map of civil aviation emergencies. By using the knowledge fusion technology, the description from different entities is unified. Then these entities are input into the target knowledge base to make the knowledge map more accurate and enrich the content of knowledge, which is of great significance to further improve the efficient decision-making when facing emergencies in the field of civil aviation and form a complete and scientific emergency management system.\u003c/p\u003e \u003cp\u003eThe construction technology of knowledge graph for civil aviation emergencies is researched in this paper. Aiming at the problems of multi-source and heterogeneous knowledge are difficult to organize and integrate effectively in civil aviation, further fusion is carried out on the basis of knowledge extraction. On the background of civil aviation emergencies, we focuses on the entity alignment method in knowledge fusion based on knowledge extraction to achieve information symmetry between entities and eliminate ambiguity. The main research content is as follows: For a large number of complex and mixed entities in civil aviation emergencies, and the entities cannot be extracted automatically and the names are long, it leads to low extraction accuracy. Thus, a Siamese-network Multiple Attribute Attention Entity Alignment (SMAAEA) model is proposed. The information between entities can be symmetrical, and the diversity of knowledge graph is improved, so as to overcome the problems of low recall rate and the contribution to entities of important attributes cannot be distinguished, which are caused by the inability to use entities to describe the details of text information. It improves the quality and confidence of the knowledge graph, and provides theoretical support for the construction of the knowledge graph in the civil aviation field.\u003c/p\u003e \u003cp\u003eThe remainder of this paper is organized as follows: In section 1 we summarize the research background and significance of this paper. Then, the process and key technology of entity alignment is introduced, and proposes the entity alignment method with SMAAEA model in section 2. In section 3, experiments are carried out to verify the proposed model. We summary our work briefly in section 4.\u003c/p\u003e"},{"header":"Related work","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eResearch status of knowledge graph\u003c/h2\u003e \u003cp\u003eThe historical evolution of knowledge graphs is broadly divided into three phases. In the first phase (late 1950s and early 1960s), the concept of Semantic Network was first proposed by M. Ross Quillian and Robert F. Simmons \u003csup\u003e2\u003c/sup\u003e, when the Semantic Network was used to achieve the storage of data information through the form of graphs, which were easier to understand and display, and the large amount of information using the Semantic Web is able to more concise and effective way to achieve information expression and storage by means of diagrams. In the second phase (mid-1980s and early 2000s), the Semantic Web has developed rapidly. The Knowledge Graph was developed from the Semantic Web to enable computers to access and use knowledge as humans do. The category of ontology, as \"an explicit specification of a conceptual model \u003csup\u003e3\u003c/sup\u003e, describes knowledge at the semantic level, and computers can process and recognize ontologies, so that information processing, flow, and interaction between humans and computers can be achieved, making human-computer interaction simpler and clearer. The third phase (early 1920s-present) is a boom phase in the development of knowledge graphs, such as YAGO\u003csup\u003e4\u003c/sup\u003e, DBpedia\u003csup\u003e5\u003c/sup\u003e, Freebase\u003csup\u003e6\u003c/sup\u003e, etc. In the academic world, the first knowledge graph construction project that adopts the combination of Chinese and English with different language features is Xlore\u003csup\u003e7\u003c/sup\u003e and OpenKE platform, which was built by Li Juanzi's team at Tsinghua University and created a system system for media event analysis and mining, entity linking, etc. based on Xlore; Fudan University and Shanghai Jiaotong University rely on Wikipedia and Baidu Encyclopedia as the basis, considering that Based on Wikipedia and Baidu, Fudan University and Shanghai Jiao Tong University created zhishi.me \u003csup\u003e8\u003c/sup\u003e and CN-pedia\u003csup\u003e9\u003c/sup\u003e Chinese knowledge graph platforms, considering the complexity of obtaining information from encyclopedias; the Institute of Computer Technology of the Chinese Academy of Sciences applied Open KN to build a prototype system of \"people cube, matter cube, and knowledge cube\"; In 2017, the cnSchema.org project\u003csup\u003e10–11\u003c/sup\u003e, which is used to maintain the schema layer of open domain knowledge graphs, was initiated by several universities in China. The main advantages of these knowledge graphs above are that they can be applied in different domains with a wide range of applications, with search functions and question and answer service systems. However, there are relatively few open data links and databases, and there is a lack of rich structured Chinese data. Meanwhile, in the commercial world, the search engines established by Sogou and Baidu are widely used. The development of these search engines has made the search method more convenient and intelligent, satisfying the needs of users and bringing a good experience at the same time.\u003c/p\u003e \u003cp\u003eIn recent years, knowledge graph technology \u003csup\u003e12–19\u003c/sup\u003e has continued to develop and improve, and has been widely used in many fields such as finance, medicine, social media, fisheries, animal health, military industry, ocean, and agriculture.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eResearch status of knowledge fusion\u003c/h2\u003e \u003cp\u003eTwo main methods, entity alignment and entity linking, are used in the knowledge fusion process of multi-source knowledge bases. Early studies by researchers on entity alignment techniques20 mainly use the relationships between entities to determine, whether the entities are successfully aligned with each other, such as rule-based entity alignment techniques, which require manual intervention to customize the rules by themselves. The literature 21proposes that the degree of similarity of attributes can be determined by the sides of a triangle. The literature 22 proposes to find matching laws in a continuous iterative learning process when there is a part of entity seeds that have already been matched. For the text information that is not annotated, this matching law is used to annotate and get more candidate matching seeds. The literature 23 evaluates entities by labeling and scoring them. The methods are supervised learning, semi-supervised learning and unsupervised learning, respectively. The literature24 proposes the use of semi-supervised learning for joint training of entities to achieve entity alignment. The literature25 proposes an entity alignment approach in Conditional Random Field (CRF), in which the distance function of similarity between entities is computed so as to adjust the learning parameters and the characteristic function to maximize their product. The most representative models are Support Vector Machine (SVM)26, integrated learning27, and decision trees28. The literature29 proposes a support vector machine based entity alignment approach under supervised learning. When there is no way to practice and analyze a large amount of data, more practice data can be obtained by active learning with continuous information exchange with the annotators. The literature30 proposed the ALIAS system to deal with the redundancy problem between similar entities through information exchange between humans and machines. As the research progresses, the Deep Learning (DP) based entity alignment model represents the features and attribute values of entities as vectors, which in turn form dense vectors to determine whether the two are the same entity between them. The literature 31 proposes a TextCNN convolutional neural network that randomly initializes the text vectors in the text to derive a feature vector representation of each text, extracts the feature representation of the text information using the convolutional layer, and outputs the final vector representation in the fully connected layer. The literature32 introduces TextRNN, which can recognize text information containing time series, input the text information into the RNN in an orderly manner to obtain the sequence feature representation, and the number of neurons in the hidden layer can be used as an abstract representation of the text information. The literature33 mentions FastText, which performs text vectorization representation of text data in the word2vec34 pre-training model, and introduces an n-grams model based on this, which is used to count the probability of occurrence of text information, and finally sums all the input feature vector representations and then averages them as the final input text vector representation, and then finally uses the softmax layers. The current research approach suffers from the problems arising from over-reliance on machine learning-based and rule-based entity alignment tasks: first, there is no way to use a large number of entities for detailed portrayal of text information for manually limited features, which results in low recall rate, and second, there is no way to distinguish the contribution of important attributes to entities. In this paper, to address the problem of information asymmetry among civil aviation emergency entities and the lack of differentiation of the importance of different attributes to entities, SMAAEA is proposed to improve the symmetry of information among entities based on knowledge extraction. By mapping entities based on different knowledge bases in the civil aviation domain to the same entity in the target knowledge base, the problem of information asymmetry among entities is overcome and the importance of different attributes to the same entity can be distinguished.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eThe general research idea of entity alignment\u003c/h2\u003e \u003cp\u003eSuppose the set of entities obtained by knowledge extraction is \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(E=\\left\\{{e}_{1},{e}_{2},\\dots {e}_{M}\\right\\}\\)\u003c/span\u003e\u003c/span\u003e, and the set of entities in the civil aviation knowledge base is \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(K=\\left\\{{k}_{1},{k}_{2},\\dots {k}_{N}\\right\\}\\)\u003c/span\u003e\u003c/span\u003e, the goal of entity alignment is to align the different entities \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{i}\\)\u003c/span\u003e\u003c/span\u003e in the set of entities \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(E\\)\u003c/span\u003e\u003c/span\u003e of the civil aviation database obtained by knowledge extraction through crawling corresponds to the different entities \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({k}_{j}\\)\u003c/span\u003e\u003c/span\u003e in the civil aviation knowledge base, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOverall idea : Firstly, we pre-process the data in the civil aviation emergency text set and the civil aviation emergency text set obtained after knowledge extraction, and then we label the processed data and train the vectors; further, we construct the Siamese BERT network multi-attribute attention alignment model, and input the processed entity vectors into the alignment model to obtain the complete serialized entity with importance distinction vector with importance distinction.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eData preprocessing\u003c/h3\u003e\n\u003cp\u003eThe pre-processing process of the data is: firstly, the civil aviation emergency text data are organized and cleaned; secondly, the data are used to create a dataset by rule matching, and finally, the entities in the dataset are annotated by automatic annotation method.\u003c/p\u003e \u003cp\u003eThe rule-based entity alignment is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The specific process is: firstly, the new entities obtained after crawling and knowledge extraction. Secondly, the entity names of the new entities are matched separately and simultaneously by a combination of exact and fuzzy matching methods to find the entities in the civil aviation database, and the new entities are compared with the original candidate entities in turn, and finally the four attribute values such as flight number, departure location, arrival location, and flight time generated by the entities are matched separately by a more detailed multi-level matching method. The specific rules are as follows.\u003c/p\u003e \u003cp\u003eFirstly, if the flight number of the paired entities can be matched, then they can be mapped to the same entity; secondly, if the flight number of the aircraft cannot be matched, then the flight departure location of the paired entities are matched, and if they can be matched, then they can be mapped to the same entity; finally, if both the flight number of the aircraft and the flight departure location cannot be Finally, if neither the flight number of the aircraft nor the departure location of the aircraft can be matched, the arrival location of the aircraft of the paired entity is matched, and if it can be matched, it means that they can be mapped as the same entity. And so on, if the flight number, flight departure location and flight arrival location cannot complete the matching, then the matching of flight time is completed, if the matching is completed, it indicates that the entities can be aligned, otherwise it is considered that the entities cannot be aligned.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe number of alignable and non-alignable entities can be obtained in the above way, which are 11119 and 48350 in order. the alignable entities, non-alignable entities and entities in the civil aviation knowledge base are counted to generate a new civil aviation dataset, with successful alignment between entities represented by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(classify=1\\)\u003c/span\u003e\u003c/span\u003e and failed alignment between entities represented by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(classify=0\\)\u003c/span\u003e\u003c/span\u003e. And each set of candidate entities with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(classify=0\\)\u003c/span\u003e\u003c/span\u003e is detected by matching technique for further alignment. The positive and negative examples in the dataset are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eStatistical table of emergency data sets in the field of civil aviation\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eTraining set\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e \u003cp\u003eTest set\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9000\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePositive example: 4500 Negative example༚4500\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003e2000\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003ePositive example: 1000\u003c/p\u003e \u003cp\u003eNegative example: 1000\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10000\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePositive example: 5000 Negative example༚5000\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11000\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePositive example: 5500 Negative example༚5500\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003eAt the same time, to better improve the quality of word separation, we match the airline name text base, flight number text base and starting place name text base in the domain dictionary, so that we can better construct a professional dictionary for the civil aviation domain. In order to align the entities more accurately, the data corpus is labeled by BIEO (Begin, Inside, End, Outside) method, which can effectively determine the boundary of the entities. Among them, B represents the first character of the entity participle of the successfully predicted category, I represents the entity participle character in the middle position of the successfully predicted category, E represents the end part character of the successfully predicted category, and O represents the predicted character that is not an entity. An example of the labeling is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eBuild Siamese network based multi-attribute attentional entity alignment model\u003c/h2\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003eBuild Siamese network based model\u003c/h2\u003e \u003cp\u003eThe Siamese network is proposed to solve the problem of similarity metrics in face verification. The core idea is that if a set of data is input and mapped into a high-dimensional feature space, then all its samples of text information can be described as pairs of high-dimensional feature information, so that the high-dimensional feature information can be used to determine the similarity degree of sample pairs. For example, Google's BERT pre-training model can match two sentences based on their semantic similarity, and it has achieved good results so far. However, if N sentences are input at the same time, it takes O(N*N) time to match the semantic similarity of each sentence one by one and find out the most similar sentence among them, which greatly increases the query time. Since the SBEART network system can handle the problem of semantic similarity between texts, we put it through the Siamese BERT network model which can also capture the connection between sentences and perform the complete representation, significantly reducing the amount of operations. Since the entities used for entity alignment in the construction of the knowledge graph of civil aviation domain are all from the civil aviation knowledge base, and the structures of these entities are similar, the above-mentioned characteristics of Siamese network can be used to encode these structurally similar entities. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, entity 1 of the Siamese network is input to one of the sub-networks for encoding to get the entity representation vector u, and entity 2 of the Siamese network is input to the other sub-networks to get the entity representation vector v. The two high-dimensional feature information representation vectors u and v are spliced, and the sigmoid activation function is used in the fully connected layer to compute the high-dimensional feature vector for Classification.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn this regard, the two-part network model structure of the Siamese network is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, and the two different neural networks of the Siamese network encode the entities correspondingly. When encoding the Bi-LSTM network, the attribute name vector representation and the attribute value vector representation containing character features are spliced and input to the Bi-LSTM network, so that not only the text information can be described but also the information contained in the encoding can be learned completely. When encoding the Bi-GRU network, an attention mechanism is introduced to assign weights to the encoding, and the resulting high-dimensional feature entity vector representation outputs the entity feature vector representation with importance.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eBuild Entity Encoding with Siamese Network\u003c/h2\u003e \u003cp\u003eIn this paper, SMAAEA model is proposed based on Siamese network, and the two parts of entity vectors in SMAAEA model are based on different encoding methods. One is to input the mixed vector of words containing character features into Bi-LSTM for encoding, which can improve the integrity of encoding by using Bi-LSTM to get the timing information in both positive and negative directions; the other is to input the shorter attribute text into Bi-GRU + Attention mechanism for encoding, and get the final The final encoding is obtained by the weight parameter in the Attention mechanism. The encoding of the two molecular networks is summed to obtain a high-dimensional feature information vector representation. The calculation is shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e):\u003c/p\u003e\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$${e}_{i}=Add\\left({e}_{i1},{e}_{i2}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\({e}_{i1}\\)\u003c/span\u003e \u003c/span\u003e denotes the Govett vector of entities output by the Bi-LSTM network, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{i2}\\)\u003c/span\u003e\u003c/span\u003e denotes the high-dimensional feature vector obtained from the output of the Bi-GRU + Attention network, and the two parts of the encoding will be elaborated separately in the following.\u003c/p\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003eBuild Bi-lstm-based Entity Encoding\u003c/h2\u003e \u003cp\u003eBased on the characteristics of the entity structure of civil aviation emergencies and considering the network characteristics of Bi-LSTM, it can not only distinguish the front-to-back order of words in, but also encode from back to front, so in the input part of the Bi-LSTM encoding layer, the attributes of the entities and different attributes are spliced for finer-grained classification as the starting input. Considering the long length of entities in civil aviation field and the existence of a large number of mixed and composite entities, firstly, a word embedding layer and a char feature layer are added to the initial coding layer, and the initial feature vector containing character information can be obtained in the embedding layer, and then the initial feature vector containing character information is used as the input of the Bi-LSTM layer. Finally, for each word vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({w}_{ij}\\)\u003c/span\u003e\u003c/span\u003e, the information in the context of the two temporal directions before and after the entity can be captured using the bi-directional LSTM, and the intensity relationship between different words can be captured. The feature vectors containing character information are stitched to obtain the new representation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\overrightarrow{{\\text{h}}_{\\text{i}}}=\\left[\\overrightarrow{{\\text{h}}_{\\text{i}\\text{j}}},\\overleftarrow{{\\text{h}}_{\\text{i}\\text{j}}}\\right]\\)\u003c/span\u003e\u003c/span\u003e, and this new representation contains bidirectional semantic relationships. The output of the model is a stitching of high-dimensional feature vectors of bidirectional semantics in textual information from the civil aviation domain, and the structure of Bi-LSTM is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe Bi-LSTM-based entity vector\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{i}\\)\u003c/span\u003e\u003c/span\u003e is calculated as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ2\" class=\"InternalRef\"\u003e2\u003c/span\u003e):\u003c/p\u003e\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$${e}_{i}=BiLSTM\\left({w}_{i1},{w}_{i2},\\dots .{w}_{iL}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eAmong them, the Embedding layer can reduce the dimensionality of the data by matrix multiplication without changing the information contained in the word vectors. Because of the similarity of some entities and the similarity of words contained in the names of two entities, the word feature vector representation is combined with the character feature vector representation in the Embedding layer, and the vectorized text information with character features is finally obtained by vectorizing the Char Embedding layer.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAmong them, the word embedding layer is used to extract the context vectors of the words in the text, and the char feature layer is used to extract the conformational feature vector and the glyph feature vector of each word, and the vectors and character feature vectors in the civil aviation text are obtained through the char feature layer and the word embedding layer, and then the char embedding layer to incorporate the character feature vectors and lexical information into the character representation, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eThe specific training process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. Firstly, the civil text subwords are extracted, and then the feature vectors of different words and characters are extracted and pre-trained. Since in one-hot coding, each word corresponds to an n-dimensional feature vector representation, the sparse matrix obtained by such a representation will cause a waste of resources, therefore, the word feature vector representation is one-to-one corresponded with the character feature vector representation in the Embeding layer to turn the sparse matrix into a dense matrix, and by adding the word vectors containing character feature information with the subvector sequence, the The vector representation of the civil aviation textualization.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eBuild Entity Encoding with Bi-GRU + Attention\u003c/h2\u003e \u003cp\u003eWe use the Bi-LSTM entity encoding method, and the final obtained entity feature representation vector representation only contains the information of attributes and attribute values, but there is no precise differentiation of the specific impact of different attribute values on each attribute. In this paper, we change the input of Char Embedding layer, instead of just inputting words, we input a shorter sentence containing attribute features, and the encoding structure is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e. Since each entity not only consists of ontology, but also contains different attributes and attribute values, the feature vector representation output by Bi-GRU is input to the attribute encoder, and the entity information of the feature vector representation is encoded to output the feature vector text information representation with attribute features. Then, the obtained text information representations are interacted with the Attention mechanism to obtain feature vector text information representations with different attributes.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eSince GRU belongs to the derivative structure of LSTM, the number of gating units used is reduced by merging the input gate control unit and forgetting gate control unit in LSTM, and the parameters are reduced accordingly, so the training speed of textualized information can be improved without changing the learned temporal information. Based on the above analysis, the entity coding based on Bi-GRU + Attention is proposed, and the attribute information is added in the coding process.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe reset gate in GRU can better capture the dependencies at shorter distances in time order, and the update gate in GRU can better capture the dependencies at longer times in time order. bi-GRU network can obtain the timing information in two opposite directions, and compared with Bi-LSTM, Bi-GRU has fewer network parameters, which can improve the training speed. In practice, the GRU performance does not differ much from the LSTM performance, and the training speed is improved. the working structure of GRU is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eThe calculation process of reset gate control unit and update gate control single gate is shown in Equations (\u003cspan refid=\"Equ3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) and (\u003cspan refid=\"Equ4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$${r}_{t}=\\sigma ({W}_{xr}\\cdot {x}_{t}+{W}_{hr}\\cdot {h}_{t-1}+{b}_{r})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$${z}_{t}=\\sigma \\left({w}_{xz}{X}_{t}+{w}_{hc}\\cdot {h}_{t-1}+{b}_{z}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eThe candidate hidden layer unit is calculated by first multiplying the output of the reset gate at the current time point by the hidden layer unit at the same time, then connecting the current time point sequence, and finally deriving the sequence unit of the candidate hidden layer by the operation on the tanh activation function, as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$${h}_{t}^{{\\prime }}={tan}h({W}_{xh}\\cdot {x}_{t}+({r}_{t}\\cdot {h}_{t-1})\\cdot {W}_{hh}+{b}_{h})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eAt moment \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(t\\)\u003c/span\u003e\u003c/span\u003e, the hidden layer state is \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({h}_{t}\\)\u003c/span\u003e\u003c/span\u003e, and the update gate control unit \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({z}_{t}\\)\u003c/span\u003e\u003c/span\u003e of the current time series is multiplied with the hidden state feature vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({h}_{t-1}\\)\u003c/span\u003e\u003c/span\u003e at the previous moment and multiplied with the hidden layer state feature vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({h}_{t}^{{\\prime }}\\)\u003c/span\u003e\u003c/span\u003e at the current moment for the integrated operation as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ6\" class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$${h}_{t}={z}_{t}\\cdot {h}_{t-1}+(1-{z}_{t})\\cdot {h}_{t}^{{\\prime }}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eFor the word text information vectors, after inputting them in the order they are in the sentence, for different words \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({w}_{ij}\\)\u003c/span\u003e\u003c/span\u003e, input to the Bi-GRU network, the text information feature vector representation containing two different directions before and after can also be obtained, and then these textualized feature vectors are stitched together to obtain the new feature vector representation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({h}_{i}=\\left[\\overrightarrow{{h}_{\\text{i}j}},\\Leftarrow {h}_{\\text{i}j}\\right]\\)\u003c/span\u003e\u003c/span\u003e. The calculation procedure is shown in equations (\u003cspan refid=\"Equ7\" class=\"InternalRef\"\u003e7\u003c/span\u003e) and (\u003cspan refid=\"Equ8\" class=\"InternalRef\"\u003e8\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$${h}_{ij}=BiGRU\\left({w}_{i1},{w}_{i2},\\dots .{w}_{iL}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$att{r}_{i}=Add\\left({h}_{i1},{h}_{i2},\\dots {h}_{iL}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eAfter the Bi-GRU network, the different attributes of each entity feature vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{i}\\)\u003c/span\u003e\u003c/span\u003e are passed through the attribute encoder to obtain \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N\\)\u003c/span\u003e\u003c/span\u003e feature vector text representations containing attribute information, namely \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(att{r}_{i1},att{r}_{i2},...,att{r}_{iN}\\)\u003c/span\u003e\u003c/span\u003e. Finally, the final encoding is obtained by integrating the feature text vectors containing attributes through the Attention mechanism. The computation process is shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ9\" class=\"InternalRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$${e}_{i}=Attention(att{r}_{i1},att{r}_{i2},att{r}_{iN})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eAttention can be specifically divided into three steps, and the specific process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn the first stage, the relationship between the two is calculated based on the Query and Key values, as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ10\" class=\"InternalRef\"\u003e10\u003c/span\u003e):\u003c/p\u003e\u003cdiv id=\"Equ10\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ10\" name=\"EquationSource\"\u003e\n$${a}_{i}=softmax\\left(si{m}_{i}\\right)=\\frac{{e}^{si{m}_{i}}}{{\\varSigma }_{j=1}^{{L}_{x}}{e}^{si{m}_{j}}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eIn the second stage, the operator of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Softmax\\)\u003c/span\u003e\u003c/span\u003e is introduced, whose purpose is to do a numerical transformation of the results of the first paragraph, calculated as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ11\" class=\"InternalRef\"\u003e11\u003c/span\u003e):\u003c/p\u003e\u003cdiv id=\"Equ11\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ11\" name=\"EquationSource\"\u003e\n$$Attention(Query,Source)={\\sum }_{i=1}^{{L}_{x}}{a}_{i}\\cdot Vlau{e}_{i}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e11\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eIn the third stage, the corresponding values are weighted by their weights to obtain the attention vector.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eBuild Classification module\u003c/h2\u003e \u003cp\u003eFor a specified set of candidate entities, such as Entity 1 and Entity 2, the first entity is first input into the Siamese network to obtain the feature vector representation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{11}\\)\u003c/span\u003e\u003c/span\u003e with completeness by Bi-LSTM coding; then the same entity vector is input into the Bi-GRU + Attention network model to obtain the feature representation vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{12}\\)\u003c/span\u003e\u003c/span\u003e with importance; finally, the feature vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{11}\\)\u003c/span\u003e\u003c/span\u003e and the feature vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{12}\\)\u003c/span\u003e\u003c/span\u003e are added together to obtain the complete high-dimensional feature representation vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{1}\\)\u003c/span\u003e\u003c/span\u003e and the feature vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{12}\\)\u003c/span\u003e\u003c/span\u003e are added to obtain the complete high-dimensional feature representation vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({e}_{1}\\)\u003c/span\u003e\u003c/span\u003e with importance representation. and so on, the same operation is performed for another entity to obtain two high-dimensional feature vector representations. The two feature vector representations are stitched in the fully connected layer, and the stitched feature vector representation is predicted. The SMAAEA model training process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig12\" class=\"InternalRef\"\u003e12\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe calculation steps are shown in Equations (\u003cspan refid=\"Equ12\" class=\"InternalRef\"\u003e12\u003c/span\u003e), (\u003cspan refid=\"Equ13\" class=\"InternalRef\"\u003e13\u003c/span\u003e), (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e), and (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ12\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ12\" name=\"EquationSource\"\u003e\n$${e}_{1}=TN({w}_{11},{w}_{12},\\dots {w}_{1L};att{r}_{11},att{r}_{12},\\dots ,att{r}_{1M})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e12\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ13\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ13\" name=\"EquationSource\"\u003e\n$${e}_{2}=TN({w}_{21},{w}_{22},\\dots {w}_{2L};att{r}_{21},att{r}_{22},\\dots ,att{r}_{2M})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e13\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ14\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ14\" name=\"EquationSource\"\u003e\n$$e=Concate\\left({e}_{1},{e}_{2}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e14\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ15\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ15\" name=\"EquationSource\"\u003e\n$$py=sigmoid\\left(e\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e15\u003c/div\u003e\u003c/div\u003e\u003cp\u003e\u003c/p\u003e \u003cp\u003eIn the above Equations, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(TN\\)\u003c/span\u003e\u003c/span\u003e is represented as Siamese network, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(w\\)\u003c/span\u003e\u003c/span\u003e is character feature vector representation, attr is attribute feature vector representation, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(att{r}_{ij}=BiGRU({w}_{i1},{w}_{i2},\\dots {w}_{is})\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(sigmod\\left(y\\right)=\\frac{1}{1+{\\text{e}}^{-y}}\\)\u003c/span\u003e\u003c/span\u003e, and the final prediction result \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(py\\)\u003c/span\u003e\u003c/span\u003e by two classification method on the high-dimensional feature vector representation. In the model encoding process, a group of entities share the weights and parameters in the model.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eBuild Entity alignment fusion module\u003c/h2\u003e \u003cp\u003eDue to the specificity of the civil aviation domain and the fact that the data used are jointly constructed by knowledge extraction and civil aviation knowledge base data, it results in certain differences in the types of entities and the names of the attributes corresponding to the entities in the data from different sources. If knowledge alignment fusion is not performed on these different representations of the same entity, it will lead to the confusion of entity links and reduce the quality of the knowledge graph, so it is necessary to further process the entities through knowledge fusion, and in this paper, the SMAAEA model is applied to build the alignment fusion module in knowledge fusion.\u003c/p\u003e \u003cp\u003eAlignment fusion is the integration of information from different sources. When fusing the external civil aviation knowledge base with the original civil aviation knowledge base, two aspects need to be considered: one is the fusion of data quality, i.e., whether the entities can be aligned, their attributes and categories are considered, and the redundancy and conflicts between entities are resolved, and then the new entities obtained are fused into the civil aviation knowledge base through the Schema layer; the other is the merging of knowledge bases in the same domain, and the fusion of historical data with similar structure of historical data are fused and the data are converted into RDF triples. The process of alignment fusion is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig13\" class=\"InternalRef\"\u003e13\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eFirst, the table of stored entities is checked to detect whether there is any entity data different from that in the original civil aviation knowledge base, and if a new entity representation is detected, then the entity is processed according to the table of stored entities. The new entities contain two parts, one part is detected from the external entity repository and the other part is from the local civil aviation entity repository, and the entities in the local civil aviation entity repository are used as candidate entities. The candidate entities are matched in multiple directions, and if the matching is successful, the candidate entities are input into the SMAAEA network model together with the new entities detected from the external knowledge base, from which it is judged whether the two describe the same object. If the entities are successfully aligned, the next aligned entities are unified in terms of attributes and input into the local civil aviation knowledge base together.\u003c/p\u003e"},{"header":"Experiment and Analysis","content":"\u003ch2\u003eExperimental dataset\u003c/h2\u003e\u003cp\u003eThe text information of civil aviation accidents used in the experiments of this paper is more than 15,000 data, which are jointly constructed by the crawled investigation tracking reports about safety accidents in aviation field released by the Civil Aviation Safety Information System of China \u003csup\u003e1\u003c/sup\u003e, the expanded dataset, and the dataset obtained by knowledge extraction. These text information are divided into 9000, 10000, and 11000 data for training, and the number of text data in the test set is 2000 data for testing. The SMAAEA model will divide the data according to 9:1 during training, and the entity labeling method is shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMethods of entity annotation\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEntity types\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEntity start\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eInside the entity\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEntity ending\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTime of incident\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-DATE\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-DATE\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-DATE\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAirlines\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-ARIL\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-ARIL\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-ARIL\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eType of aircraft\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-FTYPE\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-FTYPE\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-FTYPE\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFlight number\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-FNUM\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-FNUM\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eI-FNUM\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRegistration number\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-REG\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-REG\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-REG\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeparture airport\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-ORIG\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-ORIG\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-ORIG\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDestination\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-DEST\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-DEST\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-DEST\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAircrew\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-CREW\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-CREW\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-CREW\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of passengers\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-PSG\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-PSG\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-PSG\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ecasualties\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-CASU\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-CASU\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-CASU\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCause of accident\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-REA\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-REA\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-REA\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccident result\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eB-RES\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eI-RES\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eE-RES\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNonentity type\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eO\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eO\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eO\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e\u003ch2\u003eExperimental environment\u003c/h2\u003e\u003ch2\u003eLoss function\u003c/h2\u003e\u003cp\u003eThe cross-entropy loss function is used to evaluate the SMAAEA network model since it can provide a measure of the variability of the probability distribution of two samples.\u003c/p\u003e\u003cp\u003eFor a set of entity pairs \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\((px,py)\\)\u003c/span\u003e\u003c/span\u003e, the candidate sample prediction results are denoted by py, and the set of prediction results taking values \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\left\\{\\text{0,1}\\right\\}\\)\u003c/span\u003e\u003c/span\u003e. A sample with actual label information quantity \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(y\\)\u003c/span\u003e\u003c/span\u003e, the probability of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(y=1\\)\u003c/span\u003e\u003c/span\u003e is denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e, and the probability of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(y=0\\)\u003c/span\u003e\u003c/span\u003e is denoted by\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\((1-p)\\)\u003c/span\u003e\u003c/span\u003e, and the formula of the cross-entropy loss function is shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ16\" class=\"InternalRef\"\u003e16\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ16\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ16\" name=\"EquationSource\"\u003e\n$${log}\\left(\\left.y\\right|p\\right)=-\\left[y*{log}\\left(p\\right)+\\left(1-y\\right){log}\\left(1-p\\right)\\right]$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e16\u003c/div\u003e\u003c/div\u003e\u003cp\u003eFor all text data, the smaller the cross-entropy loss is, the better the SMAAEA model proposed in this paper is proved to be, and the loss function of the whole model is expressed as the average of all sample loss functions, as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ17\" class=\"InternalRef\"\u003e17\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ17\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ17\" name=\"EquationSource\"\u003e\n$$loss=\\frac{1}{m}\\sum _{i}^{m}\\text{log}\\left(\\left.{y}_{i}\\right|{p}_{i}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e17\u003c/div\u003e\u003c/div\u003e\u003ch2\u003eParameter setting\u003c/h2\u003e\u003cp\u003eThe parameter settings for this experiment are shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e:\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter settings\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eParameters\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSetting values\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emax_length\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e512\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emax_attribute\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebatch_size\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elearning_rate\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1e-3\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eepoch\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003edropout\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eoptimizer\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdam\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e\u003ch2\u003eExperimental environment\u003c/h2\u003e\u003cp\u003eThe experimental environment for this experiment is shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eExperimental Environment parameter settings\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"2\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExperimental environment\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eParameter setting\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSystem environment\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWindows 10\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGPU parameters\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e24 cores、128G、1080Ti\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePython version\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3.6.7\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKeras version\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.2.5\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTensorFlow version\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.14.0\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e\u003ch2\u003eExperimental evaluation index\u003c/h2\u003e\u003ch2\u003ePerformance metrics\u003c/h2\u003e\u003cp\u003eIn this paper, the harmonic mean F1 is used to measure the precision and recall, and is used as the performance index of the SMAAEA model.\u003c/p\u003e\u003cp\u003eThe accuracy rate P represents the proportion of the real results in the sample that are correct to the predicted results that are positive, as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ18\" class=\"InternalRef\"\u003e18\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ18\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ18\" name=\"EquationSource\"\u003e\n$$P=\\frac{TP}{TP+FP}=\\frac{\\begin{array}{c}The EquationNumber of instances where the true class is positive \\\\ and the model discrimination is also positive\\end{array}}{\\text{T}\\text{h}\\text{e} \\text{n}\\text{u}\\text{m}\\text{b}\\text{e}\\text{r} \\text{o}\\text{f} \\text{i}\\text{n}\\text{s}\\text{t}\\text{a}\\text{n}\\text{c}\\text{e}\\text{s} \\text{w}\\text{h}\\text{e}\\text{r}\\text{e} \\text{t}\\text{h}\\text{e} \\text{p}\\text{r}\\text{e}\\text{d}\\text{i}\\text{c}\\text{t}\\text{i}\\text{o}\\text{n} \\text{r}\\text{e}\\text{s}\\text{u}\\text{l}\\text{t} \\text{i}\\text{s} \\text{p}\\text{o}\\text{s}\\text{i}\\text{t}\\text{i}\\text{v}\\text{e}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e18\u003c/div\u003e\u003c/div\u003e\u003cp\u003eRecall rate R represents the proportion of predicted correct results to actual correct results in the sample, as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ19\" class=\"InternalRef\"\u003e19\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ19\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ19\" name=\"EquationSource\"\u003e\n$$R=\\frac{TP}{TP+FN}=\\frac{\\begin{array}{c}The EquationNumber of instances where the true class is positive \\\\ and the model discrimination is positive\\end{array}}{\\text{T}\\text{h}\\text{e} \\text{n}\\text{u}\\text{m}\\text{b}\\text{e}\\text{r} \\text{o}\\text{f} \\text{i}\\text{n}\\text{s}\\text{t}\\text{a}\\text{n}\\text{c}\\text{e}\\text{s} \\text{t}\\text{h}\\text{a}\\text{t} \\text{a}\\text{r}\\text{e} \\text{t}\\text{r}\\text{u}\\text{l}\\text{y} \\text{p}\\text{o}\\text{s}\\text{i}\\text{t}\\text{i}\\text{v}\\text{e}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e19\u003c/div\u003e\u003c/div\u003e\u003cp\u003eThe harmonic mean F1 is calculated as shown in Eq.\u0026nbsp;(\u003cspan refid=\"Equ20\" class=\"InternalRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e\u003cdiv id=\"Equ20\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ20\" name=\"EquationSource\"\u003e\n$$F1=\\frac{2\\text{*}P\\text{*}R}{P+R}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e20\u003c/div\u003e\u003c/div\u003e\u003ch2\u003eComparison model\u003c/h2\u003e\u003cp\u003eExperiments are conducted on the constructed dataset to compare several typical neural network models with the SMAAEA network model proposed in this paper, and the models are described as follows.\u003c/p\u003e\u003cp\u003ea. SVM: The support vector machine based model mentioned in reference 26.\u003c/p\u003e\u003cp\u003eb. CNN: The initialized feature vector representation of the character and word mixed features in the civil aviation text was used as the input, the number of convolution kernels was set to 3, 4, and 5, respectively, and the textual vector features were extracted to obtain the high-dimensional feature representation vector and input to the fully connected layer.\u003c/p\u003e \u003cp\u003ec. RNN: The input part is the same as the CNN network model, the initialized feature representation vector is obtained and input to the hidden layer to extract the feature information of the vector, and the output part is the same as the CNN model.\u003c/p\u003e\u003cp\u003ed. LSTM: The input and output parts are the same as those of the RNN network model, but different from the RNN network when extracting information, LSTM can extract information in the front and back directions.\u003c/p\u003e\u003cp\u003ee. NSMAAEA: The network model proposed in this paper is not used for training.\u003c/p\u003e \u003cp\u003e(\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e)Ablation experiment\u003c/p\u003e\u003cp\u003eIn order to verify the enhancement effect of the character feature word pre-training module, the Attention module and the Bi-LSTM module in the SMAAEA model on the whole model, ablation experiments were conducted, and the comparative models used in the experiments are explained as follows.\u003c/p\u003e\u003cp\u003ea. SMAAEA-char embedding: no pre-training containing character feature word vectors;\u003c/p\u003e\u003cp\u003eb. SMAAEA-attention: without Bi-GRU and Attention-based entity coding methods;\u003c/p\u003e\u003cp\u003ec. SMAAEA-bilstm: no Bi-LSTM-based entity coding method is used.\u003c/p\u003e\u003cp\u003eIn order to verify the enhancement effects of character feature word pre-training module, Attention module and Bi-LSTM module in the SMAAEA model on the entire model, ablation experiments were conducted. The comparison model used in the experiment was explained as follows.\u003c/p\u003e\u003cp\u003ea. SMAAEA-char embedding: Do not carry out the pre-training of the character feature word vector;\u003c/p\u003e\u003cp\u003eb. SMAAEA-attention: The entity coding method based on Bi-GRU and Attention is not adopted.\u003c/p\u003e\u003cp\u003ec. SMAAEA-bilstm: The entity coding method based on Bi-LSTM is not adopted.\u003c/p\u003e\u003ch2\u003eSimulation experiment and performance analysis\u003c/h2\u003e\u003cp\u003eFirst, the SMAAEA model was trained on the training set, and then 2000 text data were selected as the test set to test the training results, and the P, R, and F1 values of the entity types were counted, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eResults of the civil aviation emergency entity alignment experiment\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEntity types\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision/%\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall/%\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1/%\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTime of incident\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e97.46\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e99.98\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e98.67\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAirlines\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e97.76\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e86.98\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e92.28\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eType of aircraft\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e91.91\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e96.36\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e94.11\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFlight number\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e87.63\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e91.58\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e88.48\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRegistration number\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e91.79\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e94.43\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e93.25\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeparture airport\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.14\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e88.46\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e91.67\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDestination\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e84.78\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e79.56\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e82.23\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAircrew\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e93.17\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.58\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e94.47\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of passengers\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.69\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.69\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e95.69\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCasualties\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e84.42\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e84.42\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e84.42\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCause of accident\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e88.58\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e85.62\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e87.03\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccident result\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.57\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e93.18\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e92.89\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOverall indicator\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.12\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e97.89\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e96.97\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e\u003cp\u003eFrom the results, it can be seen that the reconciled mean F1 for the time of the incident is the highest, and the reconciled mean F1 values for the number of crew and passengers are 95.69% and 94.47%, respectively, which are only 2.5% and 2.23% lower than the overall index. The main reason is that the entity types of the number of passengers and crew are specific values and the entity types are relatively simple. The entity type of incident time is date type is also specific number, and the expression form of these entity types is fixed and simpler compared with text-based data, which is not easy to generate ambiguity. The R-value of destination entity alignment is relatively low, with a difference of 18.33% compared to the overall index, which is a large difference. The main reason is that the entity names are relatively long, there is similarity in the glyph names, and the semantics are more complex and prone to ambiguity. Since the number of casualties on the civil aviation event dataset is prone to ambiguity, its F1 value is also relatively smaller.\u003c/p\u003e\u003cp\u003eIn order to verify the feasibility of SMAAEA network model, this paper will conduct a comprehensive analysis of SMAAEA network model from three aspects: reference experiments, comparison experiments and stability experiments. The first group is the reference experiment, in which the SMAAEA model is compared with machine learning models and other classical neural network models to prove the effectiveness of SMAAEA model in improving recall and accuracy; the second group is the ablation experiment, in which SMAAEA model is ablated with different parts to analyze the specific roles of different modules and prove the effectiveness of each part of the module; the third is the stability experiment. Five different sets of training data are selected respectively, and the SMAAEA model is compared with its simplified model to verify the reliability and stability of the model.\u003c/p\u003e\u003cp\u003e \u003cem\u003eReference experiment\u003c/em\u003e \u003c/p\u003e\u003cp\u003eTraining different comparison model networks and recording the F1 values after every 5 rounds of iterations of the dataset, training found that the training effect grows faster in rounds 0–15, and the growth of the model training effect tends to be stable after the 20th round, training found that the F1 values are optimal after 25 rounds, and the effect of the 30th round is not significant with the 25th round. Comparing the F1 value performance of SMAAEA model with other models under different training rounds, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig14\" class=\"InternalRef\"\u003e14\u003c/span\u003e.\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSMAAEA and comparative model experimental results\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision/%\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall/%\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1/%\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e94.56\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e82.34\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e87.08\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNN\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.79\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e93.20\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e95.10\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRNN\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.38\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e93.20\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e92.91\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLSTM\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.99\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e89.90\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e93.42\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNSMAAEA\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.59\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e93.38\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e94.83\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMAAEA\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.12\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e97.89\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e96.97\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e\u003cp\u003ea. The experimental results in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e show that the SMAAEA model achieves an accuracy rate of 96.12% and a recall rate of 97.89% compared with machine learning models and other neural network models on the civil aviation emergencies dataset, both of which are higher than other comparison models. The SVM-based entity alignment method has the lowest F1 value compared with the F1 values of other neural network models, mainly because only the same attributes can exchange information, and there are restrictions on the exchange, and information cannot be exchanged between different attributes, resulting in an F1 value of only 87.08%. the SMAAEA model and other comparison models can enable the exchange of information between various attributes of the same entity, thus improving the The SMAAEA model and other comparative models are able to exchange information among the attributes of the same entity, thus improving the recall rate and achieving better experimental results.\u003c/p\u003e\u003cp\u003eb. From the experimental results, it can be concluded that the F1 value of the proposed SMAAEA model is improved by 1.87%, 4.06%, and 3.55% compared with CNN, RNN, and LSTM, respectively. The effectiveness of the SMAAEA model proposed in this thesis can be proved.\u003c/p\u003e\u003cp\u003ec. The number of parameters of NSMAAEA model is 3424337, and the number of parameters of SMAAEA model is 1796374. The difference between the two is whether the Siamese network is used or not. This can prove that the SMAAEA network model is beneficial to achieve the entity alignment task and realize the information exchange and connection between entities.\u003c/p\u003e\u003cp\u003ed. From the comparison results of the whole experiment, the F1 value of SMAAEA, the model proposed in this paper, is higher compared with other models. The reason for this is mainly that the two sub-networks of the Siamese network distinguish the importance of the same entity according to the integrity of encoding and different attributes, respectively, and combine the high-dimensional feature vector representations obtained from the two sub-networks to learn entity encoding from multiple perspectives, thus improving the recall rate of the model.\u003c/p\u003e\u003ch2\u003eAblation experiment\u003c/h2\u003e\u003cp\u003eTraining different simplified model networks and recording the F1 values after every 5 iterations of the dataset, training found that the training effect grows faster in rounds 0–10, and the growth of the model training effect tends to be stable after round 15. The F1 values were found to be optimal after 25 rounds, and the effect of the 30th round was not significant with the 25th round. Comparing the F1 value performance of SMAAEA model with its simplified model under different training rounds, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig15\" class=\"InternalRef\"\u003e15\u003c/span\u003e.\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSMAAEA and simplified model experimental results\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMAAEA-char embedding\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e95.89\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e94.90\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e95.97\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMAAEA-attention\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.82\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.40\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e94.01\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMAAEA-bilstm\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e97.34\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e88.30\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e92.59\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMAAEA\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e96.12\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e97.89\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e96.97\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e\u003cp\u003ea. In Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e,By comparing and analyzing the SMAAEA-embedding model with the SMAAEA model in this paper, we find that the F1 value of model SMAAEA is improved by 0.75% compared with model SAAEA-embedding, which proves that adding the Emedding layer to pre-train the word feature representation vector and character feature vector improves the experimental results very effective.\u003c/p\u003e\u003cp\u003eb. The SMAAEA-attention model is compared with the SMAAEA model, i.e., the attention mechanism is removed from the SMAAEA model and then compared with the SMAAEA model. the F1 value of the SMAAEA model is improved by 2.71%. It can be seen that the introduction of the attention mechanism can make the model pay attention to relatively more important attributes in the learning process, so that the important attributes can play a greater role in the classification process, thus improving the effectiveness of the model.\u003c/p\u003e\u003cp\u003ec. The SMAAEA-bilstm model and SMAAEA model are compared and analyzed, i.e., the Bi-LSTM network is removed from the SMAAEA model, and then it is compared with the SMAAEA model. The F1 value of model SMAAEA increased by 4.13% compared with model SMAAEA-bilstm. This is mainly due to the fact that Bi-LSTM can learn the information in both directions before and after the feature representation vector, which ensures that the integrity of the encoding can be learned, thus improving the learning quality of the model.\u003c/p\u003e\u003cp\u003ed. By comparing and analyzing the above simplified model with the SMAAEA model, the introduction of the Bi-LSTM modular model has the best improvement of 4.13%, which proves the important role of the Bi-LSTM network model for coding integrity learning.\u003c/p\u003e\u003cp\u003eStability experiment\u003c/p\u003e\u003cp\u003eIn order to verify the stability of SMAAEA model, five groups of text data with different numbers of entries are selected for the experiment, and the comparison models are the three simplified models mentioned in the ablation experiment, and the number of text data set for training is 9000, 9500, 10000, 10500 and 11000, and the number of text data used for testing is 1000. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig16\" class=\"InternalRef\"\u003e16\u003c/span\u003e, the horizontal coordinates in the figure indicate the number of text data trained and the vertical coordinates indicate the F1 values corresponding to the different models.\u003c/p\u003e\u003cp\u003eAs can be seen from Fig.\u0026nbsp;\u003cspan refid=\"Fig16\" class=\"InternalRef\"\u003e16\u003c/span\u003e, when the number of text data is 10000, the F1 values of SMAAEA-embedding model and SMAAEA model are close to each other, and SMAAEA-embedding is slightly higher than SMAAEA model, but the difference between their F1 values is only about 0.04%, and in other cases with different number of text data, the F1 values of SMAAEA model are all different. The F1 values of SMAAEA model are higher than the other models, which proves that SMAAEA model shows better stability and reliability for different training items of text information.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eKnowledge fusion is an important foundation for the construction of the knowledge graph of civil aviation emergencies. In this paper, we propose an entity alignment method based on knowledge extraction to address the special characteristics of civil aviation domain and the problems of long and complex entities in civil aviation domain, the difficulty of fast entity extraction and the low extraction accuracy in the current knowledge fusion process. It is verified and analyzed by reference experiments, comparison experiments and stability experiments.\u003c/p\u003e \u003cp\u003eThis paper firstly analyzes the entity alignment task based on civil aviation emergencies and carries out the general design. Secondly, in constructing the dataset, the rule-based entity alignment method is combined with the annotation of entities through lexicon matching to improve the accuracy and recall rate of entity alignment. Then SMAAEA network model is proposed to investigate the entity alignment task. Due to the special characteristics of civil aviation domain, entities in the domain may have similar glyph structures or the same radicals, and entity alignment is relatively difficult. In the process of entity coding and improving the pre-training model, a char feature layer is introduced to obtain a word mixture feature vector containing character features. And the obtained character-word mixed feature vector representation is input to the Bi-LSTM layer and the shorter attribute sentences are input to Bi-GRU\u0026thinsp;+\u0026thinsp;Attention respectively, and finally the high-dimensional vector representation is obtained. After entity alignment, knowledge fusion is performed to fuse the new knowledge with the knowledge in the target civil aviation knowledge base to further improve the quality of the knowledge spectrum graph. Finally, the feasibility and reliability of the SMAAEA model proposed in this thesis are verified by designing experiments and analyzing the results. In the future work, we consider whether feature matching and partition indexing of entities can be performed in the process of entity alignment to make entity alignment more efficient and accurate, but this method requires a large amount of computational resources and is usually applied to the construction of general domain knowledge graphs, and the method requires high data quality and needs to be calculated by various types of similarity functions, and the next step can be to apply the The next step is to apply the structural relationship between entities in this method to the civil aviation knowledge base.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eAuthor contributions statement\u003c/h2\u003e \u003cp\u003eJ.W. designed the study and writed the main manuscript text. J.Q. performed the experiment and analyzed the results, Z.Z collected the data and analyzed the results.X.D. prepared fgures and tables, and all authors reviewed the manuscript.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eCompeting interests\u003c/strong\u003e \u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e \u003c/p\u003e\n \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003eData availibility\u003c/h2\u003e \u003cp\u003eThe datasets used and analysed during the current study will be available from the corresponding author on reasonable request\u003c/p\u003e \u003c/div\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eWorld Accident Investigation and Tracking[EB/OL]Aviation Safety Information System of CAAC.[2018-09]. http://safety.caac.gov.cn /index /initpage.act.\u003c/li\u003e\n\u003cli\u003eSowa J F. Principles of semantic networks: Explorations in the representation of knowledge[M]. Morgan Kaufmann, 2014.\u003c/li\u003e\n\u003cli\u003eGruber T R. Toward principles for the design of ontologies used for knowledge sharing[J]. International journal of human-computer studies, 1995, 43(5-6): 907-928.\u003c/li\u003e\n\u003cli\u003eSuchanek F M, Kasneci G, Weikum G. Yago: a core of semantic knowledge[C]//Proceedings of the 16th international conference on World Wide Web. 2007: 697-706.\u003c/li\u003e\n\u003cli\u003eBizer C, Lehmann J, Kobilarov G, et al. Dbpedia-a crystallization point for the web of data[J]. Journal of web semantics, 2009, 7(3): 154-165.\u003c/li\u003e\n\u003cli\u003eBollacker K, Evans C, Paritosh P, et al. Freebase: a collaboratively created graph database for structuring human knowledge[C]//Proceedings of the 2008 ACM SIGMOD international conference on Management of data. 2008: 1247-1250.\u003c/li\u003e\n\u003cli\u003eWang Z, Li J, Wang Z, et al. XLore: A Large-scale English-Chinese Bilingual Knowledge Graph[C]//International semantic web conference (Posters \u0026amp; Demos). 2013, 1035: 121-124.\u003c/li\u003e\n\u003cli\u003eNiu X, Sun X, Wang H, et al. Zhishi. me-weaving chinese linking open data[C]//International Semantic Web Conference. Springer, Berlin, Heidelberg, 2021: 205-220.\u003c/li\u003e\n\u003cli\u003eXu B, Xu Y, Liang J, et al. CN-DBpedia: A never-ending Chinese knowledge extraction system[C]//International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems. Springer, Cham, 2017: 428-438.\u003c/li\u003e\n\u003cli\u003eLiDing. Schema[EB/OL][2019,3,25].https//github.com/cnschema/cnschema/wiki/Schema.\u003c/li\u003e\n\u003cli\u003eKejriwal M. Domain-specific knowledge graph construction[M]. Cham: Springer International Publishing, 2019.\u003c/li\u003e\n\u003cli\u003ePlank B, Goldberg Y. Multilingual part-of-speech tagging with bidirectional long short-term memory models and auxiliary loss[J]. arXiv preprint arXiv:1604.05529, 2016.\u003c/li\u003e\n\u003cli\u003eBengio Y, Sen\u0026eacute;cal J S. Adaptive importance sampling to accelerate training of a neural probabilistic language model[J]. IEEE Transactions on Neural Networks, 2008, 19(4): 713-722.\u003c/li\u003e\n\u003cli\u003eCastro M J, Prat F. New directions in connectionist language modeling[C]//International Work-Conference on Artificial Neural Networks. Springer, Berlin, Heidelberg, 2021: 598-605.\u003c/li\u003e\n\u003cli\u003eMikolov T, Karafi\u0026aacute;t M, Burget L, et al. Recurrent neural network based language model[C]//Interspeech. 2010, 2(3): 1045-1048.\u003c/li\u003e\n\u003cli\u003eMikolov T, Kombrink S, Burget L, et al. Extensions of recurrent neural network language model[C]//2011 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2011: 5528-5531.\u003c/li\u003e\n\u003cli\u003eCho K, Van Merri\u0026euml;nboer B, Gulcehre C, et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation[J]. arXiv preprint arXiv:1406.1078, 2014.\u003c/li\u003e\n\u003cli\u003eChopra S, Hadsell R, Le Cun Y. Learning a similarity metric discriminatively, with application to face verification[C]//2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u0026apos;05). IEEE, 2005, 1: 539-546.\u003c/li\u003e\n\u003cli\u003eNiu X, Sun X, Wang H, et al. me-weaving chinese linking open data[C]//International Semantic Web Conference. Springer, Berlin, Heidelberg, 2021: 205-220. \u003c/li\u003e\n\u003cli\u003eLin H L, Wang Y Z, Jia Y T, et al. Network big data oriented knowledge fusion methods: A survey[J]. Chinese Journal of Computers, 2017, 40(1): 1-27.\u003c/li\u003e\n\u003cli\u003eNgomo A C N, Auer S. LIMES\u0026mdash;a time-efficient approach for large-scale link discovery on the web of data[C]//Twenty-Second International Joint Conference on Artificial Intelligence. 2011: 905-912.\u003c/li\u003e\n\u003cli\u003eNiu X, Rong S, Wang H, et al. An effective rule miner for instance matching in a web of data[C]//Proceedings of the 21st ACM international conference on Information and knowledge management. 2022: 1085-1094.\u003c/li\u003e\n\u003cli\u003eWang X, Liu K, He S, et al. Multi-source knowledge bases entity alignment by leveraging semantic tags[J]. Chinese Journal of Computers, 2017, 40(3): 701-711.\u003c/li\u003e\n\u003cli\u003eZhang Weili, Huang Tinglei, Liang Xiao Entity alignment of encyclopedia knowledge base based on semi supervised collaborative training [J] Computer and Modernization, 2017 (12): 88-93.\u003c/li\u003e\n\u003cli\u003e[Yuan Man, Cao Yang, Chen Ping. Research on the Standard Vocabulary Reference Model in Construction of Educational Knowledge Graph [J]. E education Research , 2022, 41 (03): 76-84.\u003c/li\u003e\n\u003cli\u003eVapnik V. The nature of statistical learning theory[M]. Springer science \u0026amp; business media, 1999.\u003c/li\u003e\n\u003cli\u003eKantardzic M. Data mining: concepts, models, methods, and algorithms[M]. John Wiley \u0026amp; Sons, 2021.\u003c/li\u003e\n\u003cli\u003eHan J, Pei J, Kamber M. Data mining: concepts and techniques[M]. Elsevier, 2021.\u003c/li\u003e\n\u003cli\u003eHu Fanghuai Research on the Construction Method of Chinese Knowledge Graph Based on Multiple Data Sources [D] Doctoral Dissertation Shanghai: Department of Computer Application Technology, East China University of Science and Technology, 2015.\u003c/li\u003e\n\u003cli\u003eSarawagi S, Bhamidipaty A. Interactive deduplication using active learning[C]//Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining. 2002: 269-278.\u003c/li\u003e\n\u003cli\u003eKim Y. Convolutional Neural Networks for Sentence Classification[C]//empirical methods in natural language processing. 2014: 1746-1751. \u003c/li\u003e\n\u003cli\u003eLiu P, Qiu X, Huang X, et al. Recurrent neural network for text classification with multi-task learning[C]//international joint conference on artificial intelligence. 2016: 2873-2879.\u003c/li\u003e\n\u003cli\u003eJoulin A, Grave E, Bojanowski P, et al. Bag of Tricks for Efficient Text Classification[C]//conference of the european chapter of the association for computational linguistics. 2017: 427-431. \u003c/li\u003e\n\u003cli\u003eMikolov T, Sutskever I, Chen K, et al. Distributed representations of words and phrases and their compositionality[C]//Advances in neural information processing systems. 2023: 3111-3119.\u003c/li\u003e\n\u003cli\u003eSong, Yu, Gui, Mingyu, et al. BG-INT: An Entity Alignment Interaction Model Based on BERT and GCN[C]//Communications in Computer and Information Science, Volume 1772, 2023:67-81.\u003c/li\u003e\n\u003cli\u003eMunne, Rumana Ferdous,Ichise, Ryutaro. Entity alignment via summary and attribute embeddings[C]Logic Journal of the IGPL, Volume 31, Issue 2, 2023:314-324.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3066877/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3066877/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eWith the growth of demand for air-space integrated transportation industry, the civil aviation transportation industry plays an increasingly important role in the national economy and transportation. The diversity and diversified development of civil aviation emergencies have greatly affected the decision-making efficiency for emergencies in civil aviation emergency management system. Constructing a knowledge graph of civil aviation and improving the reliability and richness of the knowledge graph have become an urgent problem to be solved. There are some problems during the construction of civil aviation domain knowledge graph such as long entity length, existing hybrid and composite entities, similarity of the font features of domain entity names, difference of information between entities, separated coding between entities, and error prone in the transmission process of coding. So we research the construction method of knowledge graph for civil aviation emergencies to resolve these problems. (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) We propose a font pre-training model and incorporate a character feature layer into the embedding layer. (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) We propose a multi-attribute attention alignment method based on Siamese network; entities with similar structure in civil aviation knowledge base are input into the font layer for pre-training; the complete semantic information of entities and the importance of different attributes to entities are learned through two sub-networks respectively to improve the integrity of civil aviation knowledge graph. The experimental results show that, compared with other models, the F1 value of the proposed model can reach 96.97%, which verifies the feasibility of our model.\u003c/p\u003e","manuscriptTitle":"Research on multi-attribute attentional civil aviation entity alignment based on Siamese network","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-06-27 20:50:07","doi":"10.21203/rs.3.rs-3066877/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-09-08T08:24:04+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-09-01T18:03:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"8d9fff10-7363-4eec-9667-a73196d9961d","date":"2023-08-27T14:16:26+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"33b1ac54-5837-4d98-87c4-3d0f3add280b","date":"2023-08-27T08:43:21+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-08-24T21:26:14+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-08-19T21:03:39+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2023-06-22T06:36:24+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-06-22T06:32:48+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2023-06-15T08:59:53+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ae5cd1d3-8145-427e-b7db-28c79e9508d7","owner":[],"postedDate":"June 27th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":22621625,"name":"Physical sciences/Mathematics and computing/Computer science"},{"id":22621626,"name":"Physical sciences/Mathematics and computing/Information technology"},{"id":22621627,"name":"Physical sciences/Engineering"}],"tags":[],"updatedAt":"2023-10-05T05:14:15+00:00","versionOfRecord":[],"versionCreatedAt":"2023-06-27 20:50:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3066877","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3066877","identity":"rs-3066877","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00