Identification of Strong Motion Record Baseline Drift Based on Bayesian Optimized Transformer Network

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Research in earthquake engineering heavily relies on strong motion observation. The quality of strong motion records directly affects the reliability of earthquake disaster prevention, rapid reporting of seismic magnitude, earthquake early warning, and other areas. Currently, the quality assurance of strong motion records often relies on basic methods such as zero-line adjustment and high-pass filtering. However, these methods often fail to satisfactorily identify and handle abnormal waveforms present in strong motion records, and their efficiency is relatively low. In this paper, a Bayesian-optimized Transformer-based approach is proposed to improve the identification of baseline drift anomalies in strong motion records. By partitioning the strong motion record data from the 1999 Chi-Chi earthquake in Taiwan, China, into two categories: high-quality records (with minimal baseline drift) and low-quality records (with significant baseline drift), we extracted data with distinct features and inputted them into the proposed model for training. Finally, the model was used to predict whether strong motion records exhibited baseline drift abnormalities. The experimental results show that after performing Bayesian optimization on the parameters of the Transformer model, this method achieves an accuracy of over 85% in baseline drift identification. It is capable of efficiently identifying a large volume of strong motion records with baseline drift within a short period of time. The model performs well in the task of classifying baseline drift in strong motion records and can be used for subsequent identification of abnormalities after baseline drift correction, enabling automation in handling baseline drift abnormal data.
Full text 101,684 characters · extracted from preprint-html · click to expand
Identification of Strong Motion Record Baseline Drift Based on Bayesian Optimized Transformer Network | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Identification of Strong Motion Record Baseline Drift Based on Bayesian Optimized Transformer Network Baofeng Zhou, Yue Yin, Maofa Wang, Runjie Zhang, Yue Zhang, Wenheng Guo This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3237271/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 09 Oct, 2024 Read the published version in Acta Geophysica → Version 1 posted 6 You are reading this latest preprint version Abstract Research in earthquake engineering heavily relies on strong motion observation. The quality of strong motion records directly affects the reliability of earthquake disaster prevention, rapid reporting of seismic magnitude, earthquake early warning, and other areas. Currently, the quality assurance of strong motion records often relies on basic methods such as zero-line adjustment and high-pass filtering. However, these methods often fail to satisfactorily identify and handle abnormal waveforms present in strong motion records, and their efficiency is relatively low. In this paper, a Bayesian-optimized Transformer-based approach is proposed to improve the identification of baseline drift anomalies in strong motion records. By partitioning the strong motion record data from the 1999 Chi-Chi earthquake in Taiwan, China, into two categories: high-quality records (with minimal baseline drift) and low-quality records (with significant baseline drift), we extracted data with distinct features and inputted them into the proposed model for training. Finally, the model was used to predict whether strong motion records exhibited baseline drift abnormalities. The experimental results show that after performing Bayesian optimization on the parameters of the Transformer model, this method achieves an accuracy of over 85% in baseline drift identification. It is capable of efficiently identifying a large volume of strong motion records with baseline drift within a short period of time. The model performs well in the task of classifying baseline drift in strong motion records and can be used for subsequent identification of abnormalities after baseline drift correction, enabling automation in handling baseline drift abnormal data. Baseline Drift Transformer Network Strong Motion Record Bayesian Optimization Sequence Classification Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Key Points A neural network utilizing the Transformer model is proposed for the recognition of baseline drift in strong motion records. The Bayesian optimization method is employed to optimize the hyperparameters of the Transformer model, thereby enhancing its overall performance. The baseline drift recognition model holds great potential as it can be employed in tasks such as baseline drift correction and sequence classification, offering broad prospects for its application. 1. Introduction Strong motion observation is an important research area in the field of earthquake engineering. It has significant applications in studying source mechanisms (Zhao and Helmberger 1994 ; Kuang et al. 2021 ), ground motion attenuation relationships (Somerville et al. 1997 ; Youngs et al. 1997 ), as well as seismic design of structures (Katsanos et al. 2010 ) and seismic resilience assessment of cities (Shang et al. 2020 ). Additionally, strong motion observation networks can be used to establish earthquake early warning systems and rapid assessment systems for disaster situations. In the past few decades, strong motion observation networks have been extensively developed worldwide, leading to the acquisition of a large number of strong motion records and the establishment of multiple strong motion databases (Pacor et al. 2011 ; Ancheta et al. 2013 ; Akkar et al. 2014 ; Dawood et al. 2016 ; Luzi et al. 2016 ). However, due to environmental noise, instrument failures, and interference in the telecommunications transmission process, the quality of strong motion records may be poor (Boore and Bommer 2005 ), which is mainly manifested in the acceleration time history, such as spikes, baseline drifts, and asymmetry (Zhou et al. 2015 ). One major problem encountered by digital accelerometers is the distortion and offset of the reference baseline, resulting in unphysical velocity and displacement, known as baseline drift (Boore and Bommer 2005 ). Many scholars have conducted related research on the baseline drift phenomenon in strong motion records and proposed many correction methods (Iwan et al. 1985 ; Boore 1999 ; Zhou et al. 2014 ). For each individual strong motion record, we apply the method of integrating acceleration to obtain velocity and displacement. Then, based on the displacement time history, we can quickly determine whether there is a baseline drift phenomenon (Boore 1999 ). However, with the significant expansion of strong motion observation networks and the substantial increase in the number of strong motion records, it has become difficult for manual selection to efficiently and quickly identify records with baseline drift from thousands of strong motion records. There is an urgent need for automated processing of strong motion data, automated identification of abnormal waveform data, in order to further improve the quality of strong motion records. In recent years, there has been significant development in the application of machine learning and deep learning techniques in the field of earthquake engineering. For instance, Wang et al. ( 2019 ) conducted research on aftershock prediction using machine learning-based support vector regression algorithms. Subsequently, they also employed deep convolutional neural networks for earthquake prediction (Shan et al. 2022 ). Xu et al. ( 2022 ) performed dynamic time history analysis of structural seismic performance using long short-term memory networks. Bellagamba et al. ( 2019 ) utilized feedforward neural networks to classify the quality of strong motion records with low seismic intensity. Regarding the quality of strong motion records, researchers have proposed numerous methods for classification and processing using machine learning and deep learning approaches. For instance, Musson(1998) proposed a data quality assessment system for three types of quality problems: reliability of intensity assessment, positional certainty or uncertainty, and accuracy of raw data. Each indicator was quantified as a binary variable and given a final quality code from 0 (best) to 7 (worst). Chen et al. ( 2019 ) developed a new denoising framework based on unsupervised machine learning techniques using the concept of autoencoding to improve the signal-to-noise ratio of many earthquake datasets and adaptively learn seismic signals from noisy observations. Moreover, they also clustered the time samples into two groups, waveform points and non-waveform points, using unsupervised machine learning algorithms to help identify seismic waveforms in microseism or earthquake data (Chen 2020 ). Meier et al. ( 2019 ) showed how modern machine learning classifiers can significantly improve real-time signal/noise discrimination. Zhu and Beroza ( 2019 ) proposed a deep neural network-based method for earthquake arrival time picking called 'PhaseNet,' which simultaneously picks the arrival times of both P and S waves. Later, W. Zhu et al. ( 2019 ) developed a deep neural network-based denoising/decomposition method called "Deep Denoiser", which can decompose input data into the interested signal and noise (defined as any non-seismic signal). Yu et al. (2022) utilized Long Short-Term Memory (LSTM) models to assess the quality of records and perform zero-baseline correction on near-fault strong motion records, achieving automated data processing. These methods have significantly advanced the automation of processing and identification of strong motion records quality. However, more efficient and targeted methods are still needed for the recognition of peculiar waveforms such as baseline drift, spikes, and asymmetry caused by quality issues in strong motion records. Recently, Chat-GPT has been a hot topic, and research in natural language processing has shown the strong capabilities of Chat-GPT (OpenAI 2023 ). The deep learning model used by Chat-GPT is Transformer, which is a powerful and pioneering model in deep learning that can adapt to different use cases such as translation, language processing, classification, etc. In this paper, we utilize the superior performance of Transformer and establish a baseline drift identification model for strong motion records based on the encoding part of Transformer. Additionally, Bayesian optimization is employed to optimize the parameters of the model. The model fully utilizes the characteristics of sequence input and classification in Transformer, enabling it to effectively learn the temporal sequence features of strong motion records and make judgments about their data quality. In the experiments, by considering the temporal nature and visual characteristics of the strong motion records, significant features are extracted and imported into the baseline drift identification model, allowing for quick determination of whether baseline drift has occurred. The baseline drift identification neural network based on Transformer significantly enhances the efficiency of quality processing for baseline drift phenomena in strong motion records. The Bayesian optimized transformer-based neural network for baseline drift identification has greatly improved the efficiency of identifying records with baseline drift in strong motion records. Based on the findings of this study, it provides a more efficient and accurate method for handling the quality of strong motion records, and also offers a new approach and methodology for the application of deep learning techniques in the field of earthquakes. We believe that this Bayesian optimized transformer-based neural network for baseline drift identification will have a positive impact on research and applications in the field of seismology. 2. Research Content This section provides an overview of the research content, which mainly consists of three parts. The first part is a systematic process description that introduces the overall architecture and operational flow of the system. The second part includes the experimental dataset, including data sources, data formats, and data processing methods. The third part describes the model used in this experiment. 2.1 System Framework As shown in Fig. 1 , the system consists of two main aspects: constructing the dataset and training the model. The strong motion records undergo data preprocessing, followed by the generation of acceleration, velocity, and displacement plots. The record is classified based on the velocity or displacement image to determine if there is baseline drift, and the strong motion records are classified into two categories: those with baseline drift and those without baseline drift. The upper and lower images in the figure respectively show velocity images without and with baseline drift. Finally, the extracted displacement sequences corresponding to the two categories of images are converted into vector form and used to construct the training, validation, and test sets, which are input into the Transformer model. The feature information obtained after training the model is used to predict whether a new dataset contains baseline drift. 2.2 Dataset To apply the Transformer model to identify baseline drift waveforms in strong motion records, data need to be processed and divided to ensure that they can be effectively input to the model for learning and classification. The data processing process is as follows: (1) Data collection: Collecting strong motion records for the study. The data used in this research is from the Chi-Chi earthquake in Taiwan, China, comprising a total of 420 strong motion records. These records are divided into training, validation, and prediction sets. (2) Data preprocessing: Preprocess strong motion records by first adjusting the acceleration time history to the zero baseline, subtracting the average value of the noise record for the first 20 seconds(Xin-xin et al. 2022 ), and then performing a Butterworth second-order causal filtering with a bandpass range of 0.01-35Hz to remove noise from the acceleration records for better feature recognition. Finally, the data is standardized to adjust the data to the same scale for easier comparison classification and feature learning in the Transformer model. (3) Classification: Classify the strong motion records into baseline drift and non-baseline drift categories. First, integrate the processed acceleration records once and twice to obtain velocity and displacement time histories. Then, based on the velocity or displacement time histories, classify whether baseline drift exists. The basis for the classification is that slight baseline shifts in acceleration can lead to significant displacement in velocity and displacement time history (Boore 2001 ), with baseline drift being particularly evident in displacement time history. Figure 2 shows the acceleration, velocity, and displacement time histories without baseline drift. In the displacement time history, there is a nearly constant residual value in the last approximately 1/3 of the duration. The velocity time history exhibits oscillations around the zero baseline at the end of strong motion (Boore 2001 ). However, as shown in Fig. 3, when baseline drift occurs, different manifestations are observed. The slight variations in the baseline of the acceleration time history can lead to a certain slope in the observed velocity time history and a diverging displacement in the final part of the displacement time history, exhibiting a continuous upward trend instead of a steady state. Comparing the two graphs, it can be observed that the occurrence of baseline drift in the acceleration time history of strong motion records can be identified by examining the last one-third portion of the displacement time history, where more pronounced baseline drift characteristics are evident. Therefore, the displacement time history can be used to classify the strong motion records. (4) Sequence processing: Considering the input requirements of the Transformer model, this study further processes the preprocessed data by segmenting it with a certain time step. The last 4000 data points of the extracted displacement data are used for model training, divided into sequences of shape 10*400. These sequences are then input into the model in the form of a three-dimensional tensor with dimensions of n*10*400. A total of 72 training sets, 24 validation sets, and 24 test sets are created. (5) Sequence encoding: Encode the split sequences to transform them into tensor format that can be processed by the Transformer model. (6) Label encoding: Encode label information to a form that the Transformer model can be processed. 2.3 Model We have adopted a Bayesian-optimized Transformer model to establish a neural network for baseline drift recognition in strong motion records. The framework of the network is shown in Fig. 4 . The entire network mainly consists of the encoder layer of the Transformer model. The input displacement sequences, after being divided and position encoded, are then fed into the Transformer module to extract image features, progressively restore details, and finally obtain the ultimate prediction results through global average pooling and a fully connected network. The core of our neural network architecture is the Transformer, a deep learning model used for sequence-to-sequence learning. The Transformer model, introduced by Google in 2017 (Vaswani et al. 2017 ), is widely used in natural language processing and computer vision. At the heart of the Transformer is the self-attention mechanism, which learns the dependencies between elements in a sequence. By using self-attention, the Transformer can address issues such as gradient vanishing and exploding commonly found in traditional recurrent neural networks. Additionally, self-attention enables parallel computation, accelerating the training process of the model. In this paper, we calculate the similarity between each element of the input to obtain a weight vector, and then compute the representation vector for each element through weighted averaging. The self-attention mechanism is implemented using a set of query (Q), key (K), and value (V) matrices. The computation is as follows: \(Attention\left(Q,K,V\right)=softmax\left(\frac{Q{K}^{T}}{\left(\sqrt{{d}_{k}}\right)}\right)V\) (1) where \(Q,K, and V\) are matrices representing queries, keys, and values, respectively, and \({d}_{k}\) is the dimension of the key vectors. \(Attention\left(Q,K,V\right)\) represents the output vector of the self-attention mechanism. The Transformer model adds a multi-head attention mechanism based on self-attention. The multi-head attention mechanism is an important component of the Transformer model that is used to learn multiple different dependencies. It applies multiple parallel self-attention "heads" to learn the projected Q, K, and V matrices. Compared to a single attention head, this improves performance. Specifically, it linearly transforms the input sequence separately, and then obtains vectors representing the dependencies between elements in the sequence by computing attention. Finally, these vectors are concatenated and linearly transformed to obtain the final representation. The representation is shown below: \(MultiHead\left(Q,K,V\right)=Concat\left(hea{d}_{1},hea{d}_{2},\dots ,hea{d}_{h}\right){W}^{O}\) (2) where, \(hea{d}_{i}=Attention\left(Q{W}_{i}^{Q},K{W}_{i}^{K},V{W}_{i}^{V}\right)\) represents the vector output by the i-th attention head. In this paper, we use two attention heads. \(Concat\) combines the outputs of multiple heads. \({W}_{i}^{Q},{W}_{i}^{K}, and {W}_{i}^{V}\) represent the linear transformation matrices of the query, key, and value for the i-th head, respectively, and \({W}^{O}\) represents the output matrix of the multi-head attention. Before inputting the sequence into the Transformer, we need to apply positional encoding to the input sequence. Strong motion records contain temporal features and exhibit a clear time order. The temporal ordering information of the input sequence is crucial. The purpose of positional encoding is to add the time position information of the input sequence to the model, so that the model can better capture the order information in the sequence. Specifically, positional encoding can be calculated using the following formula: \(P{E}_{\left(pos,2i\right)}=sin\left(pos/{10000}^{2i/{d}_{model}}\right)\) (3) \(P{E}_{\left(pos,2i+1\right)}=cos\left(pos/{10000}^{2i/{d}_{model}}\right)\) (4) Where \(P{E}_{\left(pos,2i\right)}\) represents the positional encoding value at position \(pos\) and dimension \(2i\) , and \(P{E}_{\left(pos,2i+1\right)}\) represents the positional encoding value at position \(pos\) and dimension \(2i+1\) . \({d}_{model}\) represents the dimension of the input sequence. In this paper, the input vector is 10*400, so the sequence dimension \({d}_{model}\) is 400, \(pos\) represents the position in the input sequence, and \(i\) represents the dimension index in the positional encoding matrix. After positional encoding, the sequences of strong motion records are passed through two Transformer encoder layers, each with two attention heads. The outputs are then normalized and connected with residual connections before being fed into a feed-forward neural network to generate the final output. In this paper, a baseline drift recognition model based on Keras is constructed and optimized using the RMSprop optimizer. Binary cross-entropy is employed as the loss function, and accuracy is used as the evaluation metric. The key parameters of the model are shown in Table 1 . Table 1 Key parameters of the Transformer model Model \(Block\) \(parameter\) Transformer \(MultiHeadAttention\) num_heads = 2; key_dim = 32 TimeDistributed ff_dim = 32; activation="relu" Dense activation="sigmoid" We used Bayesian optimization to optimize the hyperparameters of the model. Bayesian optimization (B. Shahriari et al. 2016 ) is a black-box optimization method, where the black-box function refers to a function without an explicit formula or expression. This algorithm builds a surrogate model, typically a Gaussian process, to estimate the objective function. It then employs Bayesian inference to update the posterior distribution of the surrogate model, effectively exploring and exploiting the function space. By optimizing the posterior distribution of the surrogate model, the Bayesian optimization algorithm can find the global optimum with relatively few iterations. In our case, we apply Bayesian optimization to optimize the vector dimensions (key_dim) of the Q, K, and V matrices in the multi-head attention layer, as well as the dimension (ff_dim) of the feed-forward neural network layer. The parameter optimization ranges are set as shown in Table 2 . The model uses the gp_minimize optimizer to perform Bayesian optimization on the hyperparameters. Through Bayesian optimization, the most suitable vector dimensions and network layers for the model are selected. Table 2 Bayesian optimization of Transformer model parameters. Model \(Parameter\) Range Transformer \(\text{k}\text{e}\text{y}\_\text{d}\text{i}\text{m}\) (16, 64) \(ff\_dim\) (16, 64) 3. Result 3.1 Evaluation Metrics To assess the performance of the Transformer model, three commonly used evaluation metrics, namely accuracy, recall, and F1 score, are employed in this study. Figures 5 and 6 show the confusion matrix, which provides a visual representation to evaluate the performance of the classification algorithm. 3.1.1 Precision Precision refers to the proportion of sequences classified as baseline drift among the total number of strong motion sequences. It is an important metric for evaluating the performance of the classifier. The formula for calculating precision is as follows: \(Precision= \frac{TP+TN}{TP+FP+TN+FN}\) (5) Where: TP represents the number of strong motion displacement sequences that have baseline drift and are classified as having baseline drift. FP represents the number of strong motion displacement sequences that do not have baseline drift but are classified as having baseline drift. TN represents the number of strong motion displacement sequences that do not have baseline drift and are classified as not having baseline drift. FN represents the number of strong motion displacement sequences that have baseline drift but are classified as not having baseline drift. 3.1.2 Recall Recall, also known as sensitivity or true positive rate, refers to the proportion of correctly classified baseline drift samples among the total number of baseline drift samples. It is calculated using the following formula: \(Recall= \frac{TP}{TP+FN}\) (6) 3.1.3 F1 score F1 is a metric that combines Precision and Recall of a model, and is calculated as follows. \(F1=\frac{2\times Precision\times Recall}{Precision+Recall}\) (7) The precision, recall rate, and F1 score of the model are calculated through the confusion matrix shown in Fig. 5 . As shown in the figure, "normal" represents normal waveforms and "abnormal" represents baseline drift waveforms. It can be seen that TP = 67; TN = 92; FP = 8; FN = 33. The precision is calculated as Precision = (67 + 92) / 200 = 0.795; Recall = 67 / (67 + 33) = 0.67; F1 = 2 * 0.795 * 0.67 / (0.795 + 0.67) = 0.727. It can be seen that the accuracy calculated by the model without Bayesian optimization is generally lower. The confusion matrix after adding Bayesian optimization is shown in Fig. 6 . the precision, Recall, and F1 score are calculated using the above formula with TP = 78, TN = 95, FP = 5, and FN = 22. The results are as follows: Precision = (78 + 95)/200 = 0.865; Recall = 78/(78 + 22) = 0.78; F1 = (2×0.865×0.78)/(0.865 + 0.78) = 0.820. The precision has significantly improved. The precision, recall, and F1 score of the model before and after Bayesian optimization are shown in Table 3 . Table 3 The precision, recall, and F1 score of the model. Model \(Precision\) \(Recall\) \(F1 score\) Transformer 0.795 0.67 0.727 Transformer by Bayes 0.865 0.78 0.820 According to the data analysis in Table 3 , the Transformer model exhibited good performance in identifying and classifying baseline drift waveforms, Particularly, after applying Bayesian optimization, the model's precision and F1 score improved to 86.5% and 82.0% respectively. This indicates that Bayesian optimization can significantly enhance the model's performance, leading to better solutions for real-world problems. Furthermore, from the table, it can be observed that the model's recall is 67%, slightly lower than the precision. This indicates that the model may have some misclassifications in identifying drift waveforms, specifically classifying normal waveforms as drift waveforms. However, after applying Bayesian optimization, the recall of the model has also improved to some extent. In summary, the data analysis in the table above shows that the Transformer model performs well in the identification and classification of baseline drift waveforms, but there is still room for further optimization. By continuously optimizing the feature extraction and classification strategy of the model, the performance of the model can be further improved to better solve practical problems. 4. Conclusion Based on the characteristics of baseline drift in strong motion records, this paper proposes a deep learning neural network model, the transformer, for identifying baseline drift in strong motion records from an image recognition perspective. By utilizing the temporal features of the time series in strong motion records, the transformer model effectively extracts features related to baseline drift and achieves high recognition efficiency even with small sample sizes. Based on the proposed method and experimental results, the following conclusions can be drawn: (1) Strong motion data, as temporal data, possess distinctive features that vary with time. These features can be effectively extracted by the transformer model, which demonstrates high accuracy in baseline drift identification. The accuracy of the model on the test set exceeds 85%, indicating its effectiveness in recognition. (2) The performance of the model is further improved by Bayesian optimization of the transformer model. This model can be embedded in the calibration of baseline drift in strong motion records to determine whether the corrected acceleration images meet the requirements. Furthermore, the development of this model holds great promise. The transformer model shows good prospects for the classification of sequences with temporal features. However, due to the limited number of samples in this study, there is still room for further optimization. By continuously improving the feature extraction and classification strategies of the model, its performance can be further enhanced to better address real-world problems. Declarations Acknowledgment This research was funded by the Scientific Research Fund of Institute of Engineering Mechanics, China Earthquake Administration (Nos.2022C05), and the Natural Science Foundation of Heilongjiang Province (LH2022E120) and the National Key Research and Development Program of China (2018YFE0109800). This support is greatly appreciated. This work was also supported by Guangxi Project of Technology Base and Special Talent (AD20325004), the National Natural Science Foundation of China (42164002). Competing interests The authors declare that there are no competing interests. References Akkar S, Sandıkkaya MA, Şenyurt M, et al (2014) Reference database for seismic ground-motion in Europe (RESORCE). Bulletin of Earthquake Engineering 12:311–339. https://doi.org/10.1007/s10518-013-9506-8 Ancheta T, Darragh R, Stewart J, et al (2013) PEER NGA-West2 Database, PEER Report 2013/03, pacific earthquake engineering research center. University of California, Berkeley B. Shahriari, K. Swersky, Z. Wang, et al (2016) Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE 104:148–175. https://doi.org/10.1109/JPROC.2015.2494218 Bellagamba X, Lee R, Bradley BA (2019) A Neural Network for Automated Quality Screening of Ground Motion Records from Small Magnitude Earthquakes. Earthquake Spectra 35:1637–1661. https://doi.org/10.1193/122118EQS292M Boore DM (1999) Effect of baseline corrections on response spectra for two recordings of the 1999 Chi-Chi, Taiwan, earthquake Boore DM (2001) Effect of Baseline Corrections on Displacements and Response Spectra for Several Recordings of the 1999 Chi-Chi, Taiwan, Earthquake. Bulletin of the Seismological Society of America 91:1199–1211. https://doi.org/10.1785/0120000703 Boore DM, Bommer JJ (2005) Processing of strong-motion accelerograms: needs, options and consequences. Soil Dynamics and Earthquake Engineering 25:93–115. https://doi.org/10.1016/j.soildyn.2004.10.007 Chen Y (2020) Automatic microseismic event picking via unsupervised machine learning. Geophysical Journal International 222:1750–1764. https://doi.org/10.1093/gji/ggaa186 Chen Y, Zhang M, Bai M, Chen W (2019) Improving the Signal‐to‐Noise Ratio of Seismological Datasets by Unsupervised Machine Learning. Seismological Research Letters 90:1552–1564. https://doi.org/10.1785/0220190028 Dawood HM, Rodriguez-Marek A, Bayless J, et al (2016) A flatfile for the KiK-net database processed using an automated protocol. Earthquake Spectra 32:1281–1302 Haiying Yu, Wenbin Wang, Quancai Xie, Yingchun Ma (2022) Zero-baseline correction method for near-fault strong motion records based on long-term and short-term memory model LSTM. Earthquake Engineering and Engineering Dynamics 42:35–42. https://doi.org/10.13197/j.eeed.2022.0405 Iwan WD, Moser MA, Peng C-Y (1985) Some observations on strong-motion earthquake measurement using a digital accelerograph. Bulletin of the Seismological Society of America 75:1225–1246. https://doi.org/10.1785/BSSA0750051225 Katsanos EI, Sextos AG, Manolis GD (2010) Selection of earthquake ground motion records: A state-of-the-art review from a structural engineering perspective. Soil Dynamics and Earthquake Engineering 30:157–169. https://doi.org/10.1016/j.soildyn.2009.10.005 Kuang W, Yuan C, Zhang J (2021) Real-time determination of earthquake focal mechanism via deep learning. Nat Commun 12:1432. https://doi.org/10.1038/s41467-021-21670-x Luzi L, Puglia R, Russo E, et al (2016) The Engineering Strong‐Motion Database: A Platform to Access Pan‐European Accelerometric Data. Seismological Research Letters 87:987–997. https://doi.org/10.1785/0220150278 Meier M-A, Ross ZE, Ramachandran A, et al (2019) Reliable Real-Time Seismic Signal/Noise Discrimination With Machine Learning. Journal of Geophysical Research: Solid Earth 124:788–800. https://doi.org/10.1029/2018JB016661 Musson RMW (1998) Intensity assignments from historical earthquake data: issues of certainty and quality. Annals of Geophysics 41:79–91. https://doi.org/10.4401/ag-3795 OpenAI R (2023) GPT-4 technical report. arXiv 2303–08774 Pacor F, Paolucci R, Ameri G, et al (2011) Italian strong motion records in ITACA: overview and record processing. Bulletin of Earthquake Engineering 9:1741–1759. https://doi.org/10.1007/s10518-011-9295-x Shan W, Zhang M, Wang M, et al (2022) EPM–DCNN: Earthquake Prediction Models Using Deep Convolutional Neural Networks. Bulletin of the Seismological Society of America 112:2933–2945. https://doi.org/10.1785/0120220058 Shang Q, Guo X, Li Q, et al (2020) A benchmark city for seismic resilience assessment. Earthquake Engineering and Engineering Vibration 19:811–826. https://doi.org/10.1007/s11803-020-0597-3 Somerville PG, Smith NF, Graves RW, Abrahamson NA (1997) Modification of Empirical Strong Ground Motion Attenuation Relations to Include the Amplitude and Duration Effects of Rupture Directivity. Seismological Research Letters 68:199–222. https://doi.org/10.1785/gssrl.68.1.199 Vaswani A, Shazeer N, Parmar N, et al (2017) Attention is All you Need. In: Guyon I, Luxburg UV, Bengio S, et al. (eds) Advances in Neural Information Processing Systems. Curran Associates, Inc. W. Zhu, S. M. Mousavi, G. C. Beroza (2019) Seismic Signal Denoising and Decomposition Using Deep Neural Networks. IEEE Transactions on Geoscience and Remote Sensing 57:9476–9488. https://doi.org/10.1109/TGRS.2019.2926772 Wang M, Shen J, Pan ZA, Han DL (2019) An improved supported vector regression algorithm with application to predict aftershocks. Journal of Seismology 23:983–993. https://doi.org/10.1007/s10950-019-09848-9 Xin-xin Y, Ye-fei R, Tadahiro K, et al (2022) The procedure of filtering the strong motion record: Denoising and filtering. 工程力学 39:320–329. https://doi.org/10.6052/j.issn.1000-4750.2021.05.S058 Xu Z, Chen J, Shen J, Xiang M (2022) Recursive long short-term memory network for predicting nonlinear structural seismic response. Engineering Structures 250:113406. https://doi.org/10.1016/j.engstruct.2021.113406 Youngs RR, Chiou S-J, Silva WJ, Humphrey JR (1997) Strong Ground Motion Attenuation Relationships for Subduction Zone Earthquakes. Seismological Research Letters 68:58–73. https://doi.org/10.1785/gssrl.68.1.58 Zhao L-S, Helmberger DV (1994) Source estimation from broadband regional seismograms. Bulletin of the Seismological Society of America 84:91–104. https://doi.org/10.1785/BSSA0840010091 Zhou B, Wang H, Xie L, Wang Y (2015) Bizarre Waveforms in Strong Motion Records. Shock and Vibration 2015:630362. https://doi.org/10.1155/2015/630362 Zhou BF, Song TS, Wen RZ, Xie LL (2014) Permanent Displacement Identification Analysis in 2011 Mw9.0 Tohoku Earthquake, Japan. Applied Mechanics and Materials 580–583:1533–1537. https://doi.org/10.4028/www.scientific.net/AMM.580-583.1533 Zhu W, Beroza GC (2019) PhaseNet: a deep-neural-network-based seismic arrival-time picking method. Geophysical Journal International 216:261–273. https://doi.org/10.1093/gji/ggy423 Cite Share Download PDF Status: Published Journal Publication published 09 Oct, 2024 Read the published version in Acta Geophysica → Version 1 posted Editorial decision: Minor revisions 05 Feb, 2024 Editor invited by journal 10 Nov, 2023 Reviewers agreed at journal 21 Sep, 2023 Reviewers invited by journal 16 Sep, 2023 Editor assigned by journal 12 Aug, 2023 First submitted to journal 05 Aug, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3237271","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":234065337,"identity":"aa25a266-f92e-4510-af2a-eb5cb5afa33b","order_by":0,"name":"Baofeng Zhou","email":"","orcid":"","institution":"China Earthquake Administration Institute of Engineering Mechanics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Baofeng","middleName":"","lastName":"Zhou","suffix":""},{"id":234065338,"identity":"6fc280a5-ca0d-472c-b9fc-6053f7a0a014","order_by":1,"name":"Yue Yin","email":"","orcid":"","institution":"China Earthquake Administration Institute of Engineering Mechanics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yue","middleName":"","lastName":"Yin","suffix":""},{"id":234065339,"identity":"2f31b2e0-72b4-4f24-969a-52256f89f4b7","order_by":2,"name":"Maofa Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9ElEQVRIiWNgGAWjYBADfgb2xoYDHyok5PiJ1SLZwHP44MMZZyyMJRuI1iKRlmzM21aRuIGQFoPjZw+/ulFjI8HfkGMmOXOeBOMGBuaHj27g03ImL80651iahMSBM2YSH7dJMJszsBkb5+DRYnYgx8w4h+1wHcPBHqAt2yTYLBt42KTxajn/Bqjl32EJ+cM8ZtK8cyR4DA4Q0nIjx/hxbtthCYNjbEDvN0hIENRif+ONGXNuX5qE4RlmYCAfkzCQbCbgF8n+HOPPOd9sJOTuPwRGZU1dfT9788PH+LQAAZsEKp8Zv3Kwkg+E1YyCUTAKRsGIBgBbrVBH7ARJqQAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0001-6517-4042","institution":"Guilin University of Electronic Technology School of Computer Science and Information Security","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Maofa","middleName":"","lastName":"Wang","suffix":""},{"id":234065340,"identity":"0a603c73-095a-4948-8e10-c8c33c202969","order_by":3,"name":"Runjie Zhang","email":"","orcid":"","institution":"Institute of Disaster Prevention","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Runjie","middleName":"","lastName":"Zhang","suffix":""},{"id":234065341,"identity":"98e5604c-5cad-4a36-bb5e-69ea089bcb4f","order_by":4,"name":"Yue Zhang","email":"","orcid":"","institution":"China Earthquake Administration Institute of Engineering Mechanics","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yue","middleName":"","lastName":"Zhang","suffix":""},{"id":234065342,"identity":"1200e838-1098-4ecf-9045-559b4a4c3923","order_by":5,"name":"Wenheng Guo","email":"","orcid":"","institution":"Institute of Disaster Prevention","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wenheng","middleName":"","lastName":"Guo","suffix":""}],"badges":[],"createdAt":"2023-08-05 11:48:04","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3237271/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3237271/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11600-024-01460-x","type":"published","date":"2024-10-09T15:57:32+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":43535260,"identity":"6618c4f1-6848-4bc7-84a2-326f87ca093a","added_by":"auto","created_at":"2023-09-22 13:49:45","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":36902,"visible":true,"origin":"","legend":"\u003cp\u003eSystem Framework\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/6408894fa66a55025b3062e4.png"},{"id":43534339,"identity":"a2025778-1c5e-4f97-8cb1-8cbc95620d80","added_by":"auto","created_at":"2023-09-22 13:41:45","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":50935,"visible":true,"origin":"","legend":"\u003cp\u003eAcceleration, velocity, and displacement time history with baseline drift\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/35b921783be4f56a36a5f435.png"},{"id":43534341,"identity":"47205f76-3c35-4eaf-8905-bb6a9162cb52","added_by":"auto","created_at":"2023-09-22 13:41:45","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":57569,"visible":true,"origin":"","legend":"\u003cp\u003eAcceleration, velocity, and displacement time history without baseline drift\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/a5aa6e9fee561f6277db4420.png"},{"id":43535261,"identity":"279918cc-0164-44a4-9acb-429cd851bd82","added_by":"auto","created_at":"2023-09-22 13:49:45","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":44207,"visible":true,"origin":"","legend":"\u003cp\u003eThe Transformer network framework for identifying baseline drift\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/32985c23c7b2c9d8fc5c4d3a.png"},{"id":43534340,"identity":"a5444afc-ef9e-40e9-acce-63d2b58939db","added_by":"auto","created_at":"2023-09-22 13:41:45","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":40344,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix of Transformer model without Bayesian optimization\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/5c83dd7cc80ba01cef6656f5.png"},{"id":43534343,"identity":"1535a0a1-36cb-4085-9744-dfa3f153295e","added_by":"auto","created_at":"2023-09-22 13:41:45","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":40831,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrix of the Bayesian optimized Transformer model\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/6d6fe73e2c349998dffcde0f.png"},{"id":66597383,"identity":"7f60a85a-a623-4f17-a2ce-c4fd47b640cf","added_by":"auto","created_at":"2024-10-14 16:10:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":626976,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3237271/v1/353f1f28-e252-4147-bb9d-2f1328a510f3.pdf"}],"financialInterests":"","formattedTitle":"Identification of Strong Motion Record Baseline Drift Based on Bayesian Optimized Transformer Network","fulltext":[{"header":"Key Points","content":"\u003cp\u003e\u003cul\u003e \u003cli\u003e \u003cp\u003eA neural network utilizing the Transformer model is proposed for the recognition of baseline drift in strong motion records.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe Bayesian optimization method is employed to optimize the hyperparameters of the Transformer model, thereby enhancing its overall performance.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe baseline drift recognition model holds great potential as it can be employed in tasks such as baseline drift correction and sequence classification, offering broad prospects for its application.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e\u003c/p\u003e"},{"header":"1. Introduction","content":"\u003cp\u003eStrong motion observation is an important research area in the field of earthquake engineering. It has significant applications in studying source mechanisms (Zhao and Helmberger \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e1994\u003c/span\u003e; Kuang et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), ground motion attenuation relationships (Somerville et al. \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e1997\u003c/span\u003e; Youngs et al. \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e1997\u003c/span\u003e), as well as seismic design of structures (Katsanos et al. \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2010\u003c/span\u003e) and seismic resilience assessment of cities (Shang et al. \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Additionally, strong motion observation networks can be used to establish earthquake early warning systems and rapid assessment systems for disaster situations. In the past few decades, strong motion observation networks have been extensively developed worldwide, leading to the acquisition of a large number of strong motion records and the establishment of multiple strong motion databases (Pacor et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2011\u003c/span\u003e; Ancheta et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Akkar et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2014\u003c/span\u003e; Dawood et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Luzi et al. \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). However, due to environmental noise, instrument failures, and interference in the telecommunications transmission process, the quality of strong motion records may be poor (Boore and Bommer \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2005\u003c/span\u003e), which is mainly manifested in the acceleration time history, such as spikes, baseline drifts, and asymmetry (Zhou et al. \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). One major problem encountered by digital accelerometers is the distortion and offset of the reference baseline, resulting in unphysical velocity and displacement, known as baseline drift (Boore and Bommer \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2005\u003c/span\u003e). Many scholars have conducted related research on the baseline drift phenomenon in strong motion records and proposed many correction methods (Iwan et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e1985\u003c/span\u003e; Boore \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e1999\u003c/span\u003e; Zhou et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). For each individual strong motion record, we apply the method of integrating acceleration to obtain velocity and displacement. Then, based on the displacement time history, we can quickly determine whether there is a baseline drift phenomenon (Boore \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e1999\u003c/span\u003e). However, with the significant expansion of strong motion observation networks and the substantial increase in the number of strong motion records, it has become difficult for manual selection to efficiently and quickly identify records with baseline drift from thousands of strong motion records. There is an urgent need for automated processing of strong motion data, automated identification of abnormal waveform data, in order to further improve the quality of strong motion records.\u003c/p\u003e \u003cp\u003eIn recent years, there has been significant development in the application of machine learning and deep learning techniques in the field of earthquake engineering. For instance, Wang et al. (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) conducted research on aftershock prediction using machine learning-based support vector regression algorithms. Subsequently, they also employed deep convolutional neural networks for earthquake prediction (Shan et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Xu et al. (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) performed dynamic time history analysis of structural seismic performance using long short-term memory networks. Bellagamba et al. (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) utilized feedforward neural networks to classify the quality of strong motion records with low seismic intensity. Regarding the quality of strong motion records, researchers have proposed numerous methods for classification and processing using machine learning and deep learning approaches. For instance, Musson(1998) proposed a data quality assessment system for three types of quality problems: reliability of intensity assessment, positional certainty or uncertainty, and accuracy of raw data. Each indicator was quantified as a binary variable and given a final quality code from 0 (best) to 7 (worst). Chen et al. (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) developed a new denoising framework based on unsupervised machine learning techniques using the concept of autoencoding to improve the signal-to-noise ratio of many earthquake datasets and adaptively learn seismic signals from noisy observations. Moreover, they also clustered the time samples into two groups, waveform points and non-waveform points, using unsupervised machine learning algorithms to help identify seismic waveforms in microseism or earthquake data (Chen \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Meier et al. (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) showed how modern machine learning classifiers can significantly improve real-time signal/noise discrimination. Zhu and Beroza (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) proposed a deep neural network-based method for earthquake arrival time picking called 'PhaseNet,' which simultaneously picks the arrival times of both P and S waves. Later, W. Zhu et al. (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) developed a deep neural network-based denoising/decomposition method called \"Deep Denoiser\", which can decompose input data into the interested signal and noise (defined as any non-seismic signal). Yu et al. (2022) utilized Long Short-Term Memory (LSTM) models to assess the quality of records and perform zero-baseline correction on near-fault strong motion records, achieving automated data processing. These methods have significantly advanced the automation of processing and identification of strong motion records quality. However, more efficient and targeted methods are still needed for the recognition of peculiar waveforms such as baseline drift, spikes, and asymmetry caused by quality issues in strong motion records.\u003c/p\u003e \u003cp\u003eRecently, Chat-GPT has been a hot topic, and research in natural language processing has shown the strong capabilities of Chat-GPT (OpenAI \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). The deep learning model used by Chat-GPT is Transformer, which is a powerful and pioneering model in deep learning that can adapt to different use cases such as translation, language processing, classification, etc. In this paper, we utilize the superior performance of Transformer and establish a baseline drift identification model for strong motion records based on the encoding part of Transformer. Additionally, Bayesian optimization is employed to optimize the parameters of the model. The model fully utilizes the characteristics of sequence input and classification in Transformer, enabling it to effectively learn the temporal sequence features of strong motion records and make judgments about their data quality. In the experiments, by considering the temporal nature and visual characteristics of the strong motion records, significant features are extracted and imported into the baseline drift identification model, allowing for quick determination of whether baseline drift has occurred. The baseline drift identification neural network based on Transformer significantly enhances the efficiency of quality processing for baseline drift phenomena in strong motion records. The Bayesian optimized transformer-based neural network for baseline drift identification has greatly improved the efficiency of identifying records with baseline drift in strong motion records. Based on the findings of this study, it provides a more efficient and accurate method for handling the quality of strong motion records, and also offers a new approach and methodology for the application of deep learning techniques in the field of earthquakes. We believe that this Bayesian optimized transformer-based neural network for baseline drift identification will have a positive impact on research and applications in the field of seismology.\u003c/p\u003e"},{"header":"2. Research Content","content":"\u003cp\u003eThis section provides an overview of the research content, which mainly consists of three parts. The first part is a systematic process description that introduces the overall architecture and operational flow of the system. The second part includes the experimental dataset, including data sources, data formats, and data processing methods. The third part describes the model used in this experiment.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 System Framework\u003c/h2\u003e \u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, the system consists of two main aspects: constructing the dataset and training the model. The strong motion records undergo data preprocessing, followed by the generation of acceleration, velocity, and displacement plots. The record is classified based on the velocity or displacement image to determine if there is baseline drift, and the strong motion records are classified into two categories: those with baseline drift and those without baseline drift. The upper and lower images in the figure respectively show velocity images without and with baseline drift. Finally, the extracted displacement sequences corresponding to the two categories of images are converted into vector form and used to construct the training, validation, and test sets, which are input into the Transformer model. The feature information obtained after training the model is used to predict whether a new dataset contains baseline drift.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Dataset\u003c/h2\u003e \u003cp\u003eTo apply the Transformer model to identify baseline drift waveforms in strong motion records, data need to be processed and divided to ensure that they can be effectively input to the model for learning and classification. The data processing process is as follows:\u003c/p\u003e \u003cp\u003e(1) Data collection: Collecting strong motion records for the study. The data used in this research is from the Chi-Chi earthquake in Taiwan, China, comprising a total of 420 strong motion records. These records are divided into training, validation, and prediction sets.\u003c/p\u003e \u003cp\u003e(2) Data preprocessing: Preprocess strong motion records by first adjusting the acceleration time history to the zero baseline, subtracting the average value of the noise record for the first 20 seconds(Xin-xin et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), and then performing a Butterworth second-order causal filtering with a bandpass range of 0.01-35Hz to remove noise from the acceleration records for better feature recognition. Finally, the data is standardized to adjust the data to the same scale for easier comparison classification and feature learning in the Transformer model.\u003c/p\u003e \u003cp\u003e(3) Classification: Classify the strong motion records into baseline drift and non-baseline drift categories. First, integrate the processed acceleration records once and twice to obtain velocity and displacement time histories. Then, based on the velocity or displacement time histories, classify whether baseline drift exists. The basis for the classification is that slight baseline shifts in acceleration can lead to significant displacement in velocity and displacement time history (Boore \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2001\u003c/span\u003e), with baseline drift being particularly evident in displacement time history. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the acceleration, velocity, and displacement time histories without baseline drift. In the displacement time history, there is a nearly constant residual value in the last approximately 1/3 of the duration. The velocity time history exhibits oscillations around the zero baseline at the end of strong motion (Boore \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2001\u003c/span\u003e). However, as shown in Fig.\u0026nbsp;3, when baseline drift occurs, different manifestations are observed. The slight variations in the baseline of the acceleration time history can lead to a certain slope in the observed velocity time history and a diverging displacement in the final part of the displacement time history, exhibiting a continuous upward trend instead of a steady state. Comparing the two graphs, it can be observed that the occurrence of baseline drift in the acceleration time history of strong motion records can be identified by examining the last one-third portion of the displacement time history, where more pronounced baseline drift characteristics are evident. Therefore, the displacement time history can be used to classify the strong motion records.\u003c/p\u003e \u003cp\u003e(4) Sequence processing: Considering the input requirements of the Transformer model, this study further processes the preprocessed data by segmenting it with a certain time step. The last 4000 data points of the extracted displacement data are used for model training, divided into sequences of shape 10*400. These sequences are then input into the model in the form of a three-dimensional tensor with dimensions of n*10*400. A total of 72 training sets, 24 validation sets, and 24 test sets are created.\u003c/p\u003e \u003cp\u003e(5) Sequence encoding: Encode the split sequences to transform them into tensor format that can be processed by the Transformer model.\u003c/p\u003e \u003cp\u003e(6) Label encoding: Encode label information to a form that the Transformer model can be processed.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Model\u003c/h2\u003e \u003cp\u003eWe have adopted a Bayesian-optimized Transformer model to establish a neural network for baseline drift recognition in strong motion records. The framework of the network is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The entire network mainly consists of the encoder layer of the Transformer model. The input displacement sequences, after being divided and position encoded, are then fed into the Transformer module to extract image features, progressively restore details, and finally obtain the ultimate prediction results through global average pooling and a fully connected network.\u003c/p\u003e \u003cp\u003eThe core of our neural network architecture is the Transformer, a deep learning model used for sequence-to-sequence learning. The Transformer model, introduced by Google in 2017 (Vaswani et al. \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), is widely used in natural language processing and computer vision. At the heart of the Transformer is the self-attention mechanism, which learns the dependencies between elements in a sequence. By using self-attention, the Transformer can address issues such as gradient vanishing and exploding commonly found in traditional recurrent neural networks. Additionally, self-attention enables parallel computation, accelerating the training process of the model. In this paper, we calculate the similarity between each element of the input to obtain a weight vector, and then compute the representation vector for each element through weighted averaging. The self-attention mechanism is implemented using a set of query (Q), key (K), and value (V) matrices. The computation is as follows:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Attention\\left(Q,K,V\\right)=softmax\\left(\\frac{Q{K}^{T}}{\\left(\\sqrt{{d}_{k}}\\right)}\\right)V\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(1)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Q,K, and V\\)\u003c/span\u003e\u003c/span\u003e are matrices representing queries, keys, and values, respectively, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({d}_{k}\\)\u003c/span\u003e\u003c/span\u003e is the dimension of the key vectors. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Attention\\left(Q,K,V\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the output vector of the self-attention mechanism.\u003c/p\u003e \u003cp\u003eThe Transformer model adds a multi-head attention mechanism based on self-attention. The multi-head attention mechanism is an important component of the Transformer model that is used to learn multiple different dependencies. It applies multiple parallel self-attention \"heads\" to learn the projected Q, K, and V matrices. Compared to a single attention head, this improves performance. Specifically, it linearly transforms the input sequence separately, and then obtains vectors representing the dependencies between elements in the sequence by computing attention. Finally, these vectors are concatenated and linearly transformed to obtain the final representation. The representation is shown below:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabb\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(MultiHead\\left(Q,K,V\\right)=Concat\\left(hea{d}_{1},hea{d}_{2},\\dots ,hea{d}_{h}\\right){W}^{O}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(2)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ewhere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(hea{d}_{i}=Attention\\left(Q{W}_{i}^{Q},K{W}_{i}^{K},V{W}_{i}^{V}\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the vector output by the i-th attention head. In this paper, we use two attention heads. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Concat\\)\u003c/span\u003e\u003c/span\u003e combines the outputs of multiple heads. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{i}^{Q},{W}_{i}^{K}, and {W}_{i}^{V}\\)\u003c/span\u003e\u003c/span\u003e represent the linear transformation matrices of the query, key, and value for the i-th head, respectively, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}^{O}\\)\u003c/span\u003e\u003c/span\u003e represents the output matrix of the multi-head attention.\u003c/p\u003e \u003cp\u003eBefore inputting the sequence into the Transformer, we need to apply positional encoding to the input sequence. Strong motion records contain temporal features and exhibit a clear time order. The temporal ordering information of the input sequence is crucial. The purpose of positional encoding is to add the time position information of the input sequence to the model, so that the model can better capture the order information in the sequence. Specifically, positional encoding can be calculated using the following formula:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabc\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(P{E}_{\\left(pos,2i\\right)}=sin\\left(pos/{10000}^{2i/{d}_{model}}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(3)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(P{E}_{\\left(pos,2i+1\\right)}=cos\\left(pos/{10000}^{2i/{d}_{model}}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(P{E}_{\\left(pos,2i\\right)}\\)\u003c/span\u003e\u003c/span\u003e represents the positional encoding value at position \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(pos\\)\u003c/span\u003e\u003c/span\u003e and dimension \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(2i\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(P{E}_{\\left(pos,2i+1\\right)}\\)\u003c/span\u003e\u003c/span\u003e represents the positional encoding value at position \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(pos\\)\u003c/span\u003e\u003c/span\u003e and dimension \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(2i+1\\)\u003c/span\u003e\u003c/span\u003e. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({d}_{model}\\)\u003c/span\u003e\u003c/span\u003e represents the dimension of the input sequence. In this paper, the input vector is 10*400, so the sequence dimension \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({d}_{model}\\)\u003c/span\u003e\u003c/span\u003e is 400, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(pos\\)\u003c/span\u003e\u003c/span\u003e represents the position in the input sequence, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003e represents the dimension index in the positional encoding matrix.\u003c/p\u003e \u003cp\u003eAfter positional encoding, the sequences of strong motion records are passed through two Transformer encoder layers, each with two attention heads. The outputs are then normalized and connected with residual connections before being fed into a feed-forward neural network to generate the final output. In this paper, a baseline drift recognition model based on Keras is constructed and optimized using the RMSprop optimizer. Binary cross-entropy is employed as the loss function, and accuracy is used as the evaluation metric. The key parameters of the model are shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eKey parameters of the Transformer model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Block\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(parameter\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eTransformer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(MultiHeadAttention\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003enum_heads\u0026thinsp;=\u0026thinsp;2; key_dim\u0026thinsp;=\u0026thinsp;32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTimeDistributed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eff_dim\u0026thinsp;=\u0026thinsp;32; activation=\"relu\"\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eactivation=\"sigmoid\"\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWe used Bayesian optimization to optimize the hyperparameters of the model. Bayesian optimization (B. Shahriari et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) is a black-box optimization method, where the black-box function refers to a function without an explicit formula or expression. This algorithm builds a surrogate model, typically a Gaussian process, to estimate the objective function. It then employs Bayesian inference to update the posterior distribution of the surrogate model, effectively exploring and exploiting the function space. By optimizing the posterior distribution of the surrogate model, the Bayesian optimization algorithm can find the global optimum with relatively few iterations. In our case, we apply Bayesian optimization to optimize the vector dimensions (key_dim) of the Q, K, and V matrices in the multi-head attention layer, as well as the dimension (ff_dim) of the feed-forward neural network layer. The parameter optimization ranges are set as shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. The model uses the gp_minimize optimizer to perform Bayesian optimization on the hyperparameters. Through Bayesian optimization, the most suitable vector dimensions and network layers for the model are selected.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBayesian optimization of Transformer model parameters.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Parameter\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRange\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eTransformer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\text{k}\\text{e}\\text{y}\\_\\text{d}\\text{i}\\text{m}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e(16, 64)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(ff\\_dim\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e(16, 64)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"3. Result","content":"\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Evaluation Metrics\u003c/h2\u003e \u003cp\u003eTo assess the performance of the Transformer model, three commonly used evaluation metrics, namely accuracy, recall, and F1 score, are employed in this study. Figures\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003e and \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003e show the confusion matrix, which provides a visual representation to evaluate the performance of the classification algorithm.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e3.1.1 Precision\u003c/h2\u003e \u003cp\u003ePrecision refers to the proportion of sequences classified as baseline drift among the total number of strong motion sequences. It is an important metric for evaluating the performance of the classifier. The formula for calculating precision is as follows:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabd\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Precision= \\frac{TP+TN}{TP+FP+TN+FN}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(5)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \n\u003cp\u003eWhere:\u003c/p\u003e\n\u003cp\u003eTP represents the number of strong motion displacement sequences that have baseline drift and are classified as having baseline drift.\u003c/p\u003e \u003cp\u003eFP represents the number of strong motion displacement sequences that do not have baseline drift but are classified as having baseline drift.\u003c/p\u003e \u003cp\u003eTN represents the number of strong motion displacement sequences that do not have baseline drift and are classified as not having baseline drift.\u003c/p\u003e \u003cp\u003eFN represents the number of strong motion displacement sequences that have baseline drift but are classified as not having baseline drift.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.1.2 Recall\u003c/h2\u003e \u003cp\u003eRecall, also known as sensitivity or true positive rate, refers to the proportion of correctly classified baseline drift samples among the total number of baseline drift samples. It is calculated using the following formula:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabe\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Recall= \\frac{TP}{TP+FN}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(6)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e3.1.3 F1 score\u003c/h2\u003e \u003cp\u003eF1 is a metric that combines Precision and Recall of a model, and is calculated as follows.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabf\" border=\"1\"\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(F1=\\frac{2\\times Precision\\times Recall}{Precision+Recall}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(7)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe precision, recall rate, and F1 score of the model are calculated through the confusion matrix shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003e. As shown in the figure, \"normal\" represents normal waveforms and \"abnormal\" represents baseline drift waveforms. It can be seen that TP\u0026thinsp;=\u0026thinsp;67; TN\u0026thinsp;=\u0026thinsp;92; FP\u0026thinsp;=\u0026thinsp;8; FN\u0026thinsp;=\u0026thinsp;33. The precision is calculated as Precision = (67\u0026thinsp;+\u0026thinsp;92) / 200\u0026thinsp;=\u0026thinsp;0.795; Recall\u0026thinsp;=\u0026thinsp;67 / (67\u0026thinsp;+\u0026thinsp;33)\u0026thinsp;=\u0026thinsp;0.67; F1\u0026thinsp;=\u0026thinsp;2 * 0.795 * 0.67 / (0.795\u0026thinsp;+\u0026thinsp;0.67)\u0026thinsp;=\u0026thinsp;0.727. It can be seen that the accuracy calculated by the model without Bayesian optimization is generally lower.\u003c/p\u003e \u003cp\u003eThe confusion matrix after adding Bayesian optimization is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003e. the precision, Recall, and F1 score are calculated using the above formula with TP\u0026thinsp;=\u0026thinsp;78, TN\u0026thinsp;=\u0026thinsp;95, FP\u0026thinsp;=\u0026thinsp;5, and FN\u0026thinsp;=\u0026thinsp;22. The results are as follows: Precision = (78\u0026thinsp;+\u0026thinsp;95)/200\u0026thinsp;=\u0026thinsp;0.865; Recall\u0026thinsp;=\u0026thinsp;78/(78\u0026thinsp;+\u0026thinsp;22)\u0026thinsp;=\u0026thinsp;0.78; F1\u0026thinsp;=\u0026thinsp;(2\u0026times;0.865\u0026times;0.78)/(0.865\u0026thinsp;+\u0026thinsp;0.78)\u0026thinsp;=\u0026thinsp;0.820. The precision has significantly improved.\u003c/p\u003e\u003cp\u003eThe precision, recall, and F1 score of the model before and after Bayesian optimization are shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eThe precision, recall, and F1 score of the model.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Precision\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Recall\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(F1 score\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTransformer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.795\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.727\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTransformer by Bayes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.865\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.820\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAccording to the data analysis in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, the Transformer model exhibited good performance in identifying and classifying baseline drift waveforms, Particularly, after applying Bayesian optimization, the model's precision and F1 score improved to 86.5% and 82.0% respectively. This indicates that Bayesian optimization can significantly enhance the model's performance, leading to better solutions for real-world problems.\u003c/p\u003e \u003cp\u003eFurthermore, from the table, it can be observed that the model's recall is 67%, slightly lower than the precision. This indicates that the model may have some misclassifications in identifying drift waveforms, specifically classifying normal waveforms as drift waveforms. However, after applying Bayesian optimization, the recall of the model has also improved to some extent.\u003c/p\u003e \u003cp\u003eIn summary, the data analysis in the table above shows that the Transformer model performs well in the identification and classification of baseline drift waveforms, but there is still room for further optimization. By continuously optimizing the feature extraction and classification strategy of the model, the performance of the model can be further improved to better solve practical problems.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Conclusion","content":"\u003cp\u003eBased on the characteristics of baseline drift in strong motion records, this paper proposes a deep learning neural network model, the transformer, for identifying baseline drift in strong motion records from an image recognition perspective. By utilizing the temporal features of the time series in strong motion records, the transformer model effectively extracts features related to baseline drift and achieves high recognition efficiency even with small sample sizes. Based on the proposed method and experimental results, the following conclusions can be drawn:\u003c/p\u003e \u003cp\u003e(1) Strong motion data, as temporal data, possess distinctive features that vary with time. These features can be effectively extracted by the transformer model, which demonstrates high accuracy in baseline drift identification. The accuracy of the model on the test set exceeds 85%, indicating its effectiveness in recognition.\u003c/p\u003e \u003cp\u003e(2) The performance of the model is further improved by Bayesian optimization of the transformer model.\u003c/p\u003e \u003cp\u003eThis model can be embedded in the calibration of baseline drift in strong motion records to determine whether the corrected acceleration images meet the requirements. Furthermore, the development of this model holds great promise. The transformer model shows good prospects for the classification of sequences with temporal features. However, due to the limited number of samples in this study, there is still room for further optimization. By continuously improving the feature extraction and classification strategies of the model, its performance can be further enhanced to better address real-world problems.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgment\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research was funded by the Scientific Research Fund of Institute of Engineering Mechanics, China Earthquake Administration (Nos.2022C05), and the Natural Science Foundation of Heilongjiang Province (LH2022E120) and the National Key Research and Development Program of China (2018YFE0109800). This support is greatly appreciated. This work was also supported by Guangxi Project of Technology Base and Special Talent (AD20325004), the National Natural Science Foundation of China (42164002).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that there are no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAkkar S, Sandıkkaya MA, Şenyurt M, et al (2014) Reference database for seismic ground-motion in Europe (RESORCE). Bulletin of Earthquake Engineering 12:311\u0026ndash;339. https://doi.org/10.1007/s10518-013-9506-8\u003c/li\u003e\n\u003cli\u003eAncheta T, Darragh R, Stewart J, et al (2013) PEER NGA-West2 Database, PEER Report 2013/03, pacific earthquake engineering research center. University of California, Berkeley\u003c/li\u003e\n\u003cli\u003eB. Shahriari, K. Swersky, Z. Wang, et al (2016) Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE 104:148\u0026ndash;175. https://doi.org/10.1109/JPROC.2015.2494218\u003c/li\u003e\n\u003cli\u003eBellagamba X, Lee R, Bradley BA (2019) A Neural Network for Automated Quality Screening of Ground Motion Records from Small Magnitude Earthquakes. Earthquake Spectra 35:1637\u0026ndash;1661. https://doi.org/10.1193/122118EQS292M\u003c/li\u003e\n\u003cli\u003eBoore DM (1999) Effect of baseline corrections on response spectra for two recordings of the 1999 Chi-Chi, Taiwan, earthquake\u003c/li\u003e\n\u003cli\u003eBoore DM (2001) Effect of Baseline Corrections on Displacements and Response Spectra for Several Recordings of the 1999 Chi-Chi, Taiwan, Earthquake. Bulletin of the Seismological Society of America 91:1199\u0026ndash;1211. https://doi.org/10.1785/0120000703\u003c/li\u003e\n\u003cli\u003eBoore DM, Bommer JJ (2005) Processing of strong-motion accelerograms: needs, options and consequences. Soil Dynamics and Earthquake Engineering 25:93\u0026ndash;115. https://doi.org/10.1016/j.soildyn.2004.10.007\u003c/li\u003e\n\u003cli\u003eChen Y (2020) Automatic microseismic event picking via unsupervised machine learning. Geophysical Journal International 222:1750\u0026ndash;1764. https://doi.org/10.1093/gji/ggaa186\u003c/li\u003e\n\u003cli\u003eChen Y, Zhang M, Bai M, Chen W (2019) Improving the Signal‐to‐Noise Ratio of Seismological Datasets by Unsupervised Machine Learning. Seismological Research Letters 90:1552\u0026ndash;1564. https://doi.org/10.1785/0220190028\u003c/li\u003e\n\u003cli\u003eDawood HM, Rodriguez-Marek A, Bayless J, et al (2016) A flatfile for the KiK-net database processed using an automated protocol. Earthquake Spectra 32:1281\u0026ndash;1302\u003c/li\u003e\n\u003cli\u003eHaiying Yu, Wenbin Wang, Quancai Xie, Yingchun Ma (2022) Zero-baseline correction method for near-fault strong motion records based on long-term and short-term memory model LSTM. Earthquake Engineering and Engineering Dynamics 42:35\u0026ndash;42. https://doi.org/10.13197/j.eeed.2022.0405\u003c/li\u003e\n\u003cli\u003eIwan WD, Moser MA, Peng C-Y (1985) Some observations on strong-motion earthquake measurement using a digital accelerograph. Bulletin of the Seismological Society of America 75:1225\u0026ndash;1246. https://doi.org/10.1785/BSSA0750051225\u003c/li\u003e\n\u003cli\u003eKatsanos EI, Sextos AG, Manolis GD (2010) Selection of earthquake ground motion records: A state-of-the-art review from a structural engineering perspective. Soil Dynamics and Earthquake Engineering 30:157\u0026ndash;169. https://doi.org/10.1016/j.soildyn.2009.10.005\u003c/li\u003e\n\u003cli\u003eKuang W, Yuan C, Zhang J (2021) Real-time determination of earthquake focal mechanism via deep learning. Nat Commun 12:1432. https://doi.org/10.1038/s41467-021-21670-x\u003c/li\u003e\n\u003cli\u003eLuzi L, Puglia R, Russo E, et al (2016) The Engineering Strong‐Motion Database: A Platform to Access Pan‐European Accelerometric Data. Seismological Research Letters 87:987\u0026ndash;997. https://doi.org/10.1785/0220150278\u003c/li\u003e\n\u003cli\u003eMeier M-A, Ross ZE, Ramachandran A, et al (2019) Reliable Real-Time Seismic Signal/Noise Discrimination With Machine Learning. Journal of Geophysical Research: Solid Earth 124:788\u0026ndash;800. https://doi.org/10.1029/2018JB016661\u003c/li\u003e\n\u003cli\u003eMusson RMW (1998) Intensity assignments from historical earthquake data: issues of certainty and quality. Annals of Geophysics 41:79\u0026ndash;91. https://doi.org/10.4401/ag-3795\u003c/li\u003e\n\u003cli\u003eOpenAI R (2023) GPT-4 technical report. arXiv 2303\u0026ndash;08774\u003c/li\u003e\n\u003cli\u003ePacor F, Paolucci R, Ameri G, et al (2011) Italian strong motion records in ITACA: overview and record processing. Bulletin of Earthquake Engineering 9:1741\u0026ndash;1759. https://doi.org/10.1007/s10518-011-9295-x\u003c/li\u003e\n\u003cli\u003eShan W, Zhang M, Wang M, et al (2022) EPM\u0026ndash;DCNN: Earthquake Prediction Models Using Deep Convolutional Neural Networks. Bulletin of the Seismological Society of America 112:2933\u0026ndash;2945. https://doi.org/10.1785/0120220058\u003c/li\u003e\n\u003cli\u003eShang Q, Guo X, Li Q, et al (2020) A benchmark city for seismic resilience assessment. Earthquake Engineering and Engineering Vibration 19:811\u0026ndash;826. https://doi.org/10.1007/s11803-020-0597-3\u003c/li\u003e\n\u003cli\u003eSomerville PG, Smith NF, Graves RW, Abrahamson NA (1997) Modification of Empirical Strong Ground Motion Attenuation Relations to Include the Amplitude and Duration Effects of Rupture Directivity. Seismological Research Letters 68:199\u0026ndash;222. https://doi.org/10.1785/gssrl.68.1.199\u003c/li\u003e\n\u003cli\u003eVaswani A, Shazeer N, Parmar N, et al (2017) Attention is All you Need. In: Guyon I, Luxburg UV, Bengio S, et al. (eds) Advances in Neural Information Processing Systems. Curran Associates, Inc.\u003c/li\u003e\n\u003cli\u003eW. Zhu, S. M. Mousavi, G. C. Beroza (2019) Seismic Signal Denoising and Decomposition Using Deep Neural Networks. IEEE Transactions on Geoscience and Remote Sensing 57:9476\u0026ndash;9488. https://doi.org/10.1109/TGRS.2019.2926772\u003c/li\u003e\n\u003cli\u003eWang M, Shen J, Pan ZA, Han DL (2019) An improved supported vector regression algorithm with application to predict aftershocks. Journal of Seismology 23:983\u0026ndash;993. https://doi.org/10.1007/s10950-019-09848-9\u003c/li\u003e\n\u003cli\u003eXin-xin Y, Ye-fei R, Tadahiro K, et al (2022) The procedure of filtering the strong motion record: Denoising and filtering. 工程力学 39:320\u0026ndash;329. https://doi.org/10.6052/j.issn.1000-4750.2021.05.S058\u003c/li\u003e\n\u003cli\u003eXu Z, Chen J, Shen J, Xiang M (2022) Recursive long short-term memory network for predicting nonlinear structural seismic response. Engineering Structures 250:113406. https://doi.org/10.1016/j.engstruct.2021.113406\u003c/li\u003e\n\u003cli\u003eYoungs RR, Chiou S-J, Silva WJ, Humphrey JR (1997) Strong Ground Motion Attenuation Relationships for Subduction Zone Earthquakes. Seismological Research Letters 68:58\u0026ndash;73. https://doi.org/10.1785/gssrl.68.1.58\u003c/li\u003e\n\u003cli\u003eZhao L-S, Helmberger DV (1994) Source estimation from broadband regional seismograms. Bulletin of the Seismological Society of America 84:91\u0026ndash;104. https://doi.org/10.1785/BSSA0840010091\u003c/li\u003e\n\u003cli\u003eZhou B, Wang H, Xie L, Wang Y (2015) Bizarre Waveforms in Strong Motion Records. Shock and Vibration 2015:630362. https://doi.org/10.1155/2015/630362\u003c/li\u003e\n\u003cli\u003eZhou BF, Song TS, Wen RZ, Xie LL (2014) Permanent Displacement Identification Analysis in 2011 Mw9.0 Tohoku Earthquake, Japan. Applied Mechanics and Materials 580\u0026ndash;583:1533\u0026ndash;1537. https://doi.org/10.4028/www.scientific.net/AMM.580-583.1533\u003c/li\u003e\n\u003cli\u003eZhu W, Beroza GC (2019) PhaseNet: a deep-neural-network-based seismic arrival-time picking method. Geophysical Journal International 216:261\u0026ndash;273. https://doi.org/10.1093/gji/ggy423\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"acta-geophysica","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"agph","sideBox":"Learn more about [Acta Geophysica](http://link.springer.com/journal/11600)","snPcode":"11600","submissionUrl":"https://www.editorialmanager.com/agph/default2.aspx","title":"Acta Geophysica","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Baseline Drift, Transformer Network, Strong Motion Record, Bayesian Optimization, Sequence Classification","lastPublishedDoi":"10.21203/rs.3.rs-3237271/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3237271/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eResearch in earthquake engineering heavily relies on strong motion observation. The quality of strong motion records directly affects the reliability of earthquake disaster prevention, rapid reporting of seismic magnitude, earthquake early warning, and other areas. Currently, the quality assurance of strong motion records often relies on basic methods such as zero-line adjustment and high-pass filtering. However, these methods often fail to satisfactorily identify and handle abnormal waveforms present in strong motion records, and their efficiency is relatively low. In this paper, a Bayesian-optimized Transformer-based approach is proposed to improve the identification of baseline drift anomalies in strong motion records. By partitioning the strong motion record data from the 1999 Chi-Chi earthquake in Taiwan, China, into two categories: high-quality records (with minimal baseline drift) and low-quality records (with significant baseline drift), we extracted data with distinct features and inputted them into the proposed model for training. Finally, the model was used to predict whether strong motion records exhibited baseline drift abnormalities. The experimental results show that after performing Bayesian optimization on the parameters of the Transformer model, this method achieves an accuracy of over 85% in baseline drift identification. It is capable of efficiently identifying a large volume of strong motion records with baseline drift within a short period of time. The model performs well in the task of classifying baseline drift in strong motion records and can be used for subsequent identification of abnormalities after baseline drift correction, enabling automation in handling baseline drift abnormal data.\u003c/p\u003e","manuscriptTitle":"Identification of Strong Motion Record Baseline Drift Based on Bayesian Optimized Transformer Network","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-09-22 13:41:40","doi":"10.21203/rs.3.rs-3237271/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Minor revisions","date":"2024-02-05T06:49:48+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"Acta Geophysica","date":"2023-11-10T10:40:12+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2023-09-21T12:03:47+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-09-16T13:16:59+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-08-12T09:08:05+00:00","index":"","fulltext":""},{"type":"submitted","content":"Acta Geophysica","date":"2023-08-05T07:47:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"acta-geophysica","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"agph","sideBox":"Learn more about [Acta Geophysica](http://link.springer.com/journal/11600)","snPcode":"11600","submissionUrl":"https://www.editorialmanager.com/agph/default2.aspx","title":"Acta Geophysica","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"f310cf2c-6b59-43d6-bd28-96f9f8989d2c","owner":[],"postedDate":"September 22nd, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-10-14T16:05:25+00:00","versionOfRecord":{"articleIdentity":"rs-3237271","link":"https://doi.org/10.1007/s11600-024-01460-x","journal":{"identity":"acta-geophysica","isVorOnly":false,"title":"Acta Geophysica"},"publishedOn":"2024-10-09 15:57:32","publishedOnDateReadable":"October 9th, 2024"},"versionCreatedAt":"2023-09-22 13:41:40","video":"","vorDoi":"10.1007/s11600-024-01460-x","vorDoiUrl":"https://doi.org/10.1007/s11600-024-01460-x","workflowStages":[]},"version":"v1","identity":"rs-3237271","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3237271","identity":"rs-3237271","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00