Ensemble empirical mode decomposition and a long short-term memory neural network for surface water quality prediction of the Xiaofu River, China | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Ensemble empirical mode decomposition and a long short-term memory neural network for surface water quality prediction of the Xiaofu River, China Lan Luo, Yanjun Zhang, Wenxun Dong, Anni Qiu, Jinglin Zhang, Liping Zhang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2116084/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Water quality prediction is an important part of water pollution prevention and control. Using a long short-term memory (LSTM) neural network to predict water quality can solve the problem that comprehensive water quality models are too complex and difficult to apply. However, as water quality time series are generally multiperiod hybrid time series, which have strongly nonlinear and nonstationary characteristics, the prediction accuracy of LSTM for water quality is not high. The ensemble empirical mode decomposition (EEMD) method can decompose the multiperiod hybrid water quality time series into several simpler single-period components. To improve the accuracy of surface water quality prediction, a water quality prediction model based on EEMD-LSTM was proposed in this paper. The water quality time series was first decomposed into several intrinsic mode function components and one residual item, and then these components were used as the input of LSTM to predict water quality. The model was trained and validated using four water quality parameters (NH 3 N, pH, DO, COD Mn ) collected from the Xiaofu River and compared with the results of a single LSTM. During the validation period, the R 2 values when using LSTM for NH 3 N, pH, DO and COD Mn were 0.567, 0.657, 0.817 and 0.693, respectively, and the R 2 values when using EEMD-LSTM for NH 3 N, pH, DO and COD Mn were 0.924, 0.965, 0.961 and 0.936, respectively. The results show that the proposed model outperforms the single LSTM model in various evaluation indicators and greatly improves the model performance in terms of the hysteresis problem. The EEMD-LSTM model has high prediction accuracy and strong generalization ability, and further development may be valuable. Water quality prediction Ensemble empirical mode decomposition Long short-term memory network Deep learning Xiaofu River Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 1. Introduction With the rapid development of the economy over the past few decades, many water bodies in China have been seriously polluted, which affects people's quality of life and the safe water quality level (Tang et al. 2022 ; Xiong et al. 2020 ). Water environment management and protection have gradually become the focus of attention. Water quality prediction is an important link in the management and protection of aquatic environments. Scientific and accurate water quality prediction can help to understand the changing laws and development trends of the water environment, provide technical support for water environmental protection and water pollution prevention and control, and improve the decision-making initiatives of management departments (Liang et al. 2020 ; Yu et al. 2022 ). Many researchers have used comprehensive water quality models to simulate and predict water quality (Bui et al. 2019 ; Deus et al. 2013 ; Kim et al. 2017 ; Rui et al. 2015 ). At present, the comprehensive water quality models that have been widely used in water quality simulation and water environment management include the Water Quality Analysis Simulation Program (WASP) (Wool et al. 2020 ), QUAL series model (Shabani et al. 2021 ), the Environmental Fluid Dynamics Code (EFDC) (Kim et al. 2017 ), Delft3D model (Mendes et al. 2021 ), etc. However, there are many water quality parameters in comprehensive water quality models, and a large amount of measured water quality data is needed to set initial conditions and boundary conditions during simulation. Comprehensive water quality models are too complex and difficult to apply, and they are always data-intensive and time-consuming to develop (Da Silva Burigato Costa et al. 2021 ; Ejigu 2021 ). In addition, the use of complex models in the absence of data reduces the reliability of water quality prediction. Therefore, although water quality models can simulate the complex dynamics of water quality variables well, water quality prediction remains difficult. Using deep learning methods for water quality prediction can solve the problem of difficult application of comprehensive water quality models, as deep learning methods can effectively establish relationships between water quality parameters without complex boundaries and initial conditions (Liang et al. 2020 ). In recent years, artificial intelligence models such as artificial neural networks (ANNs) have been gradually applied to hydrological process analysis (Kourgialas et al. 2015 ; Yang et al. 2020 ; Zema et al. 2020 ) and water quality prediction (Kim et al. 2017 ; Palani et al. 2008 ; Seo et al. 2016 ). However, sequential order information is not reflected in the ANN training process, and ANNs do not perform well in nonlinear simulations (An et al. 2020 ). To overcome the shortcomings of ANNs, researchers have proposed recurrent neural networks (RNNs) and long short-term memory (LSTM) networks (Hochreiter and Schmidhuber 1997 ; Rumelhart et al. 1986 ). LSTM is an improved network structure proposed on the theoretical basis of RNNs. The network effectively overcomes the long-term dependence and easy gradient disappearance problems in RNNs and has better long-term and short-term memory function. Since LSTM was proposed, some researchers have applied it to the field of water quality modeling. For example, Zheng et al. ( 2021 ) used LSTM to effectively predict the concentration of chlorophyll-a and the outbreak of harmful algal blooms in a water body and provide another method for water resource management. Liang et al. ( 2020 ) found that LSTM could achieve the prediction accuracy of a comprehensive water quality model (such as EFDC). When the types of water quality data available are relatively simple, LSTM can be an effective tool for water quality prediction. However, due to the influence of hydrometeorological factors and human factors, water quality time series are nonlinear and nonstationary, so the prediction accuracy of LSTM for surface water quality is not high (Eze et al. 2021 ; Zhou et al. 2021 ; An et al. 2020 ). Surface water quality time series are generally multiperiod hybrid time series. According to the different periods, the water quality time series can be divided into high-frequency components (period of 1–10 days) and low-frequency components (period > 10 days). The main factors affecting the high-frequency components are sudden pollution and discontinuous nonpoint source pollution. These factors are closely related to the physical and chemical properties of pollutants, water quality, temperature, hydraulic conditions and other factors (Tant et al. 2015 ). The changing trend of these factors is large, which has a great influence on the accuracy of water quality prediction. The main factors affecting the low-frequency components are climate change, constant point source pollution, endogenous pollution and so on. The changing trend of these factors is relatively stable. Therefore, the main difficulty in water quality prediction is accurately predicting the high-frequency components in water quality time series. However, in existing studies, predicting the high-frequency components and low-frequency components separately when using the LSTM model for water quality prediction has rarely been considered, and the fluctuation term of the water quality series cannot be accurately predicted. Signal decomposition techniques can decompose the original water quality time series into a set of components with specific meanings and provide more detailed information. When predicting water quality, we can focus on high-frequency components to enhance the details and reduce the impact of interference information with signal decomposition techniques. Generally, the residence time of pollutants in water is 5–10 days. When the period of decomposed components is consistent with the degradation cycle of pollutants, the prediction accuracy of LSTM is likely to be improved. To overcome the limitations of a single LSTM method, LSTM can be combined with signal decomposition techniques to improve the accuracy of water quality prediction. Among signal decomposition algorithms, empirical mode decomposition (EMD) is widely used due to its orthogonality and convergence. It is easier to apply than wavelet decomposition. Huang et al. ( 1998 ) proposed EMD, which is a data-adaptive time frequency analysis method for nonlinear and nonstationary time series. EMD decomposes the original sequence into multiple intrinsic mode functions (IMFs) and residuals to reduce the complexity of the sequence. However, EMD has limitations such as modal confounding and end effects. Zhaohua and Norden ( 2009 ) proposed an improved empirical mode decomposition algorithm, EEMD, which addressed the modal confounding problem of EMD. EEMD can effectively reflect the nature of the original signal and has been widely used in many fields in recent years (An et al. 2020 ). For example, Wang et al. ( 2020 ) used EEMD to extract the oscillation period and the trend of runoff series and analyzed the relationship between runoff and climate phenomenon indicators. Niu et al. ( 2019 ) used EEMD to decompose the original monthly flow series, combined the improved gravitational search algorithm (IGSA) and extreme learning machine (ELM) for hydrological prediction, and successfully predicted the monthly runoff of the Three Gorges. Huan et al. ( 2018 ) proposed a combined prediction model based on EEMD and a least squares support vector machine (LSSVM), which had high prediction accuracy and strong generalization ability for dissolved oxygen (DO). From previous research, we know that EEMD can decompose the original water quality series into components arranged from high frequency to low frequency, and the time period of the high-frequency components is likely to be consistent with the degradation cycle of pollutants. Therefore, EEMD is suitable for decomposing water quality series into several components for water quality prediction using LSTM separately. To acquire better prediction performance of the surface water quality, a hybrid water quality prediction model based on the ensemble empirical mode decomposition method and long short-term memory neural network is proposed in this paper. The original water quality time series is decomposed into high-frequency and low-frequency components by EEMD, and the details in the time series are enlarged so that the fluctuation degree of the subsequence is more stable than that of the original series, which greatly reduces the data complexity. Then, each subsequence is predicted by LSTM separately so that the high-frequency components that have a greater impact on water quality changes are focused on. Finally, the prediction results of different components are aggregated to obtain the water quality prediction results. 2. Study Area And Data 2.1 Study area The study area selected for this research is part of the Xiaofu River in Shandong Province, China (Fig. 1 ). The study area is a temperate monsoon climate zone, with the same period of rain and heat, strong seasonal rainfall, and approximately 70% of annual precipitation falls during the flood season (June to September). The Xiaofu River is a first-class tributary on the right bank of the Xiaoqing River (Huan et al. 2018 ). The total length of the river is 136 km. The average gradient of the river is 1.8/1000. The Xiaofu River basin is located at 36°25′N ~ 37°07′N, 117°42′E~118°08′E. The watershed is 40 km wide from east to west and 76 km long from north to south, and the watershed area is 1705 km 2 . The main tributaries are the Fanyang River, Banyang River, Mansi River, Gan River, Zhulong West River and so on (Ding et al. 2022 ). Since the 1980s, the Xiaofu River has been used as a sewage channel for factories, mines, enterprises and residents along the river. In addition, rainfall is relatively low, so the water pollution of the Xiaofu River is relatively serious (Zhang et al. 2000 ). The lack of water resources upstream of the Xiaofu River and the impact of sluice gates and dam impoundments have led to poor water connectivity, poor self-purification ability, and fragile aquatic ecosystems. In recent years, a series of water environment improvement projects have been carried out in the Xiaofu River basin, and the quality of water resources is generally good, but the overall situation of the water environment is still not satisfactory. Predicting the water quality of the Xiaofu River can help to design water environment treatment plans. 2.2 Data sources In this paper, water quality data from the Zhangzhouluqiao Provincial Control Station along the Xiaofu River (36°48′19″N, 117°56′08″E) are taken as the research object. The quality of the water taken from this station is poor, and there is great room for improvement. The main water quality indicators monitored are based on the Environmental Quality Standards for Surface Water (GB3838-2002) and include chemical oxygen demand (COD), ammonia nitrogen (NH 3 N), permanganate index (COD Mn ), pH, dissolved oxygen (DO), electrical conductivity, turbidity and water temperature. The data were collected every 24 hours from April 13, 2019, to April 12, 2021. There are a total of 1096 groups of data, which fully reflect the periodic changes in water quality. According to the water quality of the Xiaofu River, pH, DO, COD Mn and NH 3 N were selected in this paper as the water quality prediction indicators. Statistical analysis was performed on the data series to check for missing data. The statistical analysis results are shown in Table 1 . Table 1 Statistical descriptions of data series Variable Name Description Average Standard Deviation Maximum Value Minimum Value Number of Missing Data pH Pondus hydrogenii 7.912 0.437 8.83 6.02 0 DO Dissolved oxygen (mg/L) 8.779 2.379 18.9 0.5 0 COD Mn Permanganate index (mg/L) 4.327 1.149 9 1.82 1 NH 3 N Ammonia nitrogen (mg/L) 0.472 0.415 5.16 0.028 1 3. Method The prediction accuracy of LSTM for multiperiod hybrid water quality time series is not high. To improve the accuracy of LSTM in predicting water quality, a surface water quality prediction model based on EEMD-LSTM is proposed. The flowchart for the EEMD-LSTM prediction model is shown in Fig. 2 . The EEMD-LSTM consists of the following steps. Step 1: Data preprocessing. The min-max normalization (MMN) method is used to normalize the original water quality series (Singh and Singh 2020 ). MMN can accelerate the speed of the gradient descent method to find the optimal solution and improve the accuracy of the prediction model. Then, the isolation forest algorithm is used to identify abnormal fluctuations, and the input and output samples are determined according to the selected sliding time window width. Step 2: Series decomposition. After the preprocessing of the original water quality time series, EEMD is used to decompose the series into multiple components that contain high-frequency and low-frequency components. The high-frequency components mainly reflect the influence of sudden pollution and discontinuous nonpoint source pollution, and the low-frequency components mainly reflect the physicochemical properties and long-term trend of surface water quality. Step 3: Period calculation. The fast fourier transform (FFT) method can reflect the periodic characteristics of signals that cannot be extracted in the time domain from the frequency domain and is a commonly used signal analysis method (Guia et al. 2015 ). The components that have a great impact on water quality changes are identified according to the significant period. Step 4: Then, independent LSTM submodels are developed for each decomposed component. When training the LSTM submodels, the mean squared error (MSE) of the training dataset is chosen as a criterion to calibrate the model, and the Adam algorithm is chosen as the optimizer. Finally, the prediction results of each submodel are aggregated to obtain the final water quality prediction results. 3.1 Ensemble empirical mode decomposition (EEMD) Huang et al. ( 1998 ) proposed a new analysis and preprocessing method for nonlinear signals, which is referred to as empirical mode decomposition. This method is suitable for dealing with nonlinear and nonstationary time series. The EMD must obey the following two rules at the same time: (1) All the extrema and zero crossing numbers must be the same or different at most by one. (2) All upper and lower envelopes must be locally symmetrical along the time axis. To solve the problem of mode mixing (i.e., decomposed IMFs that contain multiple frequencies), Zhaohua and Norden ( 2009 ) proposed an ensemble empirical mode decomposition method. EEMD utilizes the sensitivity of the signal-to-noise, first adding Gaussian white noise to the original signal to match the signals of different frequencies to the corresponding time scale and then implementing the EMD process. Given an original signal \(x\left(t\right)\) , the specific process of EEMD is as follows: (1) Add Gaussian white noise to the original signal, $${x}^{i}\left(t\right)=x\left(t\right)-{n}^{i}\left(t\right)$$ 1 where \(i\) represents the number of times Gaussian white noise is added. (2) Decompose the mixed signal \({x}^{i}\left(t\right)\) by EMD into IMFs \({C}_{j}^{i}\left(t\right)\) , ( \(j\) = 1, 2, …, n) and residual \({r}^{i}\left(t\right)\) . $${x}^{i}\left(t\right)=\sum _{j=1}^{n}{C}_{j}^{i}\left(t\right)+{r}^{i}\left(t\right)$$ 2 where \({C}_{j}^{i}\left(t\right)\) represents the \(j\) th IMF component obtained by decomposing the \(i\) th mixed signal. (3) Repeat the above steps \(N\) times with different Gaussian white noise each time and find the corresponding IMFs. (4) Average the summation of corresponding decomposed IMFs \(N\) times to eliminate the influence of the added white noise on the original signal. $$\stackrel{-}{{C}_{j}\left(t\right)}=\frac{1}{N}\sum _{j=1}^{n}{C}_{j}^{i}\left(t\right)$$ 3 where \({C}_{j}^{i}\left(t\right)\) represents the \(j\) th IMF component. Finally, after being decomposed by EEMD, the original signal \(x\left(t\right)\) can be expressed as: $$x\left(t\right)=\sum _{j=1}^{N}\stackrel{-}{{c}_{j}\left(t\right)}+r\left(t\right), i=\text{1,2}, \dots ,N$$ 4 3.2 Long short-term memory (LSTM) A long short-term memory network is an improved network structure proposed on the basis of RNNs that effectively overcomes the long-term dependence problem and gradient vanishing problem of RNNs (Hochreiter and Schmidhuber 1997 ). LSTM is suitable for processing and predicting events with long time intervals and delays in time series (An et al. 2020 ). LSTM introduces gates, which can selectively remove or add information. The LSTM cell mainly includes four gate structures: forget gate, input gate, update gate and output gate (Hochreiter and Schmidhuber 1997 ). The function of the forget gate is to forget the irrelevant state information of the previous moment. The input gate determines what information can enter the memory cell at the current moment. The output gate determines the output of the complex network. The memory unit of LSTM can use these three gate structures to screen long-term and short-term memory information. The general architecture of the LSTM cell is shown in Fig. 3 . The key to LSTM is the transmission of the cell state, which controls the information passed into the network through the combination of three gates and determines the cell state. In Fig. 3 , \({X}_{t}\) represents the input of the network at time \(t\) , \({h}_{t}\) represents the output of the network at time \(t\) , and \({C}_{t}\) represents the cell state at time \(t\) . $${f}_{t}=\sigma ({W}_{f}*\left[{h}_{t-1}, {X}_{t}\right]+{b}_{f})$$ 5 The operation ‘*’ represents the elementwise multiplication of the vectors. \({i}_{t}=\sigma ({W}_{i}*\left[{h}_{t-1}, {X}_{t}\right]+{b}_{i})\) (6) \({o}_{t}=\sigma ({W}_{o}*\left[{h}_{t-1}, {X}_{t}\right]+{b}_{o})\) (7) \(\tilde{{C}_{t}}=\text{t}\text{a}\text{n}\text{h}({W}_{c}*\left[{h}_{t-1}, {X}_{t}\right]+{b}_{c})\) (8) \({C}_{t}={f}_{t}*{C}_{t-1}+{i}_{t}*\tilde{{C}_{t}}\) (9) where \(\sigma\) is the logistic sigmoid function ( \(\sigma \left(x\right)=\frac{1}{1+{e}^{-x}}\) ), \({W}_{f}\) , \({W}_{i}\) , \({W}_{o}\) and \({W}_{c}\) represent the weight matrices of the forget gate, the input gate, the output gate and the tanh layer, respectively, \({b}_{f}\) , \({b}_{i}\) , \({b}_{o}\) and \({b}_{c}\) represent the bias vectors of the forget gate, the input gate, the output gate and the tanh layer ( \(\text{tanh}\left(x\right)=\frac{1-{e}^{-2x}}{1+{e}^{-x}}\) ), respectively, \({f}_{t}\) , \({i}_{t}\) and \({o}_{t}\) represent the output of the forget gate, the input gate and the output gate at time \(t\) , respectively, and \(\tilde{{C}_{t}}\) is an update vector for the cell state. Finally, the output \({h}_{t}\) of the memory cell is obtained through the hyperbolic tangent activation function tanh. $${h}_{t}={O}_{t}*\text{t}\text{a}\text{n}\text{h}\left({C}_{t}\right)$$ 10 LSTM is suitable for processing and predicting time series data due to its good ability to deal with the long-term dependence problem on time series data and the problem of gradient disappearance. 3.3 Data preprocessing 3.3.1 Data normalization Data normalization is an important data preprocessing step that can accelerate the speed of the gradient descent method to find the optimal solution and improve the accuracy of the forecasting model. A large amount of unscaled data will slow the learning speed of the artificial neural network and the convergence speed of the model. Since LSTM is very sensitive to fluctuations in time series data and to capturing the trends in time series data, the data need to be normalized before being fed to the neural network (ArunKumar et al. 2021 ). Original data are normalized using the min-max normalization (MMN) method, which linearly scales unnormalized data to predefined lower and upper bounds (Singh and Singh 2020 ). The equation is given as follows: $${x}_{n}=\frac{x-{x}_{min}}{{x}_{max}-{x}_{min}}$$ 11 where \(x\) represents the original time series data, \({x}_{n}\) represents the normalized time series data, \({x}_{min}\) represents the minimum value of the time series data, and \({x}_{max}\) represents the maximum value of the time series data. The min-max normalization method scales the data between 0 and 1. 3.3.2 Outlier detection The real-time monitoring data of water quality are usually unprocessed raw data. Weather factors such as strong wind and heavy rainfall may affect the results of real-time water quality monitoring, and problems such as abnormal monitoring equipment or manual input errors will lead to missing values, abnormal values or noise in the original data. Abnormal values will affect the accuracy of the model prediction. Certain methods are used to identify these outliers and deal with them. The characteristics of abnormal data are as follows: (1) they represent a small proportion of the sample data; and (2) they have significantly different properties compared with normal sample data. Liu et al. ( 2008 ) proposed the isolation forest algorithm and applied it to data outlier detection. The isolation forest algorithm has a linear time complexity and high accuracy and is a neural network algorithm that meets the requirements of big data processing. Any outlier detection method requires an anomaly score, and the calculation equation of the search path length of the isolation forest is as follows: $$c\left(n\right)=2H\left(n-1\right)-\left(\frac{2\left(n-1\right)}{n}\right)$$ 12 where \(n\) is the number of samples, \(H\left(i\right)\) is the harmonic number and can be estimated by \(ln\left(i\right)+ \xi\) (Euler’s constant), and \(c\left(n\right)\) is the average path length of the binary search tree. By normalizing the length of the isolated binary tree, a number between 0 and 1 can be obtained as the abnormal score of the detected sample. The anomaly score \(s\) of an instance \(x\) is defined as: $$s(x, n)={2}^{\frac{E\left(h\left(x\right)\right)}{c\left(n\right)}}$$ 13 where \(h\left(x\right)\) represents the path length from the root node to the x node, and \(E\left(h\right(x\left)\right)\) is the average of the path lengths of all the isolated trees in the isolated forest for the sample point \(x\) . When the anomaly score is larger, the sample point is more likely to be an outlier. Based on the anomaly score s , we can make the following assessments (Liu et al. 2008 ): (1) If the anomaly score is very close to 1, then the data are definitely anomalies. (2) If the anomaly score is much smaller than 0.5, then it is safe to regard the data as normal instances. (3) If all the anomaly scores are approximately 0.5, then there are no distinct outliers in the sample. 3.4 Performance evaluation To objectively and comprehensively evaluate the prediction performance of each model, four different evaluation indicators are selected: root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and determination coefficient (R 2 ). RMSE is sensitive to errors that are evident in the experimental data. MAE is the average value of absolute error and can truly reflect the state of the model's error in prediction. MAPE is the expected value of the absolute error and percentage of the true value. The smaller the RMSE, MAE, and MAEP are, the more accurate the prediction result and the better the model effect. The value of the determination coefficient R 2 is between 0 and 1, and the closer to 1 the value is, the better the model’s prediction ability of the regression effect. Generally, if the coefficient of determination exceeds 0.8, the model is considered to have a high goodness of fit. The specific calculation equation of each loss function is as follows: \(RMSE=\sqrt{\frac{1}{N}\sum _{i=1}^{N}{({y}_{i}-{y}_{i}^{*})}^{2}}\) (14) \(MAE=\frac{1}{N}\sum _{i=1}^{N}{|y}_{i}-{y}_{i}^{*}|\) (15) \(MAPE=\frac{1}{N}\sum _{i=1}^{N}\left|\frac{{y}_{i}-{y}_{i}^{*}}{{y}_{i}}\right|\times 100\%\) (16) \({R}^{2}=1-\frac{\sum _{i=1}^{N}{({y}_{i}-{y}_{i}^{*})}^{2}}{\sum _{i=1}^{N}{({y}_{i}-\stackrel{-}{{y}_{i}})}^{2}}\) (17) where \(N\) is the number of samples, \({y}_{i}\) is the measured value, \({y}_{i}^{*}\) is the predicted value, and \(\stackrel{-}{{y}_{i}}\) is the average value of the measured data. 4. Results 4.1 Data preprocessing Since there were few missing data (< 10%) in the water quality time series, the mean smoothing method was used to fill in the missing part of the data; the missing data were replaced by the average value of the two adjacent data on the left and right of the missing data. The min-max normalization method was used to convert the original values into values between [0, 1]. The normalization results are shown in Fig. 4 . This figure shows that the water quality parameters have apparent fluctuations. The isolated forest algorithm described above was used to identify abnormal fluctuations, such as some data jumps in the original series of water quality parameters (the maximum abnormal sample ratio was set to 0.025), and the outliers were marked. The outlier identification results are shown in Fig. 5 . Considering the small number of outliers and the large difference between an outlier and its adjacent values, outliers were directly removed from the original series, and the average value of the data on both sides of an outlier was used to fill missing values. This figure shows that compared with the original series, obvious outliers in the denoised water quality time series were removed. However, this time series is still complex and has obvious nonstationary and nonlinear characteristics from the overall trend. Measures are still needed to reduce the complexity of the water quality time series. 4.2 EEMD decomposition results After data preprocessing, the water quality time series was decomposed by the EEMD method. The ensemble number was set to 100, and the standard deviation of Gaussian white noise \({n}^{i}\left(t\right)\) was 0.05 (Ren et al. 2015 ; Liu et al. 2022 ). The EEMD results of each water quality parameter are shown in Fig. 6 . The NH 3 N, pH and DO time series were decomposed into eight IMFs and one residual item Res and arranged in the order of frequency from high to low. The COD Mn time series was decomposed into seven IMFs and one residual item Res. The first four IMF components fluctuate greatly, among which IMF1 has the strongest nonlinearity, the largest amplitude and the highest frequency. The residual item can represent the long-term trend of the time series (Liu et al. 2022 ; Ren et al. 2015 ). As illustrated in Fig. 6 , the residual items of the NH 3 N and COD Mn time series have obvious declining trends, indicating that the water environment control measures of the Xiaofu River have achieved certain results in recent years. Table 2 The period of IMF components for water quality parameters Variable Name Period (day) IMF1 IMF2 IMF3 IMF4 IMF5 IMF6 IMF7 IMF8 NH 3 N 3 7 38 41 152 356 534 534 pH 3 7 22 53 89 356 534 534 DO 5 9 20 42 97 356 356 534 COD Mn 4 8 12 59 66 356 356 - Table 2 shows that the period of the first two IMF components is 3–9 days, which corresponds to the number of days that the pollutants are naturally degraded in the water body. Therefore, IMF1 and IMF2 may represent the fluctuation of water quality due to water affected by sudden pollution, discontinuous nonpoint source pollution and so on. IMF3-IMF5 mainly reflect the seasonal changes in water quality, and IMF6-IMF8 mainly reflect the interannual changes in water quality. The seasonal and interannual changes in the water quality series are relatively stable, but the fluctuations caused by sudden pollution and discontinuous nonpoint source pollution are large and complex. Therefore, to obtain more accurate water quality prediction results, it is necessary to accurately simulate the high-frequency components. The EEMD method is able to separate the high-frequency components, enhance the details and transform nonlinear water quality series into several relatively simple and stationary time series that help to improve prediction results. 4.3 Model training and parameter optimization In this paper, we chose the MSE of the training dataset as a criterion to calibrate the model and chose the Adam algorithm as the optimizer. The Adam algorithm can solve the problems of a disappearing learning rate and slow convergence property of the error term. It can optimize the performance of the model and has lower running costs with high computational efficiency and less running memory (Diederik and Jimmy 2014 ). The Adam algorithm was adopted to train the model multiple times and update the parameters continuously. When the error between the actual value and the predicted value meets the accuracy requirements, the model was saved. The hyperparameters of the LSTM model were finally determined. The number of neurons was 50, the number of epochs for each training was 100, and the batch size was 16. In general, the larger the batch size is, the faster the training. However, if the batch size is too large, the network easily converges to the local optimum (Xiang et al. 2020 ). Different sliding time window widths n impact the output of the model. In this paper, the water quality time series of the corresponding time width was divided from the dataset as the input sample, and one time step water quality value after the sliding window was used as the output sample. Taking n = 4 as an example, its dynamic modeling process is shown in Fig. 7 . To improve the prediction accuracy of the model, the model performance is compared under different sliding time window widths, and the results are shown in Table 3 . This result illustrates that the optimal sliding time window widths for NH 3 N, pH, DO and COD Mn are 5, 5, 8 and 7, respectively. Table 3 The LSTM model performance under different sliding time window widths Water quality indicator Sliding time window width RMSE (mg/L) MAE (mg/L) MAPE (%) R 2 NH 3 N 4 0.096 0.071 67.387 0.423 5 0.089 0.057 31.901 0.783 6 0.089 0.060 42.082 0.727 7 0.089 0.059 39.477 0.746 8 0.093 0.067 60.726 0.545 pH 4 0.080 0.049 1.787 0.656 5 0.078 0.045 1.425 0.741 6 0.078 0.046 1.521 0.722 7 0.078 0.046 1.558 0.721 8 0.087 0.059 1.908 0.656 DO 4 0.590 0.420 7.831 0.769 5 0.587 0.424 7.600 0.772 6 0.594 0.434 7.741 0.763 7 0.591 0.429 7.630 0.769 8 0.588 0.422 7.628 0.777 COD Mn 4 0.246 0.167 11.041 0.748 5 0.247 0.168 12.646 0.724 6 0.244 0.165 11.538 0.743 7 0.249 0.170 10.615 0.752 8 0.243 0.165 11.701 0.744 4.4 Water quality prediction by EEMD-LSTM The water quality data were divided into a training period and validation period; the first 85% of the data were from the training period, and the last 15% of the data were from the validation period. After the model is trained, the learning situation of the model can be judged by the loss curve. If the loss curve declines smoothly or continues to decline at the end of the training period, it indicates that there is an underfitting phenomenon. If the loss curve continues to decline but begins to rise at a certain point or there is an upward trend in the fluctuation, it means that there is an overfitting phenomenon. When the loss values of the model in the training period and the validation period decrease and become stable at the same time, the model training effect is good and can be used for water quality prediction. To fully verify the performance of EEMD-LSTM, single LSTM and EEMD-LSTM were used to predict water quality parameters using the same data as input. The prediction results of LSTM and EEMD-LSTM are shown in Fig. 8 , and their performance metrics results are listed in Table 4 . It can be seen from Fig. 8 that although LSTM can predict the trend of water quality changes, the error between observed and predicted values is large, and the prediction accuracy of details and jump points is insufficient. EEMD-LSTM can more accurately predict the detailed changes and greatly improve the model performance in terms of the hysteresis problem. It is also evident in Table 4 that the EEMD-LSTM model outperforms LSTM in water quality time series prediction. Compared with LSTM, the prediction accuracy of EEMD-LSTM on the four evaluation indicators of RMSE, MAE, MAPE and R 2 has been improved. The RMSE, MAE, and MAPE of NH 3 N are decreased by 80.0%, 82.6%, and 93.7%, respectively, and R 2 is increased by 63.0%. The RMSE, MAE, and MAPE of pH are decreased by 71.3%, 74.3%, and 82.4%, respectively, and R 2 is increased by 46.9%. The RMSE, MAE, and MAPE of DO are decreased by 78.2%, 80.4%, and 78.8%, respectively, and R 2 is increased by 17.6%. The RMSE, MAE, and MAPE of COD Mn are decreased by 69.8%, 73.9%, and 84.1%, respectively, and R 2 is increased by 35.1%. These indicators illustrate that the EEMD method can better extract essential features of the water quality time series and reduce the interference of random factors. They also indicate that the prediction performance of the model is greatly improved with the EEMD method. Figure 8 also shows that compared with the single LSTM model, the predicted values of EEMD-LSTM are closer to the observed values in the extreme value prediction. A scatter plot of the observed and predicted values of the two models during the validation period is shown in Fig. 9 The scatter plot intuitively shows that the EEMD-LSTM prediction results are closer to the observed value and have better performance. Table 4 Model performance comparison of LSTM and EEMD-LSTM Model Water quality indicator Training Validation RMSE (mg/L) MAE (mg/L) MAPE (%) R 2 RMSE (mg/L) MAE (mg/L) MAPE (%) R 2 LSTM NH 3 N 0.169 0.111 37.694 0.754 0.110 0.109 50.381 0.567 pH 0.136 0.080 1.032 0.872 0.122 0.113 1.554 0.657 DO 1.151 0.826 10.807 0.733 1.027 0.820 4.685 0.817 COD Mn 0.457 0.314 7.239 0.811 0.440 0.326 13.990 0.693 EEMD-LSTM NH 3 N 0.077 0.050 5.419 0.950 0.022 0.019 3.150 0.924 pH 0.047 0.032 0.321 0.988 0.035 0.029 0.273 0.965 DO 0.531 0.355 2.245 0.945 0.224 0.161 0.994 0.961 COD Mn 0.189 0.131 2.756 0.969 0.133 0.085 2.219 0.936 In addition, the reason why EEMD-LSTM improves water quality prediction performance is further discussed. There are seasonal changes, interannual changes and short-term fluctuations in surface water quality parameters. The subsequences obtained by decomposing the original water quality sequence can more clearly show the seasonal periodic changes, interannual periodic changes and short-term fluctuations and reduce the complexity of the input data, which is beneficial to the learning and training of the model. At the same time, the high-frequency components IMF1 and IMF2 decomposed by the EEMD method can reflect the fluctuations in the water quality series caused by sudden pollution, and the prediction of these components separately can effectively improve the prediction accuracy. 5. Discussion LSTM has achieved high accuracy prediction results in applications of many fields. However, the prediction accuracy of water quality is not satisfactory, as water quality series are generally multiperiod hybrid time series that have strongly nonlinear and nonstationary characteristics, and LSTM is not suitable for predicting multiperiod hybrid time series. In this paper, we introduced the EEMD method to decompose the water quality time series into several simpler single-period components. The EEMD method can decompose the original water quality series into some components arranged from high frequency to low frequency. Among the IMFs decomposed by EEMD, IMF1 and IMF2 reflect the changing process of sudden pollutants discharged into surface water, and these components have great impacts on the accuracy of water quality prediction. Predicting these high-frequency components separately can improve the accuracy in predicting extreme values and the overall performance of the model. Therefore, the predicted values of EEMD-LSTM are closer to the observed values in the extreme value prediction, and the whole prediction accuracy of the EEMD-LSTM model is also improved compared with the single LSTM model. The EEMD-LSTM model has achieved good results in the time series prediction of water quality. The MAE, MAPE and RMSE of EEMD-LSTM for DO are 0.161, 0.994 and 0.224, respectively. The performance predictors of other water quality parameters have also achieved high accuracy. Li et al. ( 2017 ) proposed a multimodal water quality prediction model called MSVR and proved that the combination of EEMD and SVR could achieve better prediction performance. The MAE, MAPE and RMSE of MSVR for DO were 0.175, 2.153 and 0.228, respectively (Li et al. 2017 ). This shows that EEMD-LSTM is reliable in predicting water quality. Limited by time and effort, only the performance of the hybrid model EEMD-LSTM was studied in this paper for water quality prediction. Subsequently, other methods to improve the performance of LSTM will be considered. The influence of different sliding time window widths on the prediction accuracy is also discussed in this paper. The optimal sliding time window width of different water quality parameters is different, which is related to the migration, transformation and degradation rates of pollutants in water. The degradation coefficients of COD Mn and NH 3 N in rivers are 0.08–0.15 and 0.2–0.44 day − 1 , respectively (Ma et al. 2014 ). Therefore, the residence times of COD Mn and NH 3 N in water are 6.7–12.5 day and 2.3-5 days. The optimal sliding time window widths for NH3N, pH, DO and COD Mn are 5, 5, 8 and 7, respectively. This indicates that the optimal sliding time window width is consistent with the degradation time of pollutants in water. This is because after pollutants are discharged into the water, the concentration of pollutants at any point in the water increases with time and then tends to the equilibrium value. As the number of predicted time steps increases, the prediction accuracy of the model will decline, so the EEMD-LSTM model can only predict short time steps at present. Water quality prediction over long time steps is still a challenging issue. 6. Conclusions To achieve highly accurate water quality prediction results, a water quality prediction model based on the combination of the EEMD method and LSTM network is proposed in this paper. The water quality monitoring data of the Xiaofu River are used as a sample for verification, and the four water quality parameters (NH 3 N, pH, DO, COD Mn ) of the Xiaofu River are predicted. The following conclusions were drawn from this study: (1) The EEMD method can decompose time series into components arranged from high frequency to low frequency. In this paper, it is used to decompose the water quality time series to obtain several single-period components, which can effectively reduce the complexity and nonlinearity of the original time series. Among all components, the high-frequency components have the greatest impact on the accuracy of water quality prediction. Predicting the high-frequency components and the low-frequency components separately when using LSTM can significantly improve model accuracy. (2) Compared with LSTM, EEMD-LSTM significantly improves the accuracy of water quality prediction and greatly improves the model performance in terms of the hysteresis problem. During the validation period, the RMSE, MAE, MAPE and R 2 of EEMD-LSTM for NH 3 N are 0.022 mg/L, 0.019 mg/L, 3.150% and 0.924, respectively. The RMSE, MAE, MAPE and R 2 of EEMD-LSTM for pH are 0.035 mg/L, 0.029 mg/L, 0.273% and 0.965, respectively. The RMSE, MAE, MAPE and R 2 of EEMD-LSTM for DO are 0.224 mg/L, 0.161 mg/L, 0.994% and 0.961, respectively. The RMSE, MAE, MAPE and R 2 of the EEMD-LSTM for COD Mn are 0.133 mg/L, 0.085 mg/L, 2.219% and 0.936, respectively. This shows that EEMD-LSTM has high prediction accuracy and strong generalization ability. In addition, the predicted values of EEMD-LSTM are closer to the observed values in the extreme value prediction. In summary, EEMD-LSTM can be an effective tool for water quality prediction. The EEMD-LSTM model can quickly and accurately predict water quality changes, which can reflect the trend of future water quality changes and can provide a basis for formulating water environment governance measures. Declarations Funding This paper was supported by the Major Program of National Natural Science Foundation of China (No. 41790431). Competing interests The authors declare no competing interests. Author Contributions Conceptualization: Lan Luo, Yanjun Zhang and Wenxun Dong; Methodology: Lan Luo, Yanjun Zhang and Anni Qiu; Formal analysis and investigation: Lan Luo, Jinglin Zhang and Liping Zhang; Writing - original draft preparation: Lan Luo; Writing - review and editing: Lan Luo and Yanjun Zhang; All authors read and approved the final manuscript. Ethical approval Not applicable. Consent to participate Not applicable. Consent to publish All authors give consent to publication. Availability of data and materials The datasets used during the current study are available from the corresponding author on request. References An L, Hao Y, Yeh TJ, Liu Y, Liu W, Zhang B (2020) Simulation of karst spring discharge using a combination of time-frequency analysis methods and long short-term memory neural networks. J Hydrol 589:125320. https://doi.org/10.1016/j.jhydrol.2020.125320 ArunKumar KE, Kalaga DV, Kumar CMS, Kawaji M, Brenza TM (2021) Forecasting of COVID-19 using deep layer Recurrent Neural Networks (RNNs) with Gated Recurrent Units (GRUs) and Long Short-Term Memory (LSTM) cells. Chaos Solitons Fractals 146:110861. https://doi.org/10.1016/j.chaos.2021.110861 Bui HH, Ha NH, Nguyen TND, Nguyen AT, Pham TTH, Kandasamy J, Tien VN (2019) Integration of SWAT and QUAL2K for water quality modeling in a data scarce basin of Cau River basin in Vietnam. Ecohydrol Hydrobiol 19(2SI):210–223. https://doi.org/10.1016/j.ecohyd.2019.03.005 Da Costa SB, Leite CM, Almeida IR, de Almeida AK, I.K (2021) Choosing an appropriate water quality model-a review. Environ Monit Assess 193. https://doi.org/10.1007/s10661-020-08786-1 Deus R, Brito D, Mateus M, Kenov I, Fornaro A, Neves R, Alves CN (2013) Impact evaluation of a pisciculture in the Tucurui reservoir (Para, Brazil) using a two-dimensional water quality model. J Hydrol 487:1–12. https://doi.org/10.1016/j.jhydrol.2013.01.022 Diederik PK, Jimmy B (2014) Adam: A method for stochastic optimization. arXiv 1412:6980. https://doi.org/10.48550/arXiv.1412.6980 Ding S, Wang F, Sun X, Ding J, Lu J (2022) Water environmental functional zoning at county level and environmental contamination carrying capacity accounting in the mainstream of Xiaofu River. Water-Sui 14(4):615. https://doi.org/10.3390/w14040615 Ejigu MT (2021) Overview of water quality modeling. Cogent Eng 8(1). https://doi.org/10.1080/23311916.2021.1891711 Eze E, Halse S, Ajmal T (2021) Developing a novel water quality prediction model for a South African aquaculture farm. Water-Sui 13(13):1782. https://doi.org/10.3390/w13131782 Guia SS, Espirito-Santo A, Paciello V, Abate F, Pietrosanto A (2015) A comparison between FFT and MCT for period measurement with an ARM Microcontroller. 2015 IEEE International Instrumentation and Measurement Technology Conference, pp. 1938–1942 Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735–1780. https://doi.org/10.1007/978-3-642-24797-2_4 Huan J, Cao W, Qin Y (2018) Prediction of dissolved oxygen in aquaculture based on EEMD and LSSVM optimized by the Bayesian evidence framework. Comput Electron Agric 150:257–265. https://doi.org/10.1016/j.compag.2018.04.022 Huang NE, Shen Z, Long SR, Wu M, Shih HH, Zheng QN, Yen NC, Tung CC, Liu HH (1998) The empirical mode decomposition and the Hilbert spectrum for nonlinear and non-stationary time series analysis. P Roy Soc A-Math Phy 454(1971):903–995. https://doi.org/10.1098/rspa.1998.0193 Kim J, Lee T, Seo D (2017) Algal bloom prediction of the lower Han River, Korea using the EFDC hydrodynamic and water quality model. Ecol Model 366:27–36. https://doi.org/10.1016/j.ecolmodel.2017.10.015 Kourgialas NN, Dokou Z, Karatzas GP (2015) Statistical analysis and ANN modeling for predicting hydrological extremes under climate change scenarios: The example of a small mediterranean agro-watershed. J Environ Manage 154:86–101. https://doi.org/10.1016/j.jenvman.2015.02.034 Li XJ, Cheng ZW, Yu QB, Bai Y, Li C (2017) Water-quality prediction using multimodal support vector regression: case study of Jialing River, China.J Environ Eng143(10) Liang Z, Zou R, Chen X, Ren T, Su H, Liu Y (2020) Simulate the forecast capacity of a complicated water quality model using the long short-term memory approach. J Hydrol 581:124432. https://doi.org/10.1016/j.jhydrol.2019.124432 Liu FT, Ting KM, Zhou Z (2008) Isolation forest. 2008 Eighth Ieee International Conference on Data Mining, pp. 413–422 Liu X, Zhang Y, Zhang Q (2022) Comparison of EEMD-ARIMA, EEMD-BP and EEMD-SVM algorithms for predicting the hourly urban water consumption. J Hydroinform 24(3):535–558. https://doi.org/10.2166/hydro.2022.146 Liu Y, Zhang Q, Song L, Chen Y (2019) Attention-based recurrent neural networks for accurate short-term and long-term dissolved oxygen prediction. Comput Electron Agric 165. https://doi.org/10.1016/j.compag.2019.104964 Ma L, Liu L, Song LL, Yan WM (2014) A study on water pollutant degradation capability affected by water diversion. J Environ Prot Ecol 15(1):39–47 Mendes J, Ruela R, Picado A, Pinheiro JP, Ribeiro AS, Pereira H, Dias JM (2021) Modeling dynamic processes of Mondego Estuary and Oacute, Bidos Lagoon using Delft3D.J Mar Sci Technol9(1) Niu W, Feng Z, Zeng M, Feng B, Min Y, Cheng C, Zhou J (2019) Forecasting reservoir monthly runoff via ensemble empirical mode decomposition and extreme learning machine optimized by an improved gravitational search algorithm. Appl Soft Comput 82. https://doi.org/10.1016/j.asoc.2019.105589 Palani S, Liong S, Tkalich P (2008) An ANN application for water quality forecasting. Mar Pollut Bull 56(9):1586–1597. https://doi.org/10.1016/j.marpolbul.2008.05.021 Qingmei M, Min L, Aiju L (2013) Spatial variation and contamination assessment of heavy metals in surface sediments of Xiaofu River. Health Environ Res 6:785–790 Ren Y, Suganthan PN, Srikanth N (2015) A comparative study of empirical mode decomposition-based short-term wind speed forecasting methods. IEEE T Sustain Energ 6(1):236–244 Rui Y, Shen D, Khalid S, Yang Z, Wang J (2015) GIS-based emergency response system for sudden water pollution accidents. Phys Chem Earth 79–82:115–121. https://doi.org/10.1016/j.pce.2015.03.001 Rumelhart DE, Hinton GE, Williams RJ (1986) Learning representations by back-propagating errors. Nature 323(6088):533–536. https://doi.org/10.1038/323533a0 Seo IW, Yun SH, Choi SY (2016) Forecasting water quality parameters by ANN model using pre-processing technique at the downstream of Cheongpyeong Dam. Procedia Eng 154:1110–1115. https://doi.org/10.1016/j.proeng.2016.07.519 Shabani A, Zhang X, Chu X, Zheng H (2021) Automatic calibration for CE-QUAL-W2 model using improved global-best harmony search algorithm. Water-Sui 13(16). https://doi.org/10.3390/w13162308 Singh D, Singh B (2020) Investigating the impact of data normalization on classification performance. Appl Soft Comput. https://doi.org/10.1016/j.asoc.2019.105524 . 97(B) Tang W, Pei Y, Zheng H, Zhao Y, Shu L, Zhang H (2022) Twenty years of China's water pollution control: Experiences and challenges. https://doi.org/10.1016/j.chemosphere.2022.133875 . Chemosphere 295 Tant CJ, Rosemond AD, Helton AM, First MR (2015) Nutrient enrichment alters the magnitude and timing of fungal, bacterial, and detritivore contributions to litter breakdown. Freshw Sci 34(4):1259–1271 Wang J, Wang X, Lei XH, Wang H, Zhang XH, You JJ, Tan QF, Liu XL (2020) Teleconnection analysis of monthly streamflow using ensemble empirical mode decomposition. J Hydrol 582. https://doi.org/10.1016/j.jhydrol.2019.124411 Wool T Jr, Ambrose RB, Martin JL, Comer A (2020) WASP 8: the next generation in the 50-year evolution of USEPA's water quality model.Water-Sui12(5) Xiang Z, Yan J, Demir I (2020) A rainfall-runoff model with LSTM-based sequence-to-sequence learning. Water Resour Res 56(1). https://doi.org/10.1029/2019WR025326 Xiong Y, Ran Y, Zhao S, Zhao H, Tian Q (2020) Remotely assessing and monitoring coastal and inland water quality in China: Progress, challenges and outlook. Crit Rev Environ Sci Technol 50(12):1266–1302. https://doi.org/10.1080/10643389.2019.1656511 Yang S, Yang D, Chen J, Santisirisomboon J, Lu W, Zhao B (2020) A physical process and machine learning combined hydrological model for daily streamflow simulations of large watersheds with limited observation data. J Hydrol 590. https://doi.org/10.1016/j.jhydrol.2020.125206 Yu J, Kim J, Li X, Jong Y, Kim K, Ryang G (2022) Water quality forecasting based on data decomposition, fuzzy clustering and deep learning neural network. Environ Pollut 303:119136. https://doi.org/10.1016/j.envpol.2022.119136 Zema DA, Lucas-Borja ME, Fotia L, Rosaci D, Sarne GML, Zimbone SM (2020) Predicting the hydrological response of a forest after wildfire and soil treatments using an Artificial Neural Network. Comput Electron Agric 170. https://doi.org/10.1016/j.compag.2020.105280 Zhang JL, Tang MG, Liu F, Zhong ZS (2000) Vulnerability analysis of groundwater pollution by mining drainage in Zibo coal mine, Shandong Province, China. International Symposium on Hydrogeology and the Environment, pp. 157–162 Zhaohua WU, Norden EH (2009) Ensemble empirical mode decomposition: a noise-assisted data analysis method. Adv Adapt Data Anal 1(1):1–41. https://doi.org/10.1142/S1793536909000047 Zheng L, Wang H, Liu C, Zhang S, Ding A, Xie E, Li J, Wang S (2021) Prediction of harmful algal blooms in large water bodies using the combined EFDC and LSTM models. J Environ Manage 295. https://doi.org/10.1016/j.jenvman.2021.113060 Zhou J, Wang J, Chen Y, Li X, Xie Y (2021) Water quality prediction method based on multi-source transfer learning for water environmental IoT system. Sensors 21(21):7271. https://doi.org/10.3390/s21217271 Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2116084","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":150716193,"identity":"e6c3fac7-c198-43bf-9bad-dca93b87064c","order_by":0,"name":"Lan Luo","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDklEQVRIie2QsUrEMBzGEwItQo6u/1LwGSIFUYTzVRoKjoWjSwbRlEK6OV/h8Bk63dxS8JY8wMEtJwdODp0EJ01ctaWjQ35D+Ajfjy8EIYfjHxIQXH0ODB5NKI+DQIm9hSklrEoZgbhKTFAXaz1DYXongWqR2BCdqRkK2nPJQgVZuOYyxs99xlrSHShaZmMGNs1kpSAPgMvTatvnrPXSG4rSfEwhptmaFVzblXp74E1LLyOKWi5HFA94IRdGacwLo8XGKsHHpEJpVyKqgTe6M4r8WfEmFfALRUBAHFaF+eSXL173Xny9Yemoctv7b3hgD+cB8U/H4f6OP+3K1/27WI4qf0Dsweb3HQ6Hw/Gbb10LYVOc0gBfAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-2008-6417","institution":"Wuhan University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Lan","middleName":"","lastName":"Luo","suffix":""},{"id":150716194,"identity":"810b7fe5-efb1-49a2-930e-f1ce7381ded6","order_by":1,"name":"Yanjun Zhang","email":"","orcid":"https://orcid.org/0000-0001-6889-9321","institution":"Wuhan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yanjun","middleName":"","lastName":"Zhang","suffix":""},{"id":150716195,"identity":"0e5a6b15-e0b0-4cc6-8f18-e3b3d3decffb","order_by":2,"name":"Wenxun Dong","email":"","orcid":"","institution":"Wuhan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wenxun","middleName":"","lastName":"Dong","suffix":""},{"id":150716196,"identity":"6fca2998-8756-465a-b596-48c8b5793c16","order_by":3,"name":"Anni Qiu","email":"","orcid":"","institution":"Wuhan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Anni","middleName":"","lastName":"Qiu","suffix":""},{"id":150716197,"identity":"b5b79274-d174-4046-b99b-3369977a7b33","order_by":4,"name":"Jinglin Zhang","email":"","orcid":"","institution":"Wuhan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jinglin","middleName":"","lastName":"Zhang","suffix":""},{"id":150716198,"identity":"14fec004-8176-437a-8ddd-d672451d2795","order_by":5,"name":"Liping Zhang","email":"","orcid":"","institution":"Wuhan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Liping","middleName":"","lastName":"Zhang","suffix":""}],"badges":[],"createdAt":"2022-09-29 10:18:42","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2116084/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2116084/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":29209680,"identity":"fd2b06d2-8877-4ca5-b4f6-118c91b74594","added_by":"auto","created_at":"2022-11-17 20:07:46","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":767042,"visible":true,"origin":"","legend":"\u003cp\u003eOverview map of the study area\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/abf4d7e0a85925ecc763109f.png"},{"id":29207295,"identity":"26167f0c-3be1-474b-9aa9-18cdea76e5ac","added_by":"auto","created_at":"2022-11-17 19:43:46","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":55532,"visible":true,"origin":"","legend":"\u003cp\u003eA schematic flowchart for EEMD-LSTM\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/444444f744798540d8927632.png"},{"id":29208993,"identity":"04866dee-e9ac-462d-a793-1848784c3d06","added_by":"auto","created_at":"2022-11-17 19:59:46","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":35599,"visible":true,"origin":"","legend":"\u003cp\u003eLSTM memory cell unit structure (Hochreiter and Schmidhuber 1997)\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/4e6f5e4faf245778d9dc38a1.png"},{"id":29207293,"identity":"dbfbfcde-a37c-4e1c-806f-afb7e4a59275","added_by":"auto","created_at":"2022-11-17 19:43:46","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":97201,"visible":true,"origin":"","legend":"\u003cp\u003eThe normalization results of the water quality data: (a) pH, (b) DO, (c) COD\u003csub\u003eMn\u003c/sub\u003e, and (d) NH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e","description":"","filename":"Fig4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/67c48fbfecd090656bebfb98.jpg"},{"id":29207289,"identity":"ec41e77e-1035-4d04-a043-7809f600a783","added_by":"auto","created_at":"2022-11-17 19:43:46","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":114498,"visible":true,"origin":"","legend":"\u003cp\u003eThe outlier detection results of (a) pH, (b) DO, (c) COD\u003csub\u003eMn\u003c/sub\u003e, and (d) NH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e","description":"","filename":"Fig5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/9277a853bb4987f685760750.jpg"},{"id":29208516,"identity":"e15ebd48-a9f4-42ef-8107-f323789d03fa","added_by":"auto","created_at":"2022-11-17 19:51:46","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":231011,"visible":true,"origin":"","legend":"\u003cp\u003eDecomposition results of the (a) NH3N, (b) pH, (c) DO, and (d) COD\u003csub\u003eMn\u003c/sub\u003e time series\u003c/p\u003e","description":"","filename":"Fig6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/da014dae3c7f20e047ee3a05.jpg"},{"id":29208995,"identity":"78866ee4-ba0b-444a-b1fe-2a8b27785564","added_by":"auto","created_at":"2022-11-17 19:59:46","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":78241,"visible":true,"origin":"","legend":"\u003cp\u003eDynamic modeling process\u003c/p\u003e","description":"","filename":"Fig7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/1553d060c4c5114180a9e673.jpg"},{"id":29208520,"identity":"60533321-2375-404f-9aa9-4d358a51a9d2","added_by":"auto","created_at":"2022-11-17 19:51:46","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":620065,"visible":true,"origin":"","legend":"\u003cp\u003eThe LSTM and EEMD-LSTM prediction results of (a) NH\u003csub\u003e3\u003c/sub\u003eN, (b) pH, (c) DO, and (d) COD\u003csub\u003eMn\u003c/sub\u003e\u003c/p\u003e","description":"","filename":"Fig8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/7f422ece3b1851645c4b61a1.jpg"},{"id":29207297,"identity":"cc9633e0-41bb-44e4-aab9-299121351bee","added_by":"auto","created_at":"2022-11-17 19:43:46","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":385406,"visible":true,"origin":"","legend":"\u003cp\u003eScatter plot of the observed and predicted values by LSTM and EEMD-LSTM in the validation period\u003c/p\u003e","description":"","filename":"Fig9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/f7439cc1785dbd1f6a30486f.jpg"},{"id":31345253,"identity":"c9e8a177-2572-4764-97e7-dbddd69d46f1","added_by":"auto","created_at":"2023-01-10 06:16:00","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1773995,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2116084/v1/08a25c51-e542-4280-a5b1-6f0c8468f622.pdf"}],"financialInterests":"","formattedTitle":"Ensemble empirical mode decomposition and a long short-term memory neural network for surface water quality prediction of the Xiaofu River, China","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eWith the rapid development of the economy over the past few decades, many water bodies in China have been seriously polluted, which affects people's quality of life and the safe water quality level (Tang et al. \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Xiong et al. \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Water environment management and protection have gradually become the focus of attention. Water quality prediction is an important link in the management and protection of aquatic environments. Scientific and accurate water quality prediction can help to understand the changing laws and development trends of the water environment, provide technical support for water environmental protection and water pollution prevention and control, and improve the decision-making initiatives of management departments (Liang et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Yu et al. \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eMany researchers have used comprehensive water quality models to simulate and predict water quality (Bui et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Deus et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Kim et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Rui et al. \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). At present, the comprehensive water quality models that have been widely used in water quality simulation and water environment management include the Water Quality Analysis Simulation Program (WASP) (Wool et al. \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), QUAL series model (Shabani et al. \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), the Environmental Fluid Dynamics Code (EFDC) (Kim et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), Delft3D model (Mendes et al. \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2021\u003c/span\u003e), etc. However, there are many water quality parameters in comprehensive water quality models, and a large amount of measured water quality data is needed to set initial conditions and boundary conditions during simulation. Comprehensive water quality models are too complex and difficult to apply, and they are always data-intensive and time-consuming to develop (Da Silva Burigato Costa et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Ejigu \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). In addition, the use of complex models in the absence of data reduces the reliability of water quality prediction. Therefore, although water quality models can simulate the complex dynamics of water quality variables well, water quality prediction remains difficult.\u003c/p\u003e \u003cp\u003eUsing deep learning methods for water quality prediction can solve the problem of difficult application of comprehensive water quality models, as deep learning methods can effectively establish relationships between water quality parameters without complex boundaries and initial conditions (Liang et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). In recent years, artificial intelligence models such as artificial neural networks (ANNs) have been gradually applied to hydrological process analysis (Kourgialas et al. \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Yang et al. \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Zema et al. \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) and water quality prediction (Kim et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Palani et al. \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Seo et al. \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). However, sequential order information is not reflected in the ANN training process, and ANNs do not perform well in nonlinear simulations (An et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). To overcome the shortcomings of ANNs, researchers have proposed recurrent neural networks (RNNs) and long short-term memory (LSTM) networks (Hochreiter and Schmidhuber \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e1997\u003c/span\u003e; Rumelhart et al. \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e1986\u003c/span\u003e). LSTM is an improved network structure proposed on the theoretical basis of RNNs. The network effectively overcomes the long-term dependence and easy gradient disappearance problems in RNNs and has better long-term and short-term memory function. Since LSTM was proposed, some researchers have applied it to the field of water quality modeling. For example, Zheng et al. (\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2021\u003c/span\u003e) used LSTM to effectively predict the concentration of chlorophyll-a and the outbreak of harmful algal blooms in a water body and provide another method for water resource management. Liang et al. (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) found that LSTM could achieve the prediction accuracy of a comprehensive water quality model (such as EFDC). When the types of water quality data available are relatively simple, LSTM can be an effective tool for water quality prediction.\u003c/p\u003e \u003cp\u003eHowever, due to the influence of hydrometeorological factors and human factors, water quality time series are nonlinear and nonstationary, so the prediction accuracy of LSTM for surface water quality is not high (Eze et al. \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Zhou et al. \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; An et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Surface water quality time series are generally multiperiod hybrid time series. According to the different periods, the water quality time series can be divided into high-frequency components (period of 1\u0026ndash;10 days) and low-frequency components (period\u0026thinsp;\u0026gt;\u0026thinsp;10 days). The main factors affecting the high-frequency components are sudden pollution and discontinuous nonpoint source pollution. These factors are closely related to the physical and chemical properties of pollutants, water quality, temperature, hydraulic conditions and other factors (Tant et al. \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). The changing trend of these factors is large, which has a great influence on the accuracy of water quality prediction. The main factors affecting the low-frequency components are climate change, constant point source pollution, endogenous pollution and so on. The changing trend of these factors is relatively stable. Therefore, the main difficulty in water quality prediction is accurately predicting the high-frequency components in water quality time series. However, in existing studies, predicting the high-frequency components and low-frequency components separately when using the LSTM model for water quality prediction has rarely been considered, and the fluctuation term of the water quality series cannot be accurately predicted. Signal decomposition techniques can decompose the original water quality time series into a set of components with specific meanings and provide more detailed information. When predicting water quality, we can focus on high-frequency components to enhance the details and reduce the impact of interference information with signal decomposition techniques. Generally, the residence time of pollutants in water is 5\u0026ndash;10 days. When the period of decomposed components is consistent with the degradation cycle of pollutants, the prediction accuracy of LSTM is likely to be improved. To overcome the limitations of a single LSTM method, LSTM can be combined with signal decomposition techniques to improve the accuracy of water quality prediction.\u003c/p\u003e \u003cp\u003eAmong signal decomposition algorithms, empirical mode decomposition (EMD) is widely used due to its orthogonality and convergence. It is easier to apply than wavelet decomposition. Huang et al. (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e1998\u003c/span\u003e) proposed EMD, which is a data-adaptive time frequency analysis method for nonlinear and nonstationary time series. EMD decomposes the original sequence into multiple intrinsic mode functions (IMFs) and residuals to reduce the complexity of the sequence. However, EMD has limitations such as modal confounding and end effects. Zhaohua and Norden (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) proposed an improved empirical mode decomposition algorithm, EEMD, which addressed the modal confounding problem of EMD. EEMD can effectively reflect the nature of the original signal and has been widely used in many fields in recent years (An et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). For example, Wang et al. (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) used EEMD to extract the oscillation period and the trend of runoff series and analyzed the relationship between runoff and climate phenomenon indicators. Niu et al. (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) used EEMD to decompose the original monthly flow series, combined the improved gravitational search algorithm (IGSA) and extreme learning machine (ELM) for hydrological prediction, and successfully predicted the monthly runoff of the Three Gorges. Huan et al. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) proposed a combined prediction model based on EEMD and a least squares support vector machine (LSSVM), which had high prediction accuracy and strong generalization ability for dissolved oxygen (DO). From previous research, we know that EEMD can decompose the original water quality series into components arranged from high frequency to low frequency, and the time period of the high-frequency components is likely to be consistent with the degradation cycle of pollutants. Therefore, EEMD is suitable for decomposing water quality series into several components for water quality prediction using LSTM separately.\u003c/p\u003e \u003cp\u003eTo acquire better prediction performance of the surface water quality, a hybrid water quality prediction model based on the ensemble empirical mode decomposition method and long short-term memory neural network is proposed in this paper. The original water quality time series is decomposed into high-frequency and low-frequency components by EEMD, and the details in the time series are enlarged so that the fluctuation degree of the subsequence is more stable than that of the original series, which greatly reduces the data complexity. Then, each subsequence is predicted by LSTM separately so that the high-frequency components that have a greater impact on water quality changes are focused on. Finally, the prediction results of different components are aggregated to obtain the water quality prediction results.\u003c/p\u003e"},{"header":"2. Study Area And Data","content":"\u003cp\u003e2.1 Study area\u003c/p\u003e\n\u003cp\u003eThe study area selected for this research is part of the Xiaofu River in Shandong Province, China (Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). The study area is a temperate monsoon climate zone, with the same period of rain and heat, strong seasonal rainfall, and approximately 70% of annual precipitation falls during the flood season (June to September). The Xiaofu River is a first-class tributary on the right bank of the Xiaoqing River (Huan et al. \u003cspan class=\"CitationRef\"\u003e2018\u003c/span\u003e). The total length of the river is 136 km. The average gradient of the river is 1.8/1000. The Xiaofu River basin is located at 36\u0026deg;25\u0026prime;N\u0026thinsp;~\u0026thinsp;37\u0026deg;07\u0026prime;N, 117\u0026deg;42\u0026prime;E~118\u0026deg;08\u0026prime;E. The watershed is 40 km wide from east to west and 76 km long from north to south, and the watershed area is 1705 km\u003csup\u003e2\u003c/sup\u003e. The main tributaries are the Fanyang River, Banyang River, Mansi River, Gan River, Zhulong West River and so on (Ding et al. \u003cspan class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eSince the 1980s, the Xiaofu River has been used as a sewage channel for factories, mines, enterprises and residents along the river. In addition, rainfall is relatively low, so the water pollution of the Xiaofu River is relatively serious (Zhang et al. \u003cspan class=\"CitationRef\"\u003e2000\u003c/span\u003e). The lack of water resources upstream of the Xiaofu River and the impact of sluice gates and dam impoundments have led to poor water connectivity, poor self-purification ability, and fragile aquatic ecosystems. In recent years, a series of water environment improvement projects have been carried out in the Xiaofu River basin, and the quality of water resources is generally good, but the overall situation of the water environment is still not satisfactory. Predicting the water quality of the Xiaofu River can help to design water environment treatment plans.\u003c/p\u003e\n\u003cp\u003e2.2 Data sources\u003c/p\u003e\n\u003cp\u003eIn this paper, water quality data from the Zhangzhouluqiao Provincial Control Station along the Xiaofu River (36\u0026deg;48\u0026prime;19\u0026Prime;N, 117\u0026deg;56\u0026prime;08\u0026Prime;E) are taken as the research object. The quality of the water taken from this station is poor, and there is great room for improvement. The main water quality indicators monitored are based on the Environmental Quality Standards for Surface Water (GB3838-2002) and include chemical oxygen demand (COD), ammonia nitrogen (NH\u003csub\u003e3\u003c/sub\u003eN), permanganate index (COD\u003csub\u003eMn\u003c/sub\u003e), pH, dissolved oxygen (DO), electrical conductivity, turbidity and water temperature. The data were collected every 24 hours from April 13, 2019, to April 12, 2021. There are a total of 1096 groups of data, which fully reflect the periodic changes in water quality. According to the water quality of the Xiaofu River, pH, DO, COD\u003csub\u003eMn\u003c/sub\u003e and NH\u003csub\u003e3\u003c/sub\u003eN were selected in this paper as the water quality prediction indicators. Statistical analysis was performed on the data series to check for missing data. The statistical analysis results are shown in Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" id=\"Tab1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eStatistical descriptions of data series\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eVariable\u003c/p\u003e\n \u003cp\u003eName\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eDescription\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAverage\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eStandard\u003c/p\u003e\n \u003cp\u003eDeviation\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMaximum\u003c/p\u003e\n \u003cp\u003eValue\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMinimum\u003c/p\u003e\n \u003cp\u003eValue\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eNumber of\u003c/p\u003e\n \u003cp\u003eMissing Data\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003epH\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ePondus hydrogenii\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.912\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.437\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8.83\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6.02\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eDO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDissolved oxygen\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8.779\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.379\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e18.9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eCOD\u003csub\u003eMn\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ePermanganate index\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4.327\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.149\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colspan=\"2\"\u003e\n \u003cp\u003eNH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eAmmonia nitrogen\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.472\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.415\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5.16\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.028\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"},{"header":"3. Method","content":"\u003cp\u003eThe prediction accuracy of LSTM for multiperiod hybrid water quality time series is not high. To improve the accuracy of LSTM in predicting water quality, a surface water quality prediction model based on EEMD-LSTM is proposed. The flowchart for the EEMD-LSTM prediction model is shown in Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003eThe EEMD-LSTM consists of the following steps.\u003c/p\u003e\n\u003cp\u003eStep 1: Data preprocessing. The min-max normalization (MMN) method is used to normalize the original water quality series (Singh and Singh \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e). MMN can accelerate the speed of the gradient descent method to find the optimal solution and improve the accuracy of the prediction model. Then, the isolation forest algorithm is used to identify abnormal fluctuations, and the input and output samples are determined according to the selected sliding time window width.\u003c/p\u003e\n\u003cp\u003eStep 2: Series decomposition. After the preprocessing of the original water quality time series, EEMD is used to decompose the series into multiple components that contain high-frequency and low-frequency components. The high-frequency components mainly reflect the influence of sudden pollution and discontinuous nonpoint source pollution, and the low-frequency components mainly reflect the physicochemical properties and long-term trend of surface water quality.\u003c/p\u003e\n\u003cp\u003eStep 3: Period calculation. The fast fourier transform (FFT) method can reflect the periodic characteristics of signals that cannot be extracted in the time domain from the frequency domain and is a commonly used signal analysis method (Guia et al. \u003cspan class=\"CitationRef\"\u003e2015\u003c/span\u003e). The components that have a great impact on water quality changes are identified according to the significant period.\u003c/p\u003e\n\u003cp\u003eStep 4: Then, independent LSTM submodels are developed for each decomposed component. When training the LSTM submodels, the mean squared error (MSE) of the training dataset is chosen as a criterion to calibrate the model, and the Adam algorithm is chosen as the optimizer. Finally, the prediction results of each submodel are aggregated to obtain the final water quality prediction results.\u003c/p\u003e\n\u003cp\u003e3.1 Ensemble empirical mode decomposition (EEMD)\u003c/p\u003e\n\u003cp\u003eHuang et al. (\u003cspan class=\"CitationRef\"\u003e1998\u003c/span\u003e) proposed a new analysis and preprocessing method for nonlinear signals, which is referred to as empirical mode decomposition. This method is suitable for dealing with nonlinear and nonstationary time series. The EMD must obey the following two rules at the same time: (1) All the extrema and zero crossing numbers must be the same or different at most by one. (2) All upper and lower envelopes must be locally symmetrical along the time axis.\u003c/p\u003e\n\u003cp\u003eTo solve the problem of mode mixing (i.e., decomposed IMFs that contain multiple frequencies), Zhaohua and Norden (\u003cspan class=\"CitationRef\"\u003e2009\u003c/span\u003e) proposed an ensemble empirical mode decomposition method. EEMD utilizes the sensitivity of the signal-to-noise, first adding Gaussian white noise to the original signal to match the signals of different frequencies to the corresponding time scale and then implementing the EMD process.\u003c/p\u003e\n\u003cp\u003eGiven an original signal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(x\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e, the specific process of EEMD is as follows:\u003c/p\u003e\n\u003cp\u003e(1) Add Gaussian white noise to the original signal,\u0026nbsp;\u003cimg src=\"data:image/jpeg;base64,/9j/4AAQSkZJRgABAQEAYABgAAD/4RD6RXhpZgAATU0AKgAAAAgABAE7AAIAAAAQAAAISodpAAQAAAABAAAIWpydAAEAAAAgAAAQ0uocAAcAAAgMAAAAPgAAAAAc6gAAAAgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAFNhY2hpbiBNYWhhcm51cgAABZADAAIAAAAUAAAQqJAEAAIAAAAUAAAQvJKRAAIAAAADMzkAAJKSAAIAAAADMzkAAOocAAcAAAgMAAAInAAAAAAc6gAAAAgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAADIwMjI6MTE6MTcgMTU6MzA6MzkAMjAyMjoxMToxNyAxNTozMDozOQAAAFMAYQBjAGgAaQBuACAATQBhAGgAYQByAG4AdQByAAAA/+ELImh0dHA6Ly9ucy5hZG9iZS5jb20veGFwLzEuMC8APD94cGFja2V0IGJlZ2luPSfvu78nIGlkPSdXNU0wTXBDZWhpSHpyZVN6TlRjemtjOWQnPz4NCjx4OnhtcG1ldGEgeG1sbnM6eD0iYWRvYmU6bnM6bWV0YS8iPjxyZGY6UkRGIHhtbG5zOnJkZj0iaHR0cDovL3d3dy53My5vcmcvMTk5OS8wMi8yMi1yZGYtc3ludGF4LW5zIyI+PHJkZjpEZXNjcmlwdGlvbiByZGY6YWJvdXQ9InV1aWQ6ZmFmNWJkZDUtYmEzZC0xMWRhLWFkMzEtZDMzZDc1MTgyZjFiIiB4bWxuczpkYz0iaHR0cDovL3B1cmwub3JnL2RjL2VsZW1lbnRzLzEuMS8iLz48cmRmOkRlc2NyaXB0aW9uIHJkZjphYm91dD0idXVpZDpmYWY1YmRkNS1iYTNkLTExZGEtYWQzMS1kMzNkNzUxODJmMWIiIHhtbG5zOnhtcD0iaHR0cDovL25zLmFkb2JlLmNvbS94YXAvMS4wLyI+PHhtcDpDcmVhdGVEYXRlPjIwMjItMTEtMTdUMTU6MzA6MzkuMzkyPC94bXA6Q3JlYXRlRGF0ZT48L3JkZjpEZXNjcmlwdGlvbj48cmRmOkRlc2NyaXB0aW9uIHJkZjphYm91dD0idXVpZDpmYWY1YmRkNS1iYTNkLTExZGEtYWQzMS1kMzNkNzUxODJmMWIiIHhtbG5zOmRjPSJodHRwOi8vcHVybC5vcmcvZGMvZWxlbWVudHMvMS4xLyI+PGRjOmNyZWF0b3I+PHJkZjpTZXEgeG1sbnM6cmRmPSJodHRwOi8vd3d3LnczLm9yZy8xOTk5LzAyLzIyLXJkZi1zeW50YXgtbnMjIj48cmRmOmxpPlNhY2hpbiBNYWhhcm51cjwvcmRmOmxpPjwvcmRmOlNlcT4NCgkJCTwvZGM6Y3JlYXRvcj48L3JkZjpEZXNjcmlwdGlvbj48L3JkZjpSREY+PC94OnhtcG1ldGE+DQogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgIAogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgIAogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgIAogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgIAogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgIAogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgIAogICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgCiAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKICAgICAgICAgICAgICAgICAgICAgICAgICAgIDw/eHBhY2tldCBlbmQ9J3cnPz7/2wBDAAcFBQYFBAcGBQYIBwcIChELCgkJChUPEAwRGBUaGRgVGBcbHichGx0lHRcYIi4iJSgpKywrGiAvMy8qMicqKyr/2wBDAQcICAoJChQLCxQqHBgcKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKioqKir/wAARCAAYAHUDASIAAhEBAxEB/8QAHwAAAQUBAQEBAQEAAAAAAAAAAAECAwQFBgcICQoL/8QAtRAAAgEDAwIEAwUFBAQAAAF9AQIDAAQRBRIhMUEGE1FhByJxFDKBkaEII0KxwRVS0fAkM2JyggkKFhcYGRolJicoKSo0NTY3ODk6Q0RFRkdISUpTVFVWV1hZWmNkZWZnaGlqc3R1dnd4eXqDhIWGh4iJipKTlJWWl5iZmqKjpKWmp6ipqrKztLW2t7i5usLDxMXGx8jJytLT1NXW19jZ2uHi4+Tl5ufo6erx8vP09fb3+Pn6/8QAHwEAAwEBAQEBAQEBAQAAAAAAAAECAwQFBgcICQoL/8QAtREAAgECBAQDBAcFBAQAAQJ3AAECAxEEBSExBhJBUQdhcRMiMoEIFEKRobHBCSMzUvAVYnLRChYkNOEl8RcYGRomJygpKjU2Nzg5OkNERUZHSElKU1RVVldYWVpjZGVmZ2hpanN0dXZ3eHl6goOEhYaHiImKkpOUlZaXmJmaoqOkpaanqKmqsrO0tba3uLm6wsPExcbHyMnK0tPU1dbX2Nna4uPk5ebn6Onq8vP09fb3+Pn6/9oADAMBAAIRAxEAPwD6Rqrqd8ml6Td38sbypawPMyRldzBVLEDcQM8dyB6kVzPxQa5XwPMun313Z3008NtaNaTGJmmlkWNMlecAuGwDzt5yMg52lQ22oePPFvh3UtRu762l0+3guLK4uJGDmRJDLIoziJWWSNfk2jKnHINAHVnW49O8MW+p61J5bPFGXCwlWeR8AIsYZzuLEAKCxycZNXYr6JvsyXH+i3FypaO2mdRIcDLDAJBxnnBNcr4it4rXxT4GsV8xbNb+cqrOW3SLaSlNzHJJ+8ck5JGetZGrQaLffHKNNVvLu3uLXS4vskUWozxNPJLM2diRuMqohXdgbfmBboDQB3Wn61b6hfXVkElt7y0IMtvOoDhGztcYJDK2DggnkEHBBA0K43WGeL4xeGfsoG6bStQW4weqK9uUz7BicfU1Kmh+JdJkjutI1RL+a4TF/bancSGHzCP9bCcMY8H/AJZgBSMfdPzEA1/EHifSPCtpBda9dG0t7idbdJfJd1Dt90MVB2g+rYHvUh1/T18SroBkm/tFrc3Ij+zSbPLBxu8zbs68Y3Z6Vx3iix03xJrmk+AdU1Fbl10y4urgyyDzZG8v7OjlR1J82V/YpnsMHwsutQ1r+0NT1yNkv7FItEl3DG6S33GVx7M8nX/ZHpQB2mq6zb6SIElWSe5un8u2tYADJM2MkAEgAAAkkkADqRVDVPG/h7RtTNhqOoeVOhQSkQSPHb7zhPNkVSkW7tvIzVBzv+MsKzZIi0F2twRxlrhRKR7/ACxfnXCeIrmC10r4meHL5l/tvWrovpts3+svUktoki8sfxbWVgcfdxzigD2iisvSNStpZJtIW58/UNLihS8GxhtZ0ypyRg5AJ4zivLLHUZPF954dR9Z1W61q8vGm1awsdSmto9LtgkmYZFhZdjK3lgb/AJ2YHnGRQB6Nq1/r1t4k0+GwOnSWdxMqPbPG7TtHjMkocMFQLxwVbdkDIJArRtdZt7nVrjTHjlt7yBfM8qZQPNjzgSIQSGXPvkZGQMiqk/hHSbnWo9VmF8bqPysY1G4EbeWcpujD7GweeQeeetZPiYvH8TPBLQAb3kvYpW7+T9nLEfTesX44oA7Giq9lf2ep2oudOu4LuAsyiWCQSKSpIYZHGQQQfQiigCxRRRQBna5osGuWCwTSSQSwyrPb3ERAeCVTlXXPHsQcggkHIJq7brMlvGt1IkswUB3jQorH1CknH0yaKKAM2w0MW+uXWsXtx9qv50ECPs2JBCDkRouTjJOWJJLHHQAAa1FFABRRRQBlavoY1G6s7+1n+yalYlvs9xs3jawAeN1yNyNgZGQcqCCCK1aKKAK9tYW1pNczW8KpLdSeZO/VpGwFBJPoAAB2A4qxRRQAVkwaGP8AhI31q+n+03KxNb2qhNqW0TMGYAZJLNtXc2edowB0oooA1qKKKAP/2Q==\" width=\"117\" height=\"24\"\u003e\u003c/p\u003e\n\u003cdiv class=\"Equation\" id=\"Equ1\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e$${x}^{i}\\left(t\\right)=x\\left(t\\right)-{n}^{i}\\left(t\\right)$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003e represents the number of times Gaussian white noise is added.\u003c/p\u003e\n\u003cp\u003e(2) Decompose the mixed signal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({x}^{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e by EMD into IMFs \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({C}_{j}^{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e, (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(j\\)\u003c/span\u003e\u003c/span\u003e= 1, 2, \u0026hellip;, n) and residual \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({r}^{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv class=\"Equation\" id=\"Equ2\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e$${x}^{i}\\left(t\\right)=\\sum _{j=1}^{n}{C}_{j}^{i}\\left(t\\right)+{r}^{i}\\left(t\\right)$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({C}_{j}^{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(j\\)\u003c/span\u003e\u003c/span\u003eth IMF component obtained by decomposing the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003eth mixed signal.\u003c/p\u003e\n\u003cp\u003e(3) Repeat the above steps \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N\\)\u003c/span\u003e\u003c/span\u003e times with different Gaussian white noise each time and find the corresponding IMFs.\u003c/p\u003e\n\u003cp\u003e(4) Average the summation of corresponding decomposed IMFs \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N\\)\u003c/span\u003e\u003c/span\u003e times to eliminate the influence of the added white noise on the original signal.\u003c/p\u003e\n\u003cdiv class=\"Equation\" id=\"Equ3\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e$$\\stackrel{-}{{C}_{j}\\left(t\\right)}=\\frac{1}{N}\\sum _{j=1}^{n}{C}_{j}^{i}\\left(t\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({C}_{j}^{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(j\\)\u003c/span\u003e\u003c/span\u003eth IMF component.\u003c/p\u003e\u003cp\u003eFinally, after being decomposed by EEMD, the original signal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(x\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e can be expressed as:\u003c/p\u003e\u003cdiv class=\"Equation\" id=\"Equ4\"\u003e\u003cdiv class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e$$x\\left(t\\right)=\\sum _{j=1}^{N}\\stackrel{-}{{c}_{j}\\left(t\\right)}+r\\left(t\\right), i=\\text{1,2}, \\dots ,N$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003cp\u003e3.2 Long short-term memory (LSTM)\u003c/p\u003e\u003cp\u003eA long short-term memory network is an improved network structure proposed on the basis of RNNs that effectively overcomes the long-term dependence problem and gradient vanishing problem of RNNs (Hochreiter and Schmidhuber \u003cspan class=\"CitationRef\"\u003e1997\u003c/span\u003e). LSTM is suitable for processing and predicting events with long time intervals and delays in time series (An et al. \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e). LSTM introduces gates, which can selectively remove or add information. The LSTM cell mainly includes four gate structures: forget gate, input gate, update gate and output gate (Hochreiter and Schmidhuber \u003cspan class=\"CitationRef\"\u003e1997\u003c/span\u003e). The function of the forget gate is to forget the irrelevant state information of the previous moment. The input gate determines what information can enter the memory cell at the current moment. The output gate determines the output of the complex network. The memory unit of LSTM can use these three gate structures to screen long-term and short-term memory information. The general architecture of the LSTM cell is shown in Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e\u003cp\u003eThe key to LSTM is the transmission of the cell state, which controls the information passed into the network through the combination of three gates and determines the cell state. In Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X}_{t}\\)\u003c/span\u003e\u003c/span\u003e represents the input of the network at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(t\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({h}_{t}\\)\u003c/span\u003e\u003c/span\u003e represents the output of the network at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(t\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({C}_{t}\\)\u003c/span\u003e\u003c/span\u003e represents the cell state at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(t\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\u003cdiv class=\"Equation\" id=\"Equ5\"\u003e\u003cdiv class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e$${f}_{t}=\\sigma ({W}_{f}*\\left[{h}_{t-1}, {X}_{t}\\right]+{b}_{f})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003cp\u003eThe operation \u0026lsquo;*\u0026rsquo; represents the elementwise multiplication of the vectors.\u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable border=\"1\" id=\"Tabb\"\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({i}_{t}=\\sigma ({W}_{i}*\\left[{h}_{t-1}, {X}_{t}\\right]+{b}_{i})\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\"\u003e\u003cp\u003e(6)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({o}_{t}=\\sigma ({W}_{o}*\\left[{h}_{t-1}, {X}_{t}\\right]+{b}_{o})\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\"\u003e\u003cp\u003e(7)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\tilde{{C}_{t}}=\\text{t}\\text{a}\\text{n}\\text{h}({W}_{c}*\\left[{h}_{t-1}, {X}_{t}\\right]+{b}_{c})\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e(8)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({C}_{t}={f}_{t}*{C}_{t-1}+{i}_{t}*\\tilde{{C}_{t}}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e(9)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e is the logistic sigmoid function (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma \\left(x\\right)=\\frac{1}{1+{e}^{-x}}\\)\u003c/span\u003e\u003c/span\u003e), \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{f}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{o}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{c}\\)\u003c/span\u003e\u003c/span\u003e represent the weight matrices of the forget gate, the input gate, the output gate and the tanh layer, respectively, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({b}_{f}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({b}_{i}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({b}_{o}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({b}_{c}\\)\u003c/span\u003e\u003c/span\u003e represent the bias vectors of the forget gate, the input gate, the output gate and the tanh layer (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\text{tanh}\\left(x\\right)=\\frac{1-{e}^{-2x}}{1+{e}^{-x}}\\)\u003c/span\u003e\u003c/span\u003e), respectively, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({f}_{t}\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({i}_{t}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({o}_{t}\\)\u003c/span\u003e\u003c/span\u003e represent the output of the forget gate, the input gate and the output gate at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(t\\)\u003c/span\u003e\u003c/span\u003e, respectively, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\tilde{{C}_{t}}\\)\u003c/span\u003e\u003c/span\u003e is an update vector for the cell state.\u003c/p\u003e\n \u003cp\u003eFinally, the output \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({h}_{t}\\)\u003c/span\u003e\u003c/span\u003e of the memory cell is obtained through the hyperbolic tangent activation function tanh.\u003c/p\u003e\n \u003cdiv class=\"Equation\" id=\"Equ6\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e$${h}_{t}={O}_{t}*\\text{t}\\text{a}\\text{n}\\text{h}\\left({C}_{t}\\right)$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003eLSTM is suitable for processing and predicting time series data due to its good ability to deal with the long-term dependence problem on time series data and the problem of gradient disappearance.\u003c/p\u003e\n \u003cp\u003e3.3 Data preprocessing\u003c/p\u003e\n \u003cdiv class=\"Section2\" id=\"Sec4\"\u003e\n \u003cp\u003e3.3.1 Data normalization\u003c/p\u003e\n \u003cp\u003eData normalization is an important data preprocessing step that can accelerate the speed of the gradient descent method to find the optimal solution and improve the accuracy of the forecasting model. A large amount of unscaled data will slow the learning speed of the artificial neural network and the convergence speed of the model. Since LSTM is very sensitive to fluctuations in time series data and to capturing the trends in time series data, the data need to be normalized before being fed to the neural network (ArunKumar et al. \u003cspan class=\"CitationRef\"\u003e2021\u003c/span\u003e). Original data are normalized using the min-max normalization (MMN) method, which linearly scales unnormalized data to predefined lower and upper bounds (Singh and Singh \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e). The equation is given as follows:\u003c/p\u003e\n \u003cdiv class=\"Equation\" id=\"Equ7\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e$${x}_{n}=\\frac{x-{x}_{min}}{{x}_{max}-{x}_{min}}$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e11\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(x\\)\u003c/span\u003e\u003c/span\u003e represents the original time series data, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({x}_{n}\\)\u003c/span\u003e\u003c/span\u003e represents the normalized time series data, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({x}_{min}\\)\u003c/span\u003e\u003c/span\u003e represents the minimum value of the time series data, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({x}_{max}\\)\u003c/span\u003e\u003c/span\u003e represents the maximum value of the time series data. The min-max normalization method scales the data between 0 and 1.\u003c/p\u003e\n \u003cdiv class=\"Section3\" id=\"Sec5\"\u003e\n \u003cp\u003e3.3.2 Outlier detection\u003c/p\u003e\n \u003cp\u003eThe real-time monitoring data of water quality are usually unprocessed raw data. Weather factors such as strong wind and heavy rainfall may affect the results of real-time water quality monitoring, and problems such as abnormal monitoring equipment or manual input errors will lead to missing values, abnormal values or noise in the original data. Abnormal values will affect the accuracy of the model prediction. Certain methods are used to identify these outliers and deal with them. The characteristics of abnormal data are as follows: (1) they represent a small proportion of the sample data; and (2) they have significantly different properties compared with normal sample data.\u003c/p\u003e\n \u003cp\u003eLiu et al. (\u003cspan class=\"CitationRef\"\u003e2008\u003c/span\u003e) proposed the isolation forest algorithm and applied it to data outlier detection. The isolation forest algorithm has a linear time complexity and high accuracy and is a neural network algorithm that meets the requirements of big data processing. Any outlier detection method requires an anomaly score, and the calculation equation of the search path length of the isolation forest is as follows:\u003c/p\u003e\n \u003cdiv class=\"Equation\" id=\"Equ8\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e$$c\\left(n\\right)=2H\\left(n-1\\right)-\\left(\\frac{2\\left(n-1\\right)}{n}\\right)$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e12\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e is the number of samples, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(H\\left(i\\right)\\)\u003c/span\u003e\u003c/span\u003e is the harmonic number and can be estimated by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(ln\\left(i\\right)+ \\xi\\)\u003c/span\u003e\u003c/span\u003e (Euler\u0026rsquo;s constant), and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\left(n\\right)\\)\u003c/span\u003e\u003c/span\u003e is the average path length of the binary search tree.\u003c/p\u003e\n \u003cp\u003eBy normalizing the length of the isolated binary tree, a number between 0 and 1 can be obtained as the abnormal score of the detected sample. The anomaly score \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(s\\)\u003c/span\u003e\u003c/span\u003e of an instance \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(x\\)\u003c/span\u003e\u003c/span\u003e is defined as:\u003c/p\u003e\n \u003cdiv class=\"Equation\" id=\"Equ9\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e$$s(x, n)={2}^{\\frac{E\\left(h\\left(x\\right)\\right)}{c\\left(n\\right)}}$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e13\u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(h\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e represents the path length from the root node to the \u003cem\u003ex\u003c/em\u003e node, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(E\\left(h\\right(x\\left)\\right)\\)\u003c/span\u003e\u003c/span\u003e is the average of the path lengths of all the isolated trees in the isolated forest for the sample point \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(x\\)\u003c/span\u003e\u003c/span\u003e. When the anomaly score is larger, the sample point is more likely to be an outlier. Based on the anomaly score \u003cem\u003es\u003c/em\u003e, we can make the following assessments (Liu et al. \u003cspan class=\"CitationRef\"\u003e2008\u003c/span\u003e):\u003c/p\u003e\n \u003cp\u003e(1) If the anomaly score is very close to 1, then the data are definitely anomalies.\u003c/p\u003e\n \u003cp\u003e(2) If the anomaly score is much smaller than 0.5, then it is safe to regard the data as normal instances.\u003c/p\u003e\n \u003cp\u003e(3) If all the anomaly scores are approximately 0.5, then there are no distinct outliers in the sample.\u003c/p\u003e\n \u003cp\u003e3.4 Performance evaluation\u003c/p\u003e\n \u003cp\u003eTo objectively and comprehensively evaluate the prediction performance of each model, four different evaluation indicators are selected: root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and determination coefficient (R\u003csup\u003e2\u003c/sup\u003e). RMSE is sensitive to errors that are evident in the experimental data. MAE is the average value of absolute error and can truly reflect the state of the model\u0026apos;s error in prediction. MAPE is the expected value of the absolute error and percentage of the true value. The smaller the RMSE, MAE, and MAEP are, the more accurate the prediction result and the better the model effect. The value of the determination coefficient R\u003csup\u003e2\u003c/sup\u003e is between 0 and 1, and the closer to 1 the value is, the better the model\u0026rsquo;s prediction ability of the regression effect. Generally, if the coefficient of determination exceeds 0.8, the model is considered to have a high goodness of fit. The specific calculation equation of each loss function is as follows:\u003c/p\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable border=\"1\" id=\"Tabc\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(RMSE=\\sqrt{\\frac{1}{N}\\sum _{i=1}^{N}{({y}_{i}-{y}_{i}^{*})}^{2}}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e(14)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(MAE=\\frac{1}{N}\\sum _{i=1}^{N}{|y}_{i}-{y}_{i}^{*}|\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e(15)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(MAPE=\\frac{1}{N}\\sum _{i=1}^{N}\\left|\\frac{{y}_{i}-{y}_{i}^{*}}{{y}_{i}}\\right|\\times 100\\%\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e(16)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}=1-\\frac{\\sum _{i=1}^{N}{({y}_{i}-{y}_{i}^{*})}^{2}}{\\sum _{i=1}^{N}{({y}_{i}-\\stackrel{-}{{y}_{i}})}^{2}}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e(17)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N\\)\u003c/span\u003e\u003c/span\u003e is the number of samples, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({y}_{i}\\)\u003c/span\u003e\u003c/span\u003e is the measured value, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({y}_{i}^{*}\\)\u003c/span\u003e\u003c/span\u003e is the predicted value, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\stackrel{-}{{y}_{i}}\\)\u003c/span\u003e\u003c/span\u003e is the average value of the measured data.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/div\u003e"},{"header":"4. Results","content":"\u003cp\u003e4.1 Data preprocessing\u003c/p\u003e\n\u003cp\u003eSince there were few missing data (\u0026lt;\u0026thinsp;10%) in the water quality time series, the mean smoothing method was used to fill in the missing part of the data; the missing data were replaced by the average value of the two adjacent data on the left and right of the missing data. The min-max normalization method was used to convert the original values into values between [0, 1]. The normalization results are shown in Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e. This figure shows that the water quality parameters have apparent fluctuations.\u003c/p\u003e\n\u003cp\u003eThe isolated forest algorithm described above was used to identify abnormal fluctuations, such as some data jumps in the original series of water quality parameters (the maximum abnormal sample ratio was set to 0.025), and the outliers were marked. The outlier identification results are shown in Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. Considering the small number of outliers and the large difference between an outlier and its adjacent values, outliers were directly removed from the original series, and the average value of the data on both sides of an outlier was used to fill missing values. This figure shows that compared with the original series, obvious outliers in the denoised water quality time series were removed. However, this time series is still complex and has obvious nonstationary and nonlinear characteristics from the overall trend. Measures are still needed to reduce the complexity of the water quality time series.\u003c/p\u003e\n\u003cp\u003e4.2 EEMD decomposition results\u003c/p\u003e\n\u003cp\u003eAfter data preprocessing, the water quality time series was decomposed by the EEMD method. The ensemble number was set to 100, and the standard deviation of Gaussian white noise \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({n}^{i}\\left(t\\right)\\)\u003c/span\u003e\u003c/span\u003e was 0.05 (Ren et al. \u003cspan class=\"CitationRef\"\u003e2015\u003c/span\u003e; Liu et al. \u003cspan class=\"CitationRef\"\u003e2022\u003c/span\u003e). The EEMD results of each water quality parameter are shown in \u003cstrong\u003eFig.\u0026nbsp;6\u003c/strong\u003e. The NH\u003csub\u003e3\u003c/sub\u003eN, pH and DO time series were decomposed into eight IMFs and one residual item Res and arranged in the order of frequency from high to low. The COD\u003csub\u003eMn\u003c/sub\u003e time series was decomposed into seven IMFs and one residual item Res.\u003c/p\u003e\n\u003cp\u003eThe first four IMF components fluctuate greatly, among which IMF1 has the strongest nonlinearity, the largest amplitude and the highest frequency. The residual item can represent the long-term trend of the time series (Liu et al. \u003cspan class=\"CitationRef\"\u003e2022\u003c/span\u003e; Ren et al. \u003cspan class=\"CitationRef\"\u003e2015\u003c/span\u003e). As illustrated in \u003cstrong\u003eFig.\u0026nbsp;6\u003c/strong\u003e, the residual items of the NH\u003csub\u003e3\u003c/sub\u003eN and COD\u003csub\u003eMn\u003c/sub\u003e time series have obvious declining trends, indicating that the water environment control measures of the Xiaofu River have achieved certain results in recent years.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003ctable border=\"1\" id=\"Tab2\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eThe period of IMF components for water quality parameters\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" rowspan=\"2\"\u003e\n \u003cp\u003eVariable\u003c/p\u003e\n \u003cp\u003eName\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colspan=\"8\"\u003e\n \u003cp\u003ePeriod (day)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF1\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF2\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF3\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF4\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF5\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF6\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF7\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eIMF8\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e41\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e152\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e356\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e534\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e534\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epH\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e53\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e89\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e356\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e534\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e534\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e356\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e356\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e534\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCOD\u003csub\u003eMn\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e59\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e66\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e356\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e356\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e-\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eTable \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e shows that the period of the first two IMF components is 3\u0026ndash;9 days, which corresponds to the number of days that the pollutants are naturally degraded in the water body. Therefore, IMF1 and IMF2 may represent the fluctuation of water quality due to water affected by sudden pollution, discontinuous nonpoint source pollution and so on. IMF3-IMF5 mainly reflect the seasonal changes in water quality, and IMF6-IMF8 mainly reflect the interannual changes in water quality. The seasonal and interannual changes in the water quality series are relatively stable, but the fluctuations caused by sudden pollution and discontinuous nonpoint source pollution are large and complex. Therefore, to obtain more accurate water quality prediction results, it is necessary to accurately simulate the high-frequency components. The EEMD method is able to separate the high-frequency components, enhance the details and transform nonlinear water quality series into several relatively simple and stationary time series that help to improve prediction results.\u003c/p\u003e\n\u003cp\u003e4.3 Model training and parameter optimization\u003c/p\u003e\n\u003cp\u003eIn this paper, we chose the MSE of the training dataset as a criterion to calibrate the model and chose the Adam algorithm as the optimizer. The Adam algorithm can solve the problems of a disappearing learning rate and slow convergence property of the error term. It can optimize the performance of the model and has lower running costs with high computational efficiency and less running memory (Diederik and Jimmy \u003cspan class=\"CitationRef\"\u003e2014\u003c/span\u003e). The Adam algorithm was adopted to train the model multiple times and update the parameters continuously. When the error between the actual value and the predicted value meets the accuracy requirements, the model was saved. The hyperparameters of the LSTM model were finally determined. The number of neurons was 50, the number of epochs for each training was 100, and the batch size was 16. In general, the larger the batch size is, the faster the training. However, if the batch size is too large, the network easily converges to the local optimum (Xiang et al. \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eDifferent sliding time window widths \u003cem\u003en\u003c/em\u003e impact the output of the model. In this paper, the water quality time series of the corresponding time width was divided from the dataset as the input sample, and one time step water quality value after the sliding window was used as the output sample. Taking \u003cem\u003en\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4 as an example, its dynamic modeling process is shown in Fig. \u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003eTo improve the prediction accuracy of the model, the model performance is compared under different sliding time window widths, and the results are shown in Table \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. This result illustrates that the optimal sliding time window widths for NH\u003csub\u003e3\u003c/sub\u003eN, pH, DO and COD\u003csub\u003eMn\u003c/sub\u003e are 5, 5, 8 and 7, respectively.\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" id=\"Tab3\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eThe LSTM model performance under different sliding time window widths\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eWater quality\u003c/p\u003e\n \u003cp\u003eindicator\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSliding time\u003c/p\u003e\n \u003cp\u003ewindow width\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMAE\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMAPE\u003c/p\u003e\n \u003cp\u003e(%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.096\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.071\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e67.387\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.423\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.089\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.057\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e31.901\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.783\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.089\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.060\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e42.082\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.727\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.089\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.059\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e39.477\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.746\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.093\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.067\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e60.726\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.545\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epH\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.080\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.049\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.787\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.656\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.078\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.045\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.425\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.741\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.078\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.046\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.521\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.722\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.078\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.046\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.558\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.721\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.087\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.059\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.908\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.656\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.590\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.420\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.831\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.769\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.587\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.424\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.600\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.772\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.594\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.434\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.741\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.763\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.591\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.429\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.630\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.769\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.588\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.422\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.628\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.777\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCOD\u003csub\u003eMn\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.246\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.167\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e11.041\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.748\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.247\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.168\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12.646\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.724\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.244\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.165\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e11.538\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.743\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.249\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.170\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e10.615\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.752\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.243\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.165\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e11.701\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.744\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003e4.4 Water quality prediction by EEMD-LSTM\u003c/p\u003e\n\u003cp\u003eThe water quality data were divided into a training period and validation period; the first 85% of the data were from the training period, and the last 15% of the data were from the validation period. After the model is trained, the learning situation of the model can be judged by the loss curve. If the loss curve declines smoothly or continues to decline at the end of the training period, it indicates that there is an underfitting phenomenon. If the loss curve continues to decline but begins to rise at a certain point or there is an upward trend in the fluctuation, it means that there is an overfitting phenomenon. When the loss values of the model in the training period and the validation period decrease and become stable at the same time, the model training effect is good and can be used for water quality prediction.\u003c/p\u003e\n\u003cp\u003eTo fully verify the performance of EEMD-LSTM, single LSTM and EEMD-LSTM were used to predict water quality parameters using the same data as input. The prediction results of LSTM and EEMD-LSTM are shown in Fig. \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e, and their performance metrics results are listed in Table \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e. It can be seen from Fig. \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e that although LSTM can predict the trend of water quality changes, the error between observed and predicted values is large, and the prediction accuracy of details and jump points is insufficient. EEMD-LSTM can more accurately predict the detailed changes and greatly improve the model performance in terms of the hysteresis problem. It is also evident in Table \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e that the EEMD-LSTM model outperforms LSTM in water quality time series prediction. Compared with LSTM, the prediction accuracy of EEMD-LSTM on the four evaluation indicators of RMSE, MAE, MAPE and R\u003csup\u003e2\u003c/sup\u003e has been improved. The RMSE, MAE, and MAPE of NH\u003csub\u003e3\u003c/sub\u003eN are decreased by 80.0%, 82.6%, and 93.7%, respectively, and R\u003csup\u003e2\u003c/sup\u003e is increased by 63.0%. The RMSE, MAE, and MAPE of pH are decreased by 71.3%, 74.3%, and 82.4%, respectively, and R\u003csup\u003e2\u003c/sup\u003e is increased by 46.9%. The RMSE, MAE, and MAPE of DO are decreased by 78.2%, 80.4%, and 78.8%, respectively, and R\u003csup\u003e2\u003c/sup\u003e is increased by 17.6%. The RMSE, MAE, and MAPE of COD\u003csub\u003eMn\u003c/sub\u003e are decreased by 69.8%, 73.9%, and 84.1%, respectively, and R\u003csup\u003e2\u003c/sup\u003e is increased by 35.1%. These indicators illustrate that the EEMD method can better extract essential features of the water quality time series and reduce the interference of random factors. They also indicate that the prediction performance of the model is greatly improved with the EEMD method. Figure \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e also shows that compared with the single LSTM model, the predicted values of EEMD-LSTM are closer to the observed values in the extreme value prediction.\u003c/p\u003e\n\u003cp\u003eA scatter plot of the observed and predicted values of the two models during the validation period is shown in Fig. \u003cspan class=\"InternalRef\"\u003e9\u003c/span\u003e The scatter plot intuitively shows that the EEMD-LSTM prediction results are closer to the observed value and have better performance.\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" id=\"Tab4\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eModel performance comparison of LSTM and EEMD-LSTM\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" rowspan=\"2\"\u003e\n \u003cp\u003eModel\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" rowspan=\"2\"\u003e\n \u003cp\u003eWater quality\u003c/p\u003e\n \u003cp\u003eindicator\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colspan=\"4\"\u003e\n \u003cp\u003eTraining\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colspan=\"4\"\u003e\n \u003cp\u003eValidation\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMAE\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMAPE\u003c/p\u003e\n \u003cp\u003e(%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMAE\u003c/p\u003e\n \u003cp\u003e(mg/L)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMAPE\u003c/p\u003e\n \u003cp\u003e(%)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLSTM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.169\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.111\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e37.694\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.754\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.110\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.109\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e50.381\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.567\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epH\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.136\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.080\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.032\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.872\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.122\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.113\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.554\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.657\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.151\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.826\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e10.807\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.733\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e1.027\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.820\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4.685\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.817\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCOD\u003csub\u003eMn\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.457\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.314\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.239\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.811\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.440\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.326\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e13.990\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.693\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eEEMD-LSTM\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNH\u003csub\u003e3\u003c/sub\u003eN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.077\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.050\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5.419\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.950\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.019\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e3.150\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.924\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003epH\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.047\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.032\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.321\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.988\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.035\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.029\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.273\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.965\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.531\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.355\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.245\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.945\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.224\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.161\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.994\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.961\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCOD\u003csub\u003eMn\u003c/sub\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.189\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.131\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.756\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.969\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.133\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.085\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.219\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.936\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eIn addition, the reason why EEMD-LSTM improves water quality prediction performance is further discussed. There are seasonal changes, interannual changes and short-term fluctuations in surface water quality parameters. The subsequences obtained by decomposing the original water quality sequence can more clearly show the seasonal periodic changes, interannual periodic changes and short-term fluctuations and reduce the complexity of the input data, which is beneficial to the learning and training of the model. At the same time, the high-frequency components IMF1 and IMF2 decomposed by the EEMD method can reflect the fluctuations in the water quality series caused by sudden pollution, and the prediction of these components separately can effectively improve the prediction accuracy.\u003c/p\u003e"},{"header":"5. Discussion","content":"\u003cp\u003eLSTM has achieved high accuracy prediction results in applications of many fields. However, the prediction accuracy of water quality is not satisfactory, as water quality series are generally multiperiod hybrid time series that have strongly nonlinear and nonstationary characteristics, and LSTM is not suitable for predicting multiperiod hybrid time series. In this paper, we introduced the EEMD method to decompose the water quality time series into several simpler single-period components. The EEMD method can decompose the original water quality series into some components arranged from high frequency to low frequency. Among the IMFs decomposed by EEMD, IMF1 and IMF2 reflect the changing process of sudden pollutants discharged into surface water, and these components have great impacts on the accuracy of water quality prediction. Predicting these high-frequency components separately can improve the accuracy in predicting extreme values and the overall performance of the model. Therefore, the predicted values of EEMD-LSTM are closer to the observed values in the extreme value prediction, and the whole prediction accuracy of the EEMD-LSTM model is also improved compared with the single LSTM model.\u003c/p\u003e \u003cp\u003eThe EEMD-LSTM model has achieved good results in the time series prediction of water quality. The MAE, MAPE and RMSE of EEMD-LSTM for DO are 0.161, 0.994 and 0.224, respectively. The performance predictors of other water quality parameters have also achieved high accuracy. Li et al. (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2017\u003c/span\u003e) proposed a multimodal water quality prediction model called MSVR and proved that the combination of EEMD and SVR could achieve better prediction performance. The MAE, MAPE and RMSE of MSVR for DO were 0.175, 2.153 and 0.228, respectively (Li et al. \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). This shows that EEMD-LSTM is reliable in predicting water quality. Limited by time and effort, only the performance of the hybrid model EEMD-LSTM was studied in this paper for water quality prediction. Subsequently, other methods to improve the performance of LSTM will be considered.\u003c/p\u003e \u003cp\u003eThe influence of different sliding time window widths on the prediction accuracy is also discussed in this paper. The optimal sliding time window width of different water quality parameters is different, which is related to the migration, transformation and degradation rates of pollutants in water. The degradation coefficients of COD\u003csub\u003eMn\u003c/sub\u003e and NH\u003csub\u003e3\u003c/sub\u003eN in rivers are 0.08\u0026ndash;0.15 and 0.2\u0026ndash;0.44 day\u003csup\u003e\u0026minus;\u0026thinsp;1\u003c/sup\u003e, respectively (Ma et al. \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). Therefore, the residence times of COD\u003csub\u003eMn\u003c/sub\u003e and NH\u003csub\u003e3\u003c/sub\u003eN in water are 6.7\u0026ndash;12.5 day and 2.3-5 days. The optimal sliding time window widths for NH3N, pH, DO and COD\u003csub\u003eMn\u003c/sub\u003e are 5, 5, 8 and 7, respectively. This indicates that the optimal sliding time window width is consistent with the degradation time of pollutants in water. This is because after pollutants are discharged into the water, the concentration of pollutants at any point in the water increases with time and then tends to the equilibrium value. As the number of predicted time steps increases, the prediction accuracy of the model will decline, so the EEMD-LSTM model can only predict short time steps at present. Water quality prediction over long time steps is still a challenging issue.\u003c/p\u003e"},{"header":"6. Conclusions","content":"\u003cp\u003eTo achieve highly accurate water quality prediction results, a water quality prediction model based on the combination of the EEMD method and LSTM network is proposed in this paper. The water quality monitoring data of the Xiaofu River are used as a sample for verification, and the four water quality parameters (NH\u003csub\u003e3\u003c/sub\u003eN, pH, DO, COD\u003csub\u003eMn\u003c/sub\u003e) of the Xiaofu River are predicted. The following conclusions were drawn from this study:\u003c/p\u003e \u003cp\u003e(1) The EEMD method can decompose time series into components arranged from high frequency to low frequency. In this paper, it is used to decompose the water quality time series to obtain several single-period components, which can effectively reduce the complexity and nonlinearity of the original time series. Among all components, the high-frequency components have the greatest impact on the accuracy of water quality prediction. Predicting the high-frequency components and the low-frequency components separately when using LSTM can significantly improve model accuracy.\u003c/p\u003e \u003cp\u003e(2) Compared with LSTM, EEMD-LSTM significantly improves the accuracy of water quality prediction and greatly improves the model performance in terms of the hysteresis problem. During the validation period, the RMSE, MAE, MAPE and R\u003csup\u003e2\u003c/sup\u003e of EEMD-LSTM for NH\u003csub\u003e3\u003c/sub\u003eN are 0.022 mg/L, 0.019 mg/L, 3.150% and 0.924, respectively. The RMSE, MAE, MAPE and R\u003csup\u003e2\u003c/sup\u003e of EEMD-LSTM for pH are 0.035 mg/L, 0.029 mg/L, 0.273% and 0.965, respectively. The RMSE, MAE, MAPE and R\u003csup\u003e2\u003c/sup\u003e of EEMD-LSTM for DO are 0.224 mg/L, 0.161 mg/L, 0.994% and 0.961, respectively. The RMSE, MAE, MAPE and R\u003csup\u003e2\u003c/sup\u003e of the EEMD-LSTM for COD\u003csub\u003eMn\u003c/sub\u003e are 0.133 mg/L, 0.085 mg/L, 2.219% and 0.936, respectively. This shows that EEMD-LSTM has high prediction accuracy and strong generalization ability. In addition, the predicted values of EEMD-LSTM are closer to the observed values in the extreme value prediction.\u003c/p\u003e \u003cp\u003eIn summary, EEMD-LSTM can be an effective tool for water quality prediction. The EEMD-LSTM model can quickly and accurately predict water quality changes, which can reflect the trend of future water quality changes and can provide a basis for formulating water environment governance measures.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u0026nbsp;\u003c/strong\u003eThis paper was supported by the Major Program of National Natural Science Foundation of China (No. 41790431).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e The authors declare no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e Conceptualization: Lan Luo, Yanjun Zhang and Wenxun Dong; Methodology: Lan Luo, Yanjun Zhang and Anni Qiu; Formal analysis and investigation: Lan Luo, Jinglin Zhang and Liping Zhang; Writing - original draft preparation: Lan Luo; Writing - review and editing: Lan Luo and Yanjun Zhang; All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical approval\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to participate\u003c/strong\u003e Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to publish\u003c/strong\u003e All authors give consent to publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e The datasets used during the current study are available from the corresponding author on request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAn L, Hao Y, Yeh TJ, Liu Y, Liu W, Zhang B (2020) Simulation of karst spring discharge using a combination of time-frequency analysis methods and long short-term memory neural networks. J Hydrol 589:125320. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jhydrol.2020.125320\u003c/span\u003e\u003cspan address=\"10.1016/j.jhydrol.2020.125320\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArunKumar KE, Kalaga DV, Kumar CMS, Kawaji M, Brenza TM (2021) Forecasting of COVID-19 using deep layer Recurrent Neural Networks (RNNs) with Gated Recurrent Units (GRUs) and Long Short-Term Memory (LSTM) cells. Chaos Solitons Fractals 146:110861. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.chaos.2021.110861\u003c/span\u003e\u003cspan address=\"10.1016/j.chaos.2021.110861\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBui HH, Ha NH, Nguyen TND, Nguyen AT, Pham TTH, Kandasamy J, Tien VN (2019) Integration of SWAT and QUAL2K for water quality modeling in a data scarce basin of Cau River basin in Vietnam. Ecohydrol Hydrobiol 19(2SI):210\u0026ndash;223. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ecohyd.2019.03.005\u003c/span\u003e\u003cspan address=\"10.1016/j.ecohyd.2019.03.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDa Costa SB, Leite CM, Almeida IR, de Almeida AK, I.K (2021) Choosing an appropriate water quality model-a review. Environ Monit Assess 193. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10661-020-08786-1\u003c/span\u003e\u003cspan address=\"10.1007/s10661-020-08786-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeus R, Brito D, Mateus M, Kenov I, Fornaro A, Neves R, Alves CN (2013) Impact evaluation of a pisciculture in the Tucurui reservoir (Para, Brazil) using a two-dimensional water quality model. J Hydrol 487:1\u0026ndash;12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jhydrol.2013.01.022\u003c/span\u003e\u003cspan address=\"10.1016/j.jhydrol.2013.01.022\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDiederik PK, Jimmy B (2014) Adam: A method for stochastic optimization. arXiv 1412:6980. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.1412.6980\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.1412.6980\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing S, Wang F, Sun X, Ding J, Lu J (2022) Water environmental functional zoning at county level and environmental contamination carrying capacity accounting in the mainstream of Xiaofu River. Water-Sui 14(4):615. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/w14040615\u003c/span\u003e\u003cspan address=\"10.3390/w14040615\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEjigu MT (2021) Overview of water quality modeling. Cogent Eng 8(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/23311916.2021.1891711\u003c/span\u003e\u003cspan address=\"10.1080/23311916.2021.1891711\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEze E, Halse S, Ajmal T (2021) Developing a novel water quality prediction model for a South African aquaculture farm. Water-Sui 13(13):1782. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/w13131782\u003c/span\u003e\u003cspan address=\"10.3390/w13131782\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuia SS, Espirito-Santo A, Paciello V, Abate F, Pietrosanto A (2015) A comparison between FFT and MCT for period measurement with an ARM Microcontroller. 2015 IEEE International Instrumentation and Measurement Technology Conference, pp.\u0026nbsp;1938\u0026ndash;1942\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735\u0026ndash;1780. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-3-642-24797-2_4\u003c/span\u003e\u003cspan address=\"10.1007/978-3-642-24797-2_4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuan J, Cao W, Qin Y (2018) Prediction of dissolved oxygen in aquaculture based on EEMD and LSSVM optimized by the Bayesian evidence framework. Comput Electron Agric 150:257\u0026ndash;265. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.compag.2018.04.022\u003c/span\u003e\u003cspan address=\"10.1016/j.compag.2018.04.022\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang NE, Shen Z, Long SR, Wu M, Shih HH, Zheng QN, Yen NC, Tung CC, Liu HH (1998) The empirical mode decomposition and the Hilbert spectrum for nonlinear and non-stationary time series analysis. P Roy Soc A-Math Phy 454(1971):903\u0026ndash;995. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1098/rspa.1998.0193\u003c/span\u003e\u003cspan address=\"10.1098/rspa.1998.0193\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim J, Lee T, Seo D (2017) Algal bloom prediction of the lower Han River, Korea using the EFDC hydrodynamic and water quality model. Ecol Model 366:27\u0026ndash;36. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ecolmodel.2017.10.015\u003c/span\u003e\u003cspan address=\"10.1016/j.ecolmodel.2017.10.015\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKourgialas NN, Dokou Z, Karatzas GP (2015) Statistical analysis and ANN modeling for predicting hydrological extremes under climate change scenarios: The example of a small mediterranean agro-watershed. J Environ Manage 154:86\u0026ndash;101. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jenvman.2015.02.034\u003c/span\u003e\u003cspan address=\"10.1016/j.jenvman.2015.02.034\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi XJ, Cheng ZW, Yu QB, Bai Y, Li C (2017) Water-quality prediction using multimodal support vector regression: case study of Jialing River, China.J Environ Eng143(10)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang Z, Zou R, Chen X, Ren T, Su H, Liu Y (2020) Simulate the forecast capacity of a complicated water quality model using the long short-term memory approach. J Hydrol 581:124432. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jhydrol.2019.124432\u003c/span\u003e\u003cspan address=\"10.1016/j.jhydrol.2019.124432\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu FT, Ting KM, Zhou Z (2008) Isolation forest. 2008 Eighth Ieee International Conference on Data Mining, pp.\u0026nbsp;413\u0026ndash;422\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu X, Zhang Y, Zhang Q (2022) Comparison of EEMD-ARIMA, EEMD-BP and EEMD-SVM algorithms for predicting the hourly urban water consumption. J Hydroinform 24(3):535\u0026ndash;558. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2166/hydro.2022.146\u003c/span\u003e\u003cspan address=\"10.2166/hydro.2022.146\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Y, Zhang Q, Song L, Chen Y (2019) Attention-based recurrent neural networks for accurate short-term and long-term dissolved oxygen prediction. Comput Electron Agric 165. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.compag.2019.104964\u003c/span\u003e\u003cspan address=\"10.1016/j.compag.2019.104964\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa L, Liu L, Song LL, Yan WM (2014) A study on water pollutant degradation capability affected by water diversion. J Environ Prot Ecol 15(1):39\u0026ndash;47\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMendes J, Ruela R, Picado A, Pinheiro JP, Ribeiro AS, Pereira H, Dias JM (2021) Modeling dynamic processes of Mondego Estuary and Oacute, Bidos Lagoon using Delft3D.J Mar Sci Technol9(1)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNiu W, Feng Z, Zeng M, Feng B, Min Y, Cheng C, Zhou J (2019) Forecasting reservoir monthly runoff via ensemble empirical mode decomposition and extreme learning machine optimized by an improved gravitational search algorithm. Appl Soft Comput 82. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.asoc.2019.105589\u003c/span\u003e\u003cspan address=\"10.1016/j.asoc.2019.105589\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePalani S, Liong S, Tkalich P (2008) An ANN application for water quality forecasting. Mar Pollut Bull 56(9):1586\u0026ndash;1597. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.marpolbul.2008.05.021\u003c/span\u003e\u003cspan address=\"10.1016/j.marpolbul.2008.05.021\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQingmei M, Min L, Aiju L (2013) Spatial variation and contamination assessment of heavy metals in surface sediments of Xiaofu River. Health Environ Res 6:785\u0026ndash;790\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRen Y, Suganthan PN, Srikanth N (2015) A comparative study of empirical mode decomposition-based short-term wind speed forecasting methods. IEEE T Sustain Energ 6(1):236\u0026ndash;244\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRui Y, Shen D, Khalid S, Yang Z, Wang J (2015) GIS-based emergency response system for sudden water pollution accidents. Phys Chem Earth 79\u0026ndash;82:115\u0026ndash;121. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.pce.2015.03.001\u003c/span\u003e\u003cspan address=\"10.1016/j.pce.2015.03.001\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRumelhart DE, Hinton GE, Williams RJ (1986) Learning representations by back-propagating errors. Nature 323(6088):533\u0026ndash;536. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/323533a0\u003c/span\u003e\u003cspan address=\"10.1038/323533a0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeo IW, Yun SH, Choi SY (2016) Forecasting water quality parameters by ANN model using pre-processing technique at the downstream of Cheongpyeong Dam. Procedia Eng 154:1110\u0026ndash;1115. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.proeng.2016.07.519\u003c/span\u003e\u003cspan address=\"10.1016/j.proeng.2016.07.519\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShabani A, Zhang X, Chu X, Zheng H (2021) Automatic calibration for CE-QUAL-W2 model using improved global-best harmony search algorithm. Water-Sui 13(16). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/w13162308\u003c/span\u003e\u003cspan address=\"10.3390/w13162308\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSingh D, Singh B (2020) Investigating the impact of data normalization on classification performance. Appl Soft Comput. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.asoc.2019.105524\u003c/span\u003e\u003cspan address=\"10.1016/j.asoc.2019.105524\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. 97(B)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang W, Pei Y, Zheng H, Zhao Y, Shu L, Zhang H (2022) Twenty years of China's water pollution control: Experiences and challenges. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.chemosphere.2022.133875\u003c/span\u003e\u003cspan address=\"10.1016/j.chemosphere.2022.133875\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Chemosphere 295\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTant CJ, Rosemond AD, Helton AM, First MR (2015) Nutrient enrichment alters the magnitude and timing of fungal, bacterial, and detritivore contributions to litter breakdown. Freshw Sci 34(4):1259\u0026ndash;1271\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang J, Wang X, Lei XH, Wang H, Zhang XH, You JJ, Tan QF, Liu XL (2020) Teleconnection analysis of monthly streamflow using ensemble empirical mode decomposition. J Hydrol 582. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jhydrol.2019.124411\u003c/span\u003e\u003cspan address=\"10.1016/j.jhydrol.2019.124411\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWool T Jr, Ambrose RB, Martin JL, Comer A (2020) WASP 8: the next generation in the 50-year evolution of USEPA's water quality model.Water-Sui12(5)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiang Z, Yan J, Demir I (2020) A rainfall-runoff model with LSTM-based sequence-to-sequence learning. Water Resour Res 56(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1029/2019WR025326\u003c/span\u003e\u003cspan address=\"10.1029/2019WR025326\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiong Y, Ran Y, Zhao S, Zhao H, Tian Q (2020) Remotely assessing and monitoring coastal and inland water quality in China: Progress, challenges and outlook. Crit Rev Environ Sci Technol 50(12):1266\u0026ndash;1302. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/10643389.2019.1656511\u003c/span\u003e\u003cspan address=\"10.1080/10643389.2019.1656511\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang S, Yang D, Chen J, Santisirisomboon J, Lu W, Zhao B (2020) A physical process and machine learning combined hydrological model for daily streamflow simulations of large watersheds with limited observation data. J Hydrol 590. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jhydrol.2020.125206\u003c/span\u003e\u003cspan address=\"10.1016/j.jhydrol.2020.125206\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu J, Kim J, Li X, Jong Y, Kim K, Ryang G (2022) Water quality forecasting based on data decomposition, fuzzy clustering and deep learning neural network. Environ Pollut 303:119136. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.envpol.2022.119136\u003c/span\u003e\u003cspan address=\"10.1016/j.envpol.2022.119136\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZema DA, Lucas-Borja ME, Fotia L, Rosaci D, Sarne GML, Zimbone SM (2020) Predicting the hydrological response of a forest after wildfire and soil treatments using an Artificial Neural Network. Comput Electron Agric 170. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.compag.2020.105280\u003c/span\u003e\u003cspan address=\"10.1016/j.compag.2020.105280\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang JL, Tang MG, Liu F, Zhong ZS (2000) Vulnerability analysis of groundwater pollution by mining drainage in Zibo coal mine, Shandong Province, China. International Symposium on Hydrogeology and the Environment, pp.\u0026nbsp;157\u0026ndash;162\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhaohua WU, Norden EH (2009) Ensemble empirical mode decomposition: a noise-assisted data analysis method. Adv Adapt Data Anal 1(1):1\u0026ndash;41. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1142/S1793536909000047\u003c/span\u003e\u003cspan address=\"10.1142/S1793536909000047\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng L, Wang H, Liu C, Zhang S, Ding A, Xie E, Li J, Wang S (2021) Prediction of harmful algal blooms in large water bodies using the combined EFDC and LSTM models. J Environ Manage 295. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jenvman.2021.113060\u003c/span\u003e\u003cspan address=\"10.1016/j.jenvman.2021.113060\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou J, Wang J, Chen Y, Li X, Xie Y (2021) Water quality prediction method based on multi-source transfer learning for water environmental IoT system. Sensors 21(21):7271. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/s21217271\u003c/span\u003e\u003cspan address=\"10.3390/s21217271\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Water quality prediction, Ensemble empirical mode decomposition, Long short-term memory network, Deep learning, Xiaofu River","lastPublishedDoi":"10.21203/rs.3.rs-2116084/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2116084/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eWater quality prediction is an important part of water pollution prevention and control. Using a long short-term memory (LSTM) neural network to predict water quality can solve the problem that comprehensive water quality models are too complex and difficult to apply. However, as water quality time series are generally multiperiod hybrid time series, which have strongly nonlinear and nonstationary characteristics, the prediction accuracy of LSTM for water quality is not high. The ensemble empirical mode decomposition (EEMD) method can decompose the multiperiod hybrid water quality time series into several simpler single-period components. To improve the accuracy of surface water quality prediction, a water quality prediction model based on EEMD-LSTM was proposed in this paper. The water quality time series was first decomposed into several intrinsic mode function components and one residual item, and then these components were used as the input of LSTM to predict water quality. The model was trained and validated using four water quality parameters (NH\u003csub\u003e3\u003c/sub\u003eN, pH, DO, COD\u003csub\u003eMn\u003c/sub\u003e) collected from the Xiaofu River and compared with the results of a single LSTM. During the validation period, the R\u003csup\u003e2\u003c/sup\u003e values when using LSTM for NH\u003csub\u003e3\u003c/sub\u003eN, pH, DO and COD\u003csub\u003eMn\u003c/sub\u003e were 0.567, 0.657, 0.817 and 0.693, respectively, and the R\u003csup\u003e2\u003c/sup\u003e values when using EEMD-LSTM for NH\u003csub\u003e3\u003c/sub\u003eN, pH, DO and COD\u003csub\u003eMn\u003c/sub\u003e were 0.924, 0.965, 0.961 and 0.936, respectively. The results show that the proposed model outperforms the single LSTM model in various evaluation indicators and greatly improves the model performance in terms of the hysteresis problem. The EEMD-LSTM model has high prediction accuracy and strong generalization ability, and further development may be valuable.\u003c/p\u003e","manuscriptTitle":"Ensemble empirical mode decomposition and a long short-term memory neural network for surface water quality prediction of the Xiaofu River, China","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-11-17 19:43:41","doi":"10.21203/rs.3.rs-2116084/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"42557232-b726-4e4c-b90a-55f2f6ef8126","owner":[],"postedDate":"November 17th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-01-10T06:15:50+00:00","versionOfRecord":[],"versionCreatedAt":"2022-11-17 19:43:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2116084","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2116084","identity":"rs-2116084","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.