Detection of urban flood inundation from traffic images using deep learning methods | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Detection of urban flood inundation from traffic images using deep learning methods pengcheng zhong, Yueyi Liu, Hang Zheng, Jianshi Zhao This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3075920/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Dec, 2023 Read the published version in Water Resources Management → Version 1 posted 4 You are reading this latest preprint version Abstract Urban hydrological monitoring is the basis for urban hydrological analysis and storm flood control. However, current monitoring of urban hydrological data is insufficient, including flood inundation depth. This limits calibration and flood early warning ability of the hydrological model. In response to this limitation, a method for evaluating the depth of urban floods based on image recognition using deep learning was established in this study. This method can identify the submerged positions of pedestrians or vehicles in the image, such as pedestrian legs and car exhaust pipes, using the object recognition model YOLOv4. The mean average precision of water depth recognition in a dataset of 1177 flood images reached 89.29%. The established method extracted on-site, real-time, and continuous water depth data from images or video data provided by existing traffic cameras. This system does not require installation of additional water gauges and thus has a low cost and immediate usability. urban flood inundation image recognition deep learning water depth Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Highlights Water depth can be calculated by recognizing the flood image using deep learning. The method can identify submerged positions of pedestrians or vehicles in the image. The precision of the YOLOv4 model for flood depth recognition can reach 89.29%. 1. Introduction Urban floods have posed increasing challenges on global sustainable developments (Hammond et al., 2015 ; Nkwunonwo et al., 2020 ). Urban flooding problems in terms of the flooding frequency and the damage caused are growing due to climate change and intensive urbanization (Chen et al.,2015; Yin et al.,2015; Jamali et al., 2018 ). For example, in China, an extreme storm with a daily rainfall of approximately 650 mm battered the city Zhengzhou in central China, which has a population of 13 million in an urban area of about 1000 km 2 (Wang et al.,2022). A total of 398 people died within a week during the storm, from 17 July to 23 July 2021, because of the rapid inundation of flooding in urban low-lying areas, including traffic tunnels and subways where passengers were trapped by flash floods (Disaster Investigation Team of the State Council, China, 2022 ). Early warning and efficient management of urban flash floods are still critical in saving lives in modern metropolises. Urban hydrology provides a basis to estimate the inundated areas and damage of urban floods (Fletcher et al., 2013 ; Ichiba et al., 2018 ; Rubinato et al., 2019 ). Urban floods are generally caused by storm runoff that cannot be completely discharged by urban drainage systems (Cristiano et al., 2019 ; Kourtis and Tsihrintzis, 2021 ). Compared with floods in mountain areas, urban floods occur faster because runoff is generated more quickly on the impervious surfaces of cities (Sohn et al., 2020 ; de Mello Silva and da Silva, 2020 ). A higher percentage of impervious surfaces in the city reduces the precipitation infiltration and at the same time increases the speed of surface water flow during a storm (Lee and Brody, 2018 ; Yu et al., 2018 ; Feng et al., 2021 ). Consequently, this allows only a limited time for warning and evacuating people possibly caught by the flash floods. Urban flash floods are also more damaging because more people and expensive infrastructures might be flooded in cities (Jamali et al., 2018 ; Singh et al.,2018). Several hydrological models were previously developed for urban flood simulations and control (Mignot et al., 2019 ; Bulti and Abebe, 2020 ; Guo et al., 2021 ). Some of them, such as Storm Water Management Model and MIKE FLOOD, are widely applied worldwide (Babaei et al., 2018 ; Bai et al., 2018 ; Li et al.,2018; Nigussie and Altunkaynak, 2019 ). The distribution of inundated areas in cities can be obtained through modeling and consequently used for urban flood early warnings (Rangari et al., 2019 ; Seenu et al., 2020 ). However, a gap exists in terms of the model calibration and validation due to a lack of measured data of floodwater depths across urban areas. The reasons for this might be the limited coverage of monitoring spots and a relatively high cost for maintaining the gauging equipment (Arshad et al., 2019 ; Moy de Vitry et al., 2019 ; Li et al., 2019 ). Generally, there are numerous and scattered waterlogging locations in urban areas because different low-lying facilities or underground infrastructures, such as subway stations and traffic tunnels, tend to be dispersed on the urban landscape in the context of intensive urban developments. It is quite costly to monitor distributed flood inundations using a traditional sensor-based system covering these scattered locations. In addition, the measurement equipment already installed in waterlogging locations, such as remote water level sensors or gauges, are threatened by possible vandalism as many passengers and a high volume of traffic flow might appear around them. This further increases the cost for equipment maintenance. As a result, monitoring networks with extensive coverage on urban flood depths are limited throughout the world (Wang et al., 2018 ). Consequently, the existing models of urban hydrology can generally only use the discharge data of underground drainage pipes or urban rivers for calibration and validation rather than the actual flood depth data (Wang et al., 2018 ; Liu et al., 2020 ). This significantly reduces the reliability of flood models. To obtain more flood depth data in urban areas, new technologies were applied in the past decade, including remote sensing (Cohen et al.2018; Shen et al., 2019 ; Rong et al., 2020 ), social media or crowdsourcing data retrieval (Ilieva and McPhearson, 2018 ; Scotti et al., 2020 ; de Vitry and Leitão, 2020 ; Kankanamge, et al., 2020 ), and object recognition from flood images (Jiang et al., 2019 ; Bhola et al.,2019; Park et al., 2021 ). However, the performances of those methods are not currently satisfactory, either in terms of accuracy or practicality. For example, the application of aerial and satellite imagery based on optical sensors is limited by vegetation canopies and cloud cover during floods (Munasinghe et al., 2018 ; DeVries et al., 2020 ; Hashemi-Beni and Gebrehiwot, 2021 ). Microwave remote sensing technology, which can penetrate cloud cover, lacks the ability to extract high-quality flood data in small-scale urban areas due to insufficient spatial resolution and consequently has low sensitivity for urban flooding (Zeng et al., 2020 ; Dubey et al., 2021 ; Galantowicz and Picton, 2021 ). Other studies tried to alleviate the urban flood data limitations by leveraging information contained in images of the event posted on social media platforms (Feng et al., 2020 ; Songchon et al., 2021 ). For example, Barz et al. established a method using content-based image retrieval technology and an artificial intelligence method to identify "flood" and "non-flood" events from a dataset of about 100,000 online images (Barz et al., 2019 ). Bînă et al. established a deep learning classifier that detected flooding events by analyzing text published by online news outlets as well as the accompanying article images (Bînă et al., 2019 ). Although these methods provided new information for urban flood monitoring, it is still difficult to obtain accurate and continuous records on floodwater levels at a specific location due to the relatively low reliability of social media texts and images that are generally posted at various locations without time consistency during floods. To increase the quality of flood images, high resolution on-site photos of water level gauges from a video surveillance system were used to extract water level figures automatically through identification technology (Lv et al., 2018 ; Jiang et al., 2020 ). For instance, Basnyat and Gangopadhyay developed a low-cost, low-power cyber-physical system prototype using a Raspberry Pi camera to detect the rising water level using image processing and text recognition techniques (Basnyat et al., 2018 ). This could pave the way for mass deployment of a flash flood detection system with minimal human intervention, but its practicality is still limited in urban areas, as on-site water level gauges and camera systems must be installed at traffic-intensive spots in the city (Yang et al., 2014 ). Given the above limitations, this study proposed an innovative method to extract water level data at inundated spots of urban areas during flash floods through an image-recognition method based on deep learning. The method recognized flood level figures from traffic photos during floods by identifying the relative position of the water surface to the reference objects in the water, including people and vehicles passing through the water or partly submerged by the water. The proposed method could provide real-time and continuous data of inundated water levels at an appointed low-lying location without installing extra engineering sensors, by analyzing images of water and traffic objectives obtained from existing traffic surveillance cameras in road networks. As traffic surveillance cameras are widespread in cities globally, this method might be practical for expanding urban flood monitoring and providing timely early warning at a low cost. 2. Materials and Methods 2.1. Data A total of 1177 images of flooded streets with pedestrians or vehicles were used for water depth identification. The image resolutions ranged from 142 × 107 pixels to 4613 × 2595 pixels. These images were obtained by Google Search using keywords that included “urban waterlogging” and “urban floods.” Using these images, a standard dataset for floodwater depth recognition was generated by the pattern analysis, statistical modeling, and computational learning (PASCAL) visual object class (VOC) framework. PASCAL is a network organization sponsored by the EU, which created public datasets of images and annotations, together with standardized evaluation software (Everingham et al., 2010 ). PASCAL VOC is one of the widely used visual datasets with a strictly normalized framework (Gauen et al., 2017 ) to identify the size, object class, and coordinates of bounding boxes in the images. It increases the efficiency of image annotation by processing the annotation information in the PASCAL VOC format. In this study, the 1177 images were annotated to 2095 objects and seven categories using the VOC format by the software LabelImg (Yakovlev A et al., 2020). The seven categories were generated according to the positions of people or vehicles submerged by flood water, as shown in Table 1 . Table 1 Categories of flood images Categories Definition Number of objects Car-none Dry surface or the water depth less than 5 cm 195 Car-pipe Water depth more than 5 cm and below the exhaust pipe of a vehicle 372 Car-handle Water depth between the vehicle exhaust pipe and the door handle 337 Car-roof Water over the vehicle’s roof 276 Per-none Dry surface or the water depth less than 5 cm 209 Per-leg Water depth more than 5 cm and below the waist of a person 561 Per-waist Water depth above a person's waist 145 Total number of objects 2095 Note: Per is short for person 2.2. Target detection algorithm The “you only look once” (YOLO) algorithm was used in this study to extract features and detect objects in the flood images using a deep convolutional neural network. YOLO is a commonly used model for target detection from images because of its fast speed of image recognition in a one-stage algorithm, which performs feature extraction, target classification, and target regression through the same convolutional neural network (CNN) (Redmon et al., 2016 ). Compared with a two-stage algorithm, such as the Faster-RCNN with multiple steps for generating the candidate area of the input image and extracting features for the image recognition through different neural networks, YOLO skips the generation of a candidate area and completes the detection in one step. This provides the YOLO algorithm with a relatively higher calculation speed and a sufficient precision (Zou et al., 2019 ). This fits the requirements of this study to detect changes of flood depths in a real-time manner. YOLO conducts object detection by calculating the probability that a bounding box in an image belongs to a pre-set category using a CNN and a regression method. First, YOLO takes the entire image as the input of the network after dividing the image into a 7×7 grid. Then, each grid is framed by two alternative bounding boxes with n categories for each box to be classified. There are five parameters of each bounding box, namely, the horizontal coordinate of the box central point ( x ), the vertical coordinate of the central point ( y ), the width ( w ) and height ( h ) of the box, and the probability that the bounding box contains the target ( c ∈[0,1]). Based on this, a variable named YOLO Head is obtained as a tensor with a shape of (7, 7, 2 × 5 + n). For example, if a dataset contains 20 categories, the output of the model will be a matrix of (7, 7, 30). In this matrix, the first two dimensions with a shape of (7 × 7) represent the number of feature grids extracted. The last dimension with a size of 30 is the result of multiplying the number of bounding boxes and the number of box parameters plus the number of probability values of the box belonging to each category, which is 2×5 + 20. The YOLO algorithm was first established and updated to its third generation by Redmon J (Redmon & Farhadi, 2018 ). The fourth generation, YOLOv4, was proposed by Alexey Bochkovskiy in 2020 (Bochkovskiy et al., 2020 ). YOLOv4 was used in this study. It is composed of three parts: the backbone, feature enhancement, and prediction. In the backbone, the model extracts three feature layers from the image using five residual networks after adjusting the image to a fixed resolution using CSPDarknet53 (C.-Y. Wang et al., 2020 ). The three feature layers are then transferred to the feature enhancement part, the neck, to obtain three prediction layers, which comprise the YOLO head, for the prediction part. The YOLO head is used for final classification and recognition of the image. In this study, three feature layers with the shapes (13, 13), (26, 26), and (52, 52) were adopted to extract the feature from an image with the shape (416, 416) through pooling technology of the CNN, as shown in Fig. 1 . Each grid in the feature layer was framed by three alternative bounding boxes with seven categories. Each bounding box has five parameters, as previously stated. Therefore, the sizes of the output YOLO heads are (13, 13, 22), (26, 26, 22), and (52, 52, 22), in which the number of the third dimension with a size of 22 is the result of multiplying the number of bounding boxes with the number of box parameters plus the number of probability values that the box belongs to each category, which is 3×5 + 7. In addition, YOLOv4 adopts the activation function Mish (Misra, 2019 ) instead of LeakyRelu (Maas et al., 2013 ) used in YOLOv3. Compared with LeakyRelu, the Mish function has a lower bound; so, it helps the model to strengthen the regularization effect. Furthermore, the non-monotonic function of Mish preserves small negative values, thus stabilizing the network gradient flow. Mish is infinitely continuous with smooth curves, which also helps to improve the quality of the results. Although the Mish function uses more computing resources than LeakyRelu functions, it improves the accuracy of the model. The activation functions are shown in Fig. 2 . 2.3. Model training and validation In this study, the input data were randomly distributed into a training set (90% of data) and a validation set (10% of data). YOLOv4 was trained by the method of transfer learning, in which the initial parameters, including the initial weights of the CNN, were transferred from the existing dataset VOC2007 (Everingham et al., 2010 ). Transfer learning is a practical scheme in the training process, which can use generic features from other larger training datasets. This reduces the time required for training the model. The VOC2007 dataset consists of about 500,000 annotated images that are divided into 20 object classes. The dataset is large enough to provide the CNN of this study with initial weights. Two hundred epochs were adopted for model training in this study. In this process, frozen learning with a learning rate of 1e-3 was used in the first 90 epochs, and regular learning with a learning rate of 1e-4 was used in the rest of the epochs. Frozen learning did not change the initial parameters of the networks in the backbone part, and only generated minor adjustments to other networks during the training (Brock et al., 2017 ). In contrast, in regular learning, the parameters of all the networks were adjusted in the model. Frozen learning can reduce the time required for training by using the existed network parameters obtained from a previous dataset such as VOC2007. However, the model would not converge after a few epochs because the backbone networks were not updated under frozen learning. Therefore, regular learning was adopted subsequently in the training process for further convergence of the model. In addition, a lower learning rate was also adopted in the regular learning to avoid overfitting. The convergence curve of the model training is shown in Fig. 3 . The decrease of training loss became stable after 70 epochs of training using frozen learning due to the unchanging structure of the network. Then, the training loss was further reduced without frozen learning until 150 epochs. 2.4. Indicators for performance assessment Mean average precision (mAP) was used in this study to evaluate the accuracy of flood image recognition. It is a metric widely applied in object detection and classification (Zou et al., 2019 ). The mAP was calculated using precision and recall indicators of the modeling results, which are given by $$Precision= \frac{TP}{TP+FP}$$ 1 $$Recall= \frac{TP}{TP+FN}$$ 2 where TP (true positive) represents the number of samples that are correctly recognized or classified, FP (false positive) represents the number of samples that are recognized incorrectly, and FN (false negative) represents the number of unrecognized results. Precision is the ratio of correctly predicted positive results to the total predicted positive results. Recall is the ratio of correctly predicted positive results to the total number of actual positive samples. Given the above indicators, the mAP can be calculated as follows: $$mAP= \frac{{\sum }_{1}^{n}AP}{n}$$ 3 $$AP= {\int }_{0}^{1}p\left(r\right)dr$$ 4 where AP (average precision) represents the area under the precision-recall curve of the prediction results referring to each category, n is the total number of categories, which is seven in this study, and r is the value of recall from Eq. 2 . p (r) represents the Precision-Recall curve. The y-axis of p ( r ) is the maximum value of the corresponding precision when the value of recall increases from 0 to 1 at a given interval, such as 0.1. 3. Results The results of flood recognition from the street images are shown in Fig. 4 . As shown in the figure, the reference objects, the vehicles and pedestrians, are framed by bounding boxes with the identified category of the water depth presented at the top of each box. In this manner, the flood water depth can be extracted by real-time image recognition using street photos or surveillance videos from traffic cameras. The precision, recall, and AP of the flood recognition from 1177 images are shown in Table 2 . The mAP of the model prediction reached 89.29%, demonstrating a satisfactory performance of the model for identifying the submerged depths of the reference objects. The AP of the predictions in the categories of Car-roof and Car-pipe was 98.81% and 98.63%, respectively. This indicated that the best performance of the flood depth recognition from the image dataset in this study was for the identification of water levels reaching the roof and the exhaust pipe of the vehicles. The AP of the predictions for the categories of Car-handle and Per-leg was 92.47% and 90.44%, respectively, representing the second-best performances of the flood depth recognition. Table 2 Results of flood depth recognition categories Category Precision (%) Recall (%) AP 50 (%) Car-roof 95.00 95.00 98.81 Car-pipe 87.50 93.33 98.63 Car-handle 82.35 90.32 92.47 Per-leg 85.94 82.09 90.44 Car-none 93.75 71.43 87.60 Per-waist 75.00 64.29 83.84 Per-none 82.35 51.85 73.21 mAP 89.29 Note: Per is short for person. The AP of seven categories and the number of labeled objects in each category are shown in Fig. 5 . There were 2095 objects in the dataset, including 1885 objects used in model training and 210 objects used for validation. The category Per-leg had the highest number of objects, with 494 objects in the training set and 67 objects in the validation set. This meant that more images showing floods with a water depth reaching the legs of pedestrians were used in this study compared to images showing other categories. Compared with the Per-leg category, the Car-pipe category had a smaller number of objects, with 342 objects in the training set and 30 objects in the validation set. However, the AP of the Car-pipe was 98.63% and was higher than the AP of the Per-leg category, which was only 90.44%. In addition, the Per-waist category had the least number of objects but the second lowest AP. The AP was not only affected by the number of objects but also by the feature extractions of the objects. This founding was supported by other studies using the dataset VOC2007, in which the objects were classified into 20 categories with various APs. In previous studies, objects with relatively clearer contours and more regular shapes, such as cars and trains, achieved higher recognition APs. Conversely, irregularly shaped objects, such as humans, cats, and dogs, had lower APs (X. Wang et al., 2013 ). This is consistent with the results of this study. When a car was used as the reference object of the flood submerged depth, the AP of the flood recognition was higher than when using the human body as the inundated reference. This is because the objects with clear contours and regular shapes more easily have features extracted in image recognition. Figure 6 shows the precision and recall of water depth predictions in the categories Car-roof and Per-none, which had the best and the worst performances of the prediction, respectively, as shown in Table 5 . In Fig. 6 , the x-axis was a pre-defined confidence threshold of the bounding box containing the object in the image to be recognized. The range of the confidence threshold was from 0.1 to 1.0 in this study. The object in the bounding box could be recognized and categorized only when the box confidence was higher than the pre-defined threshold. Therefore, the precision increased with an increasing threshold. For example, the predicted category of the bounding box was highly accurate when precisely matching its annotation if the confidence threshold was close to 1. However, in this case, most of the bounding boxes failed to pass the high threshold, causing a very low value of the recall. The recall decreased when the threshold increased. As shown in Fig. 6 , the recall of the water depth prediction corresponding to the Car-roof category was 95.00%, which was almost double the 51.85% recall corresponding to the Per-none category when the confidence threshold was 0.5. This is because the posture of the vehicle tends to be more consistent than that of the pedestrian, who might have multiple postures, including standing, bending over, and squatting. The variations of the object posture also cause difficulty with recognition. This finding was consistent with a previous study (Giannakeris et al., 2018 ), which found that the recall of recognizing a fire in an image was lower than that of recognizing a flood in an image, because the variable shape of the fire made it more difficult to be recognized accurately. 4. Discussion 4.1. Performances of frozen learning and regular learning The performances of frozen learning and regular learning on inundated depth recognition were analyzed by the changes of the mAP and the loss in the learning process using the different learning methods. The mAP and the loss are shown at every 50 epochs during the learning process in Table 3 . At 200 epochs, the mAP of regular learning was higher than that of frozen learning. The loss of regular learning was less than that of frozen learning. This shows that the performance of regular learning was better than that of frozen learning. In addition, a further improvement of the model performance was found by using a combined learning method, which used frozen learning in the first 90 epochs and regular learning in the rest of the epochs. Table 3 mAP and loss across different epochs using frozen learning and regular learning Epochs Frozen learning Regular learning Combined learning mAP (%) Loss mAP (%) Loss mAP (%) Loss 50 74.94 43.44 69.44 41.40 78.33 43.62 100 73.28 43.48 72.34 38.61 79.92 33.82 150 73.73 43.42 73.51 36.41 83.14 28.36 200 73.35 42.47 75. 57 34.61 83.35 27.81 The changes in the mAP and the loss are shown in Fig. 7 . The mAP of the frozen learning was higher than that of the regular learning within the first 100 epochs, after which the mAP of the regular learning continued to rise and increased more than that of frozen learning. Moreover, the mAP of the combined learning was higher than that of the frozen learning and regular learning after 50 epochs. In addition, the loss of the combined learning decreased and became less than that of the frozen learning and regular learning after about 70 epochs. Both the mAP and the loss of the combined learning were stable without further decreases or increases at 200 epochs, showing that there was no overfitting in the learning process. Given this, in the context of transfer learning, the frozen learning with fixed parameters transferred from the existing dataset had a better performance at the early stage of the learning process. However, regular learning, which altered the transferred parameters, had a better performance at the end of the learning process. More importantly, the performance of the combined learning was better than that of using frozen learning and regular learning separately. 4.2. Impacts of training technologies Three technologies for image processing and model training were applied in the deep learning process to analyze the impacts of these technologies on improving the performance of image recognition on the flood depths. These technologies include Mosaic (Bochkovskiy et al., 2020 ), label smoothing (Szegedy et al., 2016 ), and cosine annealing (Loshchilov & Hutter, 2016 ), which are currently widely used in image recognition. Mosaic is a technique for image augmentation in which four images are combined through cropping and blending, making full use of the data (Hao & Zhili, 2020 ). Label smoothing is a regularization technique that introduces noise for the labels (Müller et al., 2019 ). Cosine annealing is a learning rate variation strategy characterized by periodic drops and resets of the learning rate during the training process. Cosine annealing helps the training process escape a local optimal solution by resetting a larger learning rate (Gotmare et al., 2018 ). The applicability of different technologies was analyzed by comparing the model mAP across application scenarios using eight individual or combined technologies, as shown in Table 4 . In scenario 1, none of above technologies was applied, and the mAP of the modeling under this scenario was only 67.58%. After using Mosaic, label smoothing, and cosine annealing separately, the mAP increased to 72.14%, 70.94%, and 70.48%, respectively. The mAP was improved the most in scenario 2 compared with scenario 1 using Mosaic. However, in scenarios 5 and 6, in which Mosaic was jointly applied with label smoothing and cosine annealing, respectively, the mAP was 71.92% and 67.53%, respectively. The mAP was reduced when Mosaic was jointly applied with label smoothing and cosine annealing compared with Mosaic used alone. Mosaic had the best performance for improving the mAP of flood depth recognition in this study. This was consistent with previous studies on YOLOv4 in which Mosaic improved the accuracy of image recognition (Bochkovskiy A, et, 2020). In contrast, the lowest mAP of 66.29% was from scenario 7 in which the label smoothing and cosine annealing were jointly applied without Mosaic. In addition, the second lowest mAP of 67.53% resulted from scenario 6 in which Mosaic and cosine annealing were jointly applied. It was found that, in this study, cosine annealing was not suitable to be used in combination with Mosaic and label smoothing because both cosine annealing with Mosaic and cosine annealing with label smoothing had lower performances than did cosine annealing alone. Table 4 mAP using different training technologies Scenario Mosaic Label Smoothing Cosine Annealing mAP (%) 1 67.58 2 √ 74.12 3 √ 70.94 4 √ 70.48 5 √ √ 71.92 6 √ √ 67.53 7 √ √ 66.29 8 √ √ √ 71.48 4.3. Impacts of activation functions The activation function can affect the performance of a deep learning model significantly in terms of the accuracy and speed of calculations. The mAP and the frames per second (FPS) were calculated in this study to evaluate the impacts of two activation functions, Mish and LeakyRelu, as shown in Fig. 2 . A higher FPS indicates a faster speed of image recognition. The results are shown in Table 5 . The YOLOv4 model using Mish generally had higher mAP results than did LeakyRelu when recognizing flood depth in images. However, the FPS of the model with Mish was 30.49, which was less than that of LeakyRelu. This indicates that Mish performed at a slower speed in image recognition, with about five fewer images per second compared to LeakyRelu in this study. This is consistent with previous studies that showed that Mish had a relatively greater loss in detection speed compared to LeakyRelu, but the accuracy was better than that of LeakyRelu (C.-Y. Wang et al., 2021 ). Table 5 mAP (%) and FPS using LeakyRelu and Mish functions mAP FPS Epochs 20 40 60 80 100 120 140 160 180 200 Mish 65.56 72.3 73.91 70.35 68.75 73.99 72.91 69.34 75.53 74.56 30.49 LeakyRelu 55.04 71.05 71.8 67.67 72.21 72.49 72.06 71.98 71.76 72.06 34.92 5. Conclusion An image recognition method for the detection of urban flood inundation was established in this study by applying the deep learning network YOLOv4 and the transfer training technique. Vehicles and pedestrians submerged by flood water in the images were adopted as reference objects for flood water depth recognition. After training of 200 epochs, the model classified images with a certain flood water level into four categories, namely, Car-none, Car-pipe, Car-handle, and Car-roof, representing that the flood water depth in the image was less than 5 cm, or that the flood level reached the exhaust pipe, the door, or the roof of a vehicle, respectively, with vehicles as the reference. Similarly, the model also classified the flood image into three categories, namely, Per-none, Per-leg, and Per-waist, representing that the flood water depth in the image was less than 5 cm, or that the flood level reached the leg or waist of a person, respectively. The results showed that the mAP of the proposed method reached 89.29%, with a recognition speed of 30.49 FPS. It was also found that the accuracy of flood depth recognition by YOLOv4 was affected by the reference object submerged by the flood. The accuracy of using a vehicle as the reference object was significantly higher than that of using a person as the reference object. This is because a vehicle generally has a regular shape and a fixed position, which makes it easier to identify. In addition, the impact on the flood depth recognition of various transfer learning techniques, image processing methods, networking training tricks, and activation functions of the CNN was discussed in the study. It was found that, in the context of flood depth recognition from images using YOLOv4, (1) the performance of the combined use of frozen learning and regular learning was better than that of using the two learning techniques separately; (2) image augmentation using Mosaic technology effectively improved the accuracy of recognition, and using Mosaic alone was better than jointly using Mosaic, label smoothing, and cosine annealing in this study; and (3) the model using the Mish activation function had a better performance than that using the LeakyRelu activation function. The proposed YOLOv4 model has the advantages of high accuracy and fast speed in recognizing the flood water depth from images using vehicles and pedestrians as reference objects. This enables the method to use the images or video data provided by existing traffic cameras to identify the flood depth during a storm, thereby obtaining on-site, real-time, and continuous water depth data. This method does not require the additional installation of a water gauge or a water depth sensing device and thus has a low cost and a high usability in practice. The flood depth data obtained by this method can be used for the verification and calibration of urban hydrological models and can also provide real-time inundation data for early warning of urban floods, making up for the insufficient coverage of flood inundation depth monitoring sites in current urban hydrological research. The proposed method could be improved in terms of recognition accuracy and reliability. For example, the flood depth was predicted in the form of a categorical grade, such as the water level to the handle of the car or to the leg of a person, rather than the specific water depth value. In addition, the method relies on reference objects, and it is difficult to identify the water depth without a reference object on the road surface. Declarations Funding : This study was financially supported by the National Natural Science Foundation of China (U2040206),(52179009) and (51909035). Competing Interests: The authors have no relevant financial or non-financial interests to disclose. Author Contributions: All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Pengcheng Zhong, Yueyi Liu, Hang Zheng, Jianshi Zhao. The first draft of the manuscript was written by Pengcheng Zhong and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript. References Hammond, M. J., Chen, A. S., Djordjević, S., Butler, D., & Mark, O. (2015). Urban flood impact assessment: A state-of-the-art review. Urban Water Journal, 12(1), 14-29. Nkwunonwo, U. C., Whitworth, M., & Baily, B. (2020). A review of the current status of flood modelling for urban flood risk management in the developing countries. Scientific African, 7, e00269. Chen, Y., Zhou, H., Zhang, H., Du, G., & Zhou, J. (2015). Urban flood risk warning under rapid urbanization. Environmental research, 139, 3-10. Yin, J., Ye, M., Yin, Z., & Xu, S. (2015). A review of advances in urban flood risk analysis over China. Stochastic Environmental Research and Risk Assessment, 29(3), 1063-1070. Jamali, B., Löwe, R., Bach, P. M., Urich, C., Arnbjerg-Nielsen, K., & Deletic, A. (2018). A rapid urban flood inundation and damage assessment model. Journal of Hydrology, 564, 1085-1098. Wang, H., Hu, Y., Guo, Y., Wu, Z., & Yan, D. (2022). Urban flood forecasting based on the coupling of numerical weather model and stormwater model: A case study of Zhengzhou city. Journal of Hydrology: Regional Studies, 39, 100985. Disaster Investigation Team of the State Council, China. (2022). Investigation report on the "7.20" heavy rainstorm disaster in Zhengzhou, Henan, China. Fletcher, T. D., Andrieu, H., & Hamel, P. (2013). Understanding, management and modelling of urban hydrology and its consequences for receiving waters: A state of the art. Advances in Water Resources, 51, 261-279. Ichiba, A., Gires, A., Tchiguirinskaia, I., Schertzer, D., Bompard, P., & Ten Veldhuis, M. C. (2018). Scale effect challenges in urban hydrology highlighted with a distributed hydrological model. Hydrology and Earth System Sciences, 22(1), 331-350. Rubinato, M., Nichols, A., Peng, Y., Zhang, J. M., Lashford, C., Cai, Y. P., ... & Tait, S. (2019). Urban and river flooding: Comparison of flood risk management approaches in the UK and China and an assessment of future knowledge needs. Water science and engineering, 12(4), 274-283. Cristiano, E., ten Veldhuis, M. C., Wright, D. B., Smith, J. A., & van de Giesen, N. (2019). The influence of rainfall and catchment critical scales on urban hydrological response sensitivity. Water Resources Research, 55(4), 3375-3390. Kourtis, I. M., & Tsihrintzis, V. A. (2021). Adaptation of urban drainage networks to climate change: A review. Science of The Total Environment, 771, 145431. Sohn, W., Kim, J. H., Li, M. H., Brown, R. D., & Jaber, F. H. (2020). How does increasing impervious surfaces affect urban flooding in response to climate variability?. Ecological Indicators, 118, 106774. de Mello Silva, C., & da Silva, G. B. L. (2020). Cumulative effect of the disconnection of impervious areas within residential lots on runoff generation and temporal patterns in a small urban area. Journal of environmental management, 253, 109719. Lee, Y., & Brody, S. D. (2018). Examining the impact of land use on flood losses in Seoul, Korea. Land use policy, 70, 500-509. Yu, H., Zhao, Y., Fu, Y., & Li, L. (2018). Spatiotemporal variance assessment of urban rainstorm waterlogging affected by impervious surface expansion: A case study of Guangzhou, China. Sustainability, 10(10), 3761. Feng, B., Zhang, Y., & Bourke, R. (2021). Urbanization impacts on flood risks based on urban growth data and coupled flood models. Natural Hazards, 106(1), 613-627. Jamali, B., Löwe, R., Bach, P. M., Urich, C., Arnbjerg-Nielsen, K., & Deletic, A. (2018). A rapid urban flood inundation and damage assessment model. Journal of Hydrology, 564, 1085-1098. Singh, P., Sinha, V. S. P., Vijhani, A., & Pahuja, N. (2018). Vulnerability assessment of urban road network from urban flood. International journal of disaster risk reduction, 28, 237-250. Mignot, E., Li, X., & Dewals, B. (2019). Experimental modelling of urban flooding: A review. Journal of Hydrology, 568, 334-342. Bulti, D. T., & Abebe, B. G. (2020). A review of flood modeling methods for urban pluvial flood application. Modeling earth systems and environment, 6(3), 1293-1302. Guo, K., Guan, M., & Yu, D. (2021). Urban surface water flood modelling–a comprehensive review of current models and future challenges. Hydrology and Earth System Sciences, 25(5), 2843-2860. Babaei, S., Ghazavi, R., & Erfanian, M. (2018). Urban flood simulation and prioritization of critical urban sub-catchments using SWMM model and PROMETHEE II approach. Physics and Chemistry of the Earth, Parts A/B/C, 105, 3-11. Bai, Y., Zhao, N., Zhang, R., & Zeng, X. (2018). Storm water management of low impact development in urban areas based on SWMM. Water, 11(1), 33. Li, J., Zhang, B., Mu, C., & Chen, L. (2018). Simulation of the hydrological and environmental effects of a sponge city based on MIKE FLOOD. Environmental earth sciences, 77(2), 1-16. Nigussie, T. A., & Altunkaynak, A. (2019). Modeling the effect of urbanization on flood risk in Ayamama Watershed, Istanbul, Turkey, using the MIKE 21 FM model. Natural Hazards, 99(2), 1031-1047. Rangari, V. A., Umamahesh, N. V., & Bhatt, C. M. (2019). Assessment of inundation risk in urban floods using HEC RAS 2D. Modeling Earth Systems and Environment, 5(4), 1839-1851. Seenu, P. Z., Venkata Rathnam, E., & Jayakumar, K. V. (2020). Visualisation of urban flood inundation using SWMM and 4D GIS. Spatial Information Research, 28(4), 459-467. Arshad, B., Ogie, R., Barthelemy, J., Pradhan, B., Verstaevel, N., & Perez, P. (2019). Computer vision and IoT-based sensors in flood monitoring and mapping: A systematic review. Sensors, 19(22), 5012. Moy de Vitry, M., Kramer, S., Wegner, J. D., & Leitão, J. P. (2019). Scalable flood level trend monitoring with surveillance cameras using a deep convolutional neural network. Hydrology and Earth System Sciences, 23(11), 4621-4634. Li, Y., Martinis, S., & Wieland, M. (2019). Urban flood mapping with an active self-learning convolutional neural network based on TerraSAR-X intensity and interferometric coherence. ISPRS Journal of Photogrammetry and Remote Sensing, 152, 178-191. Wang, R. Q., Mao, H., Wang, Y., Rae, C., & Shaw, W. (2018). Hyper-resolution monitoring of urban flooding with social media and crowdsourcing data. Computers & Geosciences, 111, 139-147. Wang, Y., Chen, A. S., Fu, G., Djordjević, S., Zhang, C., & Savić, D. A. (2018). An integrated framework for high-resolution urban flood modelling considering multiple information sources and urban features. Environmental modelling & software, 107, 85-95. Liu, J., Shao, W., Xiang, C., Mei, C., & Li, Z. (2020). Uncertainties of urban flood modeling: Influence of parameters for different underlying surfaces. Environmental research, 182, 108929. Cohen, S., Brakenridge, G. R., Kettner, A., Bates, B., Nelson, J., McDonald, R., ... & Zhang, J. (2018). Estimating floodwater depths from flood inundation maps and topography. JAWRA Journal of the American Water Resources Association, 54(4), 847-858. Shen, X., Wang, D., Mao, K., Anagnostou, E., & Hong, Y. (2019). Inundation extent mapping by synthetic aperture radar: A review. Remote Sensing, 11(7), 879. Rong, Y., Zhang, T., Zheng, Y., Hu, C., Peng, L., & Feng, P. (2020). Three-dimensional urban flood inundation simulation based on digital aerial photogrammetry. Journal of Hydrology, 584, 124308. Ilieva, R. T., & McPhearson, T. (2018). Social-media data for urban sustainability. Nature Sustainability, 1(10), 553-565. Scotti, V., Giannini, M., & Cioffi, F. (2020). Enhanced flood mapping using synthetic aperture radar (SAR) images, hydraulic modelling, and social media: A case study of Hurricane Harvey (Houston, TX). Journal of Flood Risk Management, 13(4), e12647. de Vitry, M. M., & Leitão, J. P. (2020). The potential of proxy water level measurements for calibrating urban pluvial flood models. Water Research, 175, 115669. Kankanamge, N., Yigitcanlar, T., Goonetilleke, A., & Kamruzzaman, M. (2020). Determining disaster severity through social media analysis: Testing the methodology with South East Queensland Flood tweets. International journal of disaster risk reduction, 42, 101360. Jiang, J., Liu, J., Cheng, C., Huang, J., & Xue, A. (2019). Automatic estimation of urban waterlogging depths from video images based on ubiquitous reference objects. Remote Sensing, 11(5), 587. Bhola, P. K., Nair, B. B., Leandro, J., Rao, S. N., & Disse, M. (2019). Flood inundation forecasts using validation data generated with the assistance of computer vision. Journal of Hydroinformatics, 21(2), 240-256. Park, S., Baek, F., Sohn, J., & Kim, H. (2021). Computer vision–based estimation of flood depth in flooded-vehicle images. Journal of Computing in Civil Engineering, 35(2), 04020072. Munasinghe, D., Cohen, S., Huang, Y. F., Tsang, Y. P., Zhang, J., & Fang, Z. (2018). Intercomparison of satellite remote sensing‐based flood inundation mapping techniques. JAWRA Journal of the American Water Resources Association, 54(4), 834-846. DeVries, B., Huang, C., Armston, J., Huang, W., Jones, J. W., & Lang, M. W. (2020). Rapid and robust monitoring of flood events using Sentinel-1 and Landsat data on the Google Earth Engine. Remote Sensing of Environment, 240, 111664. Hashemi-Beni, L., & Gebrehiwot, A. A. (2021). Flood extent mapping: An integrated method using deep learning and region growing using UAV optical data. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 2127-2135. Zeng, Z., Gan, Y., Kettner, A. J., Yang, Q., Zeng, C., Brakenridge, G. R., & Hong, Y. (2020). Towards high resolution flood monitoring: An integrated methodology using passive microwave brightness temperatures and Sentinel synthetic aperture radar imagery. Journal of Hydrology, 582, 124377. Dubey, A. K., Kumar, P., Chembolu, V., Dutta, S., Singh, R. P., & Rajawat, A. S. (2021). Flood modeling of a large transboundary river using WRF-Hydro and microwave remote sensing. Journal of Hydrology, 598, 126391. Galantowicz, J. F., & Picton, J. (2021). Flood Mapping with Passive Microwave Remote Sensing: Current Capabilities and Directions for Future Development. In Earth Observation for Flood Applications (pp. 39-60). Elsevier.Feng et al., 2020; Songchon et al., 2021. Feng, Y., Brenner, C., & Sester, M. (2020). Flood severity mapping from Volunteered Geographic Information by interpreting water level from images containing people: A case study of Hurricane Harvey. ISPRS Journal of Photogrammetry and Remote Sensing, 169, 301-319. Songchon, C., Wright, G., & Beevers, L. (2021). Quality assessment of crowdsourced social media data for urban flood management. Computers, Environment and Urban Systems, 90, 101690. Barz, B., Schröter, K., Münch, M., Yang, B., Unger, A., Dransch, D., & Denzler, J. (2019). Enhancing flood impact analysis using interactive retrieval of social media images. arXiv preprint arXiv:1908.03361. Bînă, D., Vlad, G. A., Onose, C., & Cercel, D. C. (2019, October). Flood severity estimation in news articles using deep learning approaches. In Proceedings of the MediaEval 2019 Workshop, Sophia Antipolis, France (pp. 27-29). Lv, Y., Gao, W., Yang, C., & Wang, N. (2018). Inundated areas extraction based on raindrop photometric model (RPM) in surveillance video. Water, 10(10), 1332. Jiang, J., Qin, C. Z., Yu, J., Cheng, C., Liu, J., & Huang, J. (2020). Obtaining urban waterlogging depths from video images using synthetic image data. Remote Sensing, 12(6), 1014. Basnyat, B., Roy, N., & Gangopadhyay, A. (2018, June). A flash flood categorization system using scene-text recognition. In 2018 IEEE International Conference on Smart Computing (SMARTCOMP) (pp. 147-154). IEEE. Yang, HC., Wang, CY. & Yang, JX. (2014). Applying image recording and identification for measuring water stages to prevent flood hazards. Nat Hazards 74, 737–754. Everingham, M., van Gool, L., Williams, C. K. I., Winn, J., & Zisserman, A. (2010). The pascal visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303–338. Gauen, K., Dailey, R., Laiman, J., Zi, Y., Asokan, N., Lu, Y.-H., Thiruvathukal, G. K., Shyu, M.-L., & Chen, S.-C. (2017). Comparison of visual datasets for machine learning. 2017 IEEE International Conference on Information Reuse and Integration (IRI), 346–355. Yakovlev, A., & Lisovychenko, O. (2020). An approach for image annotation automatization for artificial intelligence models learning. Адаптивні Системи Автоматичного Управління, 1(36), 32–40. Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 779–788. Zou, Z., Shi, Z., Guo, Y., & Ye, J. (2019). Object detection in 20 years: A survey. ArXiv Preprint ArXiv:1905.05055. Redmon, J., & Farhadi, A. (2018). Yolov3: An incremental improvement. ArXiv Preprint ArXiv:1804.02767. Bochkovskiy, A., Wang, C.-Y., & Liao, H.-Y. M. (2020). Yolov4: Optimal speed and accuracy of object detection. ArXiv Preprint ArXiv:2004.10934. Wang, C.-Y., Liao, H.-Y. M., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., & Yeh, I.-H. (2020). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 390–391. Misra, D. (2019). Mish: A self regularized non-monotonic neural activation function. ArXiv Preprint ArXiv:1908.08681, 4(2), 10–48550. Maas, A. L., Hannun, A. Y., & Ng, A. Y. (2013). Rectifier nonlinearities improve neural network acoustic models. Proc. Icml, 30(1), 3. Brock, A., Lim, T., Ritchie, J. M., & Weston, N. (2017). Freezeout: Accelerate training by progressively freezing layers. ArXiv Preprint ArXiv:1706.04983. Wang, X., Yang, M., Zhu, S., & Lin, Y. (2013). Regionlets for generic object detection. Proceedings of the IEEE International Conference on Computer Vision, 17–24. Giannakeris, P., Avgerinakis, K., Karakostas, A., Vrochidis, S., & Kompatsiaris, I. (2018). People and vehicles in danger-A fire and flood detection system in social media. 2018 IEEE 13th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP), 1–5. Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2016). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2818–2826. Loshchilov, I., & Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. ArXiv Preprint ArXiv:1608.03983. Hao, W., & Zhili, S. (2020). Improved mosaic: Algorithms for more complex images. Journal of Physics: Conference Series, 1684(1), 012094. Müller, R., Kornblith, S., & Hinton, G. E. (2019). When does label smoothing help? Advances in Neural Information Processing Systems, 32. Gotmare, A., Keskar, N. S., Xiong, C., & Socher, R. (2018). A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation. ArXiv Preprint ArXiv:1810.13243. Wang, C.-Y., Bochkovskiy, A., & Liao, H.-Y. M. (2021). Scaled-yolov4: Scaling cross stage partial network. Proceedings of the IEEE/Cvf Conference on Computer Vision and Pattern Recognition, 13029–13038. Cite Share Download PDF Status: Published Journal Publication published 02 Dec, 2023 Read the published version in Water Resources Management → Version 1 posted Reviewers agreed at journal 22 Jul, 2023 Reviewers invited by journal 22 Jul, 2023 Editor assigned by journal 25 Jun, 2023 First submitted to journal 24 Jun, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3075920","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":220707097,"identity":"001babf8-bbba-40d3-b478-e902b4f93fd5","order_by":0,"name":"pengcheng zhong","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1klEQVRIiWNgGAWjYBACCQhlI8cPYUnIEKklIc1YcgYDYwOQz0OslkOJBjfAWhgIa5Fs7zF8XPjjQILx7ebjj27UWPAwsB8+ugGfFmmeM8bGMxLu5JndOZbYnHMM6DCetLQb+LTISaSlSfMkPCs2u5Fj2JzDBtQiwWOGX4v8M5CWw4mbZ4C0/CNCi7QE8zGwlg0SQC25bURokexJPmwMdL2xxI20xNm5fRI8bIT8InH8YONjHhtgVM5IPvA551udHD/74WN4tWACNtKUj4JRMApGwSjABgAGSUS5+mR0DQAAAABJRU5ErkJggg==","orcid":"","institution":"Dongguan University of Technology","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"pengcheng","middleName":"","lastName":"zhong","suffix":""},{"id":220707098,"identity":"d763095e-5ce8-4138-8e1c-de753ebab73f","order_by":1,"name":"Yueyi Liu","email":"","orcid":"","institution":"Dongguan University of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yueyi","middleName":"","lastName":"Liu","suffix":""},{"id":220707099,"identity":"bda01f47-851c-4eeb-8bd4-fa2bfc1c24e6","order_by":2,"name":"Hang Zheng","email":"","orcid":"","institution":"Dongguan University of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hang","middleName":"","lastName":"Zheng","suffix":""},{"id":220707100,"identity":"336a1772-51a3-435b-895d-b48271fdd803","order_by":3,"name":"Jianshi Zhao","email":"","orcid":"","institution":"Tsinghua University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jianshi","middleName":"","lastName":"Zhao","suffix":""}],"badges":[],"createdAt":"2023-06-17 12:47:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3075920/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3075920/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11269-023-03669-9","type":"published","date":"2023-12-02T15:00:53+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":40620349,"identity":"8bfde3ad-76fd-4d3f-b32e-b7b9662033ef","added_by":"auto","created_at":"2023-07-26 17:33:08","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":162544,"visible":true,"origin":"","legend":"\u003cp\u003eNetwork of \"You Only See Once\" Version 4 (YOLOv4)\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/5bdf39e003dd5ca6748c6f3a.png"},{"id":40619729,"identity":"07913dee-bf77-40c4-be07-f6b420adb391","added_by":"auto","created_at":"2023-07-26 17:25:07","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":332843,"visible":true,"origin":"","legend":"\u003cp\u003eFunctions Mish and LeakyRelu\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/6c0087162841aac1425b7a76.jpg"},{"id":40620348,"identity":"10fee881-6a90-41f4-8b9f-52267bab4257","added_by":"auto","created_at":"2023-07-26 17:33:07","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":39837,"visible":true,"origin":"","legend":"\u003cp\u003eConvergence curve of the model training\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/b9cdd99c887b0dd6b336514c.jpg"},{"id":40619735,"identity":"b92dafaf-8411-4fc8-a2bf-63a11148e88b","added_by":"auto","created_at":"2023-07-26 17:25:08","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":883480,"visible":true,"origin":"","legend":"\u003cp\u003eExamples of flood depth recognitions\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/1bf4a5b28ba8382c1f48b0d7.jpg"},{"id":40620350,"identity":"f0e2013f-c11d-46ae-857d-6a0174f05eed","added_by":"auto","created_at":"2023-07-26 17:33:08","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":614682,"visible":true,"origin":"","legend":"\u003cp\u003eAP and number of labeled objects of the predictions\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/65e50ef7ab262e03aa9fc3d0.jpg"},{"id":40619731,"identity":"68e449d1-6c77-4170-a2b4-05873a792168","added_by":"auto","created_at":"2023-07-26 17:25:07","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":221372,"visible":true,"origin":"","legend":"\u003cp\u003eChanges in precision and recall of the predictions\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/f8d3c4fde624ebae9820b005.png"},{"id":40619732,"identity":"f79c7488-ef90-4389-a26c-7de95b26d346","added_by":"auto","created_at":"2023-07-26 17:25:08","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":843267,"visible":true,"origin":"","legend":"\u003cp\u003eChanges of the AP and loss across different epochs using various learning methods\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/8b7b7c5b4d766e64f9825ce6.jpg"},{"id":47561006,"identity":"a226524a-7935-4ac2-a17d-898f9454e165","added_by":"auto","created_at":"2023-12-04 15:06:36","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1094628,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3075920/v1/00c17492-2784-4400-a685-13003f92de02.pdf"}],"financialInterests":"","formattedTitle":"Detection of urban flood inundation from traffic images using deep learning methods","fulltext":[{"header":"Highlights","content":"\u003cp\u003eWater depth can be calculated by recognizing the flood image using deep learning.\u003c/p\u003e\n\u003cp\u003eThe method can identify submerged positions of pedestrians or vehicles in the image.\u003c/p\u003e\n\u003cp\u003eThe precision of the YOLOv4 model for flood depth recognition can reach 89.29%.\u003c/p\u003e"},{"header":"1. Introduction","content":"\u003cp\u003eUrban floods have posed increasing challenges on global sustainable developments (Hammond et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Nkwunonwo et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Urban flooding problems in terms of the flooding frequency and the damage caused are growing due to climate change and intensive urbanization (Chen et al.,2015; Yin et al.,2015; Jamali et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). For example, in China, an extreme storm with a daily rainfall of approximately 650 mm battered the city Zhengzhou in central China, which has a population of 13\u0026nbsp;million in an urban area of about 1000 km\u003csup\u003e2\u003c/sup\u003e (Wang et al.,2022). A total of 398 people died within a week during the storm, from 17 July to 23 July 2021, because of the rapid inundation of flooding in urban low-lying areas, including traffic tunnels and subways where passengers were trapped by flash floods (Disaster Investigation Team of the State Council, China, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). Early warning and efficient management of urban flash floods are still critical in saving lives in modern metropolises.\u003c/p\u003e \u003cp\u003eUrban hydrology provides a basis to estimate the inundated areas and damage of urban floods (Fletcher et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Ichiba et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Rubinato et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Urban floods are generally caused by storm runoff that cannot be completely discharged by urban drainage systems (Cristiano et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Kourtis and Tsihrintzis, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Compared with floods in mountain areas, urban floods occur faster because runoff is generated more quickly on the impervious surfaces of cities (Sohn et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; de Mello Silva and da Silva, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). A higher percentage of impervious surfaces in the city reduces the precipitation infiltration and at the same time increases the speed of surface water flow during a storm (Lee and Brody, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Yu et al., \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Feng et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Consequently, this allows only a limited time for warning and evacuating people possibly caught by the flash floods. Urban flash floods are also more damaging because more people and expensive infrastructures might be flooded in cities (Jamali et al., \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Singh et al.,2018).\u003c/p\u003e \u003cp\u003eSeveral hydrological models were previously developed for urban flood simulations and control (Mignot et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Bulti and Abebe, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Guo et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Some of them, such as Storm Water Management Model and MIKE FLOOD, are widely applied worldwide (Babaei et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Bai et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Li et al.,2018; Nigussie and Altunkaynak, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The distribution of inundated areas in cities can be obtained through modeling and consequently used for urban flood early warnings (Rangari et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Seenu et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). However, a gap exists in terms of the model calibration and validation due to a lack of measured data of floodwater depths across urban areas. The reasons for this might be the limited coverage of monitoring spots and a relatively high cost for maintaining the gauging equipment (Arshad et al., \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Moy de Vitry et al., \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Li et al., \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Generally, there are numerous and scattered waterlogging locations in urban areas because different low-lying facilities or underground infrastructures, such as subway stations and traffic tunnels, tend to be dispersed on the urban landscape in the context of intensive urban developments. It is quite costly to monitor distributed flood inundations using a traditional sensor-based system covering these scattered locations. In addition, the measurement equipment already installed in waterlogging locations, such as remote water level sensors or gauges, are threatened by possible vandalism as many passengers and a high volume of traffic flow might appear around them. This further increases the cost for equipment maintenance. As a result, monitoring networks with extensive coverage on urban flood depths are limited throughout the world (Wang et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Consequently, the existing models of urban hydrology can generally only use the discharge data of underground drainage pipes or urban rivers for calibration and validation rather than the actual flood depth data (Wang et al., \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Liu et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). This significantly reduces the reliability of flood models.\u003c/p\u003e \u003cp\u003eTo obtain more flood depth data in urban areas, new technologies were applied in the past decade, including remote sensing (Cohen et al.2018; Shen et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Rong et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), social media or crowdsourcing data retrieval (Ilieva and McPhearson, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Scotti et al., \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; de Vitry and Leit\u0026atilde;o, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Kankanamge, et al., \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2020\u003c/span\u003e), and object recognition from flood images (Jiang et al., \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Bhola et al.,2019; Park et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). However, the performances of those methods are not currently satisfactory, either in terms of accuracy or practicality. For example, the application of aerial and satellite imagery based on optical sensors is limited by vegetation canopies and cloud cover during floods (Munasinghe et al., \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; DeVries et al., \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Hashemi-Beni and Gebrehiwot, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Microwave remote sensing technology, which can penetrate cloud cover, lacks the ability to extract high-quality flood data in small-scale urban areas due to insufficient spatial resolution and consequently has low sensitivity for urban flooding (Zeng et al., \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Dubey et al., \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Galantowicz and Picton, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eOther studies tried to alleviate the urban flood data limitations by leveraging information contained in images of the event posted on social media platforms (Feng et al., \u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Songchon et al., \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). For example, Barz et al. established a method using content-based image retrieval technology and an artificial intelligence method to identify \"flood\" and \"non-flood\" events from a dataset of about 100,000 online images (Barz et al., \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). B\u0026icirc;nă et al. established a deep learning classifier that detected flooding events by analyzing text published by online news outlets as well as the accompanying article images (B\u0026icirc;nă et al., \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Although these methods provided new information for urban flood monitoring, it is still difficult to obtain accurate and continuous records on floodwater levels at a specific location due to the relatively low reliability of social media texts and images that are generally posted at various locations without time consistency during floods. To increase the quality of flood images, high resolution on-site photos of water level gauges from a video surveillance system were used to extract water level figures automatically through identification technology (Lv et al., \u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Jiang et al., \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). For instance, Basnyat and Gangopadhyay developed a low-cost, low-power cyber-physical system prototype using a Raspberry Pi camera to detect the rising water level using image processing and text recognition techniques (Basnyat et al., \u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). This could pave the way for mass deployment of a flash flood detection system with minimal human intervention, but its practicality is still limited in urban areas, as on-site water level gauges and camera systems must be installed at traffic-intensive spots in the city (Yang et al., \u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e2014\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eGiven the above limitations, this study proposed an innovative method to extract water level data at inundated spots of urban areas during flash floods through an image-recognition method based on deep learning. The method recognized flood level figures from traffic photos during floods by identifying the relative position of the water surface to the reference objects in the water, including people and vehicles passing through the water or partly submerged by the water. The proposed method could provide real-time and continuous data of inundated water levels at an appointed low-lying location without installing extra engineering sensors, by analyzing images of water and traffic objectives obtained from existing traffic surveillance cameras in road networks. As traffic surveillance cameras are widespread in cities globally, this method might be practical for expanding urban flood monitoring and providing timely early warning at a low cost.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003e2.1. Data\u003c/h2\u003e\n\u003cp\u003eA total of 1177 images of flooded streets with pedestrians or vehicles were used for water depth identification. The image resolutions ranged from 142 \u0026times; 107 pixels to 4613 \u0026times; 2595 pixels. These images were obtained by Google Search using keywords that included \u0026ldquo;urban waterlogging\u0026rdquo; and \u0026ldquo;urban floods.\u0026rdquo; Using these images, a standard dataset for floodwater depth recognition was generated by the pattern analysis, statistical modeling, and computational learning (PASCAL) visual object class (VOC) framework. PASCAL is a network organization sponsored by the EU, which created public datasets of images and annotations, together with standardized evaluation software (Everingham et al., \u003cspan class=\"CitationRef\"\u003e2010\u003c/span\u003e). PASCAL VOC is one of the widely used visual datasets with a strictly normalized framework (Gauen et al., \u003cspan class=\"CitationRef\"\u003e2017\u003c/span\u003e) to identify the size, object class, and coordinates of bounding boxes in the images. It increases the efficiency of image annotation by processing the annotation information in the PASCAL VOC format. In this study, the 1177 images were annotated to 2095 objects and seven categories using the VOC format by the software LabelImg (Yakovlev A et al., 2020). The seven categories were generated according to the positions of people or vehicles submerged by flood water, as shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eCategories of flood images\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eCategories\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eDefinition\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eNumber of objects\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-none\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDry surface or the water depth less than 5 cm\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e195\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-pipe\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eWater depth more than 5 cm and below the exhaust pipe of a vehicle\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e372\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-handle\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eWater depth between the vehicle exhaust pipe and the door handle\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e337\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-roof\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eWater over the vehicle\u0026rsquo;s roof\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e276\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePer-none\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDry surface or the water depth less than 5 cm\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e209\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePer-leg\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eWater depth more than 5 cm and below the waist of a person\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e561\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePer-waist\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eWater depth above a person's waist\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e145\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eTotal number of objects\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2095\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003ctfoot\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"3\"\u003eNote: Per is short for person\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tfoot\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n\u003ch2\u003e2.2. Target detection algorithm\u003c/h2\u003e\n\u003cp\u003eThe \u0026ldquo;you only look once\u0026rdquo; (YOLO) algorithm was used in this study to extract features and detect objects in the flood images using a deep convolutional neural network. YOLO is a commonly used model for target detection from images because of its fast speed of image recognition in a one-stage algorithm, which performs feature extraction, target classification, and target regression through the same convolutional neural network (CNN) (Redmon et al., \u003cspan class=\"CitationRef\"\u003e2016\u003c/span\u003e). Compared with a two-stage algorithm, such as the Faster-RCNN with multiple steps for generating the candidate area of the input image and extracting features for the image recognition through different neural networks, YOLO skips the generation of a candidate area and completes the detection in one step. This provides the YOLO algorithm with a relatively higher calculation speed and a sufficient precision (Zou et al., \u003cspan class=\"CitationRef\"\u003e2019\u003c/span\u003e). This fits the requirements of this study to detect changes of flood depths in a real-time manner.\u003c/p\u003e\n\u003cp\u003eYOLO conducts object detection by calculating the probability that a bounding box in an image belongs to a pre-set category using a CNN and a regression method. First, YOLO takes the entire image as the input of the network after dividing the image into a 7\u0026times;7 grid. Then, each grid is framed by two alternative bounding boxes with \u003cstrong\u003en\u003c/strong\u003e categories for each box to be classified. There are five parameters of each bounding box, namely, the horizontal coordinate of the box central point (\u003cstrong\u003ex\u003c/strong\u003e), the vertical coordinate of the central point (\u003cstrong\u003ey\u003c/strong\u003e), the width (\u003cstrong\u003ew\u003c/strong\u003e) and height (\u003cstrong\u003eh\u003c/strong\u003e) of the box, and the probability that the bounding box contains the target (\u003cstrong\u003ec\u003c/strong\u003e\u0026isin;[0,1]). Based on this, a variable named YOLO Head is obtained as a tensor with a shape of (7, 7, 2 \u0026times; 5\u0026thinsp;+\u0026thinsp;n). For example, if a dataset contains 20 categories, the output of the model will be a matrix of (7, 7, 30). In this matrix, the first two dimensions with a shape of (7 \u0026times; 7) represent the number of feature grids extracted. The last dimension with a size of 30 is the result of multiplying the number of bounding boxes and the number of box parameters plus the number of probability values of the box belonging to each category, which is 2\u0026times;5\u0026thinsp;+\u0026thinsp;20.\u003c/p\u003e\n\u003cp\u003eThe YOLO algorithm was first established and updated to its third generation by Redmon J (Redmon \u0026amp; Farhadi, \u003cspan class=\"CitationRef\"\u003e2018\u003c/span\u003e). The fourth generation, YOLOv4, was proposed by Alexey Bochkovskiy in 2020 (Bochkovskiy et al., \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e). YOLOv4 was used in this study. It is composed of three parts: the backbone, feature enhancement, and prediction. In the backbone, the model extracts three feature layers from the image using five residual networks after adjusting the image to a fixed resolution using CSPDarknet53 (C.-Y. Wang et al., \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e). The three feature layers are then transferred to the feature enhancement part, the neck, to obtain three prediction layers, which comprise the YOLO head, for the prediction part. The YOLO head is used for final classification and recognition of the image.\u003c/p\u003e\n\u003cp\u003eIn this study, three feature layers with the shapes (13, 13), (26, 26), and (52, 52) were adopted to extract the feature from an image with the shape (416, 416) through pooling technology of the CNN, as shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e. Each grid in the feature layer was framed by three alternative bounding boxes with seven categories. Each bounding box has five parameters, as previously stated. Therefore, the sizes of the output YOLO heads are (13, 13, 22), (26, 26, 22), and (52, 52, 22), in which the number of the third dimension with a size of 22 is the result of multiplying the number of bounding boxes with the number of box parameters plus the number of probability values that the box belongs to each category, which is 3\u0026times;5\u0026thinsp;+\u0026thinsp;7.\u003c/p\u003e\n\u003cp\u003eIn addition, YOLOv4 adopts the activation function Mish (Misra, \u003cspan class=\"CitationRef\"\u003e2019\u003c/span\u003e) instead of LeakyRelu (Maas et al., \u003cspan class=\"CitationRef\"\u003e2013\u003c/span\u003e) used in YOLOv3. Compared with LeakyRelu, the Mish function has a lower bound; so, it helps the model to strengthen the regularization effect. Furthermore, the non-monotonic function of Mish preserves small negative values, thus stabilizing the network gradient flow. Mish is infinitely continuous with smooth curves, which also helps to improve the quality of the results. Although the Mish function uses more computing resources than LeakyRelu functions, it improves the accuracy of the model. The activation functions are shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\n\u003ch2\u003e2.3. Model training and validation\u003c/h2\u003e\n\u003cp\u003eIn this study, the input data were randomly distributed into a training set (90% of data) and a validation set (10% of data). YOLOv4 was trained by the method of transfer learning, in which the initial parameters, including the initial weights of the CNN, were transferred from the existing dataset VOC2007 (Everingham et al., \u003cspan class=\"CitationRef\"\u003e2010\u003c/span\u003e). Transfer learning is a practical scheme in the training process, which can use generic features from other larger training datasets. This reduces the time required for training the model. The VOC2007 dataset consists of about 500,000 annotated images that are divided into 20 object classes. The dataset is large enough to provide the CNN of this study with initial weights.\u003c/p\u003e\n\u003cp\u003eTwo hundred epochs were adopted for model training in this study. In this process, frozen learning with a learning rate of 1e-3 was used in the first 90 epochs, and regular learning with a learning rate of 1e-4 was used in the rest of the epochs. Frozen learning did not change the initial parameters of the networks in the backbone part, and only generated minor adjustments to other networks during the training (Brock et al., \u003cspan class=\"CitationRef\"\u003e2017\u003c/span\u003e). In contrast, in regular learning, the parameters of all the networks were adjusted in the model. Frozen learning can reduce the time required for training by using the existed network parameters obtained from a previous dataset such as VOC2007. However, the model would not converge after a few epochs because the backbone networks were not updated under frozen learning. Therefore, regular learning was adopted subsequently in the training process for further convergence of the model. In addition, a lower learning rate was also adopted in the regular learning to avoid overfitting.\u003c/p\u003e\n\u003cp\u003eThe convergence curve of the model training is shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. The decrease of training loss became stable after 70 epochs of training using frozen learning due to the unchanging structure of the network. Then, the training loss was further reduced without frozen learning until 150 epochs.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003e2.4. Indicators for performance assessment\u003c/h2\u003e\n\u003cp\u003eMean average precision (mAP) was used in this study to evaluate the accuracy of flood image recognition. It is a metric widely applied in object detection and classification (Zou et al., \u003cspan class=\"CitationRef\"\u003e2019\u003c/span\u003e). The mAP was calculated using precision and recall indicators of the modeling results, which are given by\u003c/p\u003e\n\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ1\" class=\"mathdisplay\"\u003e$$Precision= \\frac{TP}{TP+FP}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ2\" class=\"mathdisplay\"\u003e$$Recall= \\frac{TP}{TP+FN}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003ewhere TP (true positive) represents the number of samples that are correctly recognized or classified, FP (false positive) represents the number of samples that are recognized incorrectly, and FN (false negative) represents the number of unrecognized results. Precision is the ratio of correctly predicted positive results to the total predicted positive results. Recall is the ratio of correctly predicted positive results to the total number of actual positive samples.\u003c/p\u003e\n\u003cp\u003eGiven the above indicators, the mAP can be calculated as follows:\u003c/p\u003e\n\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ3\" class=\"mathdisplay\"\u003e$$mAP= \\frac{{\\sum }_{1}^{n}AP}{n}$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ4\" class=\"mathdisplay\"\u003e$$AP= {\\int }_{0}^{1}p\\left(r\\right)dr$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003ewhere \u003cem\u003eAP\u003c/em\u003e (average precision) represents the area under the precision-recall curve of the prediction results referring to each category, \u003cem\u003en\u003c/em\u003e is the total number of categories, which is seven in this study, and \u003cem\u003er\u003c/em\u003e is the value of recall from Eq.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. \u003cem\u003ep\u003c/em\u003e (r) represents the Precision-Recall curve. The y-axis of \u003cem\u003ep\u003c/em\u003e(\u003cem\u003er\u003c/em\u003e) is the maximum value of the corresponding precision when the value of recall increases from 0 to 1 at a given interval, such as 0.1.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"3. Results","content":"\u003cp\u003eThe results of flood recognition from the street images are shown in \u003cstrong\u003eFig.\u0026nbsp;4\u003c/strong\u003e. As shown in the figure, the reference objects, the vehicles and pedestrians, are framed by bounding boxes with the identified category of the water depth presented at the top of each box. In this manner, the flood water depth can be extracted by real-time image recognition using street photos or surveillance videos from traffic cameras.\u003c/p\u003e\n\u003cp\u003eThe precision, recall, and AP of the flood recognition from 1177 images are shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. The mAP of the model prediction reached 89.29%, demonstrating a satisfactory performance of the model for identifying the submerged depths of the reference objects. The AP of the predictions in the categories of Car-roof and Car-pipe was 98.81% and 98.63%, respectively. This indicated that the best performance of the flood depth recognition from the image dataset in this study was for the identification of water levels reaching the roof and the exhaust pipe of the vehicles. The AP of the predictions for the categories of Car-handle and Per-leg was 92.47% and 90.44%, respectively, representing the second-best performances of the flood depth recognition.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eResults of flood depth recognition categories\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eCategory\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePrecision\u003c/p\u003e\n\u003cp\u003e(%)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eRecall\u003c/p\u003e\n\u003cp\u003e(%)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAP\u003csub\u003e50\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e(%)\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-roof\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e95.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e95.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e98.81\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-pipe\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e87.50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e93.33\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e98.63\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-handle\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e82.35\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e90.32\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e92.47\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePer-leg\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e85.94\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e82.09\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e90.44\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCar-none\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e93.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.43\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e87.60\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePer-waist\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e75.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e64.29\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e83.84\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePer-none\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e82.35\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e51.85\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.21\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"3\" align=\"left\"\u003e\n\u003cp\u003emAP\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e89.29\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003ctfoot\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"4\"\u003eNote: Per is short for person.\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tfoot\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe AP of seven categories and the number of labeled objects in each category are shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. There were 2095 objects in the dataset, including 1885 objects used in model training and 210 objects used for validation. The category Per-leg had the highest number of objects, with 494 objects in the training set and 67 objects in the validation set. This meant that more images showing floods with a water depth reaching the legs of pedestrians were used in this study compared to images showing other categories. Compared with the Per-leg category, the Car-pipe category had a smaller number of objects, with 342 objects in the training set and 30 objects in the validation set. However, the AP of the Car-pipe was 98.63% and was higher than the AP of the Per-leg category, which was only 90.44%. In addition, the Per-waist category had the least number of objects but the second lowest AP. The AP was not only affected by the number of objects but also by the feature extractions of the objects. This founding was supported by other studies using the dataset VOC2007, in which the objects were classified into 20 categories with various APs. In previous studies, objects with relatively clearer contours and more regular shapes, such as cars and trains, achieved higher recognition APs. Conversely, irregularly shaped objects, such as humans, cats, and dogs, had lower APs (X. Wang et al., \u003cspan class=\"CitationRef\"\u003e2013\u003c/span\u003e). This is consistent with the results of this study. When a car was used as the reference object of the flood submerged depth, the AP of the flood recognition was higher than when using the human body as the inundated reference. This is because the objects with clear contours and regular shapes more easily have features extracted in image recognition.\u003c/p\u003e\n\u003cp\u003eFigure \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e shows the precision and recall of water depth predictions in the categories Car-roof and Per-none, which had the best and the worst performances of the prediction, respectively, as shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. In Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, the x-axis was a pre-defined confidence threshold of the bounding box containing the object in the image to be recognized. The range of the confidence threshold was from 0.1 to 1.0 in this study. The object in the bounding box could be recognized and categorized only when the box confidence was higher than the pre-defined threshold. Therefore, the precision increased with an increasing threshold. For example, the predicted category of the bounding box was highly accurate when precisely matching its annotation if the confidence threshold was close to 1. However, in this case, most of the bounding boxes failed to pass the high threshold, causing a very low value of the recall. The recall decreased when the threshold increased.\u003c/p\u003e\n\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, the recall of the water depth prediction corresponding to the Car-roof category was 95.00%, which was almost double the 51.85% recall corresponding to the Per-none category when the confidence threshold was 0.5. This is because the posture of the vehicle tends to be more consistent than that of the pedestrian, who might have multiple postures, including standing, bending over, and squatting. The variations of the object posture also cause difficulty with recognition. This finding was consistent with a previous study (Giannakeris et al., \u003cspan class=\"CitationRef\"\u003e2018\u003c/span\u003e), which found that the recall of recognizing a fire in an image was lower than that of recognizing a flood in an image, because the variable shape of the fire made it more difficult to be recognized accurately.\u003c/p\u003e"},{"header":"4. Discussion","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n\u003ch2\u003e4.1. Performances of frozen learning and regular learning\u003c/h2\u003e\n\u003cp\u003eThe performances of frozen learning and regular learning on inundated depth recognition were analyzed by the changes of the mAP and the loss in the learning process using the different learning methods. The mAP and the loss are shown at every 50 epochs during the learning process in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e. At 200 epochs, the mAP of regular learning was higher than that of frozen learning. The loss of regular learning was less than that of frozen learning. This shows that the performance of regular learning was better than that of frozen learning. In addition, a further improvement of the model performance was found by using a combined learning method, which used frozen learning in the first 90 epochs and regular learning in the rest of the epochs.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003emAP and loss across different epochs using frozen learning and regular learning\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eEpochs\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eFrozen learning\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eRegular learning\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eCombined learning\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003emAP (%)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLoss\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003emAP (%)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLoss\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003emAP (%)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLoss\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e74.94\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e43.44\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e69.44\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e41.40\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e78.33\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e43.62\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.28\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e43.48\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.34\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e38.61\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e79.92\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e33.82\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e150\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e43.42\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.51\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e36.41\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e83.14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e28.36\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e200\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.35\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e42.47\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e75. 57\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e34.61\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e83.35\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e27.81\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe changes in the mAP and the loss are shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e. The mAP of the frozen learning was higher than that of the regular learning within the first 100 epochs, after which the mAP of the regular learning continued to rise and increased more than that of frozen learning. Moreover, the mAP of the combined learning was higher than that of the frozen learning and regular learning after 50 epochs. In addition, the loss of the combined learning decreased and became less than that of the frozen learning and regular learning after about 70 epochs. Both the mAP and the loss of the combined learning were stable without further decreases or increases at 200 epochs, showing that there was no overfitting in the learning process. Given this, in the context of transfer learning, the frozen learning with fixed parameters transferred from the existing dataset had a better performance at the early stage of the learning process. However, regular learning, which altered the transferred parameters, had a better performance at the end of the learning process. More importantly, the performance of the combined learning was better than that of using frozen learning and regular learning separately.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n\u003ch2\u003e4.2. Impacts of training technologies\u003c/h2\u003e\n\u003cp\u003eThree technologies for image processing and model training were applied in the deep learning process to analyze the impacts of these technologies on improving the performance of image recognition on the flood depths. These technologies include Mosaic (Bochkovskiy et al., \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e), label smoothing (Szegedy et al., \u003cspan class=\"CitationRef\"\u003e2016\u003c/span\u003e), and cosine annealing (Loshchilov \u0026amp; Hutter, \u003cspan class=\"CitationRef\"\u003e2016\u003c/span\u003e), which are currently widely used in image recognition. Mosaic is a technique for image augmentation in which four images are combined through cropping and blending, making full use of the data (Hao \u0026amp; Zhili, \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e). Label smoothing is a regularization technique that introduces noise for the labels (M\u0026uuml;ller et al., \u003cspan class=\"CitationRef\"\u003e2019\u003c/span\u003e). Cosine annealing is a learning rate variation strategy characterized by periodic drops and resets of the learning rate during the training process. Cosine annealing helps the training process escape a local optimal solution by resetting a larger learning rate (Gotmare et al., \u003cspan class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eThe applicability of different technologies was analyzed by comparing the model mAP across application scenarios using eight individual or combined technologies, as shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e. In scenario 1, none of above technologies was applied, and the mAP of the modeling under this scenario was only 67.58%. After using Mosaic, label smoothing, and cosine annealing separately, the mAP increased to 72.14%, 70.94%, and 70.48%, respectively. The mAP was improved the most in scenario 2 compared with scenario 1 using Mosaic. However, in scenarios 5 and 6, in which Mosaic was jointly applied with label smoothing and cosine annealing, respectively, the mAP was 71.92% and 67.53%, respectively. The mAP was reduced when Mosaic was jointly applied with label smoothing and cosine annealing compared with Mosaic used alone. Mosaic had the best performance for improving the mAP of flood depth recognition in this study. This was consistent with previous studies on YOLOv4 in which Mosaic improved the accuracy of image recognition (Bochkovskiy A, et, 2020).\u003c/p\u003e\n\u003cp\u003eIn contrast, the lowest mAP of 66.29% was from scenario 7 in which the label smoothing and cosine annealing were jointly applied without Mosaic. In addition, the second lowest mAP of 67.53% resulted from scenario 6 in which Mosaic and cosine annealing were jointly applied. It was found that, in this study, cosine annealing was not suitable to be used in combination with Mosaic and label smoothing because both cosine annealing with Mosaic and cosine annealing with label smoothing had lower performances than did cosine annealing alone.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab4\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003emAP using different training technologies\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eScenario\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eMosaic\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eLabel Smoothing\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eCosine Annealing\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003emAP (%)\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e67.58\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e74.12\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e70.94\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e70.48\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.92\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e67.53\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e66.29\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026radic;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.48\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n\u003ch2\u003e4.3. Impacts of activation functions\u003c/h2\u003e\n\u003cp\u003eThe activation function can affect the performance of a deep learning model significantly in terms of the accuracy and speed of calculations. The mAP and the frames per second (FPS) were calculated in this study to evaluate the impacts of two activation functions, Mish and LeakyRelu, as shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. A higher FPS indicates a faster speed of image recognition. The results are shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. The YOLOv4 model using Mish generally had higher mAP results than did LeakyRelu when recognizing flood depth in images. However, the FPS of the model with Mish was 30.49, which was less than that of LeakyRelu. This indicates that Mish performed at a slower speed in image recognition, with about five fewer images per second compared to LeakyRelu in this study. This is consistent with previous studies that showed that Mish had a relatively greater loss in detection speed compared to LeakyRelu, but the accuracy was better than that of LeakyRelu (C.-Y. Wang et al., \u003cspan class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab5\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003emAP (%) and FPS using LeakyRelu and Mish functions\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n\u003cth colspan=\"10\" align=\"left\"\u003e\n\u003cp\u003emAP\u003c/p\u003e\n\u003c/th\u003e\n\u003cth rowspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eFPS\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eEpochs\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e20\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e40\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e60\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e80\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e100\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e120\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e140\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e160\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e180\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e200\u003c/strong\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMish\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e65.56\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e70.35\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e68.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e73.99\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.91\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e69.34\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e75.53\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e74.56\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e30.49\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLeakyRelu\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e55.04\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e67.67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.21\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.49\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.06\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.98\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e72.06\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e34.92\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eAn image recognition method for the detection of urban flood inundation was established in this study by applying the deep learning network YOLOv4 and the transfer training technique. Vehicles and pedestrians submerged by flood water in the images were adopted as reference objects for flood water depth recognition. After training of 200 epochs, the model classified images with a certain flood water level into four categories, namely, Car-none, Car-pipe, Car-handle, and Car-roof, representing that the flood water depth in the image was less than 5 cm, or that the flood level reached the exhaust pipe, the door, or the roof of a vehicle, respectively, with vehicles as the reference. Similarly, the model also classified the flood image into three categories, namely, Per-none, Per-leg, and Per-waist, representing that the flood water depth in the image was less than 5 cm, or that the flood level reached the leg or waist of a person, respectively. The results showed that the mAP of the proposed method reached 89.29%, with a recognition speed of 30.49 FPS.\u003c/p\u003e \u003cp\u003eIt was also found that the accuracy of flood depth recognition by YOLOv4 was affected by the reference object submerged by the flood. The accuracy of using a vehicle as the reference object was significantly higher than that of using a person as the reference object. This is because a vehicle generally has a regular shape and a fixed position, which makes it easier to identify. In addition, the impact on the flood depth recognition of various transfer learning techniques, image processing methods, networking training tricks, and activation functions of the CNN was discussed in the study. It was found that, in the context of flood depth recognition from images using YOLOv4, (1) the performance of the combined use of frozen learning and regular learning was better than that of using the two learning techniques separately; (2) image augmentation using Mosaic technology effectively improved the accuracy of recognition, and using Mosaic alone was better than jointly using Mosaic, label smoothing, and cosine annealing in this study; and (3) the model using the Mish activation function had a better performance than that using the LeakyRelu activation function.\u003c/p\u003e \u003cp\u003eThe proposed YOLOv4 model has the advantages of high accuracy and fast speed in recognizing the flood water depth from images using vehicles and pedestrians as reference objects. This enables the method to use the images or video data provided by existing traffic cameras to identify the flood depth during a storm, thereby obtaining on-site, real-time, and continuous water depth data. This method does not require the additional installation of a water gauge or a water depth sensing device and thus has a low cost and a high usability in practice.\u003c/p\u003e \u003cp\u003eThe flood depth data obtained by this method can be used for the verification and calibration of urban hydrological models and can also provide real-time inundation data for early warning of urban floods, making up for the insufficient coverage of flood inundation depth monitoring sites in current urban hydrological research. The proposed method could be improved in terms of recognition accuracy and reliability. For example, the flood depth was predicted in the form of a categorical grade, such as the water level to the handle of the car or to the leg of a person, rather than the specific water depth value. In addition, the method relies on reference objects, and it is difficult to identify the water depth without a reference object on the road surface.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003cstrong\u003e:\u003c/strong\u003eThis study was financially supported by the National Natural Science Foundation of China (U2040206),(52179009) and (51909035).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e The authors have no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Pengcheng Zhong, Yueyi Liu, Hang Zheng, Jianshi Zhao. The first draft of the manuscript was written by Pengcheng Zhong and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eHammond, M. J., Chen, A. S., Djordjević, S., Butler, D., \u0026amp; Mark, O. (2015). Urban flood impact assessment: A state-of-the-art review. Urban Water Journal, 12(1), 14-29.\u003c/li\u003e\n\u003cli\u003eNkwunonwo, U. C., Whitworth, M., \u0026amp; Baily, B. (2020). A review of the current status of flood modelling for urban flood risk management in the developing countries. Scientific African, 7, e00269.\u003c/li\u003e\n\u003cli\u003eChen, Y., Zhou, H., Zhang, H., Du, G., \u0026amp; Zhou, J. (2015). Urban flood risk warning under rapid urbanization. Environmental research, 139, 3-10.\u003c/li\u003e\n\u003cli\u003eYin, J., Ye, M., Yin, Z., \u0026amp; Xu, S. (2015). A review of advances in urban flood risk analysis over China. Stochastic Environmental Research and Risk Assessment, 29(3), 1063-1070.\u003c/li\u003e\n\u003cli\u003eJamali, B., L\u0026ouml;we, R., Bach, P. M., Urich, C., Arnbjerg-Nielsen, K., \u0026amp; Deletic, A. (2018). A rapid urban flood inundation and damage assessment model. Journal of Hydrology, 564, 1085-1098.\u003c/li\u003e\n\u003cli\u003eWang, H., Hu, Y., Guo, Y., Wu, Z., \u0026amp; Yan, D. (2022). Urban flood forecasting based on the coupling of numerical weather model and stormwater model: A case study of Zhengzhou city. Journal of Hydrology: Regional Studies, 39, 100985.\u003c/li\u003e\n\u003cli\u003eDisaster Investigation Team of the State Council, China. (2022). Investigation report on the \u0026quot;7.20\u0026quot; heavy rainstorm disaster in Zhengzhou, Henan, China. \u003c/li\u003e\n\u003cli\u003eFletcher, T. D., Andrieu, H., \u0026amp; Hamel, P. (2013). Understanding, management and modelling of urban hydrology and its consequences for receiving waters: A state of the art. Advances in Water Resources, 51, 261-279.\u003c/li\u003e\n\u003cli\u003eIchiba, A., Gires, A., Tchiguirinskaia, I., Schertzer, D., Bompard, P., \u0026amp; Ten Veldhuis, M. C. (2018). Scale effect challenges in urban hydrology highlighted with a distributed hydrological model. Hydrology and Earth System Sciences, 22(1), 331-350.\u003c/li\u003e\n\u003cli\u003eRubinato, M., Nichols, A., Peng, Y., Zhang, J. M., Lashford, C., Cai, Y. P., ... \u0026amp; Tait, S. (2019). Urban and river flooding: Comparison of flood risk management approaches in the UK and China and an assessment of future knowledge needs. Water science and engineering, 12(4), 274-283.\u003c/li\u003e\n\u003cli\u003eCristiano, E., ten Veldhuis, M. C., Wright, D. B., Smith, J. A., \u0026amp; van de Giesen, N. (2019). The influence of rainfall and catchment critical scales on urban hydrological response sensitivity. Water Resources Research, 55(4), 3375-3390.\u003c/li\u003e\n\u003cli\u003eKourtis, I. M., \u0026amp; Tsihrintzis, V. A. (2021). Adaptation of urban drainage networks to climate change: A review. Science of The Total Environment, 771, 145431.\u003c/li\u003e\n\u003cli\u003eSohn, W., Kim, J. H., Li, M. H., Brown, R. D., \u0026amp; Jaber, F. H. (2020). How does increasing impervious surfaces affect urban flooding in response to climate variability?. Ecological Indicators, 118, 106774.\u003c/li\u003e\n\u003cli\u003ede Mello Silva, C., \u0026amp; da Silva, G. B. L. (2020). Cumulative effect of the disconnection of impervious areas within residential lots on runoff generation and temporal patterns in a small urban area. Journal of environmental management, 253, 109719.\u003c/li\u003e\n\u003cli\u003eLee, Y., \u0026amp; Brody, S. D. (2018). Examining the impact of land use on flood losses in Seoul, Korea. Land use policy, 70, 500-509.\u003c/li\u003e\n\u003cli\u003eYu, H., Zhao, Y., Fu, Y., \u0026amp; Li, L. (2018). Spatiotemporal variance assessment of urban rainstorm waterlogging affected by impervious surface expansion: A case study of Guangzhou, China. Sustainability, 10(10), 3761.\u003c/li\u003e\n\u003cli\u003eFeng, B., Zhang, Y., \u0026amp; Bourke, R. (2021). Urbanization impacts on flood risks based on urban growth data and coupled flood models. Natural Hazards, 106(1), 613-627.\u003c/li\u003e\n\u003cli\u003eJamali, B., L\u0026ouml;we, R., Bach, P. M., Urich, C., Arnbjerg-Nielsen, K., \u0026amp; Deletic, A. (2018). A rapid urban flood inundation and damage assessment model. Journal of Hydrology, 564, 1085-1098.\u003c/li\u003e\n\u003cli\u003eSingh, P., Sinha, V. S. P., Vijhani, A., \u0026amp; Pahuja, N. (2018). Vulnerability assessment of urban road network from urban flood. International journal of disaster risk reduction, 28, 237-250.\u003c/li\u003e\n\u003cli\u003eMignot, E., Li, X., \u0026amp; Dewals, B. (2019). Experimental modelling of urban flooding: A review. Journal of Hydrology, 568, 334-342.\u003c/li\u003e\n\u003cli\u003eBulti, D. T., \u0026amp; Abebe, B. G. (2020). A review of flood modeling methods for urban pluvial flood application. Modeling earth systems and environment, 6(3), 1293-1302.\u003c/li\u003e\n\u003cli\u003eGuo, K., Guan, M., \u0026amp; Yu, D. (2021). Urban surface water flood modelling\u0026ndash;a comprehensive review of current models and future challenges. Hydrology and Earth System Sciences, 25(5), 2843-2860.\u003c/li\u003e\n\u003cli\u003eBabaei, S., Ghazavi, R., \u0026amp; Erfanian, M. (2018). Urban flood simulation and prioritization of critical urban sub-catchments using SWMM model and PROMETHEE II approach. Physics and Chemistry of the Earth, Parts A/B/C, 105, 3-11.\u003c/li\u003e\n\u003cli\u003eBai, Y., Zhao, N., Zhang, R., \u0026amp; Zeng, X. (2018). Storm water management of low impact development in urban areas based on SWMM. Water, 11(1), 33.\u003c/li\u003e\n\u003cli\u003eLi, J., Zhang, B., Mu, C., \u0026amp; Chen, L. (2018). Simulation of the hydrological and environmental effects of a sponge city based on MIKE FLOOD. Environmental earth sciences, 77(2), 1-16.\u003c/li\u003e\n\u003cli\u003eNigussie, T. A., \u0026amp; Altunkaynak, A. (2019). Modeling the effect of urbanization on flood risk in Ayamama Watershed, Istanbul, Turkey, using the MIKE 21 FM model. Natural Hazards, 99(2), 1031-1047.\u003c/li\u003e\n\u003cli\u003eRangari, V. A., Umamahesh, N. V., \u0026amp; Bhatt, C. M. (2019). Assessment of inundation risk in urban floods using HEC RAS 2D. Modeling Earth Systems and Environment, 5(4), 1839-1851.\u003c/li\u003e\n\u003cli\u003eSeenu, P. Z., Venkata Rathnam, E., \u0026amp; Jayakumar, K. V. (2020). Visualisation of urban flood inundation using SWMM and 4D GIS. Spatial Information Research, 28(4), 459-467.\u003c/li\u003e\n\u003cli\u003eArshad, B., Ogie, R., Barthelemy, J., Pradhan, B., Verstaevel, N., \u0026amp; Perez, P. (2019). Computer vision and IoT-based sensors in flood monitoring and mapping: A systematic review. Sensors, 19(22), 5012.\u003c/li\u003e\n\u003cli\u003eMoy de Vitry, M., Kramer, S., Wegner, J. D., \u0026amp; Leit\u0026atilde;o, J. P. (2019). Scalable flood level trend monitoring with surveillance cameras using a deep convolutional neural network. Hydrology and Earth System Sciences, 23(11), 4621-4634.\u003c/li\u003e\n\u003cli\u003eLi, Y., Martinis, S., \u0026amp; Wieland, M. (2019). Urban flood mapping with an active self-learning convolutional neural network based on TerraSAR-X intensity and interferometric coherence. ISPRS Journal of Photogrammetry and Remote Sensing, 152, 178-191.\u003c/li\u003e\n\u003cli\u003eWang, R. Q., Mao, H., Wang, Y., Rae, C., \u0026amp; Shaw, W. (2018). Hyper-resolution monitoring of urban flooding with social media and crowdsourcing data. Computers \u0026amp; Geosciences, 111, 139-147.\u003c/li\u003e\n\u003cli\u003eWang, Y., Chen, A. S., Fu, G., Djordjević, S., Zhang, C., \u0026amp; Savić, D. A. (2018). An integrated framework for high-resolution urban flood modelling considering multiple information sources and urban features. Environmental modelling \u0026amp; software, 107, 85-95.\u003c/li\u003e\n\u003cli\u003eLiu, J., Shao, W., Xiang, C., Mei, C., \u0026amp; Li, Z. (2020). Uncertainties of urban flood modeling: Influence of parameters for different underlying surfaces. Environmental research, 182, 108929.\u003c/li\u003e\n\u003cli\u003eCohen, S., Brakenridge, G. R., Kettner, A., Bates, B., Nelson, J., McDonald, R., ... \u0026amp; Zhang, J. (2018). Estimating floodwater depths from flood inundation maps and topography. JAWRA Journal of the American Water Resources Association, 54(4), 847-858.\u003c/li\u003e\n\u003cli\u003eShen, X., Wang, D., Mao, K., Anagnostou, E., \u0026amp; Hong, Y. (2019). Inundation extent mapping by synthetic aperture radar: A review. Remote Sensing, 11(7), 879.\u003c/li\u003e\n\u003cli\u003eRong, Y., Zhang, T., Zheng, Y., Hu, C., Peng, L., \u0026amp; Feng, P. (2020). Three-dimensional urban flood inundation simulation based on digital aerial photogrammetry. Journal of Hydrology, 584, 124308.\u003c/li\u003e\n\u003cli\u003eIlieva, R. T., \u0026amp; McPhearson, T. (2018). Social-media data for urban sustainability. Nature Sustainability, 1(10), 553-565.\u003c/li\u003e\n\u003cli\u003eScotti, V., Giannini, M., \u0026amp; Cioffi, F. (2020). Enhanced flood mapping using synthetic aperture radar (SAR) images, hydraulic modelling, and social media: A case study of Hurricane Harvey (Houston, TX). Journal of Flood Risk Management, 13(4), e12647.\u003c/li\u003e\n\u003cli\u003ede Vitry, M. M., \u0026amp; Leit\u0026atilde;o, J. P. (2020). The potential of proxy water level measurements for calibrating urban pluvial flood models. Water Research, 175, 115669.\u003c/li\u003e\n\u003cli\u003eKankanamge, N., Yigitcanlar, T., Goonetilleke, A., \u0026amp; Kamruzzaman, M. (2020). Determining disaster severity through social media analysis: Testing the methodology with South East Queensland Flood tweets. International journal of disaster risk reduction, 42, 101360.\u003c/li\u003e\n\u003cli\u003eJiang, J., Liu, J., Cheng, C., Huang, J., \u0026amp; Xue, A. (2019). Automatic estimation of urban waterlogging depths from video images based on ubiquitous reference objects. Remote Sensing, 11(5), 587.\u003c/li\u003e\n\u003cli\u003eBhola, P. K., Nair, B. B., Leandro, J., Rao, S. N., \u0026amp; Disse, M. (2019). Flood inundation forecasts using validation data generated with the assistance of computer vision. Journal of Hydroinformatics, 21(2), 240-256.\u003c/li\u003e\n\u003cli\u003ePark, S., Baek, F., Sohn, J., \u0026amp; Kim, H. (2021). Computer vision\u0026ndash;based estimation of flood depth in flooded-vehicle images. Journal of Computing in Civil Engineering, 35(2), 04020072.\u003c/li\u003e\n\u003cli\u003eMunasinghe, D., Cohen, S., Huang, Y. F., Tsang, Y. P., Zhang, J., \u0026amp; Fang, Z. (2018). Intercomparison of satellite remote sensing‐based flood inundation mapping techniques. JAWRA Journal of the American Water Resources Association, 54(4), 834-846.\u003c/li\u003e\n\u003cli\u003eDeVries, B., Huang, C., Armston, J., Huang, W., Jones, J. W., \u0026amp; Lang, M. W. (2020). Rapid and robust monitoring of flood events using Sentinel-1 and Landsat data on the Google Earth Engine. Remote Sensing of Environment, 240, 111664.\u003c/li\u003e\n\u003cli\u003eHashemi-Beni, L., \u0026amp; Gebrehiwot, A. A. (2021). Flood extent mapping: An integrated method using deep learning and region growing using UAV optical data. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 2127-2135.\u003c/li\u003e\n\u003cli\u003eZeng, Z., Gan, Y., Kettner, A. J., Yang, Q., Zeng, C., Brakenridge, G. R., \u0026amp; Hong, Y. (2020). Towards high resolution flood monitoring: An integrated methodology using passive microwave brightness temperatures and Sentinel synthetic aperture radar imagery. Journal of Hydrology, 582, 124377.\u003c/li\u003e\n\u003cli\u003eDubey, A. K., Kumar, P., Chembolu, V., Dutta, S., Singh, R. P., \u0026amp; Rajawat, A. S. (2021). Flood modeling of a large transboundary river using WRF-Hydro and microwave remote sensing. Journal of Hydrology, 598, 126391.\u003c/li\u003e\n\u003cli\u003eGalantowicz, J. F., \u0026amp; Picton, J. (2021). Flood Mapping with Passive Microwave Remote Sensing: Current Capabilities and Directions for Future Development. In Earth Observation for Flood Applications (pp. 39-60). Elsevier.Feng et al., 2020; Songchon et al., 2021.\u003c/li\u003e\n\u003cli\u003eFeng, Y., Brenner, C., \u0026amp; Sester, M. (2020). Flood severity mapping from Volunteered Geographic Information by interpreting water level from images containing people: A case study of Hurricane Harvey. ISPRS Journal of Photogrammetry and Remote Sensing, 169, 301-319.\u003c/li\u003e\n\u003cli\u003eSongchon, C., Wright, G., \u0026amp; Beevers, L. (2021). Quality assessment of crowdsourced social media data for urban flood management. Computers, Environment and Urban Systems, 90, 101690.\u003c/li\u003e\n\u003cli\u003eBarz, B., Schr\u0026ouml;ter, K., M\u0026uuml;nch, M., Yang, B., Unger, A., Dransch, D., \u0026amp; Denzler, J. (2019). Enhancing flood impact analysis using interactive retrieval of social media images. arXiv preprint arXiv:1908.03361.\u003c/li\u003e\n\u003cli\u003eB\u0026icirc;nă, D., Vlad, G. A., Onose, C., \u0026amp; Cercel, D. C. (2019, October). Flood severity estimation in news articles using deep learning approaches. In Proceedings of the MediaEval 2019 Workshop, Sophia Antipolis, France (pp. 27-29).\u003c/li\u003e\n\u003cli\u003eLv, Y., Gao, W., Yang, C., \u0026amp; Wang, N. (2018). Inundated areas extraction based on raindrop photometric model (RPM) in surveillance video. Water, 10(10), 1332.\u003c/li\u003e\n\u003cli\u003eJiang, J., Qin, C. Z., Yu, J., Cheng, C., Liu, J., \u0026amp; Huang, J. (2020). Obtaining urban waterlogging depths from video images using synthetic image data. Remote Sensing, 12(6), 1014.\u003c/li\u003e\n\u003cli\u003eBasnyat, B., Roy, N., \u0026amp; Gangopadhyay, A. (2018, June). A flash flood categorization system using scene-text recognition. In 2018 IEEE International Conference on Smart Computing (SMARTCOMP) (pp. 147-154). IEEE.\u003c/li\u003e\n\u003cli\u003eYang, HC., Wang, CY. \u0026amp; Yang, JX. (2014). Applying image recording and identification for measuring water stages to prevent flood hazards. Nat Hazards 74, 737\u0026ndash;754.\u003c/li\u003e\n\u003cli\u003eEveringham, M., van Gool, L., Williams, C. K. I., Winn, J., \u0026amp; Zisserman, A. (2010). The pascal visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303\u0026ndash;338. \u003c/li\u003e\n\u003cli\u003eGauen, K., Dailey, R., Laiman, J., Zi, Y., Asokan, N., Lu, Y.-H., Thiruvathukal, G. K., Shyu, M.-L., \u0026amp; Chen, S.-C. (2017). Comparison of visual datasets for machine learning. 2017 IEEE International Conference on Information Reuse and Integration (IRI), 346\u0026ndash;355.\u003c/li\u003e\n\u003cli\u003eYakovlev, A., \u0026amp; Lisovychenko, O. (2020). An approach for image annotation automatization for artificial intelligence models learning. Адаптивні Системи Автоматичного Управління, 1(36), 32\u0026ndash;40.\u003c/li\u003e\n\u003cli\u003eRedmon, J., Divvala, S., Girshick, R., \u0026amp; Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 779\u0026ndash;788.\u003c/li\u003e\n\u003cli\u003eZou, Z., Shi, Z., Guo, Y., \u0026amp; Ye, J. (2019). Object detection in 20 years: A survey. ArXiv Preprint ArXiv:1905.05055.\u003c/li\u003e\n\u003cli\u003eRedmon, J., \u0026amp; Farhadi, A. (2018). Yolov3: An incremental improvement. ArXiv Preprint ArXiv:1804.02767.\u003c/li\u003e\n\u003cli\u003eBochkovskiy, A., Wang, C.-Y., \u0026amp; Liao, H.-Y. M. (2020). Yolov4: Optimal speed and accuracy of object detection. ArXiv Preprint ArXiv:2004.10934.\u003c/li\u003e\n\u003cli\u003eWang, C.-Y., Liao, H.-Y. M., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., \u0026amp; Yeh, I.-H. (2020). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 390\u0026ndash;391.\u003c/li\u003e\n\u003cli\u003eMisra, D. (2019). Mish: A self regularized non-monotonic neural activation function. ArXiv Preprint ArXiv:1908.08681, 4(2), 10\u0026ndash;48550.\u003c/li\u003e\n\u003cli\u003eMaas, A. L., Hannun, A. Y., \u0026amp; Ng, A. Y. (2013). Rectifier nonlinearities improve neural network acoustic models. Proc. Icml, 30(1), 3.\u003c/li\u003e\n\u003cli\u003eBrock, A., Lim, T., Ritchie, J. M., \u0026amp; Weston, N. (2017). Freezeout: Accelerate training by progressively freezing layers. ArXiv Preprint ArXiv:1706.04983.\u003c/li\u003e\n\u003cli\u003eWang, X., Yang, M., Zhu, S., \u0026amp; Lin, Y. (2013). Regionlets for generic object detection. Proceedings of the IEEE International Conference on Computer Vision, 17\u0026ndash;24.\u003c/li\u003e\n\u003cli\u003eGiannakeris, P., Avgerinakis, K., Karakostas, A., Vrochidis, S., \u0026amp; Kompatsiaris, I. (2018). People and vehicles in danger-A fire and flood detection system in social media. 2018 IEEE 13th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP), 1\u0026ndash;5.\u003c/li\u003e\n\u003cli\u003eSzegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., \u0026amp; Wojna, Z. (2016). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2818\u0026ndash;2826.\u003c/li\u003e\n\u003cli\u003eLoshchilov, I., \u0026amp; Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. ArXiv Preprint ArXiv:1608.03983.\u003c/li\u003e\n\u003cli\u003eHao, W., \u0026amp; Zhili, S. (2020). Improved mosaic: Algorithms for more complex images. Journal of Physics: Conference Series, 1684(1), 012094.\u003c/li\u003e\n\u003cli\u003eM\u0026uuml;ller, R., Kornblith, S., \u0026amp; Hinton, G. E. (2019). When does label smoothing help? Advances in Neural Information Processing Systems, 32.\u003c/li\u003e\n\u003cli\u003eGotmare, A., Keskar, N. S., Xiong, C., \u0026amp; Socher, R. (2018). A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation. ArXiv Preprint ArXiv:1810.13243.\u003c/li\u003e\n\u003cli\u003eWang, C.-Y., Bochkovskiy, A., \u0026amp; Liao, H.-Y. M. (2021). Scaled-yolov4: Scaling cross stage partial network. Proceedings of the IEEE/Cvf Conference on Computer Vision and Pattern Recognition, 13029\u0026ndash;13038.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"water-resources-management","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"warm","sideBox":"Learn more about [Water Resources Management](https://www.springer.com/journal/11269)","snPcode":"11269","submissionUrl":"https://submission.nature.com/new-submission/11269/3","title":"Water Resources Management","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"urban flood inundation, image recognition, deep learning, water depth","lastPublishedDoi":"10.21203/rs.3.rs-3075920/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3075920/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eUrban hydrological monitoring is the basis for urban hydrological analysis and storm flood control. However, current monitoring of urban hydrological data is insufficient, including flood inundation depth. This limits calibration and flood early warning ability of the hydrological model. In response to this limitation, a method for evaluating the depth of urban floods based on image recognition using deep learning was established in this study. This method can identify the submerged positions of pedestrians or vehicles in the image, such as pedestrian legs and car exhaust pipes, using the object recognition model YOLOv4. The mean average precision of water depth recognition in a dataset of 1177 flood images reached 89.29%. The established method extracted on-site, real-time, and continuous water depth data from images or video data provided by existing traffic cameras. This system does not require installation of additional water gauges and thus has a low cost and immediate usability.\u003c/p\u003e","manuscriptTitle":"Detection of urban flood inundation from traffic images using deep learning methods","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-07-26 17:25:03","doi":"10.21203/rs.3.rs-3075920/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"","date":"2023-07-22T12:14:32+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-07-22T12:13:27+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-06-26T00:45:39+00:00","index":"","fulltext":""},{"type":"submitted","content":"Water Resources Management","date":"2023-06-25T03:40:14+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"water-resources-management","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"warm","sideBox":"Learn more about [Water Resources Management](https://www.springer.com/journal/11269)","snPcode":"11269","submissionUrl":"https://submission.nature.com/new-submission/11269/3","title":"Water Resources Management","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"ddc85ea8-9352-489b-8c00-9b5ce01071d0","owner":[],"postedDate":"July 26th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-12-04T15:02:54+00:00","versionOfRecord":{"articleIdentity":"rs-3075920","link":"https://doi.org/10.1007/s11269-023-03669-9","journal":{"identity":"water-resources-management","isVorOnly":false,"title":"Water Resources Management"},"publishedOn":"2023-12-02 15:00:53","publishedOnDateReadable":"December 2nd, 2023"},"versionCreatedAt":"2023-07-26 17:25:03","video":"","vorDoi":"10.1007/s11269-023-03669-9","vorDoiUrl":"https://doi.org/10.1007/s11269-023-03669-9","workflowStages":[]},"version":"v1","identity":"rs-3075920","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3075920","identity":"rs-3075920","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.