Bird detection Algorithm Incorporating Attention Mechanism | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Article Bird detection Algorithm Incorporating Attention Mechanism Yuanqing Liang, Bin Wang, Houxin Huang, Hai Pang, Xiang Yue This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3319901/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The safety of the substation is related to the stability of social order and people's daily lives, and the habitat and reproduction of birds can cause serious safety accidents in the power system. In this paper, to solve the problem of low accuracy rate when the YOLOv5l model is applied to the bird-repelling robot in the substation for detection, a C3ECA-YOLOv5l algorithm is proposed to accurately detect the four common bird species near the substation in real time: pigeon, magpie, sparrow and swallow. Four attention modules—Squeeze-and-Excitation (SE), Convolutional Block Attention Module (CBAM), an efficient channel attention module (ECA), and Coordinate Attention (CA)—were added to the backbone network at different times—after the C3-3 network layer, before the SPPF network layer, and in the C3 network layer (C3-3, C3-6, C3-9, and C3-3)—to determine the best network detection performance option. After comparing the network mean average precision rates (mAP @0.5 ), we incorporated the ECA attention module into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) as the final test method. In the validation set, the mAP @0.5 of the C3ECA-YOLOv5l network was 94.7%, which, after incorporating the SE, CBAM, ECA, and CA attention modules before the SPPF network layer following the C3-3 network layer of the backbone, resulted in mean average precisions of 92.9%, 92.0%, 91.8%, and 93.1%, respectively, indicating a decrease of 1.8%, 2.7%, 2.9%, and 1.6%, respectively. Incorporating the SE, CBAM, and CA attention modules into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) resulted in mean average precision rates of 93.5%, 94.1%, and 93.4%, respectively, which were 1.2%, 0.6%, and 1.3% lower than that obtained for the C3ECA-YOLOv5l model. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 1 Introduction The stability of the electricity system affects the daily lives of thousands of households.[ 33 , 40 ] A substation is a place where the voltage and current of electrical energy in the power system is transformed, concentrated and distributed.[ 19 ] The power plant transmits the electric energy to the substation through high-voltage transmission lines, and the substation reduces the voltage of the electric energy and transmits it from the distribution network to the users' homes. The normal operation of the substation is essential for the smooth supply of electricity. However, many birds often gather around the substation[ 2 , 7 , 10 , 22 ], they will use their claws and beaks to damage the insulation skin, some birds such as sparrows, magpies, etc. will build nests in the vicinity, which will easily lead to short-circuit capacitor bank, insulation failure and other circuit failures in the long run.[ 4 , 31 , 35 ] Bird removal by human labor is expensive and ineffective. Furthermore, driving birds in a timely and efficient manner is not possible.[ 13 ] Physical bird repellent, biochemical bird repellent, and complete bird repellent[ 11 , 20 , 21 ]are other popular bird repellent techniques in addition to manual bird repellent. The most popular physical bird deterrents include voice broadcasting, colorful windmills, acoustic emitters, bird repellent thorns[ 5 , 27 ], and bird repellent vehicles. Granules, powder, and water are the main biochemical bird repellents in use today. Traditional bird repellant techniques, on the other hand, are unable to reliably detect birds near substations, and prolonged aimless effort would result in resource waste and decreased efficiency. With the continuous development of deep learning technology, its application in the field of target detection has become more and more extensive. Siriani et al.[ 26 ] linked the modified YOLO v4 algorithm with the Kalman filter-based bird tracking algorithm to monitor the activities of domestic chickens in farms, and the experiments showed that the accuracy rate was as high as 99.9%, but their dataset was relatively single, and the dataset needs to be expanded to conduct further experiments on the number and types of datasets. Liu et al.[ 15 ] based on YOLO v4 algorithm to identify the dead chicks in the poultry house, the experimental results show that its precision rate reached 95.24%, and designed a dead chick removal system can effectively solve the virus spread in the poultry house. The special signal processing mechanism of the brain for human eyesight provides the inspiration for attentional mechanisms. Human vision swiftly analyzes an entire image to identify the target area to focus on, thereby creating a focus of attention to learn more about the topic of focus while ignoring irrelevant information. This attention mechanism has been used extensively in image recognition over the past few years and has emerged as a fundamental method in deep learning for acquiring essential information from input images while suppressing irrelevant features. Wang et al. [ 32 ]proposed the efficient channel attention network (ECA-Net) for deep convolutional neural networks. By replacing the fully connected layer of the SE module with a 1 × 1 dynamic convolutional kernel, the ECA module improves the SE attention mechanism proposed by Hu et al.[ 9 ] by maintaining a direct relationship between channels and weights. Wang et al. developed a residual attention network to increase the effectiveness of target feature extraction. Zang et al.[ 37 ] used a modified YOLOv5s network for spike number detection and decided to insert the global attention mechanism (GAM)[ 18 ] into the Head section to enhance feature fusion and the ECA attention module into the C3 module of the backbone network to enhance feature extraction. The YOLOv5s network is now more accurate, and the impact of occlusion and overlap on detection outcomes is more clearly understood. Cao et al.[ 1 ] first designed the YOLOv5s network for lightweighting, and then inserted ECA modules into the backbone network to achieve coal and gangue detection. To compare the effects of different attention mechanisms on model accuracy and lightweighting performance, ECA, CA, CBAM, and GAM modules were inserted into the backbone network, respectively. The results show that the method incorporating ECA modules achieves the best detection accuracy and effectively detects objects with different shapes, sizes and surface features. Studies to increase the model's accuracy and robustness by adding an attention mechanism module also include the multi-scale YOLOv5s model proposed by Li et al. [ 12 ]for the detection of small target objects, the improved YOLOv5s model proposed by Qi et al. [ 23 ]for the detection of tomato virus disease, the improved YOLOv5s model proposed by Li et al. [ 14 ]for the detection of wheat ears detection, a Feature Pyramid Attention Network (FPANet) by Zhai et al.[ 38 ] for an accurate and effective count of the crowd. Zheng et al.[ 39 ] enhanced the ability of the backbone network to extract features by improving the Swin Transfer attention mechanism. In summary, the integration of the attention module into the YOLOv5 backbone network enables accurate detection of objects that are subject to occlusion, overlap, or have different surface characteristics. Based on the research of domestic and international scholars, this paper proposes a C3ECA-YOLOv5l network for common bird detection. The specific work is summarized as follows: (1) The 768 original photos of the dataset were improved using Open CV software and extended to a total of 6,451 images in the enhanced dataset. (2) To determine the best method for improving the performance of the network, four attention mechanisms—SE, CBAM, CA, and ECA—were added after the concentrated-comprehensive convolution (C3)-3 network layer of the feature extraction network, before the spatial pyramid pooling-fast (SPPF) network layer, and in the C3 network layer (C3-3, C3-6, C3-9, and C3-3) of the backbone feature extraction network. (3) An improved C3ECA-YOLOv5l algorithm was suggested for efficiently and accurately detecting four frequent bird species near the substation: pigeons, magpies, sparrows, and swallows. 2. Materials and methods 2.1 Data acquisition Four common birds—pigeons, sparrows, magpies, and swallows—and a total of 768 photographs were included in the dataset for this study. The dataset's whole photo collection was sourced from the Baidu Encyclopedia. The photographs of several birds in the dataset are displayed in Fig. 1 . 2.2 Image preprocessing LabelImg software[ 29 ] was used to label the pigeons, sparrows, magpies, and swallows.in the dataset; label boxes were added and the labeled files were saved. The photos were tagged using the LabelImg software, as shown in Fig. 2 . Data enhancement reduces overfitting during model training, boosts model robustness, and increases the number of images required for model generalization. To simulate the impact of weather variations on detection findings, the brightness of the original images was modified using Open CV software, which sharpened the original picture, emphasized the edges of bird, and reduced the background in the labeled box. Open CV software was also used to add Gaussian noise, Gaussian blur, and mirroring operations to the original image to simulate the blur that could result from detecting a moving object, capturing a scene from a distance, and inaccurate camera focus. The original images and image enhancements of the 5,376 images were randomly rotated in a 2:8 ratio with rotation angles of 30°, 90°, and 180°. Random rotation was employed to simulate the impact of the device on the detection, resulting from changes in the angle. After several image enhancement operations, the original dataset was expanded to include 6,451 photographs. The location coordinate information of the bird features in the annotation files of the other data augmentation methods were the same, except for the mirroring and random rotation operations. Based on the 6:2:2 ratio, the images in the dataset were split into training, testing, and validation sets. In Fig. 3 , the data-enhanced and original images are displayed. Table 1 lists the specifics of the dataset structure. Table 1 Dataset structure Ripeness Number of original images Training (image enhancement) Validation (image enhancement) Test (image enhancement) pigeons sparrows magpies swallows 202 158 184 224 1018 796 930 1129 339 265 309 376 339 265 309 276 total 768 3873 1289 1289 2.3 Construction of experimenting platform The tests were run on a Windows 10 computer with an Nvidia GeForce RTX 3080 graphics card and an Intel(R) Xeon(R) Platinum 8255C CPU running at 2.50GHz with 12 virtual CPUs. PyTorch 1.8.1 and Compute Unified Device Architecture (CUDA) 11.1 were used as the deep learning framework, and cuDNN version 8.0.5 served as the acceleration. The learning rate, batch size, momentum, weight decay, and number of iterations were all set to 0.01; the photo resolution was set to 640 pixels. All models in this experiment were run in the same environment configuration and under the same hyperparameter settings. 2.4Four common birds detection based on C3ECA-YOLOv5l Redmon et al.[ 24 ] proposed a target identification technique known as YOLO (V1-V3), a multiscale predictive detection algorithm that can recognize visual features of various sizes. YOLOv5 builds upon YOLOv3, with the four network models of YOLOv5— YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x—available on GitHub. Multiple depth and width parameters are used to modify the depth and width of the networks, even though the overall design of the networks is the same. Similar to the channel and layer control factors in EfficientNet[ 28 ], YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x increase in depth and width from left to right. The backbone and head constitute the network structure of YOLOv5 (version 6.1). The head, which can be subdivided into the neck, can also be detected. In this experiment, YOLOv5l, with depth and width multiples of 1.0, was selected as the network model. The four attention mechanisms—SE, CBAM, ECA, and CA—were incorporated separately after the C3-3 network layer of the feature extraction network, before the SPPF network layer, and in the C3 network layer of the backbone feature extraction network (C3-3, C3-6, C3-9, and C3-3) to identify detection performance improvement options for YOLOv5l. The experimental results demonstrate that the algorithm performs best when the ECA attention mechanism is incorporated into the C3 network layer. Figure 4 depicts the network structure of C3ECA-YOLOv5l. Mosaic data improvement employs CutMix data enhancement[ 36 ] as a model to randomly scale, crop, arrange, and merge four images to increase the number of detection targets and enhance data diversity. However, the batch size (training parameters) cannot be increased because the GPU constraints are resolved through mosaic data improvement. The three modules that constitute the backbone feature extraction network are focus, cross-stage partial (CSP) network, and SPPF. The focus module was replaced with a 6 × 6 convolutional layer using YOLOv5 (version 6.1). A 6 × 6 convolutional layer is more effective than employing the focus module for some GPUs currently in use, considering the same results are obtained with less computational effort. YOLOv5l has two types of CSPs[ 30 ], with the head section using direct connections and the CSPs of the backbone network connected by residual connections[ 6 ].The structure of the CSP network is shown in Fig. 4 . In this study, the CSP network of the backbone network in the C3ECA-YOLOv5l algorithm was improved by including the ECA attention mechanism in the bottleneck module to increase the feature-extraction capabilities of the backbone network. The SPP module in YOLO v5 (version 6.1) was replaced with the SPPF layer, as shown in Fig. 4 . The SPFF layer accelerates the network by cascading three 5 × 5 pooling kernels while achieving the same outcome as the SPP module. The neck and detection components comprise the head module. The neck module is composed of a path aggregation network (PANet)[ 17 ], which fuses data from the bottom and top paths that have passed through the feature pyramid network (top-to-bottom path fusion), thereby improving the neck module’s capacity to fuse features. The detection module receives feature data after the neck module and outputs 80 × 80, 40 × 40, and 20 × 20 feature maps to detect targets of different sizes. The bounding box loss and non-maximum suppression are included in the prediction process of the detection module. The precision of target detection was increased using Loss CLOU , which is used by YOLOv5l to determine the distance between real and predicted boxes. This method effectively addresses the problem of inaccurate intersection over union (IOU) computations. \(Los{s}_{CLOU}\) was calculated as shown in Eq. 1 . $$Los{s_{CLOU}}=1 - IOU+\frac{{({\rho ^2}({b^2},{b^{gt}}))}}{{{c^2}+\alpha v}}$$ 1 $$\alpha =\frac{v}{{\left( {1 - IOU+v} \right)}}$$ 2 $$v=\frac{4}{{{\pi ^2}}}{(\arctan \frac{{{w^{gt}}}}{{{h^{gt}}}} - \arctan \frac{w}{h})^2}$$ 3 where IOU denotes the intersection to merge ratio between the true and predicted values; \({\rho }^{2}\left(b,{b}^{gt}\right)\) denotes the Euclidean distance between the center points of the predicted and true target bounding boxes; c denotes the diagonal distance of the smallest external rectangle between the predicted and true target bounding boxes; \(\alpha v\) denotes the ratio of length to width. 2.5 ECA attention mechanism The ECA is a lightweight plug-and-play channel attention module, where channel-dimensionality reduction is performed without altering the size of the input feature map. Wang et al. believe that the dimensionality reduction operation in the SE attention mechanism disrupts the direct correspondence between channels and weights, thereby negatively influencing the prediction outcome. Therefore, in ECA-Net, the corresponding operation is performed using a 1 × 1 dynamic convolution kernel rather than a fully connected layer to learn the channel dimension information in the SE attention mechanism. The dynamic convolution kernel implies that the size of the convolution kernel adapts to the change in the feature map size through a function, with the relationship between the size of the convolution kernel and number of channels expressed in Eq. 4 . Figure 5 depicts the structural layout of the ECA module. $$k=\psi \left( C \right)=|\frac{{{{\log }_2}(C)}}{\gamma }+\frac{b}{\gamma }{|_{odd'}}$$ 4 where \(k\) is the convolutional kernel size; \(C\) denotes the number of channels \(; {\Vert }_{{odd}^{,}}\) restricts \(k\) by taking only odd values; \(\gamma\) and \(b\) are set to 2 and 1, respectively in the original paper, and jointly change the number of channels \(C\) and convolution kernel size \(k\) . Figure 4 , which depicts the structure of the C3ECA-YOLOv5l network, illustrates the addition of the ECA attention module to the C3 network layer (C3-3, C3-6, C3-9, C3-3) of the backbone feature network in this study, to improve the network’s ability to extract features. The attention module was added after the 1 × 1 and 3 × 3 convolution kernels and then summed with the feature maps connected by residuals. This strategy was experimentally demonstrated to obtain a mAP @0.5 of 94.7%, a 4.0% improvement over that of YOLOv5l, with the outcomes reflecting this improvement. 2.6 Evaluation indicators of the model The model was evaluated using four indices: precision, recall, mAP, and inference time. Eqs. 5 – 7 can be used to determine the values of precision, recall, and mean average precision such that the higher the value, the better the detection and the more reliable and effective the performance of the model. $$precision=\frac{{TP}}{{\left( {TP+FP} \right)}} \times 100\%$$ 5 $$recall=\frac{{TP}}{{\left( {TP+FN} \right)}} \times 100\%$$ 6 $$mAP=\sum\nolimits_{{i=1}}^{c} {\frac{{AP\left( i \right)}}{C}}$$ 7 where true positive (TP) indicates that the actual sample is positive, and the prediction is also positive; false positive (FP) indicates that the actual sample is negative, but the prediction is positive; false negative (FN) indicates that the actual sample is positive, but the prediction is negative. Average precision (AP) indicates the average mean precision rate; the higher the AP value, the better the performance of the model; mAP is the average AP of the four categories of bird; C indicates the number of categories. 3 Results and Discussion 3.1 Results of birds ripeness detection based on C3ECA-YOLOv5l The performance of the improved C3ECA-YOLOv5l algorithm was investigated using photographs from the validation set. The method achieved precision and recall of 92.1% and 90.4%, respectively, on the experimental data. Figure 6 shows examples of four frequent bird species near the substation, including pigeon, magpie, swallow, and sparrow which were recognized by the C3ECA-YOLOv5l algorithm. Figure 7 shows the mean average precision of four common birds, in conclusion, the C3ECA-YOLOv5l algorithm exhibited respectable detection performance and robustness for four common birds, as shown in Fig. 6 and Fig. 7 . 3.2 Comparison of C3ECA-YOLOv5l algorithm with other detection methods This study evaluates the detection performance of the C3ECA-YOLOv5l method with respect to four detection algorithms—Faster R-CNN[ 25 ], SSD[ 16 ], YOLOv5l, and YOLOv3-SPP[ 3 ]—utilizing the precision, recall and mAP @0.5 . Table 2 presents a complete breakdown of the comparison statistics for the five algorithms. Figure 8 shows the graphs of the mAP @0.5 for each of the five algorithms, along with the scores for precision, recall, mAP @0.5 , mAP @ [0.5−0.95] , and F1-curve. Table 2 Comparison of the five algorithms Method Precision (%) Recall (%) mAP @0.5 (%) C3ECA-YOLOv5l 92.1 90.4 94.7 YOLOv5l 91.1 90.1 90.7 SSD 71.4 76.0 77.6 Faster R-CNN 87.5 87.1 89.8 YOLOv3-SPP 80.6 81.2 83.3 As seen in Table 2 , the C3ECA-YOLOv5l algorithm clearly outperforms the YOLOv5l algorithm in terms of precision, recall, and mAP @0.5 by 1.0%, 0.3%, and 4.0%, respectively, the SSD algorithm by 20.7%, 14.4%, and 17.1%, respectively, the Faster R-CNN algorithm by 4.6%, 3.3%, and 4.9%, respectively, and the YOLOv3-SPP by 11.6%, 9.2%, and 11.4%, respectively. In conclusion, the mean average precision rates were higher than the corresponding values for the other four algorithms. The C3ECA-YOLOv5l algorithm outperformed the other models in terms of precision, recall, and mAP @0.5 and exhibited good detection performance. In Fig. 8 . mAP @0.5 graphs of the five algorithms and Performance Parameters Radar Chart, mAP variation is plotted as a function of epoch for five algorithms: C3ECA-YOLOv5l, Faster R-CNN, SSD, YOLOv3-SPP, and YOLOv5l, with the plots of precision, recall, mAP @0.5 , mAP @[0.5−0.95] , and F1-curve scores compared in a radar chart. The mAP scores of C3ECA-YOLOv5l gradually surpass the corresponding values of the other four networks, and the fluctuations during convergence are not large and gradually smooth out. In the radar plot, the area enclosed by the five values of the C3ECA-YOLOv5l network is larger than the corresponding values of the Faster R-CNN, SSD, YOLOv3-SPP, and YOLOv5l networks, with the areas of the remaining algorithms all included in the area enclosed by the C3ECA-YOLOv5l network, implying that the precision, recall, mAP @0.5 , mAP @[0.5−0.95] , and F1-curve scores of C3ECA-YOLOv5l network are larger than the corresponding values of the other network models. Overall, the detection performance of the C3ECA-YOLOv5l network is better than those of Faster R-CNN, SSD, YOLOv3-SPP, and YOLOv5l. The formula for F1 is provided in Eq. 8 : $$F1=\frac{{2 \times precision \times recall}}{{\left( {precision+recall} \right)}}$$ 8 3.3 Effects of incorporating attention mechanisms on model performance This study uses SE, CBAM[ 34 ], ECA, and CA[ 8 ] to investigate the impact of the attention mechanism on the model and determine the method that would most improve the detection performance of the algorithm. Figure 9 depicts attention mechanisms added to the C3-3 network layer of the feature extraction network before the SPPF network layer. After validation, the mean average precision rates of incorporating the SE, CBAM, ECA, and CA attention modules before the SPPF network layer and after the C3-3 network layer of the backbone network are 92.9%, 92.0%, 91.8%, and 93.1%, respectively, 2.2%, 1.3%, 1.1%, and 2.4% higher than those of the YOLOv5l network. Figure 10 depicts the network structure of the attention mechanism incorporated into the C3 network layer of the backbone feature extraction network (C3-3, C3-6, C3-9, and C3-3). The mean average precision rates of incorporating the SE, CBAM, CA, and ECA attention modules into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) experimentally are 93.5%, 94.1%, 93.4%, and 94.7%, respectively, providing improvements of 2.8%, 3.4%, 2.7%, and 4.0%, respectively, relative to the YOLOv5l network. The mean average precision of the two techniques is listed in Tables 4 and 5 . The C3ECA-YOLOv5l, which has the highest mAP @0.5 score, was chosen as the final improved algorithm after comparing the mAP values. 3.3.1 Effect of adding attention mechanism on mAP score As seen in Table 3 and Fig. 11 the mAP @0.5 score of C3ECA-YOLOv5l is 94.7%, which is 1.8%, 2.7%, 2.9%, and 1.6% higher than those obtained by incorporating the SE, CBAM, ECA, and CA attention mechanisms before the SPPF network layer following the C3-3 network layer of the backbone network. The mAP @0.5 scores of the four algorithms relative to the YOLOv5l algorithm are 2.2%, 1.3%, 1.1%, and 2%, respectively. Clearly, adding all four attention mechanisms before the SPPF network layer and following the C3-3 network layer of the backbone network improves the YOLOv5l network; however, adding the ECA attention mechanism to the C3 network layer (C3-3, C3-6, C3-9, and C3-3) provides a greater advantage in terms of mAP score. Table 3 mAP scores of the six algorithms Method C3ECA- YOLOv5l YOLOv5l SE- YOLOv5l ECA- YOLOv5l CBAM- YOLOv5l CA- YOLOv5l mAP @0.5 (%) 94.7 90.7 92.9 91.8 92.0 93.1 In Fig. 11 , the plot of mAP @0.5 of C3ECA-YOLOv5l is compared with that of YOLOv5l, with SE, CBAM, CA, and ECA attention modules incorporated before the SPPF network layer and after the C3-3 network layer of the YOLOv5l to determine the combination that achieves the best detection performance during training. 3.3.2 Effect of including attention mechanisms in C3 layer (C3-3, C3-6, C3-9, C3-3) on algorithm's mAP score. As seen in Table 4 , the mAP @0.5 scores after incorporating the attention mechanism into the C3 network layer are generally higher than the corresponding scores obtained after incorporating the attention mechanism after the C3-3 network layer of the backbone network and before the SPPF network layer. The performance of incorporating the ECA attention mechanism surpasses that of incorporating the SE, CBAM, and CA attention mechanisms in the C3 network layer (C3-3, C3-6, C3-9, C3-3) by 1.2%, 0.6%, and 1.3%, respectively. The mAP @0.5 scores of the four algorithms improved by 4.0%, 2.8%, 2.7%, and 3.4%, respectively, relative to the YOLOv5l network. In conclusion, YOLOv5l, which incorporates the ECA attention mechanism in the C3 network layer, is more conducive to the accurate detection of four common birds. Table 4 mAP scores of the five algorithms Method C3ECA- YOLOv5l YOLOv5l C3SE- YOLOv5l C3CA- YOLOv5l C3CBAM- YOLOv5l mAP @0.5 (%) 94.7 90.7 93.5 93.4 94.1 3.4 Exploring the effect of different image resolutions on detection In this experiment, training input image sizes of 416 × 416, 640 × 640, and 1024 × 1024 pixels were utilized to examine the impact of image resolution on the detection outcome of the C3ECA-YOLOv5l algorithm. The mAP @0.5 and inference times of the various resolution images are detailed in Table 5 . Table 5 mAP @0.5 and inference time of images of various resolutions Resolutions (pixels) [email protected] (%) Inference time (ms) 416 × 416 pixels 640 × 640 pixels 1024 × 1024 pixels 92.8 94.7 94.9 9.0 10.0 16.7 In Table 5 , when the image resolution increases from 416 × 416 pixels to 640 × 640 pixels, the inference time increases by 1 ms. However, mAP @0.5 increases by 1.9% and the model has a more significant increase in mean average precision at the cost of increasing the inference time. However, when the image resolution is increased from 640 × 640 pixels to 1024 × 1024 pixels, the inference time increases by 6.7 ms and the mAP @0.5 increases by only 0.2%. Therefore, a 640 × 640 image was used to train the C3ECA-YOLOv5l network. 3.5 Evaluation of C3ECA-YOLOv5l and YOLOv5l algorithms' inference times The inference times of both C3ECA-YOLOv5l and YOLOv5l algorithms and other performance metrics were compared, with the results presented in Table 6 . The inference time of C3ECA-YOLOv5l is 10.0 ms in Table 4 , 2 ms longer than that of YOLOv5l. In addition, compared to the YOLOv5l algorithm, the memory footprint grew by 3.4 MB. Compared with YOLOv5l, the C3ECA-YOLOv5l algorithm has a greater mAP increase (4.0%). Thus, the C3ECA-YOLOv5l method can still ensure real-time detection of birds when all parameters are combined. 4 Conclusion This paper proposes a method for detecting four frequent bird species near the substation based on the C3ECA-YOLOv5l model. Four attention modules—SE, CBAM, ECA, and CA—were added to the backbone network at different times—after the C3-3 network layer, before the SPPF network layer, and in the C3 network layer (C3-3, C3-6, C3-9, and C3-3)—to determine the best network detection performance. The improved C3ECA-YOLOv5l network had a mAP @0.5 rate of 94.7%, a 4.0% improvement in mAP @0.5 compared to that of the original YOLOv5l (90.7%) on the validation set. The mAP @0.5 rates of integrating the SE, CBAM, ECA, and CA attention modules before the SPPF network layer and following the C3-3 network layer of the backbone network were 92.9%, 92.0%, 91.8%, and 93.1%, reflecting improvements of 2.2%, 1.3%, 1.1%, and 2.4%, respectively, compared with that of the YOLOv5l network and reductions of 1.8%, 2.7%, 2.9%, and 1.6%, respectively, compared with that of the C3ECA-YOLOv5l network. The mAP @0.5 rates for integrating the SE, CBAM, and CA attention modules into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) were 93.5%, 94.1%, and 93.4%, respectively, indicating improvements over that of the YOLOv5l network by 2.8%, 3.4%, and 2.7%, respectively and a reduction over that of the C3ECA-YOLOv5l network by 1.2%, 0.6%, and 1.3%, respectively. The ECA attention mechanism in the C3 network layers (C3-3, C3-6, C3-9, and C3-3) was chosen as the final approach for the experiment after comparing the mean average precision rates (mAP @0.5 ). The mean average precision rate of the C3ECA-YOLOv5l network was higher than that of the YOLOv5l, SSD, Faster R-CNN, and YOLOv3-SPP networks by 4.0%, 17.1%, 4.9%, and 11.4%, respectively, indicating the improvement in the precision of bird’s detection. The model detection time was 2 ms longer than that of YOLOv5l but could still ensure real-time detection of four frequent bird species near the substation. We will continue to refine the algorithm in subsequent studies to identify methods to increase the mean detection precision rate and simplify the network structure of C3ECA-YOLOv5l to obtain a lightweight network. Declarations Author Contributions: Conceptualization, Yuanqing Liang; methodology, Yuanqing Liang; software, Bin Wang; validation, Bin Wang, Houxin Huang and Hai Pang; formal analysis, Houxin Huang; investigation, Hai Pang; resources, Hai Pang; data curation, Bin Wang; writing—original draft preparation, Yuanqing Liang; writing—review and editing, Yuanqing Liang; visualization, Bin Wang; supervision, Xiang Yue; project administration, Xiang Yue; funding acquisition, Xiang Yue. All authors have read and agreed to the published version of the manuscript. Acknowledgements This work was supported in part by Guangxi Power Grid Company Science and Technology Project under Grant040100KK52210007 Data availability If data are required, please contact the corresponding author. Conflict of interest: We have no affiliations with any organization with a direct or indirect financial interest in the subject matter discussed in the manuscript. References Cao Z, Fang L, Li Z, et al. (2023) Lightweight Target Detection for Coal and Gangue Based on Improved Yolov5s. In Processes. MDPI, 11(4): 1268. https://doi.org/10.3390/pr11041268 Chattopadhyay A, Ukil A, Jap D, Bhasin S (2018) Toward threat of implementation attacks on substation security: case study on fault detection and isolation. IEEE Trans Industr Inform 14(6):2442–2451.https://doi.org/10.1109/TII.2017.2770096 Dong S, Ma Y, Li C (2021) Implementation of detection system of grassland degradation indicator grass species based on YOLOv3-SPP algorithm. In Journal of Physics: Conference Series. IOP Publishing, 1738(1): 012051. http://dx.doi.org/10.1088/1742-6596/1738/1/012051 Fang Q, Canbing L (2011) Design of transmission line solar ultrasonic birds repeller. In: 2011 IEEE Power Engineering and Automation Conference, pp. 217–220.https://doi.org/10.1109/PEAM.2011.6134839 Feng D, Lin S, He Z, Sun X, Wang Z (2018) Failure risk interval estimation of traction power supply equipment considering the impact ofmultiple factors. IEEE Trans Transp Electrif 4(2):389–398.https://doi.org/10.1109/TTE.2017.2784959 He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016, 770-778. https://doi.org/10.1109/CVPR.2016.90 He Z, Wu H, Hu X (2021) Analysis on bird damage accident of overhead transmission lines in Ningxia region and optimization Design of Insulating Grading Ring. In: 2021 international conference on electrical materials and power equipment (ICEMPE). IEEE, pp 1–4.https://doi.org/10.1109/ICEMPE51623.2021.9509213 Hou Q, Zhou D, Feng J (2021) Coordinate attention for efficient mobile network design. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13713-13722. https://doi.org/10.1109/CVPR46437.2021.01350 Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2018, 7132- 7141. https://doi.org/10.1109/CVPR.2018.00745 Jie T, Sheng-chao J, Le W (2018) Analysis and prevention of bird hazard barriers on transmission line in Guangxi power grid. In: 2018 13th IEEE Conference on Industrial Electronics and Applications (ICIEA),Wuhan, pp. 270–274. https://doi.org/10.1109/ICIEA.2018.8397727 Langåker H-A, Kjerkreit H, Syversen CL, Moore RJD, Holhjem ØH, Jensen I, Morrison A, Transeth AA, Kvien O, Berg G, Olsen TA, Hatlestad A, Negård T, Broch R, Johnsen JE (2021) An autonomous drone-based system for inspection of electrical substations. Int J Adv Robot Syst 18(2):17298814211002973 Li A, Sun S, Zhang Z, Feng M, Wu C, Li W (2023) A Multi-Scale Traffic Object Detection Algorithm for Road Scenes Based on Improved YOLOv5. In Electronics. MDPI, 12: 878. http://dx.doi.org/10.3390/electronics12040878 Li G, Gao L, Fan X et al (2018) The design of fixed bird-repellent fitting for eliminating bird damage in substations. In: 2018 2nd IEEE Conference on Energy Internet and Energy System Integration (EI2),pp. 1–5. https://doi.org/10.1109/EI2.2018.8582423 Li R, Wu Y (2022) Improved YOLO v5 Wheat Ear Detection Algorithm Based on Attention Mechanism. In Electronics. MDPI, 11: 1673. http://dx.doi.org/10.3390/electronics11111673 Liu H W, Chen C H, Tsai Y C, et al.(2021) Identifying images of dead chickens with a chicken removal system integrated with a deep learning algorithm[J]. Sensors, 21(11): 3579.https://doi.org/10.3390/s21113579 Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu CY, Berg AC (2016) Ssd: Single shot multibox detector. In Computer Vision–ECCV 14th European Conference, Amsterdam, The Netherlands, October 11–14, Proceedings, Part I 14 ,21-37. Springer International Publishing. https://doi.org/10.1007/978-3-319-46448-02 Liu, S., Qi, L., Qin, H., Shi, J., Jia, J (2018) Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 8759-8768. https://doi.org/10.1109/CVPR.2018.00913 Liu, Y., Shao, Z., & Hoffmann, N. (2021). Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions. https://doi.org/10.48550/arXiv.2112.05561 Ma J, Zhu M, Cai X, Li Y (2019) Dc substation for dc grid—part i: comparative evaluation ofdc substation configurations. IEEE Trans Power Electron 34(10):9719–9731.http://dx.doi.org/10.1109/TPEL.2019.2895043 Mo J, Chen Y, Zhang Y et al (2020) Design and improvement of anti-bird devices for transmission line towers. In: 2020 7th International Conference on Information, Cybernetics, and Computational SocialSystems (ICCSS), pp.808–813. https://doi.org/10.1109/ICCSS52145.2020.9336861 Muminov A, Jeon YC, Na D et al (2017) Development of a solar powered bird repeller system witheffective bird scarer sounds. In: 2017 International Conference on Information Science andCommunications Technologies (ICISCT), pp. 1–4. https://doi.org/10.1109/ICISCT.2017.8188587 Pan H, Zhou F, Ma Y, Wen G (2021) A Bird-caused Damage Risk Assessment System for Power Grid Based on Intelligent Data Platform. In: 2021 IEEE sustainable power and energy conference (iSPEC). IEEE,pp 2559–2564.https://doi.org/10.1109/iSPEC53008.2021.9735826 Qi J, Liu X, Liu K et al. (2022) An improved YOLOv5 model based on visual attention mechanism: Application to recognition of tomato virus disease. In Computers and Electronics in Agriculture. Elsevier. https://doi.org/10.1016/j.compag.2022.106780 Redmon J, Divvala S, Girshick R, et al (2016) You Only Look Once: Unified, Real-Time Object Detection. IEEE Conference on Computer Vision and Pattern Recognition 779–788. https://doi.org/10.48550/arXiv.1506.02640 Ren S, He K, Girshick R, et al (2017) Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE T. Pattern Anal. 39 (6), 1137–1149. https://doi.org/10.1109/tpami.2016.2577031 Siriani A L R, Kodaira V, Mehdizadeh S A, et al. Detection and tracking of chickens in low-light images using YOLO network and Kalman filter[J]. Neural Computing and Applications, 2022, 34(24): 21987-21997.http://dx.doi.org/10.1007/s00521-022-07664-w Sundararajan R, Burnham J, Carlton R, Cherney EA, Couret G, Eldridge KT, Farzaneh M, Frazier SD, Gorur RS, Harness R, Shaffner D, Siegel S, Varner J (2004) Preventive measures to reduce bird related power outages-part ii: streamers and contamination. IEEE Trans Power Deliv 19(4):1848–1853.https://doi.org/10.1109/TPWRD.2003.822522 Tan M, Le Q (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning. PMLR ,6105-6114. https://doi.org/10.48550/arXiv.1905.11946 Tzutalin D, (2015) LabelImg.Git code. https://github.com/tzutalin/labelImg. Wang C, Liao HYM, Yeh IH, et al (2020) CSPNet: A New Backbone that can Enhance Learning Capability of CNN. IEEE Conference on Computer Vision and Pattern Recognition 1571–1580. https://doi.org/10.1109/CVPRW50498.2020.00203 Wang H, Wang S, Deng C et al (2018) Study on the flashover characteristics ofbird droppings along 110kv composite insulator. In: 2018 International Conference on Power System Technology (POWERCON),China, 6 Nov-8 Nov pp. 2929–2933.https://doi.org/10.1109/POWERCON.2018.8602005 Wang Q, Wu B, Zhu P, Li P, Zuo W, Hu Q (2020) ECA-Net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition ,11534-11542. https://doi.org/10.1109/CVPR42600.2020.01155 Wen M, Li Y, Xie X et al (2020) Key factors for efficient consumption of renewable energy in a provincial power grid in southern China. CSEE J Power Energy Syst 6(3):554–562. https://doi.org/10.17775/CSEEJPES.2019.01970 Woo S, Park J, Lee JY, Kweon IS (2018) Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV). 3-19. https://doi.org/10.48550/arXiv.1807.06521 Xiao D, Zhou C, Ma Q, Lei J, du X (2020) Wearable intelligent warning system for approaching high-voltage electrical equipment. IEEE Trans Instrum Meas 69(12):9389–9397.https://doi.org/10.1109/TIM.2020.3001696 Yun S, Han D, Oh SJ, et al (2019) CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features. IEEE International Conference on Computer Vision 6022–6031. https://doi.org/10.1109/ICCV.2019.00612 Zang H, Wang Y, Ru L, et al. (2022) Detection method of wheat spike improved YOLOv5s based on the attention mechanism. In Frontiers in Plant Science. Frontiers, 13: 993244. https://doi.org/10.3389/fpls.2022.993244 Zhai, W., et al. (2023). FPANet: Feature Pyramid Attention Network for Crowd Counting. Applied Intelligence. https://doi.org/10.1007/s10489-023-04499-3 Zheng, W., et al. (2023). A Stage-Adaptive Selective Network with Position Awareness for Semantic Segmentation of LULC Remote Sensing Images. Remote Sensing, 15(11), 2811. https://doi.org/10.3390/rs15112811 Zhou M, Yan J, Zhou X (2020) Real-time online analysis of power grid. CSEE J Power Energy Syst 6(1):236–238.https://doi.org/10.17775/CSEEJPES.2019.02840 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3319901","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":233636928,"identity":"dcf06993-396a-428d-a059-bef0e2d09934","order_by":0,"name":"Yuanqing Liang","email":"","orcid":"","institution":"Nanning Power Supply Bureau, Guangxi Power Grid Co Ltd","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yuanqing","middleName":"","lastName":"Liang","suffix":""},{"id":233636929,"identity":"21479fe9-1c58-481c-8c75-a0fe8135fa85","order_by":1,"name":"Bin Wang","email":"","orcid":"","institution":"Nanning Power Supply Bureau, Guangxi Power Grid Co Ltd","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bin","middleName":"","lastName":"Wang","suffix":""},{"id":233636930,"identity":"5ec0079d-310a-4d34-a521-fd6af2caa016","order_by":2,"name":"Houxin Huang","email":"","orcid":"","institution":"Nanning Power Supply Bureau, Guangxi Power Grid Co Ltd","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Houxin","middleName":"","lastName":"Huang","suffix":""},{"id":233636931,"identity":"c8d65c8f-77b1-412b-9585-6f4d3e2c0005","order_by":3,"name":"Hai Pang","email":"","orcid":"","institution":"Nanning Power Supply Bureau, Guangxi Power Grid Co Ltd","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hai","middleName":"","lastName":"Pang","suffix":""},{"id":233636932,"identity":"6be63571-2f9f-4756-9a85-75ef6adf3a72","order_by":4,"name":"Xiang Yue","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6ElEQVRIiWNgGAWjYLCCBCDmh7IZG4jWItlAkhYQMDhArBaDG8nHJB7U3LHbfP6M4WceBhvZDQeYnz3AryUt2SDh2LPkbTdyjKV5GNKMNxxgMzfAryXH8EEC2+Fksxs8BkAthxM3HOBhk8CvJf/DgYR/h5ON+88Y/+Zh+E+MlhzGB4lth+0MGHLMgLYcIKxF8swzY4PEvsMJEjfSyiznGCQbzzzMZoZXC9/x5GeSP74dtufvP7z5xpsKO9m+483P8GpROAChExsYOIDhBAoqZnzqgUC+AULbMzCwPyCgdhSMglEwCkYqAACtsk5zqKVFiQAAAABJRU5ErkJggg==","orcid":"","institution":"Shenyang Agricultural University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Xiang","middleName":"","lastName":"Yue","suffix":""}],"badges":[],"createdAt":"2023-09-02 13:44:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3319901/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3319901/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":43495778,"identity":"4e22218d-10a5-4e4b-a602-9233f042f4fc","added_by":"auto","created_at":"2023-09-21 16:36:15","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":907384,"visible":true,"origin":"","legend":"\u003cp\u003ePhotographs of different bird species in the dataset. They are magpies, doves, swallows, and sparrows, in that sequence, from A to D. The images were saved in JPG format with a 1280 x 720-pixel resolution.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/b746b765848317ecddbbf365.png"},{"id":43497803,"identity":"5348c9e7-b2fc-44b5-b01d-3301989af008","added_by":"auto","created_at":"2023-09-21 16:52:15","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":468742,"visible":true,"origin":"","legend":"\u003cp\u003eLabeling of images with LabelImg software.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/085578c9f9a7165d7a1372cf.png"},{"id":43497261,"identity":"ec3bb415-28f2-40d7-a922-b7478fec8329","added_by":"auto","created_at":"2023-09-21 16:44:15","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":689237,"visible":true,"origin":"","legend":"\u003cp\u003eData enhancement. (A) original image; original image following brightness alteration through (B) brightening and (C) darkening; (D) original image with Gaussian noise added; THE original image with Gaussian blur added; (F) mirrored version of original image; (G) sharpened version of original image; original image rotated by (H) 30° following brightness adjustment and (I) 180° after the addition of Gaussian noise.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/fd614dd3e800235f68ba66a0.png"},{"id":43497263,"identity":"d17ef6b9-c9db-4ab4-a462-d89a55f8a720","added_by":"auto","created_at":"2023-09-21 16:44:15","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":203355,"visible":true,"origin":"","legend":"\u003cp\u003eStructure of C3ECA-YOLOv5l\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/24d2e33d3db036352335d372.png"},{"id":43497804,"identity":"e0906d94-332b-4712-8445-57d75bfc4af8","added_by":"auto","created_at":"2023-09-21 16:52:15","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":186986,"visible":true,"origin":"","legend":"\u003cp\u003eStructure of the ECA module\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/7c73cd58b0c428e7f8f5c876.png"},{"id":43495775,"identity":"6cf96a33-3337-4ab1-a039-7ffde3d10964","added_by":"auto","created_at":"2023-09-21 16:36:15","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1159907,"visible":true,"origin":"","legend":"\u003cp\u003eDetection results of four frequent bird species near the substation. A-D: Effect of detecting pigeons E-H: Effect of detecting magpies I-L: Effect of detecting sparrows M-P: Effect of detecting swallows.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/fe5490d9350faa6f0989e68a.png"},{"id":43497259,"identity":"3febfbfc-f53f-4da3-b738-84e5b3585be7","added_by":"auto","created_at":"2023-09-21 16:44:15","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":126552,"visible":true,"origin":"","legend":"\u003cp\u003emAP of four frequent bird species near the substation; and “all classes” represents the mean average precision. (\u003ca href=\"mailto:
[email protected]\"\u003emAP\u003csub\
[email protected]\u003c/sub\u003e\u003c/a\u003e).\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/601d61b509ab2f2a4161476e.png"},{"id":43495776,"identity":"cd510089-0815-451a-9a5d-ecb609e93f01","added_by":"auto","created_at":"2023-09-21 16:36:15","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":220316,"visible":true,"origin":"","legend":"\u003cp\u003emAP\u003csub\
[email protected]\u003c/sub\u003e graphs of the five algorithms and Performance Parameters Radar Chart\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/ff49cab2a13e19cda29ae227.png"},{"id":43495771,"identity":"28fa5781-3362-4e04-acf5-efc48225ea3b","added_by":"auto","created_at":"2023-09-21 16:36:15","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":50998,"visible":true,"origin":"","legend":"\u003cp\u003eNetwork structure with attention mechanism incorporated before the SPPF network layer and after the C3-3 network layer of the backbone feature extraction network.\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/a3f17e85cc412072db979cfe.png"},{"id":43495773,"identity":"d0e16bbc-4604-4ab0-b037-1a5b9793c36b","added_by":"auto","created_at":"2023-09-21 16:36:15","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":67923,"visible":true,"origin":"","legend":"\u003cp\u003eStructure incorporating the attention mechanism in C3 network layer (C3-3, C3-6, C3-9, C3-3).\u003c/p\u003e","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/7a6d16cbba175331f6cceafe.png"},{"id":43495780,"identity":"d0352c78-58b7-4729-9968-61449ac1edd8","added_by":"auto","created_at":"2023-09-21 16:36:15","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":90492,"visible":true,"origin":"","legend":"\u003cp\u003ePlots of mAP scores for the six algorithms\u003c/p\u003e","description":"","filename":"floatimage11.png","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/d40002f63ad0bd9c7111568e.png"},{"id":53576925,"identity":"eb5fff8a-e0ac-48fc-9d58-614c971aa498","added_by":"auto","created_at":"2024-03-27 16:45:13","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5694586,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3319901/v1/1a491efa-c2c2-4015-86f1-e506aecbe95c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Bird detection Algorithm Incorporating Attention Mechanism","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eThe stability of the electricity system affects the daily lives of thousands of households.[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e] A substation is a place where the voltage and current of electrical energy in the power system is transformed, concentrated and distributed.[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] The power plant transmits the electric energy to the substation through high-voltage transmission lines, and the substation reduces the voltage of the electric energy and transmits it from the distribution network to the users' homes. The normal operation of the substation is essential for the smooth supply of electricity. However, many birds often gather around the substation[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], they will use their claws and beaks to damage the insulation skin, some birds such as sparrows, magpies, etc. will build nests in the vicinity, which will easily lead to short-circuit capacitor bank, insulation failure and other circuit failures in the long run.[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eBird removal by human labor is expensive and ineffective. Furthermore, driving birds in a timely and efficient manner is not possible.[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] Physical bird repellent, biochemical bird repellent, and complete bird repellent[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]are other popular bird repellent techniques in addition to manual bird repellent. The most popular physical bird deterrents include voice broadcasting, colorful windmills, acoustic emitters, bird repellent thorns[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], and bird repellent vehicles. Granules, powder, and water are the main biochemical bird repellents in use today. Traditional bird repellant techniques, on the other hand, are unable to reliably detect birds near substations, and prolonged aimless effort would result in resource waste and decreased efficiency.\u003c/p\u003e \u003cp\u003eWith the continuous development of deep learning technology, its application in the field of target detection has become more and more extensive. Siriani et al.[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] linked the modified YOLO v4 algorithm with the Kalman filter-based bird tracking algorithm to monitor the activities of domestic chickens in farms, and the experiments showed that the accuracy rate was as high as 99.9%, but their dataset was relatively single, and the dataset needs to be expanded to conduct further experiments on the number and types of datasets. Liu et al.[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] based on YOLO v4 algorithm to identify the dead chicks in the poultry house, the experimental results show that its precision rate reached 95.24%, and designed a dead chick removal system can effectively solve the virus spread in the poultry house.\u003c/p\u003e \u003cp\u003eThe special signal processing mechanism of the brain for human eyesight provides the inspiration for attentional mechanisms. Human vision swiftly analyzes an entire image to identify the target area to focus on, thereby creating a focus of attention to learn more about the topic of focus while ignoring irrelevant information. This attention mechanism has been used extensively in image recognition over the past few years and has emerged as a fundamental method in deep learning for acquiring essential information from input images while suppressing irrelevant features. Wang et al. [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]proposed the efficient channel attention network (ECA-Net) for deep convolutional neural networks. By replacing the fully connected layer of the SE module with a 1 \u0026times; 1 dynamic convolutional kernel, the ECA module improves the SE attention mechanism proposed by Hu et al.[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] by maintaining a direct relationship between channels and weights. Wang et al. developed a residual attention network to increase the effectiveness of target feature extraction. Zang et al.[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e] used a modified YOLOv5s network for spike number detection and decided to insert the global attention mechanism (GAM)[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] into the Head section to enhance feature fusion and the ECA attention module into the C3 module of the backbone network to enhance feature extraction. The YOLOv5s network is now more accurate, and the impact of occlusion and overlap on detection outcomes is more clearly understood. Cao et al.[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] first designed the YOLOv5s network for lightweighting, and then inserted ECA modules into the backbone network to achieve coal and gangue detection. To compare the effects of different attention mechanisms on model accuracy and lightweighting performance, ECA, CA, CBAM, and GAM modules were inserted into the backbone network, respectively. The results show that the method incorporating ECA modules achieves the best detection accuracy and effectively detects objects with different shapes, sizes and surface features. Studies to increase the model's accuracy and robustness by adding an attention mechanism module also include the multi-scale YOLOv5s model proposed by Li et al. [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]for the detection of small target objects, the improved YOLOv5s model proposed by Qi et al. [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]for the detection of tomato virus disease, the improved YOLOv5s model proposed by Li et al. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]for the detection of wheat ears detection, a Feature Pyramid Attention Network (FPANet) by Zhai et al.[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] for an accurate and effective count of the crowd. Zheng et al.[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e] enhanced the ability of the backbone network to extract features by improving the Swin Transfer attention mechanism.\u003c/p\u003e \u003cp\u003eIn summary, the integration of the attention module into the YOLOv5 backbone network enables accurate detection of objects that are subject to occlusion, overlap, or have different surface characteristics. Based on the research of domestic and international scholars, this paper proposes a C3ECA-YOLOv5l network for common bird detection. The specific work is summarized as follows:\u003c/p\u003e \u003cp\u003e(1) The 768 original photos of the dataset were improved using Open CV software and extended to a total of 6,451 images in the enhanced dataset.\u003c/p\u003e \u003cp\u003e(2) To determine the best method for improving the performance of the network, four attention mechanisms\u0026mdash;SE, CBAM, CA, and ECA\u0026mdash;were added after the concentrated-comprehensive convolution (C3)-3 network layer of the feature extraction network, before the spatial pyramid pooling-fast (SPPF) network layer, and in the C3 network layer (C3-3, C3-6, C3-9, and C3-3) of the backbone feature extraction network.\u003c/p\u003e \u003cp\u003e(3) An improved C3ECA-YOLOv5l algorithm was suggested for efficiently and accurately detecting four frequent bird species near the substation: pigeons, magpies, sparrows, and swallows.\u003c/p\u003e"},{"header":"2. Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data acquisition\u003c/h2\u003e \u003cp\u003eFour common birds\u0026mdash;pigeons, sparrows, magpies, and swallows\u0026mdash;and a total of 768 photographs were included in the dataset for this study. The dataset's whole photo collection was sourced from the Baidu Encyclopedia. The photographs of several birds in the dataset are displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Image preprocessing\u003c/h2\u003e \u003cp\u003eLabelImg software[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] was used to label the pigeons, sparrows, magpies, and swallows.in the dataset; label boxes were added and the labeled files were saved. The photos were tagged using the LabelImg software, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eData enhancement reduces overfitting during model training, boosts model robustness, and increases the number of images required for model generalization. To simulate the impact of weather variations on detection findings, the brightness of the original images was modified using Open CV software, which sharpened the original picture, emphasized the edges of bird, and reduced the background in the labeled box. Open CV software was also used to add Gaussian noise, Gaussian blur, and mirroring operations to the original image to simulate the blur that could result from detecting a moving object, capturing a scene from a distance, and inaccurate camera focus. The original images and image enhancements of the 5,376 images were randomly rotated in a 2:8 ratio with rotation angles of 30\u0026deg;, 90\u0026deg;, and 180\u0026deg;. Random rotation was employed to simulate the impact of the device on the detection, resulting from changes in the angle. After several image enhancement operations, the original dataset was expanded to include 6,451 photographs. The location coordinate information of the bird features in the annotation files of the other data augmentation methods were the same, except for the mirroring and random rotation operations. Based on the 6:2:2 ratio, the images in the dataset were split into training, testing, and validation sets. In Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, the data-enhanced and original images are displayed. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e lists the specifics of the dataset structure.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDataset structure\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRipeness\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of\u003c/p\u003e \u003cp\u003eoriginal images\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTraining\u003c/p\u003e \u003cp\u003e(image enhancement)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eValidation\u003c/p\u003e \u003cp\u003e(image enhancement)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTest\u003c/p\u003e \u003cp\u003e(image enhancement)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003epigeons\u003c/p\u003e \u003cp\u003esparrows\u003c/p\u003e \u003cp\u003emagpies\u003c/p\u003e \u003cp\u003eswallows\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e202\u003c/p\u003e \u003cp\u003e158\u003c/p\u003e \u003cp\u003e184\u003c/p\u003e \u003cp\u003e224\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1018\u003c/p\u003e \u003cp\u003e796\u003c/p\u003e \u003cp\u003e930\u003c/p\u003e \u003cp\u003e1129\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e339\u003c/p\u003e \u003cp\u003e265\u003c/p\u003e \u003cp\u003e309\u003c/p\u003e \u003cp\u003e376\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e339\u003c/p\u003e \u003cp\u003e265\u003c/p\u003e \u003cp\u003e309\u003c/p\u003e \u003cp\u003e276\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003etotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e768\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3873\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1289\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1289\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Construction of experimenting platform\u003c/h2\u003e \u003cp\u003eThe tests were run on a Windows 10 computer with an Nvidia GeForce RTX 3080 graphics card and an Intel(R) Xeon(R) Platinum 8255C CPU running at 2.50GHz with 12 virtual CPUs. PyTorch 1.8.1 and Compute Unified Device Architecture (CUDA) 11.1 were used as the deep learning framework, and cuDNN version 8.0.5 served as the acceleration. The learning rate, batch size, momentum, weight decay, and number of iterations were all set to 0.01; the photo resolution was set to 640 pixels. All models in this experiment were run in the same environment configuration and under the same hyperparameter settings.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4Four common birds detection based on C3ECA-YOLOv5l\u003c/h2\u003e \u003cp\u003eRedmon et al.[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] proposed a target identification technique known as YOLO (V1-V3), a multiscale predictive detection algorithm that can recognize visual features of various sizes. YOLOv5 builds upon YOLOv3, with the four network models of YOLOv5\u0026mdash; YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x\u0026mdash;available on GitHub. Multiple depth and width parameters are used to modify the depth and width of the networks, even though the overall design of the networks is the same. Similar to the channel and layer control factors in EfficientNet[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x increase in depth and width from left to right.\u003c/p\u003e \u003cp\u003eThe backbone and head constitute the network structure of YOLOv5 (version 6.1). The head, which can be subdivided into the neck, can also be detected. In this experiment, YOLOv5l, with depth and width multiples of 1.0, was selected as the network model. The four attention mechanisms\u0026mdash;SE, CBAM, ECA, and CA\u0026mdash;were incorporated separately after the C3-3 network layer of the feature extraction network, before the SPPF network layer, and in the C3 network layer of the backbone feature extraction network (C3-3, C3-6, C3-9, and C3-3) to identify detection performance improvement options for YOLOv5l. The experimental results demonstrate that the algorithm performs best when the ECA attention mechanism is incorporated into the C3 network layer. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e depicts the network structure of C3ECA-YOLOv5l.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eMosaic data improvement employs CutMix data enhancement[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e] as a model to randomly scale, crop, arrange, and merge four images to increase the number of detection targets and enhance data diversity. However, the batch size (training parameters) cannot be increased because the GPU constraints are resolved through mosaic data improvement. The three modules that constitute the backbone feature extraction network are focus, cross-stage partial (CSP) network, and SPPF. The focus module was replaced with a 6 \u0026times; 6 convolutional layer using YOLOv5 (version 6.1). A 6 \u0026times; 6 convolutional layer is more effective than employing the focus module for some GPUs currently in use, considering the same results are obtained with less computational effort. YOLOv5l has two types of CSPs[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], with the head section using direct connections and the CSPs of the backbone network connected by residual connections[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].The structure of the CSP network is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. In this study, the CSP network of the backbone network in the C3ECA-YOLOv5l algorithm was improved by including the ECA attention mechanism in the bottleneck module to increase the feature-extraction capabilities of the backbone network. The SPP module in YOLO v5 (version 6.1) was replaced with the SPPF layer, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The SPFF layer accelerates the network by cascading three 5 \u0026times; 5 pooling kernels while achieving the same outcome as the SPP module. The neck and detection components comprise the head module. The neck module is composed of a path aggregation network (PANet)[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], which fuses data from the bottom and top paths that have passed through the feature pyramid network (top-to-bottom path fusion), thereby improving the neck module\u0026rsquo;s capacity to fuse features. The detection module receives feature data after the neck module and outputs 80 \u0026times; 80, 40 \u0026times; 40, and 20 \u0026times; 20 feature maps to detect targets of different sizes. The bounding box loss and non-maximum suppression are included in the prediction process of the detection module. The precision of target detection was increased using Loss\u003csub\u003eCLOU\u003c/sub\u003e, which is used by YOLOv5l to determine the distance between real and predicted boxes. This method effectively addresses the problem of inaccurate intersection over union (IOU) computations. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(Los{s}_{CLOU}\\)\u003c/span\u003e\u003c/span\u003e was calculated as shown in Eq.\u0026nbsp;\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$Los{s_{CLOU}}=1 - IOU+\\frac{{({\\rho ^2}({b^2},{b^{gt}}))}}{{{c^2}+\\alpha v}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\alpha =\\frac{v}{{\\left( {1 - IOU+v} \\right)}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$v=\\frac{4}{{{\\pi ^2}}}{(\\arctan \\frac{{{w^{gt}}}}{{{h^{gt}}}} - \\arctan \\frac{w}{h})^2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere IOU denotes the intersection to merge ratio between the true and predicted values; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\rho }^{2}\\left(b,{b}^{gt}\\right)\\)\u003c/span\u003e\u003c/span\u003e denotes the Euclidean distance between the center points of the predicted and true target bounding boxes; c denotes the diagonal distance of the smallest external rectangle between the predicted and true target bounding boxes; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\alpha v\\)\u003c/span\u003e\u003c/span\u003e denotes the ratio of length to width.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 ECA attention mechanism\u003c/h2\u003e \u003cp\u003eThe ECA is a lightweight plug-and-play channel attention module, where channel-dimensionality reduction is performed without altering the size of the input feature map. Wang et al. believe that the dimensionality reduction operation in the SE attention mechanism disrupts the direct correspondence between channels and weights, thereby negatively influencing the prediction outcome. Therefore, in ECA-Net, the corresponding operation is performed using a 1 \u0026times; 1 dynamic convolution kernel rather than a fully connected layer to learn the channel dimension information in the SE attention mechanism. The dynamic convolution kernel implies that the size of the convolution kernel adapts to the change in the feature map size through a function, with the relationship between the size of the convolution kernel and number of channels expressed in Eq.\u0026nbsp;\u003cspan refid=\"Equ4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e depicts the structural layout of the ECA module.\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$k=\\psi \\left( C \\right)=|\\frac{{{{\\log }_2}(C)}}{\\gamma }+\\frac{b}{\\gamma }{|_{odd\u0026#039;}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e is the convolutional kernel size; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(C\\)\u003c/span\u003e\u003c/span\u003e denotes the number of channels\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(; {\\Vert }_{{odd}^{,}}\\)\u003c/span\u003e\u003c/span\u003e restricts \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e by taking only odd values; \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\gamma\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(b\\)\u003c/span\u003e\u003c/span\u003e are set to 2 and 1, respectively in the original paper, and jointly change the number of channels \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(C\\)\u003c/span\u003e\u003c/span\u003e and convolution kernel size \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, which depicts the structure of the C3ECA-YOLOv5l network, illustrates the addition of the ECA attention module to the C3 network layer (C3-3, C3-6, C3-9, C3-3) of the backbone feature network in this study, to improve the network\u0026rsquo;s ability to extract features. The attention module was added after the 1 \u0026times; 1 and 3 \u0026times; 3 convolution kernels and then summed with the feature maps connected by residuals. This strategy was experimentally demonstrated to obtain a mAP\u003csub\
[email protected]\u003c/sub\u003e of 94.7%, a 4.0% improvement over that of YOLOv5l, with the outcomes reflecting this improvement.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Evaluation indicators of the model\u003c/h2\u003e \u003cp\u003eThe model was evaluated using four indices: precision, recall, mAP, and inference time. Eqs.\u0026nbsp;\u003cspan refid=\"Equ5\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Equ7\" class=\"InternalRef\"\u003e7\u003c/span\u003e can be used to determine the values of precision, recall, and mean average precision such that the higher the value, the better the detection and the more reliable and effective the performance of the model.\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$precision=\\frac{{TP}}{{\\left( {TP+FP} \\right)}} \\times 100\\%$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$recall=\\frac{{TP}}{{\\left( {TP+FN} \\right)}} \\times 100\\%$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$mAP=\\sum\\nolimits_{{i=1}}^{c} {\\frac{{AP\\left( i \\right)}}{C}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere true positive (TP) indicates that the actual sample is positive, and the prediction is also positive; false positive (FP) indicates that the actual sample is negative, but the prediction is positive; false negative (FN) indicates that the actual sample is positive, but the prediction is negative. Average precision (AP) indicates the average mean precision rate; the higher the AP value, the better the performance of the model; mAP is the average AP of the four categories of bird; C indicates the number of categories.\u003c/p\u003e \u003c/div\u003e"},{"header":"3 Results and Discussion","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Results of birds ripeness detection based on C3ECA-YOLOv5l\u003c/h2\u003e \u003cp\u003eThe performance of the improved C3ECA-YOLOv5l algorithm was investigated using photographs from the validation set. The method achieved precision and recall of 92.1% and 90.4%, respectively, on the experimental data. Figure\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows examples of four frequent bird species near the substation, including pigeon, magpie, swallow, and sparrow which were recognized by the C3ECA-YOLOv5l algorithm. Figure\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e shows the mean average precision of four common birds, in conclusion, the C3ECA-YOLOv5l algorithm exhibited respectable detection performance and robustness for four common birds, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Comparison of C3ECA-YOLOv5l algorithm with other detection methods\u003c/h2\u003e \u003cp\u003eThis study evaluates the detection performance of the C3ECA-YOLOv5l method with respect to four detection algorithms\u0026mdash;Faster R-CNN[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], SSD[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], YOLOv5l, and YOLOv3-SPP[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]\u0026mdash;utilizing the precision, recall and mAP\u003csub\
[email protected]\u003c/sub\u003e.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents a complete breakdown of the comparison statistics for the five algorithms. Figure\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e shows the graphs of the mAP\u003csub\
[email protected]\u003c/sub\u003e for each of the five algorithms, along with the scores for precision, recall, mAP\u003csub\
[email protected]\u003c/sub\u003e, mAP\u003csub\u003e@ [0.5\u0026minus;0.95]\u003c/sub\u003e, and F1-curve.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the five algorithms\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003emAP\u003csub\
[email protected]\u003c/sub\u003e (%)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC3ECA-YOLOv5l\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e92.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e90.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e94.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYOLOv5l\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e91.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e90.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e90.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSSD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e71.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e76.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e77.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFaster R-CNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e87.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e87.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e89.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYOLOv3-SPP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e80.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e81.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e83.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAs seen in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, the C3ECA-YOLOv5l algorithm clearly outperforms the YOLOv5l algorithm in terms of precision, recall, and mAP\u003csub\
[email protected]\u003c/sub\u003e by 1.0%, 0.3%, and 4.0%, respectively, the SSD algorithm by 20.7%, 14.4%, and 17.1%, respectively, the Faster R-CNN algorithm by 4.6%, 3.3%, and 4.9%, respectively, and the YOLOv3-SPP by 11.6%, 9.2%, and 11.4%, respectively. In conclusion, the mean average precision rates were higher than the corresponding values for the other four algorithms. The C3ECA-YOLOv5l algorithm outperformed the other models in terms of precision, recall, and mAP\u003csub\
[email protected]\u003c/sub\u003e and exhibited good detection performance.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. mAP\u003csub\
[email protected]\u003c/sub\u003e graphs of the five algorithms and Performance Parameters Radar Chart, mAP variation is plotted as a function of epoch for five algorithms: C3ECA-YOLOv5l, Faster R-CNN, SSD, YOLOv3-SPP, and YOLOv5l, with the plots of precision, recall, mAP\u003csub\
[email protected]\u003c/sub\u003e, mAP\u003csub\u003e@[0.5\u0026minus;0.95]\u003c/sub\u003e, and F1-curve scores compared in a radar chart. The mAP scores of C3ECA-YOLOv5l gradually surpass the corresponding values of the other four networks, and the fluctuations during convergence are not large and gradually smooth out. In the radar plot, the area enclosed by the five values of the C3ECA-YOLOv5l network is larger than the corresponding values of the Faster R-CNN, SSD, YOLOv3-SPP, and YOLOv5l networks, with the areas of the remaining algorithms all included in the area enclosed by the C3ECA-YOLOv5l network, implying that the precision, recall, mAP\u003csub\
[email protected]\u003c/sub\u003e, mAP\u003csub\u003e@[0.5\u0026minus;0.95]\u003c/sub\u003e, and F1-curve scores of C3ECA-YOLOv5l network are larger than the corresponding values of the other network models. Overall, the detection performance of the C3ECA-YOLOv5l network is better than those of Faster R-CNN, SSD, YOLOv3-SPP, and YOLOv5l. The formula for F1 is provided in Eq.\u0026nbsp;\u003cspan refid=\"Equ8\" class=\"InternalRef\"\u003e8\u003c/span\u003e:\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$F1=\\frac{{2 \\times precision \\times recall}}{{\\left( {precision+recall} \\right)}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Effects of incorporating attention mechanisms on model performance\u003c/h2\u003e \u003cp\u003eThis study uses SE, CBAM[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], ECA, and CA[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] to investigate the impact of the attention mechanism on the model and determine the method that would most improve the detection performance of the algorithm. Figure\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e depicts attention mechanisms added to the C3-3 network layer of the feature extraction network before the SPPF network layer. After validation, the mean average precision rates of incorporating the SE, CBAM, ECA, and CA attention modules before the SPPF network layer and after the C3-3 network layer of the backbone network are 92.9%, 92.0%, 91.8%, and 93.1%, respectively, 2.2%, 1.3%, 1.1%, and 2.4% higher than those of the YOLOv5l network.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e depicts the network structure of the attention mechanism incorporated into the C3 network layer of the backbone feature extraction network (C3-3, C3-6, C3-9, and C3-3). The mean average precision rates of incorporating the SE, CBAM, CA, and ECA attention modules into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) experimentally are 93.5%, 94.1%, 93.4%, and 94.7%, respectively, providing improvements of 2.8%, 3.4%, 2.7%, and 4.0%, respectively, relative to the YOLOv5l network. The mean average precision of the two techniques is listed in Tables\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. The C3ECA-YOLOv5l, which has the highest mAP\u003csub\
[email protected]\u003c/sub\u003e score, was chosen as the final improved algorithm after comparing the mAP values.\u003c/p\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e3.3.1 Effect of adding attention mechanism on mAP score\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAs seen in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e the mAP\u003csub\
[email protected]\u003c/sub\u003e score of C3ECA-YOLOv5l is 94.7%, which is 1.8%, 2.7%, 2.9%, and 1.6% higher than those obtained by incorporating the SE, CBAM, ECA, and CA attention mechanisms before the SPPF network layer following the C3-3 network layer of the backbone network. The mAP\u003csub\
[email protected]\u003c/sub\u003e scores of the four algorithms relative to the YOLOv5l algorithm are 2.2%, 1.3%, 1.1%, and 2%, respectively. Clearly, adding all four attention mechanisms before the SPPF network layer and following the C3-3 network layer of the backbone network improves the YOLOv5l network; however, adding the ECA attention mechanism to the C3 network layer (C3-3, C3-6, C3-9, and C3-3) provides a greater advantage in terms of mAP score.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003emAP scores of the six algorithms\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eC3ECA- YOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSE-\u003c/p\u003e \u003cp\u003eYOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eECA- YOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eCBAM- YOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eCA-\u003c/p\u003e \u003cp\u003eYOLOv5l\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003emAP\u003c/b\u003e\u003csub\u003e\u003cb\
[email protected]\u003c/b\u003e\u003c/sub\u003e \u003cb\u003e(%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e94.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e90.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e92.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e91.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e92.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e93.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e, the plot of mAP\u003csub\
[email protected]\u003c/sub\u003e of C3ECA-YOLOv5l is compared with that of YOLOv5l, with SE, CBAM, CA, and ECA attention modules incorporated before the SPPF network layer and after the C3-3 network layer of the YOLOv5l to determine the combination that achieves the best detection performance during training.\u003c/p\u003e \u003cp\u003e \u003cb\u003e3.3.2 Effect of including attention mechanisms in C3 layer (C3-3, C3-6, C3-9, C3-3) on algorithm's mAP score.\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAs seen in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, the mAP\u003csub\
[email protected]\u003c/sub\u003e scores after incorporating the attention mechanism into the C3 network layer are generally higher than the corresponding scores obtained after incorporating the attention mechanism after the C3-3 network layer of the backbone network and before the SPPF network layer. The performance of incorporating the ECA attention mechanism surpasses that of incorporating the SE, CBAM, and CA attention mechanisms in the C3 network layer (C3-3, C3-6, C3-9, C3-3) by 1.2%, 0.6%, and 1.3%, respectively. The mAP\u003csub\
[email protected]\u003c/sub\u003e scores of the four algorithms improved by 4.0%, 2.8%, 2.7%, and 3.4%, respectively, relative to the YOLOv5l network. In conclusion, YOLOv5l, which incorporates the ECA attention mechanism in the C3 network layer, is more conducive to the accurate detection of four common birds.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003emAP scores of the five algorithms\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eC3ECA- YOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eC3SE-\u003c/p\u003e \u003cp\u003eYOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eC3CA- YOLOv5l\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eC3CBAM- YOLOv5l\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emAP\u003csub\
[email protected]\u003c/sub\u003e (%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e94.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e90.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e93.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e93.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e94.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Exploring the effect of different image resolutions on detection\u003c/h2\u003e \u003cp\u003eIn this experiment, training input image sizes of 416 \u0026times; 416, 640 \u0026times; 640, and 1024 \u0026times; 1024 pixels were utilized to examine the impact of image resolution on the detection outcome of the C3ECA-YOLOv5l algorithm. The mAP\u003csub\
[email protected]\u003c/sub\u003e and inference times of the various resolution images are detailed in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003emAP\u003csub\
[email protected]\u003c/sub\u003e and inference time of images of various resolutions\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eResolutions (pixels)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\
[email protected] (%)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eInference time (ms)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e416 \u0026times; 416 pixels\u003c/p\u003e \u003cp\u003e640 \u0026times; 640 pixels\u003c/p\u003e \u003cp\u003e1024 \u0026times; 1024 pixels\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e92.8\u003c/p\u003e \u003cp\u003e94.7\u003c/p\u003e \u003cp\u003e94.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9.0\u003c/p\u003e \u003cp\u003e10.0\u003c/p\u003e \u003cp\u003e16.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIn Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, when the image resolution increases from 416 \u0026times; 416 pixels to 640 \u0026times; 640 pixels, the inference time increases by 1 ms. However, mAP\u003csub\
[email protected]\u003c/sub\u003e increases by 1.9% and the model has a more significant increase in mean average precision at the cost of increasing the inference time. However, when the image resolution is increased from 640 \u0026times; 640 pixels to 1024 \u0026times; 1024 pixels, the inference time increases by 6.7 ms and the mAP\u003csub\
[email protected]\u003c/sub\u003e increases by only 0.2%. Therefore, a 640 \u0026times; 640 image was used to train the C3ECA-YOLOv5l network.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Evaluation of C3ECA-YOLOv5l and YOLOv5l algorithms' inference times\u003c/h2\u003e \u003cp\u003eThe inference times of both C3ECA-YOLOv5l and YOLOv5l algorithms and other performance metrics were compared, with the results presented in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e\u003cimg src=\"https://myfiles.space/user_files/122228_c8a1650c59388082/122228_custom_files/img1695283737.png\"\u003e\u003c/p\u003e\n\u003cp\u003eThe inference time of C3ECA-YOLOv5l is 10.0 ms in Table \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e ms longer than that of YOLOv5l. In addition, compared to the YOLOv5l algorithm, the memory footprint grew by 3.4 MB. Compared with YOLOv5l, the C3ECA-YOLOv5l algorithm has a greater mAP increase (4.0%). Thus, the C3ECA-YOLOv5l method can still ensure real-time detection of birds when all parameters are combined.\u003c/p\u003e"},{"header":"4 Conclusion","content":"\u003cp\u003eThis paper proposes a method for detecting four frequent bird species near the substation based on the C3ECA-YOLOv5l model. Four attention modules\u0026mdash;SE, CBAM, ECA, and CA\u0026mdash;were added to the backbone network at different times\u0026mdash;after the C3-3 network layer, before the SPPF network layer, and in the C3 network layer (C3-3, C3-6, C3-9, and C3-3)\u0026mdash;to determine the best network detection performance. The improved C3ECA-YOLOv5l network had a mAP\u003csub\
[email protected]\u003c/sub\u003e rate of 94.7%, a 4.0% improvement in mAP\u003csub\
[email protected]\u003c/sub\u003e compared to that of the original YOLOv5l (90.7%) on the validation set. The mAP\u003csub\
[email protected]\u003c/sub\u003e rates of integrating the SE, CBAM, ECA, and CA attention modules before the SPPF network layer and following the C3-3 network layer of the backbone network were 92.9%, 92.0%, 91.8%, and 93.1%, reflecting improvements of 2.2%, 1.3%, 1.1%, and 2.4%, respectively, compared with that of the YOLOv5l network and reductions of 1.8%, 2.7%, 2.9%, and 1.6%, respectively, compared with that of the C3ECA-YOLOv5l network. The mAP\u003csub\
[email protected]\u003c/sub\u003e rates for integrating the SE, CBAM, and CA attention modules into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) were 93.5%, 94.1%, and 93.4%, respectively, indicating improvements over that of the YOLOv5l network by 2.8%, 3.4%, and 2.7%, respectively and a reduction over that of the C3ECA-YOLOv5l network by 1.2%, 0.6%, and 1.3%, respectively. The ECA attention mechanism in the C3 network layers (C3-3, C3-6, C3-9, and C3-3) was chosen as the final approach for the experiment after comparing the mean average precision rates (mAP\u003csub\
[email protected]\u003c/sub\u003e). The mean average precision rate of the C3ECA-YOLOv5l network was higher than that of the YOLOv5l, SSD, Faster R-CNN, and YOLOv3-SPP networks by 4.0%, 17.1%, 4.9%, and 11.4%, respectively, indicating the improvement in the precision of bird\u0026rsquo;s detection. The model detection time was 2 ms longer than that of YOLOv5l but could still ensure real-time detection of four frequent bird species near the substation. We will continue to refine the algorithm in subsequent studies to identify methods to increase the mean detection precision rate and simplify the network structure of C3ECA-YOLOv5l to obtain a lightweight network.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization, Yuanqing Liang; methodology, Yuanqing Liang; software, Bin Wang; validation, Bin Wang, Houxin Huang and Hai Pang; formal analysis, Houxin Huang; investigation, Hai Pang; resources, Hai Pang; data curation, Bin Wang; writing\u0026mdash;original draft preparation, Yuanqing Liang; writing\u0026mdash;review and editing, Yuanqing Liang; visualization, Bin Wang; supervision, Xiang Yue; project administration, Xiang Yue; funding acquisition, Xiang Yue. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported in part by Guangxi Power Grid Company Science and Technology Project under Grant040100KK52210007\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIf data are required, please contact the corresponding author.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest:\u0026nbsp;\u003c/strong\u003eWe have no affiliations with any organization with a direct or indirect financial interest in the subject matter discussed in the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eCao Z, Fang L, Li Z, et al. (2023) Lightweight Target Detection for Coal and Gangue Based on Improved Yolov5s. In Processes. MDPI, 11(4): 1268. https://doi.org/10.3390/pr11041268\u003c/li\u003e\n\u003cli\u003eChattopadhyay A, Ukil A, Jap D, Bhasin S (2018) Toward threat of implementation attacks on substation security: case study on fault detection and isolation. IEEE Trans Industr Inform 14(6):2442\u0026ndash;2451.https://doi.org/10.1109/TII.2017.2770096\u003c/li\u003e\n\u003cli\u003eDong S, Ma Y, Li C (2021) Implementation of detection system of grassland degradation indicator grass species based on YOLOv3-SPP algorithm. In Journal of Physics: Conference Series. IOP Publishing, 1738(1): 012051. http://dx.doi.org/10.1088/1742-6596/1738/1/012051\u003c/li\u003e\n\u003cli\u003eFang Q, Canbing L (2011) Design of transmission line solar ultrasonic birds repeller. In: 2011 IEEE Power Engineering and Automation Conference, pp. 217\u0026ndash;220.https://doi.org/10.1109/PEAM.2011.6134839\u003c/li\u003e\n\u003cli\u003eFeng D, Lin S, He Z, Sun X, Wang Z (2018) Failure risk interval estimation of traction power supply equipment considering the impact ofmultiple factors. IEEE Trans Transp Electrif 4(2):389\u0026ndash;398.https://doi.org/10.1109/TTE.2017.2784959\u003c/li\u003e\n\u003cli\u003eHe K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016, 770-778. https://doi.org/10.1109/CVPR.2016.90\u003c/li\u003e\n\u003cli\u003eHe Z, Wu H, Hu X (2021) Analysis on bird damage accident of overhead transmission lines in Ningxia region and optimization Design of Insulating Grading Ring. In: 2021 international conference on electrical materials and power equipment (ICEMPE). IEEE, pp 1\u0026ndash;4.https://doi.org/10.1109/ICEMPE51623.2021.9509213\u003c/li\u003e\n\u003cli\u003eHou Q, Zhou D, Feng J (2021) Coordinate attention for efficient mobile network design. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13713-13722. https://doi.org/10.1109/CVPR46437.2021.01350\u003c/li\u003e\n\u003cli\u003eHu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2018, 7132- 7141. https://doi.org/10.1109/CVPR.2018.00745\u003c/li\u003e\n\u003cli\u003eJie T, Sheng-chao J, Le W (2018) Analysis and prevention of bird hazard barriers on transmission line in Guangxi power grid. In: 2018 13th IEEE Conference on Industrial Electronics and Applications (ICIEA),Wuhan, pp. 270\u0026ndash;274. https://doi.org/10.1109/ICIEA.2018.8397727\u003c/li\u003e\n\u003cli\u003eLang\u0026aring;ker H-A, Kjerkreit H, Syversen CL, Moore RJD, Holhjem \u0026Oslash;H, Jensen I, Morrison A, Transeth AA, Kvien O, Berg G, Olsen TA, Hatlestad A, Neg\u0026aring;rd T, Broch R, Johnsen JE (2021) An autonomous drone-based system for inspection of electrical substations. Int J Adv Robot Syst 18(2):17298814211002973\u003c/li\u003e\n\u003cli\u003eLi A, Sun S, Zhang Z, Feng M, Wu C, Li W (2023) A Multi-Scale Traffic Object Detection Algorithm for Road Scenes Based on Improved YOLOv5. In Electronics. MDPI, 12: 878. http://dx.doi.org/10.3390/electronics12040878\u003c/li\u003e\n\u003cli\u003eLi G, Gao L, Fan X et al (2018) The design of fixed bird-repellent fitting for eliminating bird damage in substations. In: 2018 2nd IEEE Conference on Energy Internet and Energy System Integration (EI2),pp. 1\u0026ndash;5. https://doi.org/10.1109/EI2.2018.8582423\u003c/li\u003e\n\u003cli\u003eLi R, Wu Y (2022) Improved YOLO v5 Wheat Ear Detection Algorithm Based on Attention Mechanism. In Electronics. MDPI, 11: 1673. http://dx.doi.org/10.3390/electronics11111673\u003c/li\u003e\n\u003cli\u003eLiu H W, Chen C H, Tsai Y C, et al.(2021) Identifying images of dead chickens with a chicken removal system integrated with a deep learning algorithm[J]. Sensors, 21(11): 3579.https://doi.org/10.3390/s21113579\u003c/li\u003e\n\u003cli\u003eLiu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu CY, Berg AC (2016) Ssd: Single shot multibox detector. In Computer Vision\u0026ndash;ECCV 14th European Conference, Amsterdam, The Netherlands, October 11\u0026ndash;14, Proceedings, Part I 14 ,21-37. Springer International Publishing. https://doi.org/10.1007/978-3-319-46448-02\u003c/li\u003e\n\u003cli\u003eLiu, S., Qi, L., Qin, H., Shi, J., Jia, J (2018) Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 8759-8768. https://doi.org/10.1109/CVPR.2018.00913\u003c/li\u003e\n\u003cli\u003eLiu, Y., Shao, Z., \u0026amp; Hoffmann, N. (2021). Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions. https://doi.org/10.48550/arXiv.2112.05561\u003c/li\u003e\n\u003cli\u003eMa J, Zhu M, Cai X, Li Y (2019) Dc substation for dc grid\u0026mdash;part i: comparative evaluation ofdc substation configurations. IEEE Trans Power Electron 34(10):9719\u0026ndash;9731.http://dx.doi.org/10.1109/TPEL.2019.2895043\u003c/li\u003e\n\u003cli\u003eMo J, Chen Y, Zhang Y et al (2020) Design and improvement of anti-bird devices for transmission line towers. In: 2020 7th International Conference on Information, Cybernetics, and Computational SocialSystems (ICCSS), pp.808\u0026ndash;813. https://doi.org/10.1109/ICCSS52145.2020.9336861\u003c/li\u003e\n\u003cli\u003eMuminov A, Jeon YC, Na D et al (2017) Development of a solar powered bird repeller system witheffective bird scarer sounds. In: 2017 International Conference on Information Science andCommunications Technologies (ICISCT), pp. 1\u0026ndash;4. https://doi.org/10.1109/ICISCT.2017.8188587\u003c/li\u003e\n\u003cli\u003ePan H, Zhou F, Ma Y, Wen G (2021) A Bird-caused Damage Risk Assessment System for Power Grid Based on Intelligent Data Platform. In: 2021 IEEE sustainable power and energy conference (iSPEC). IEEE,pp 2559\u0026ndash;2564.https://doi.org/10.1109/iSPEC53008.2021.9735826\u003c/li\u003e\n\u003cli\u003eQi J, Liu X, Liu K et al. (2022) An improved YOLOv5 model based on visual attention mechanism: Application to recognition of tomato virus disease. In Computers and Electronics in Agriculture. Elsevier. https://doi.org/10.1016/j.compag.2022.106780\u003c/li\u003e\n\u003cli\u003eRedmon J, Divvala S, Girshick R, et al (2016) You Only Look Once: Unified, Real-Time Object Detection. IEEE Conference on Computer Vision and Pattern Recognition 779\u0026ndash;788. https://doi.org/10.48550/arXiv.1506.02640\u003c/li\u003e\n\u003cli\u003eRen S, He K, Girshick R, et al (2017) Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE T. Pattern Anal. 39 (6), 1137\u0026ndash;1149. https://doi.org/10.1109/tpami.2016.2577031\u003c/li\u003e\n\u003cli\u003eSiriani A L R, Kodaira V, Mehdizadeh S A, et al. Detection and tracking of chickens in low-light images using YOLO network and Kalman filter[J]. Neural Computing and Applications, 2022, 34(24): 21987-21997.http://dx.doi.org/10.1007/s00521-022-07664-w\u003c/li\u003e\n\u003cli\u003eSundararajan R, Burnham J, Carlton R, Cherney EA, Couret G, Eldridge KT, Farzaneh M, Frazier SD, Gorur RS, Harness R, Shaffner D, Siegel S, Varner J (2004) Preventive measures to reduce bird related power outages-part ii: streamers and contamination. IEEE Trans Power Deliv 19(4):1848\u0026ndash;1853.https://doi.org/10.1109/TPWRD.2003.822522\u003c/li\u003e\n\u003cli\u003eTan M, Le Q (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning. PMLR ,6105-6114. https://doi.org/10.48550/arXiv.1905.11946 \u003c/li\u003e\n\u003cli\u003eTzutalin D, (2015) LabelImg.Git code. https://github.com/tzutalin/labelImg. \u003c/li\u003e\n\u003cli\u003eWang C, Liao HYM, Yeh IH, et al (2020) CSPNet: A New Backbone that can Enhance Learning Capability of CNN. IEEE Conference on Computer Vision and Pattern Recognition 1571\u0026ndash;1580. https://doi.org/10.1109/CVPRW50498.2020.00203\u003c/li\u003e\n\u003cli\u003eWang H, Wang S, Deng C et al (2018) Study on the flashover characteristics ofbird droppings along 110kv composite insulator. In: 2018 International Conference on Power System Technology (POWERCON),China, 6 Nov-8 Nov pp. 2929\u0026ndash;2933.https://doi.org/10.1109/POWERCON.2018.8602005\u003c/li\u003e\n\u003cli\u003eWang Q, Wu B, Zhu P, Li P, Zuo W, Hu Q (2020) ECA-Net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition ,11534-11542. https://doi.org/10.1109/CVPR42600.2020.01155\u003c/li\u003e\n\u003cli\u003eWen M, Li Y, Xie X et al (2020) Key factors for efficient consumption of renewable energy in a provincial power grid in southern China. CSEE J Power Energy Syst 6(3):554\u0026ndash;562. https://doi.org/10.17775/CSEEJPES.2019.01970\u003c/li\u003e\n\u003cli\u003eWoo S, Park J, Lee JY, Kweon IS (2018) Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV). 3-19. https://doi.org/10.48550/arXiv.1807.06521\u003c/li\u003e\n\u003cli\u003eXiao D, Zhou C, Ma Q, Lei J, du X (2020) Wearable intelligent warning system for approaching high-voltage electrical equipment. IEEE Trans Instrum Meas 69(12):9389\u0026ndash;9397.https://doi.org/10.1109/TIM.2020.3001696\u003c/li\u003e\n\u003cli\u003eYun S, Han D, Oh SJ, et al (2019) CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features. IEEE International Conference on Computer Vision 6022\u0026ndash;6031. https://doi.org/10.1109/ICCV.2019.00612\u003c/li\u003e\n\u003cli\u003eZang H, Wang Y, Ru L, et al. (2022) Detection method of wheat spike improved YOLOv5s based on the attention mechanism. In Frontiers in Plant Science. Frontiers, 13: 993244. https://doi.org/10.3389/fpls.2022.993244\u003c/li\u003e\n\u003cli\u003eZhai, W., et al. (2023). FPANet: Feature Pyramid Attention Network for Crowd Counting. Applied Intelligence. https://doi.org/10.1007/s10489-023-04499-3\u003c/li\u003e\n\u003cli\u003eZheng, W., et al. (2023). A Stage-Adaptive Selective Network with Position Awareness for Semantic Segmentation of LULC Remote Sensing Images. Remote Sensing, 15(11), 2811. https://doi.org/10.3390/rs15112811\u003c/li\u003e\n\u003cli\u003eZhou M, Yan J, Zhou X (2020) Real-time online analysis of power grid. CSEE J Power Energy Syst 6(1):236\u0026ndash;238.https://doi.org/10.17775/CSEEJPES.2019.02840\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3319901/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3319901/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe safety of the substation is related to the stability of social order and people's daily lives, and the habitat and reproduction of birds can cause serious safety accidents in the power system. In this paper, to solve the problem of low accuracy rate when the YOLOv5l model is applied to the bird-repelling robot in the substation for detection, a C3ECA-YOLOv5l algorithm is proposed to accurately detect the four common bird species near the substation in real time: pigeon, magpie, sparrow and swallow. Four attention modules\u0026mdash;Squeeze-and-Excitation (SE), Convolutional Block Attention Module (CBAM), an efficient channel attention module (ECA), and Coordinate Attention (CA)\u0026mdash;were added to the backbone network at different times\u0026mdash;after the C3-3 network layer, before the SPPF network layer, and in the C3 network layer (C3-3, C3-6, C3-9, and C3-3)\u0026mdash;to determine the best network detection performance option. After comparing the network mean average precision rates (mAP\u003csub\
[email protected]\u003c/sub\u003e), we incorporated the ECA attention module into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) as the final test method. In the validation set, the mAP\u003csub\
[email protected]\u003c/sub\u003e of the C3ECA-YOLOv5l network was 94.7%, which, after incorporating the SE, CBAM, ECA, and CA attention modules before the SPPF network layer following the C3-3 network layer of the backbone, resulted in mean average precisions of 92.9%, 92.0%, 91.8%, and 93.1%, respectively, indicating a decrease of 1.8%, 2.7%, 2.9%, and 1.6%, respectively. Incorporating the SE, CBAM, and CA attention modules into the C3 network layer (C3-3, C3-6, C3-9, and C3-3) resulted in mean average precision rates of 93.5%, 94.1%, and 93.4%, respectively, which were 1.2%, 0.6%, and 1.3% lower than that obtained for the C3ECA-YOLOv5l model.\u003c/p\u003e","manuscriptTitle":"Bird detection Algorithm Incorporating Attention Mechanism","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-09-21 16:36:10","doi":"10.21203/rs.3.rs-3319901/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"141e234d-fac4-4f3c-ba70-09a1bc211a63","owner":[],"postedDate":"September 21st, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-03-27T16:44:45+00:00","versionOfRecord":[],"versionCreatedAt":"2023-09-21 16:36:10","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3319901","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3319901","identity":"rs-3319901","version":["v1"]},"buildId":"GqpaHPwrfC8PjnIFayRh5","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.