BirdNET can be as good as experts for acoustic bird monitoring in a European city

preprint OA: closed CC-BY-ND-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

BirdNET has become a leading tool for recognising bird species in audio recordings. However, its applicability in ecological research has been questioned over the sometimes large number of species falsely identified. Using species-specific confidence thresholds has been identified as a powerful approach to solving this issue. However, determining these thresholds is time and resource-consuming. While optimising the parameter setting of the algorithm could be an alternative strategy, the effect of parameter settings on the algorithm’s performance is not well understood. Here, we compared the species identification of BirdNET against expert identification using an acoustic dataset comprising 930 minutes of recordings collected in Munich, Germany. The performance of BirdNET was evaluated using three performance metrics: precision, recall, and F1-score. The metrics were calculated using 24 combinations of the parameters: week, sensitivity, and overlap at four temporal aggregations (pooling of data across time intervals) to also test the effects of recording length. We found that BirdNET closely matched expert identification, particularly when given more data (higher temporal aggregation, F1 score = 0.84) and when including the parameters week of the year, a suitable sensitivity, and an overlap of one to two seconds. Thus, while there are still limitations, using appropriate parameter settings and recording durations, BirdNET yields results comparable to experts without the need for time-consuming estimation of species-specific thresholds. This approach offers reliable presence-absence data in a fast and efficient way while species-specific thresholds are not readily available.
Full text 38,572 characters · extracted from oa-pdf · 10 sections · click to expand

Abstract

18 BirdNET has become a leading tool for recognising bird species in audio recordings. However, its applicability in 19 ecological research has been questioned over the sometimes large number of species falsely identified. Using 20 species-specific confidence thresholds has been identified as a powerful approach to solving this issue. However, 21 determining these thresholds is time and resource-consuming. While optimising the parameter setting of the 22 algorithm could be an alternative strategy, the effect of parameter settings on the algorithm's performance is not well 23 understood. Here, we compared the species identification of BirdNET against expert identification using an acoustic 24 dataset comprising 930 minutes of recordings collected in Munich, Germany. The performance of BirdNET was 25 evaluated using three performance metrics: precision, recall, and F1-score. The metrics were calculated using 24 26 combinations of the parameters: week, sensitivity, and overlap at four temporal aggregations (pooling of data across 27 time intervals) to also test the effects of recording length. We found that BirdNET closely matched expert 28 identification, particularly when given more data (higher temporal aggregation, F1 score = 0.84) and when including 29 the parameters week of the year, a suitable sensitivity, and an overlap of one to two seconds. Thus, while there are 30 still limitations, using appropriate parameter settings and recording durations, BirdNET yields results comparable to 31 experts without the need for time-consuming estimation of species-specific thresholds. This approach offers reliable 32 presence-absence data in a fast and efficient way while species-specific thresholds are not readily available. 33 34

Introduction

35 Monitoring bird communities traditionally relies on field observations, requiring experts to manually identify 36 species through visual or auditory cues. While effective, this process is time-intensive, subject to observer bias [1,2], 37 and limited in spatial and temporal coverage [3]. Passive acoustic monitoring (PAM) has emerged as an alternative, 38 allowing continuous data collection across multiple sites simultaneously. However, the bottleneck in analysing the 39 vast amounts of audio data generated through PAM has limited its practical application in ecological research, as 40 manual species identification in recordings remains equally time-consuming and requires specialised expertise [4]. 41 Recent advancements in machine learning have transformed acoustic data analysis, making automated species 42 identification increasingly accessible to researchers without computer science expertise. Among these tools, 43 BirdNET [5] has emerged as a leader tool. The current version, 2.4, has global coverage of over 6,500 avian and 44 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 3 non-avian classes [6]. In Germany, BirdNET covers 407 of the 527 bird species tracked by the German 45 Ornithologists Society, including rare and vagrant species [7]. This extensive coverage, coupled with its open-source 46 nature and user-friendly interface, has driven BirdNET's rapid adoption in both industry applications and scientific 47 research [e.g. 5–7]. 48 Despite its growing popularity, integrating BirdNET in ecological research has been questioned over the sometimes 49 large number of species falsely identified. Using species-specific confidence thresholds has been identified as a 50 powerful approach to solving this issue [6]. As such, many BirdNET studies have focused on optimising confidence 51 thresholds [e.g. 10,11]. However, determining these thresholds is time and resource-consuming. Opt imising the 52 parameter setting of the algorithm could be an alternative strategy to improve suboptimal classifications and increase 53 the reliability of ecological metrics derived from BirdNET analyses. Yet the effect of parameter settings—such as 54 overlap, sensitivity, and week of the year—on BirdNET’s performance is not well understood [9], especially in an 55 urban environment. 56 Here, we used expertly identified acoustic recordings collected in Munich, Germany, to test BirdNET for bird 57 species classification. We aim to determine whether BirdNET can provide species lists comparable to an expert 58 ornithologist in an urban environment. More specifically, we a) assess the impact of varying BirdNET parameters on 59 classification performance, b) examine how different recording length influence output, and c) compare BirdNET’s 60 performance to expert annotations in terms of species richness and identification accuracy. Based on these findings, 61 we provide practical recommendations for parameter settings and validation approaches that optimize BirdNET 62 performance in urban acoustic surveys. 63

Methods

64 Acoustic recording 65 We placed a single Frontier Labs BAR on the roof of a housing complex in Munich, Laim, between May and 66 October 2021. We recorded in week-long blocks, with a minimum of one week between recording periods. We 67 recorded one minute every 10 minutes from two hours before sunrise to three hours after sunrise [13] to keep the 68 amount of manual identification required to a manageable level resulting in a total of 15.5 hours of recordings over 69 60 days. Recordings were taken at a sample rate of 48kHz, a bit depth of 16 and a gain of 40dB. 70 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 4 Species Identification 71 Two experts (A. Fairbairn, J.S. Burmeister) identified all bird vocalisations in each recording visually and aurally 72 using Kaleidoscope Pro version 5.6.8 to view the spectrograms and listen to the recordings [14] resulting in a list of 73 species per each one-minute recording. Next, we ran BirdNET analyser v2.3 on the same recordings, producing a list 74 of BirdNET species detections for each one-minute recording. 75 Analysis 76 Parameter effects 77 We tested four parameter settings. First, BirdNET can use the week of the year of recording in conjunction with the 78 location to filter what species are likely to occur at a location at that time of the year using eBird [15] species lists. 79 We ran all analyses with and without week included. Second, the detection sensitivity (range 0.5 to 1.5, default 1.0) 80 affects how sensitive BirdNET is to faint or background vocalisations. We ran all analyses with three sensitivity 81 levels (0.5, 1.0, 1.5). Third, BirdNET works on three-second audio segments for analysis. The overlap determines 82 how many seconds of the previous segment is “overlapped” (default 0.0s). We ran all analyses with four levels of 83 overlap (0, 1, 2, 2.9 seconds). Finally, BirdNET provides a confidence level, i.e., the confidence that BirdNET has in 84 its own predictions [6], for each detection. Setting a minimum confidence level in BirdNET causes all detections 85 with a lower confidence to be removed from the results. Thus, to identify the best settings, including the best 86 minimum confidence level, we ran all analyses with the default minimum confidence level (0.1) to get a full list of 87 detections that we could filter afterwards for the analyses. 88 Temporal aggregation 89 To test how different temporal resolutions (i.e., short versus long recording periods) affect BirdNET's performance, 90 we aggregated both our reference data and BirdNET results to four different temporal scales: minute (no 91 aggregation), day, week, and the entire dataset (Table 1). For BirdNET results, we first filtered the raw detections 92 based on confidence thresholds and other parameters before aggregation. For each temporal scale, we recorded only 93 species presence data, ignoring repeated detections of the same species. At the minute resolution, if a species was 94 detected multiple times within the same minute, it was recorded as a single presence. At the day resolution, if a 95 species was detected in any minute during that day, it was recorded as a single presence for the entire day. At the 96 week resolution, any species detected at any point during the week was recorded as a single presence for that week. 97 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 5 For the entire dataset, each species was recorded as either present or absent overall. This presence-based aggregation 98 was applied to both the expert identifications and filtered BirdNET outputs before making any comparisons. Using 99 the settings that produced the best overall F1 scores, we then calculated separate F1 scores for each individual 100 minute, day, and week to assess how performance variability changes across different temporal resolutions. 101 Table 1. Breakdown of how the data was aggregated to get the different temporal resolutions. resolution length total comparisons minute 1 minute 930 day 31 minutes 30 week 155 minutes* 5 dataset 930 minutes 1 * Two weeks varied slightly by the number of recordings. 155 minutes is the average 102 BirdNET vs expert 103 To compare BirdNET's output to expert identifications, we calculated three metrics for each temporal resolution: 104 true positives, defined as species correctly detected as present by BirdNET; false positives, defined as species 105 incorrectly reported as present by BirdNET but not confirmed by experts; and false negatives, defined as species 106 confirmed present by experts but not detected by BirdNET (Fig. 1). We did not calculate true negatives, as there is 107 no meaningful "absence" class in this context—only species that may have been missed by either the expert or 108 BirdNET. Based on these values, we calculated three commonly used machine learning evaluation metrics: 109 precision, recall, and F1 score, to assess BirdNET's performance. Precision answers the question "How reliable are 110 BirdNET's identifications?" by measuring the proportion of BirdNET's species detections that were correct (Eq. 1). 111 A high precision means that when BirdNET identifies a species, it's likely to be accurate. Recall answers the 112 question "How comprehensive is BirdNET's coverage?" by measuring the proportion of actually present species (as 113 identified by experts) that BirdNET successfully detected (Eq. 2). A high recall means BirdNET is capturing most of 114 the species present in recordings. Because there is typically a trade-off between precision and recall—improving one 115 may reduce the other—we use the F1 score (Eq. 3) as a balanced performance metric. The F1 score represents the 116 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 6 harmonic mean of precision and recall, providing a single value that is high only when both precision and recall are 117 high. This offers an integrated measure of BirdNET's overall effectiveness at correctly identifying bird species while 118 minimising both false identifications and missed species. Importantly, these metrics were calculated at each 119 temporal resolution, meaning we compared the complete species list for each time unit (minute, day, week, or entire 120 dataset) generated by BirdNET to that identified by the expert rather than evaluating individual detections. 121 122 Figure 1. Multiclass confusion matrix showing the definitions of true positive, false negative, and false positive. 123 True negative is light grey because we did not calculate it. 124 125 /g1842/g1870/g1857/g1855/g1861/g1871/g1861/g1867/g1866 /g3404 /g1846/g1870/g1873/g1857 /g1842/g1867/g1871/g1861/g1872/g1861/g1874/g1857 /g1846/g1870/g1873/g1857 /g1842/g1867/g1871/g1861/g1872/g1861/g1874/g1857 /g3397 /g1832/g1853/g1864/g1871/g1857 /g1842/g1867/g1871/g1861/g1872/g1861/g1874/g1857 (1) 126 /g1844/g1857/g1855/g1853/g1864/g1864 /g3404 /g1846/g1870/g1873/g1857 /g1842/g1867/g1871/g1861/g1872/g1861/g1874/g1857 /g1846/g1870/g1873/g1857 /g1842/g1867/g1871/g1861/g1872/g1861/g1874/g1857 /g3397 /g1832/g1853/g1864/g1871/g1857 /g1840/g1857/g1859/g1853/g1872/g1861/g1874/g1857 (2) 127 /g18321 /g3404 2 /g3400 /g1842/g1870/g1857/g1855/g1861/g1871/g1861/g1867/g1866 /g3400 /g1844/g1857/g1855/g1853/g1864/g1864 /g1842/g1870/g1857/g1855/g1861/g1871/g1861/g1867/g1866 /g3397 /g1844/g1857/g1855/g1853/g1864/g1864 (3) 128 Test of parameter settings on BirdNet performance 129 To understand the effects of each parameter, we fitted a linear mixed-effects model with F1 score as the response 130 variable. Fixed effects included the main effects of aggregation level, sensitivity, overlap, and minimum confidence 131 (minconf²), as well as all pairwise interactions among them. A random intercept was included for each run to 132 account for repeated runs of BirdNET. Models were fitted using the lme function from the nlme package [16] in R 133 [17] (Eq 4; Supplementary Table S1). 134 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 7 /g18581 /g1871/g1855/g1867/g1870/g1857 ~ /g1864/g1857/g1874/g1857/g1864 /g3397 /g1871/g1857/g1866/g1871/g1861/g1872/g1861/g1874/g1861/g1872/g1877 /g3397 /g1867/g1874/g1857/g1870/g1864/g1853/g1868 /g3397 /g1865/g1861/g1866/g1855/g1867/g1866/g1858 /g2870/g3397 /g1864/g1857/g1874/g1857/g1864: /g1871/g1857/g1866/g1871/g1861/g1872/g1861/g1874/g1861/g1872/g1877 /g3397 /g1864/g1857/g1874/g1857/g1864: /g1867/g1874/g1857/g1870/g1864/g1853/g1868 /g3397 /g1864/g1857/g1874/g1857/g1864: /g1865/g1861/g1866/g1855/g1867/g1866/g1858 /g2870/g3397 /g1871/g1857/g1866/g1871/g1861/g1872/g1861/g1874/g1861/g1872/g1877: /g1867/g1874/g1857/g1870/g1864/g1853/g1868 /g3397 /g1871/g1857/g1866/g1871/g1861/g1872/g1861/g1874/g1861/g1872/g1877: /g1865/g1861/g1866/g1855/g1867/g1866/g1858 /g2870/g3397 /g1867/g1874/g1857/g1870/g1864/g1853/g1868: /g1865/g1861/g1866/g1855/g1867/g1866/g1858 /g2870, /g1870/g1853/g1866/g1856/g1867/g1865 /g34041 | /g1870 /g1873 /g1866 (4) 135 Confirmation test 136 While in the previous steps, we assume that the expert identification is perfect, incorrect identification or missed 137 species may also occur in the species lists based on expert identification, despite our best efforts. Therefore, we 138 conducted a confirmation test assuming errors can occur within expert and BirdNET identification. Using the 139 parameter values that produced the best result for the dataset resolution (highest F1 score), we manually checked a 140 portion of the BirdNET detections. Following Sethi et. al. 2021 [18], we sorted the BirdNET results by species and 141 randomly selected up to 50 results for each species. For species with fewer than 50 detections, we reviewed all 142 available detections. Each selected detection was re-examined by listening to the audio and confirming whether it 143 was correctly identified. We additionally examined the confidence ranges of any false negative species (species 144 identified by the experts but missed by BirdNET when using the best parameter settings) to determine what the 145 impact of lowering the confidence threshold would be. 146 147 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 8

Results

148 We recorded 930 minutes over five weeks and expertly identified 9466 vocalisations of a total of 23 species in the 149 recordings. The most commonly identified species were Short-toed Treecreeper Certhia brachydactyla (n=2894), 150 Blackbird Turdus merula (n=1078), Common Swift Apus apus (n=817) and Great Tit Parus major (n=798; 151 Supplementary Table S2). With default settings, and including week of the year, BirdNET identified 93 species with 152 the most frequently identified being the European Robin Erithacus rubecula (n=2890), the Blackbird (n=1820), the 153 Great Tit (n=1264), and the Short-toed Treecreeper (n=1037; Supplementary Table S3). 154 BirdNET vs expert 155 When comparing BirdNET with expert identification, temporal aggregation and all tested parameters (minimum 156 confidence, week, sensitivity, and overlap) significantly affected BirdNET performance (F1 scores; Fig. 2, 157 Supplementary Table S4). Including the week of the year consistently provided better results (Supplement Fig. S1). 158 Minimum confidence had the strongest effect on F1 score (p < 0.001) and significantly interacted with temporal 159 aggregation, sensitivity, and overlap (Fig. 2, Supplement Table S1). As aggregation increased from minute 160 resolution to the whole dataset, maximum F1 scores improved from 0.62 to 0.84, respectively, when using optimal, 161 i.e. those providing the best F1 scores, combinations of minimum confidence, sensitivity, and overlap (Fig. 2, Table 162 2). With the exception of high confidence levels for the minute and day aggregations, decreasing overlap provided 163 higher maximum F1 scores (Fig. 2, Supplement Table S4). Default BirdNET parameter settings performed poorly 164 when considering F1 scores (Table 2). However, default settings provided higher recall, maximising the number of 165 true positives while inflating false positives. 166 167 Figure 2. Predicted F1 score as a function of data aggregation level, overlap, sensitivity and minimum confidence 168 from a liner mixed-effects model based on 1,944 comparison tests between BirdNET and an expert ornithologist. 169 Each panel shows the variation in predicted F1 scores across each aggregation level, minimum confidence and 170 overlap (left) or sensitivity (right). 171 172 Table 2. Best settings resulted in the highest F1 scores for each temporal aggregation level from 1,944 BirdNET and expert .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 9 identification comparisons. aggregation week sensitivity overlap minConf precision recall F1 score true positives false positives false negatives minute* yes 1 0 0.1 0.35 0.69 0.46 1196 2267 546 minute yes 0.5 2 0.56 0.70 0.56 0.62 975 426 767 minute yes 1 2 0.54 0.70 0.56 0.62 975 426 767 minute yes 1.5 2 0.52 0.70 0.56 0.62 975 426 767 day* yes 1 0 0.1 0.38 0.88 0.53 328 538 44 day yes 0.5 1 0.57 0.73 0.75 0.74 278 102 94 week* yes 1 0 0.1 0.34 0.96 0.50 92 181 4 week yes 0.5 1 0.83 0.77 0.74 0.76 71 21 25 week yes 1 1 0.74 0.77 0.74 0.76 71 21 25 week yes 1.5 1 0.63 0.77 0.74 0.76 71 21 25 dataset* yes 1 0 0.1 0.25 1.00 0.40 23 70 0 dataset yes 1.5 0 0.79 0.90 0.78 0.84 18 2 5 dataset yes 1.5 0 0.8 0.90 0.78 0.84 18 2 5 * Default BirdNET settings with week included. Bold are the best settings we recommend 173 Confirmation test 174 In our confirmation test, conducted using the best parameter settings identified for data aggregated across the entire 175 dataset, BirdNET detected 20 species. Among these were two species— Delichon urbicum (Western house martin) 176 and Turdus philomelos (Song thrush)—that were not present in the expert identifications and were flagged as false 177 positives. However, after manually reviewing the audio clips associated with these detections, we confirmed that all 178 14 detections of Song thrush were, in fact, correct. Only the single detection of Western house martin remained a 179 true false positive (Fig. 3). With these adjustments, the F1 score for the full-dataset resolution was revised to 0.86 180 (Precision = 0.95, Recall = 0.792). The expert identification included an additional five species missed by BirdNET. 181 It is important to note that these false negatives were not addressed in the confirmation test, as they represent species 182 that were not detected by the model. In our check of the false negative species, we found that their confidence scores 183 ranged from a max of 0.37 to a max of 0.68 (Supplementary Table S5). As such, a lower minimum confidence 184 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 10 threshold would have included some of these species, but at the cost of additional false positives. Lowering the 185 minimum confidence to 0.54, for example, reduced the false negatives to 2, but at the cost of an additional 9 false 186 positive species, which would result in an overall lower F1 score. 187 188 Figure 3. The proportion of BirdNET detections from the dataset resolution that were manually checked by an 189 expert and determined to be correct or incorrect. A maximum of 50 random detections per species were checked. If a 190 species was detected less than 50 times, all detections were checked. White numbers are the number of detections 191 checked (if less than 50, only that number were available). The red label denotes the species missed in our expert 192 identification that BirdNET identified. 193 194

Discussion

195 Our study adds to the growing body of work evaluating the performance of BirdNET and is to our knowledge one of 196 the first to systematically investigate the effect of varying BirdNET parameters. While current acoustic monitoring 197 practices recommend short recordings during peak activity periods [13,19], BirdNET eliminates the need to listen to 198 entire recordings, allowing for longer and more frequent data collection. As highlighted by our temporal resolution 199

Results

and as others have found [20], BirdNET performs better with longer recordings, which also enhances 200 monitoring by capturing species with varying activity patterns [21]. Adjusting the parameters, more specifically, by 201 including the week of the year, increasing the overlap to one or two seconds, and using a higher than default 202 confidence threshold produced better results, especially when using short recordings as we did here. When provided 203 with the correct settings and sufficient data (e.g., aggregated over longer periods or longer recordings), BirdNET 204 performed nearly as well as an expert and, in our case, even detected a species missed by the expert. However, it did 205 miss five species identified by the experts. Nevertheless, we show that BirdNET can be used to monitor birds in 206 acoustically complex environments such as cities. 207 While the best settings for BirdNET parameters varied depending on the temporal resolution level, we can deduce 208 some generalisations for running BirdNET. Since our best-performing parameter settings always included the week 209 of the year, we recommend general monitoring to include the week of the year and location. Overlap had a more 210 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 11 significant impact on performance than sensitivity. We recommend using an overlap between one and two seconds 211 for short one to five-minute recording schemes. With longer or continuous recordings, an overlap may not be 212 necessary as vocalisations that may be missed or unidentifiable from being cut are likely to occur again. Default 213 sensitivity (1.0) generally produced the same results as higher or lower sensitivities while maintaining higher 214 minimum confidence levels. We, therefore, recommend using the default sensitivity. For short recordings of 30 215 minutes or less, we suggest filtering by a minimum confidence of 0.54 or higher. For longer recordings, a higher 216 minimum confidence yields the most reliable results, although such a high threshold may exclude some low-217 confidence but valid detections. 218 Despite its performance, we recommend that the results of BirdNET be validated, especially when using very short 219 recordings. While research goals will dictate the amount of validation necessary, we recommend a few quick 220

Methods

to ensure the best results. If the researcher is familiar with what is likely to occur on their study site, only 221 manually checking unlikely or uncommon species is likely to suffice. If confirming the presence of the species at a 222 location (i.e. producing a species list) is the research goal, validation can be done easily by manually checking the 223 top results for each species/site, as only one valid detection is needed to confirm occurrence [9]. Additionally, 224 removing singletons or doubletons and checking only infrequently detected species is likely to provide more 225 accurate species lists than just the raw output of BirdNET. As our confirmation test showed, had we lowered the 226 confidence level to 0.54, removed the singletons and checked the species that were detected 10 or fewer times, we 227 would have had a species list closer to that of the expert, missing only two species. This highlights how simple 228 validation and filtering steps can significantly improve agreement between expert and automated methods. 229 We recognise that current practice recommends using individual species thresholds as model performance can vary 230 greatly between species [6,12]. We think it is important to let the research question dictate which method should be 231 used. Creating individual species scores requires a significant upfront investment in time and requires expert 232 knowledge of the different vocalisations a species can make. Further, the creation of species-specific thresholds 233 assumes that a dataset contains enough detections across the confidence range, although a universal number of 234 detections has yet to be determined. Research questions for which the number of valid vocalisations is important, 235 e.g. studies investigating activity patterns [10], could benefit from species-specific thresholds as they maintain a 236 greater number of detections. For example, Tseng et al. (2025) [12] found that using individual species scores 237 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 12 retained a much larger number of detections (70 ± 37%) than a universal threshold (17 ± 14%) as we use here. Still, 238 it has yet to be determined if these thresholds are transferrable across time (e.g., season, time of day, years) and 239 space (e.g., regions and habitats). Until standardised species-specific thresholds become available, for short-term 240 studies or rapid biodiversity assessments, universal thresholds with simple validation procedures may be sufficient 241 and more resource-efficient. Therefore, it is important to consider the aims of a project when deciding if a universal 242 threshold is adequate or if individual thresholds should be calculated. 243 Our study provides additional support for BirdNET as a practical tool for species identification, particularly in urban 244 environments, but also highlights some limitations that require careful considerati on. We show that appropriate 245 recording strategies and utilising or adjusting key parameters—such as including week of the year, increasing 246 overlap for short recordings, and using a higher minimum confidence threshold—can substantially improve 247 detection performance. It should be noted that our results represent a single site in a southern German city and 248

Results

from different regions or environments may vary. While we acknowledge that the universal confidence 249 thresholds we propose may not suit all research contexts, with basic validation or filtering (e.g., checking 250 uncommon or infrequent species), they can still yield ecologically useful results. The universal threshold approach 251 offers reliable presence-absence data in a fast and efficient way as long as species-specific thresholds are not readily 252 available. BirdNET effectively overcomes some of the limitations of conventional ornithological sampling methods, 253 thus positioning it as a valuable asset in the ongoing quest for comprehensive and efficient biodiversity monitoring 254 practices, offering new research opportunities in ecology and ornithology. 255 256

Acknowledgements

257 We thank the German Research Foundation (DFG Research Training Group 2679 - Urban Green Infrastructure) for 258 funding this research. We would also like to thank Julia Windl and Lisa Maier for assisting with identifying the bird 259 recordings. Additionally, we would like to thank Münchener Wohnen for permitting our recording. 260 261 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 13

References

262 1. Scher CL, Clark JS. Species traits and observer behaviors that bias data assimilation and how to accommodate 263 them. Ecological Applications. 2023;33: e2815. doi:10.1002/eap.2815 264 2. Harris JBC, Haskell DG. Simulated Birdwatchers’ Playback Affects the Behavior of Two Tropical Birds. 265 PLOS ONE. 2013;8: e77902. doi:10.1371/journal.pone.0077902 266 3. Kułaga K, Budka M. Bird species detection by an observer and an autonomous sound recorder in two 267 different environments: Forest and farmland. Pérez-García JM, editor. PLoS ONE. 2019;14: e0211970. 268 doi:10.1371/journal.pone.0211970 269 4. Hoefer S, McKnight ,Donald T., Allen-Ankins ,Slade, Nordberg ,Eric J., and Schwarzkopf L. Passive acoustic 270 monitoring in terrestrial vertebrates: a review. Bioacoustics. 2023;32: 506–531. 271 doi:10.1080/09524622.2023.2209052 272 5. Kahl S, Wood CM, Eibl M, Klinck H. BirdNET: A deep learning solution for avian diversity monitoring. Ecol 273 Inform. 2021;61. doi:10.1016/j.ecoinf.2021.101236 274 6. Wood CM, Kahl S. Guidelines for appropriate use of BirdNET scores and other detector outputs. J Ornithol. 275 2024;165: 777–782. doi:10.1007/s10336-024-02144-5 276 7. Barthel PH, Krüger T. Liste der Vögel Deutschlands: Version 3.2. Deutsche Ornithologen-Gesellschaft e.V.; 277 2019. Available: http://www.do-278 g.de/fileadmin/Barthel___Krueger_2019_Liste_der_Voegel_Deutschlands_3.2_DO-G.pdf 279 8. Sethi SS, Fossøy F, Cretois B, Rosten CM. Management relevant applications of acoustic monitoring for 280 Norwegian nature – The Sound of Norway. Norsk institutt for naturforskning (NINA); 2021. Available: 281 https://hdl.handle.net/11250/2832294 282 9. Pérez /i1Granados C. BirdNET: applications, performance, pitfalls and future opportunities. Ibis. 2023;165: 283 1068–1075. doi:10.1111/ibi.13193 284 10. Amorós-Ausina D, Schuchmann K-L, Marques MI, Pérez-Granados C. Living Together, Singing Together: 285 Revealing Similar Patterns of Vocal Activity in Two Tropical Songbirds Applying BirdNET. Sensors. 286 2024;24: 5780. doi:10.3390/s24175780 287 11. David Funosas, Luc Barbaro, Laura Sch illé, Arnaud Elger, Bastien Castagneyrol, Maxime Cauchoix. 288 Assessing the potential of BirdNET to infer European bird communities from large-scale ecoacoustic data. 289 bioRxiv. 2023; 2023.12.06.570351. doi:10.1101/2023.12.06.570351 290 12. Tseng S, Hodder DP, Otter KA. Setting BirdNET confidence thresholds: species-specific vs. universal 291 approaches. J Ornithol. 2025 [cited 7 Apr 2025]. doi:10.1007/s10336-025-02260-w 292 13. Abrahams C. Bird bioacoustic surveys – Developing a standard protocol. In Practice. 2018: 20–23. 293 14. Wildlife Acoustics. Kaleidoscope Pro 5. Wildlife Acoustics; 2021. Available: 294 https://www.wildlifeacoustics.com/products/kaleidoscope-pro 295 15. Sullivan BL, Wood CL, Iliff MJ, Bonney RE, Fink D, Kelling S. eBird: A citizen-based bird observation 296 network in the biological sciences. Biological Conservation. 2009;142: 2282–2292. 297 doi:10.1016/j.biocon.2009.05.006 298 16. Pinheiro J, Bates D, DebRoy S, Sarkar D, R Core Team. nlme: Linear and nonlinear mixed effects models. 299 2021. Available: https://CRAN.R-project.org/package=nlme 300 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint 14 17. R Core Team. R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for 301 Statistical Computing; 2023. Available: https://www.R-project.org/ 302 18. Sethi SS, Fossøy, F, Cretois, B, Rosten, C. M. Management relevant applications of acoustic monitoring for 303 Norwegian nature – The Sound of Norway. Norwegian Institute for Nature Research; 2021 p. 37. Report No.: 304 NINA Report 2064. 305 19. Metcalf O, Abrahams C, Ashington B, Baker E, Bradfer-Lawrence T, Browning E, et al. Good practice 306 guidelines for long-term ecoacoustic monitoring in the UK. The UK Acoustics Network; 2023 Feb pp. 1–82. 307 Available: https://acoustics.ac.uk/ 308 20. Cole JS, Michel NL, Emerson SA, Siegel RB. Automated bird sound classifications of long-duration 309 recordings produce occupancy model outputs similar to manually annotated data. Ornithol Appl. 2022;124: 1–310 15. doi:10.1093/ornithapp/duac003 311 21. Robbins CS. Effect of time of day on bird activity. C. John Ralph, J. Michael Scott, editors. Studies in Avian 312 Biology. 1981;6: 275–286. 313 314 .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint .CC-BY-ND 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted May 16, 2025. ; https://doi.org/10.1101/2024.09.17.613451doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-ND-4.0