Petal to the metal: The slow road to automating large-scale phenology labeling for herbarium specimens

preprint OA: closed CC-BY-NC-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

ABSTRACT Herbarium specimens represent critical historical records of plant phenology, yet automating annotation of reproductive structures remains challenging given the diversity of floral morphologies, specimen age and quality, and image quality. Here, we present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community. After testing multiple strategies for generating training data, we found in-house expert-curated annotations were essential for producing reliable results. Expert validation found relatively strong accuracy for detecting present floral structures, but still had moderately high false negative rates. Applying the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million records labeled with flowers present. However, only 2.9 million of these contained complete metadata necessary for downstream phenology research, highlighting the need for full label digitization efforts. Still, this dataset represents a large compilation of historical herbarium-derived phenology records available as a resource for the phenology community. We end by demonstrating how integrating these machine-labeled records into Phenobase, a publicly-available phenology database, expands taxonomic and temporal coverage for large-scale phenological analyses, and discuss remaining challenges and next steps.
Full text 64,966 characters · extracted from oa-pdf · 7 sections · click to expand

Abstract

14 Herbarium specimens represent critical historical records of plant phenology, yet automating 15 annotation of reproductive structures remains challenging given the diversity of floral 16 morphologies, specimen age and quality, and image quality. Here, we present a machine 17 learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium 18 specimens and deliver these data to the phenology research community. After testing multiple 19 strategies for generating training data, we found in-house expert-curated annotations were 20 essential for producing reliable results. Expert validation found relatively strong accuracy for 21 detecting present floral structures, but still had moderately high false negative rates. Applying 22 the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million 23 records labeled with flowers present. However, only 2.9 million of these contained complete 24 metadata necessary for downstream phenology research, highlighting the need for full label 25 digitization efforts. Still, this dataset represents a large compilation of historical herbarium-26 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 2 derived phenology records available as a resource for the phenology community. We end by 27 demonstrating how integrating these machine-labeled records into Phenobase, a publicly-28 available phenology database, expands taxonomic and temporal coverage for large-scale 29 phenological analyses, and discuss remaining challenges and next steps. 30 31 Key Words: computer vision; data integration; herbarium specimens; machine learning; plant 32 phenology; plant flowering 33 34 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 3

Introduction

35 Natural life cycle events, such as plant flowering, fruiting, and leaf-out, occur annually on 36 seasonally-defined schedules shaped by environmental cues. The timing and sequence of 37 these events, which vary depending on species and environment, underpin many biological and 38 human processes, including species interactions (Kudo et al., 2004), carbon cycling 39 (Richardson et al., 2013), resource acquisition (Deacy et al., 2017), agriculture (Yadav et al., 40 2023), tourism (Nagai et al., 2019), and allergy management (Manangan et al., 2023). Because 41 these seasonal changes are critical and fundamentally linked with many socio-ecological 42 processes, a whole branch of inquiry, phenology, is devoted to understanding their drivers and 43 consequences. 44 A fundamental question in phenology is how different species and ecosystems have 45 responded and will continue to respond to accelerating human-driven environmental change. 46 The rate and direction of these phenological shifts are of particular interest as they are often 47 considered one of the first and best indicators of broader ecological disruption (Menzel et al., 48 2006). For example, such shifts can interfere with key interactions that require phenological 49 synchrony, such as pollination by interacting insects, and can have cascading consequences for 50 species’ population health, ecosystems, and people (Visser et al., 2019). Thus, access to data 51 capturing the timing of key life stages across time, space, and taxa is increasingly of interest 52 and importance across many scientific and applied disciplines. 53 Concerted phenology monitoring has been a key means to generate data needed to 54 document phenological change. However, many regions across the globe don’t have phenology 55 monitoring programs, and in areas that do, the temporal extent is often limited. For example, the 56 USA National Phenology Network began coalescing observational monitoring in 2007 57 (Crimmins et al., 2022), but national-scale phenology data before this date are limited and often 58 of short duration (with some exceptions, see for example, Hough’s efforts from 1851-1859; 59 Guralnick et al., 2025). Beyond direct phenology monitoring, there are new resources, such as 60 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 4 iNaturalist, producing digital vouchers where phenological information can be secondarily 61 derived (Dinnage et al., 2025). Such contemporary contributory science platforms, like 62 iNaturalist, harness community members across the globe to generate large quantities of 63 observations linked to images and now include means to annotate phenological states. While 64 these resources are helping to further fill spatial and taxonomic data gaps and providing new 65 opportunities for phenology research at scale, they are even more limited in temporal extent 66 compared to monitoring data. 67 A critical component for defining and predicting phenological change is understanding 68 phenology and its trends over decadal to century-level timescales. Herbaria, libraries that house 69 collections of pressed plant vouchers from the field, contain specimens dating back sometimes 70 hundreds of years and have the potential to offer valuable historical phenology data, 71 complementing contemporary monitoring initiatives (Davis et al., 2015; Willis et al., 2017). 72 Recent efforts to image and digitize herbarium specimens have allowed for broader use of these 73 plant collections (Soltis 2017), with information on over 108 million angiosperm specimens (~40 74 million with associated images) available on the Global Biodiversity Information Facility (GBIF) 75 as of October 2025. Ideally, each of these herbarium specimens contains valuable information 76 about the location, collection date, species identity, and occasionally the phenological phase of 77 that plant individual. For the vast majority of plants lacking phenology annotations, it requires 78 significant effort to classify which, if any, phenological phases are visibly present on specimens. 79 Automating phenology annotation for digitized herbarium specimens is a clear next step, 80 one that would unlock vast amounts of historical data and enable phenological change research 81 across diverse taxa and environments. Recent machine learning efforts to classify phenological 82 phases in photos of live plants have been successful (Dinnage et al., 2025); however, 83 herbarium specimens present a unique set of challenges. Dried and pressed plant specimens 84 often lose color and structure over time, reproductive structures can be lost during processing, 85 some taxa have structures that are undetectable from digitized images alone or without 86 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 5 dissection, and specimens contain non-plant material such as labels and herbarium stamps that 87 may interfere with machine detection (Pearson et al., 2020). Additionally, accuracy of machine 88 phenology annotation is highly dependent on the annotator who creates training data used to 89 calibrate models, and curating a reliable and taxonomically-diverse training dataset requires 90 extensive effort and careful consideration. 91 Several works have applied machine learning to herbarium specimens, with greater 92 success in species identification than phenological recognition tasks (Hussein et al., 2022). 93 Phenological efforts have often focused on small, curated species and image sets (Goëau et al., 94 2020; Younis et al., 2020) and many have reported challenges annotating specific phenological 95 structures such as flowers (Lorieul et al., 2019; Younis et al., 2020). More recent efforts have 96 successfully addressed specific phenology research questions using models trained by 97 compositing many datasets from multiple different sources (Williamson et al., 2025). Our aim 98 here, however, was not to produce a standalone research product or answer a particular 99 question, but to develop a trustable, usable, and integratable data resource for the broader 100 research community. Doing so required careful curation of our own training data, which allowed 101 us to leverage information gathered during training to reduce noise in the final data product, as 102 we discuss belowt. 103 Here, we present an automated phenology annotation pipeline that accurately detects 104 flowers present on digitized herbarium specimens. With the goal of maximizing geographic and 105 taxonomic coverage across the angiosperm tree of life, this effort unlocks vast quantities of 106 previously limited historical phenology information. We outline the challenges, successes, 107 complexities and strategic decisions involved in developing the pipeline, emphasizing the need 108 for high-quality training data generated by botanically-knowledgeable individuals and rigorous 109 filtering on multiple dimensions both pre- and post-modeling. We discuss the trade off between 110 strict data filtering for improved model accuracy and data retention to maximize the volume of 111 downstream historical records annotated with phenology information. We highlight a surprising 112 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 6 and disappointing discovery that the majority of specimens lack metadata sufficient for 113 phenological annotation, which we discovered after all the rest of the effort. Finally, we 114 demonstrate the vast scale of previously-unavailable historical phenology data that are now 115 available to stakeholders, and discuss the integration of these data into Phenobase 116 (https://phenobase.netlify.app/), an online database of millions of research-ready, community-117 sourced phenology data from field images and from structured in situ phenology monitoring 118 programs. 119 120

Methods

121 Gathering specimen image and metadata 122 We opted to gather data for this work from GBIF, which provides access to billions of species 123 occurrence records and, in many cases, associated images. We searched GBIF in mid-2023 for 124 all “Magnoliopsida” and “Liliopsida”, using filters to select the basis of record as “preserved 125 specimen” and media type as “images”. This produced 33,691,939 occurrence records with 126 35,314,854 corresponding multimedia records. We did not filter out records missing coordinates 127 or location information so that when metadata of a record is updated in the future, its associated 128 phenological information will have already been extracted. We downloaded this dataset and 129 associated images in October 2024 and, using a nearest pixel filter, programmatically resized 130 each image to 1024 pixels along the shorter edge to reduce storage and satisfy GPU memory 131 limitations. Because services are often down and media links can change, the number of 132 images downloaded is a subset of the total number of records in the original GBIF dataset. In 133 sum, after contending with link issues, removing small or low quality images, and images 134 associated with taxa that we decided to filter out (described below), we collected 22,447,563 135 images for downstream labeling and model training. 136 137 Labeling training data 138 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 7 A challenging aspect of developing a reliable automated pipeline for phenophase annotation on 139 herbarium specimens across the angiosperm tree of life was generating enough high-quality 140 training data. Our goal was to produce a trusted historical phenology data resource for broad 141 use, and we therefore deliberately chose to manually generate our own training data rather than 142 utilizing data from multiple external sources. As mentioned above, manually scoring pressed 143 plant specimens for phenophases is time consuming and difficult, and a key challenge is 144 balancing the need for botanical expertise with the capacity of expert annotators. Other 145 initiatives have employed volunteers, undergraduate students, and paid non-experts with 146 relative success (Brenskelle et al., 2020; Zhou et al., 2018), however these focused on a few 147 key taxa with easily-distinguishable structures. Limiting annotations to a few taxa at the genus 148 or family level allows for more in-depth, targeted training and familiarity with target taxa which 149 likely lead to better results. We were unsure whether employing non-experts would be 150 successful when scoring thousands of diverse taxa. 151 We first tested whether a paid undergraduate student could produce accurate 152 phenophase annotations across all flowering plants. We employed a student to score a random 153 selection of specimens for presence or absence of open flowers. If the student was uncertain or 154 unable to make an informed decision about the digitized image, they had the option to mark 155 “unknown”. We then also had a botanically-knowledgeable individual validate these annotations 156 to assess their accuracy. As reported in Results, we found the undergraduate’s open flower 157 annotations were not at a level usable for training a model. Based on these results and other 158 work that suggests that one of the most important variables in annotation accuracy is simply the 159 annotator themselves (Brenskelle et al., 2020), we made the conscious choice to employ one in-160 house botanically-knowledgeable individual to generate all training data. This decision greatly 161 reduced our capacity to generate high quantities of training data, but ensured high-quality and 162 consistency, which was our priority. 163 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 8 In addition to the first set of specimens annotated to validate the undergraduate 164 annotations, we annotated a second additional set of specimens for open flowers to increase 165 training data. While compiling the second set, we randomly selected specimens in families 166 underrepresented in the first set in an effort to promote taxonomic evenness. Additional records 167 were added to the training data iteratively through repeated bouts of image validation 168 throughout model training. 169 170 Filtering training data 171 To ensure accurate and reliable results, we opted to filter out taxa pre-model-training that were 172 challenging to manually annotate. First, we removed genera that had more than 5 records in the 173 training dataset and an “unknown” annotation rate greater than or equal to 25%. After removing 174 these genera, we removed families using the same criteria. Many of these removed taxa were 175 those with small reproductive structures, such as Chenopodiaceae, or those with flowers and 176 fruits that are difficult to distinguish from each other in images alone, like Poaceae. These taxa 177 were removed permanently from our pipeline, both from training and downstream data. We 178 made this choice strategically to balance data quality and quantity. Throughout this process, we 179 found it critical to acknowledge and respond to limitations, both in human expertise, machine 180 capabilities, and overall capacity. Final lists of removed genera and families are available in 181 Supporting Information. 182 183 Model training and post-training thresholding 184 Model Ensembling Approach 185 One challenge with training models on all herbarium data is the vast and complex floral diversity 186 across the whole of the group, challenging models to properly abstract features that can apply 187 broadly and that are not easily confusable with other objects. We realized single models often 188 performed modestly well, but that ensembled models might perform better than any single one 189 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 9 given this variability. We tried several different neural networks, architectures, and model 190 parameters and ultimately determined that an ensemble modeling approach, combining three 191 models, was most effective based on downstream validation statistics. Note that all three 192 models were fine tuned versions of models that were previously pretrained on ImageNet data 193 (Deng et al., 2009). 194 Despite differences across the three models (Effnet, ViTMSE and VITCE, summarized in 195 Table 1), all training followed the same general steps. First, expert-labeled records were split 196 into 60% training, 20% testing, and 20% validation sets. After low-quality images were filtered 197 out, the images were resized to match the size requirements of the specific model (Table 1). 198 Images in the training dataset were augmented using random horizontal and vertical flips and 199 “auto augment” (Cubuk et al., 2018), which uses multiple techniques including random rotation, 200 color jitter, sharpness adjustment, and image shear. Next, all models were color adjusted to use 201 the ImageNet color mean and standard deviation. During training, each model was saved as to 202 five checkpoints The final best model was chosen as the checkpoint with the best F1 score on 203 the holdout validation dataset. The F1 score is the harmonic mean of precision (how many 204 positive model predictions are actually positive?) and recall (how many actual positives did the 205 model detect?). We chose the F1 score because it is particularly useful for assessing predictive 206 performance for the positive class (“flowers present”), which is the focus of this effort. This 207 overall approach is standard for model selection and ensures comparable scoring across model 208 approaches (e.g. EffNet CNN versus ViT) and that models are tuned, selected, and ensembled 209 based on their best validation performance. 210 We ran models using a single A100 GPU on the University of Florida’s HiPerGator. Each 211 model was trained for at least 200 epochs, and we saved five checkpoints during training. At the 212 end of every epoch, we evaluated model performance on a held-out validation set that was 213 never used for backpropagation, which we used to assure model generalization and select the 214 best checkpoint. We used several metrics for evaluating the model, including F1 score, 215 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 10 accuracy, validation set loss, precision, and recall. After training, each checkpoint from all three 216 models was post-processed to determine the probability thresholds that maximized accuracy. 217 We divided the range [0.0–1.0] into three regions defined by two thresholds. Images with a 218 score above the high threshold were considered a positive indication of the trait, images with 219 scores below the low threshold were considered a negative indication of the trait. Images with 220 scores between the high and low thresholds were classified as “equivocal”, meaning the model 221 could not decide if the trait was present or absent. We adjusted those two thresholds until the 222 images in the test (holdout) dataset accuracy was at the maximum. We also guarded against 223 equivocal records being more than 30% of the test dataset. We used a batch size of 32, the 224 initial learning rate was 5e-5 with a weight decay parameter of 0.01, and an AdamW optimizer. 225 For each image, each model generated a final outcome of positive, negative, or 226 equivocal. To ensemble our three models, we used a two-thirds “majority rule” approach, 227 focusing in particular on predictions of “flowers present” (e.g., where at least two models of the 228 three predicted present with high certainty). Since herbarium specimens often contain only a 229 portion of a plant individual and absence on the whole plant cannot be inferred, we focused on 230 the accuracy of “present” labels and only retained images machine-labeled as “present” in 231 downstream data. All other cases were scored “flowers not present” or “equivocal”. Note that the 232 ensemble itself can be undecided (equivocal), for instance with [positive, negative, equivocal] or 233 two or more equivocal votes from the three model outputs. 234 235 Model validation 236 We validated ensemble results both using held-out test data and through additional 237 expert validation. For the expert validation, we machine-labeled a random subset of images and 238 validated 600 specimens machine-labeled as “present” for flowers and 300 specimens machine-239 labeled as “absent” for flowers. Consistent with the Plant Phenology Ontology (Stucky et al., 240 2018), “flowers” here refers to any floral structure, including both flower buds and open flowers. 241 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 11 Although the models were originally trained on “open flowers”, we opted to conduct validation at 242 the broader “flower” level, as these tasks are difficult and we felt “flower” better reflected the 243 features detected by the ensemble models. To further understand instances where the model 244 consensus was incorrect, we anecdotally looked at false positives and false negatives to 245 understand general patterns in the mistakes (see Results). 246 247 Machine labeling and integration with other sources 248 After model validation, the remaining unlabeled downloaded images were passed through the 249 ensemble model pipeline. We retained only images labeled as “present” by at least two out of 250 three models and data fields were standardized for integration in Phenobase, which has a 251 defined set of required and optional fields. As we discuss below, we dropped records that we 252 labelled but were missing these key fields that are required for phenology research. Key fields 253 here include: scientific name, event date, and decimal latitude and longitude. The remaining 254 data were made publicly available through the Phenobase web portal 255 (https://phenobase.netlify.app/). 256 257 Evaluating specimens missing critical fields 258 A key challenge and major limitation during this process was that a large number of specimens 259 did not have date or location information included in their metadata. As these two pieces of 260 information are critical for meaningful phenology data, this resulted in the loss of substantial 261 data. To further investigate this issue, we randomly selected 100 specimens labeled as “flowers 262 present” that were missing both date and location information. We manually annotated these 263 specimens for whether or not they included specific (day, month and year) date information and 264 georeferencable locality information on the specimen label. 265 266

Results

267 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 12 Labeling training data 268 We worked with an undergraduate student to annotate a total of 3,525 specimens for open 269 flowers, providing them with resources for the task and asking them to use online resources 270 when needed. When validating these annotations, the undergraduate and one of our team 271 members (lead author EG), who served as a botanical expert, were in agreement 69% of the 272 time. Given this result, we chose to not include the undergraduate annotations in our final 273 training dataset. EG then generated a total of 9,004 annotations from individual specimens for 274 open flowers. After removing 49 genera and an additional seven families (Appendix S1; see 275 Supporting Information with this article), 1660 genera and 261 families remained. Our final 276 training dataset contained 7,205 specimen records, 4319 annotated as open flowers present 277 and 2886 annotated as open flowers absent. Once expert-labeled records were split, 4324 278 images were included in the training set, 1442 in the testing set, and 1439 in the validation set. 279 280 Model fitting and validation 281 We utilized three models (Table 1) and ensembled their results to predict the presence of 282 flowers. Our training model accuracy scores across all three models varied between 88.3% and 283 92.5%. In general, the vision transformers models (ViT) performed better (92.2% and 92.5% 284 accuracy) than the EffNet model (88.3% success). Overall, from the full ensemble, validation on 285 the held-out test data found 96.3% accuracy for detecting the presence of flowers and 86.8% 286 accuracy for detecting the absence of flowers (Appendix S2). However, expert validation on 287 records not in the initial training, testing, or validation processes found the ensemble models to 288 be 90.5% accurate for detecting the presence of flowers and 52.3% accurate for the absence of 289 flowers (Table 2). The low accuracy for detecting absences was expected and likely due both to 290 expert-validated data being more diverse than training data and that we were less concerned 291 with training for absences in our overall process. Further investigation of the false positives from 292 the expert validation found that the majority of false positives were fruits mistaken for flowers, 293 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 13 although some sheets had no reproductive structures present and one sheet had a fruit 294 illustration (Figure 1). Investigation into false negatives from the expert validation found a variety 295 of reasons for the mistakes, many explainable, including situations where flowers weren’t 296 attached to the main plant, flowers that may have been confused as fruits, and flowers that were 297 very challenging to see (Figure 2). 298 299 Machine labeling and integration with other sources 300 After downloading and filtering out problematic images (images that could not be downloaded, 301 images that cannot be read, or images that were either less than 10 KB or greater than 32 MB) 302 and images of hard-to-annotate genera and families (Appendix S1), we were left with 303 22,447,563 images of digitized herbarium specimens. These images were fed through the 304 ensemble modeling pipeline and annotated for presence or absence of flowers. Of these 305 records, 11,131,820 were labeled as flowers present (meaning at least ⅔ models scored as 306 present), but only 2,913,353 of these contained valid latitude, longitude, and date information 307 (Figure 3). 5,748,788 images were scored as flowers absent (meaning at least ⅔ models scored 308 as absent), and 5,566,955 were equivocal (meaning there was no consensus among the 309 models). Only records scored as flowers present and with the needed metadata were retained 310 and integrated into Phenobase (https://phenobase.netlify.app/). 311 312 Data coverage and evaluating specimens missing critical information 313 The 2,913,35 records annotated as flowers present represented 383 plant families and 8,883 314 genera (Figure 4). They represent 320 years of collecting, with the earliest record dated in 1501 315 (Figure 5). We evaluated a small subset of the 8,218,467 records that were missing locality and 316 date information to determine exactly what fields were missing and if those issues could be 317 resolved with more complete digitization of labels. Out of the 100 random specimens that were 318 missing date and latitude/longitude information, 65 had a date (specific day and year) on the 319 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 14 specimen label while 35 did not; 85 had a locality that is georeferenceable and localizable to a 320 unit lower than country, while 15 did not. A total of 57 out of 100 had both a georeferenceable 321 specific locality and date information on the label. Only 44 of the 100 records had images 322 retrievable from the original URL, while for 56 records the image URL failed. 323 324

Discussion

325 While machine-automated annotation is a clear next step in addressing bottlenecks and 326 delivering large quantities of valuable data to the research community, some annotation tasks 327 are harder than others. Further, tasks that are relatively straightforward in one context, such as 328 annotating flowers in photographs of live plants, can be much more challenging in other 329 contexts, such as attempting to do the same for dried, pressed, and aged plants on herbarium 330 sheets. Other groups have produced herbarium data using machine labeling approaches to 331 answer their own research questions about machine learning, certain taxa, or regions, and 332 many showcase reasonably high precision and recall (Goëau et al., 2020; Lorieul et al., 2019; 333 Williamson et al., 2025; Younis et al., 2020). Here, however, our focus was on creating a 334 reliable resource at the broadest possible scale and for the phenology research community as a 335 whole; we were not focused on answering specific research questions. Our first priority was that 336 stakeholders could trust, understand, and easily use these data. A key challenge for us was to 337 develop a set of solutions that we felt met a reasonable bar for trustworthiness, which lead to 338 rigorous data filtering, and to clearly communicate where these solutions may fall short. Here we 339 are not wholly concerned with the size of the data product or any other metric of success. 340 Rather, creating trustworthy data in this context meant carefully vetting the quality of training 341 data, evaluating validation statistics, and being thoughtful about the data structure and 342 documentation made available to stakeholders so they can best use these data and understand 343 their value and limitations. 344 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 15 We also caution that the most appealing solutions for annotating flowers may be difficult 345 at the taxonomic, spatial and temporal scale we attempted here. Our initial hope when 346 generating training data was that we could produce flower annotations utilizing a team of 347 people, likely undergraduate students, with limited botanical training. While this approach might 348 work when focusing on a small subset of plants, such as a genus or family, when scaling to all 349 angiosperms we found it untenable. A vast, complicated diversity of inflorescence morphologies 350 interacts with herbarium sheet taphonomic processes to make it especially challenging to detect 351 flower presence, even for an expert botanist. We argue that the ideal solution for creating high 352 quality annotations is delegating annotation tasks to botanists working on taxonomic groups 353 where they have expertise, but this is also impractical given capacity and logistical constraints. 354 Our solution here was using one expert with ample practice scoring sheets, which has both 355 benefits and drawbacks. One major benefit is that there is likely more consistency and higher 356 quality with this method than a crowdsourced solution. This is especially important in dealing 357 with the many cases where flower presence was uncertain, which was critical for removing 358 certain families and genera from downstream annotation labeling. However, a major drawback 359 in using only one annotator to generate training data is that this method is much less scalable, 360 generating significantly fewer annotated records to use for training. Finally, more focused 361 expertise in certain taxonomic groups may have produced less “unknowns” and higher quality 362 training data for those groups that could lead to inclusion in downstream modeling. 363 A key step in our process was making thoughtful, informed decisions about removing 364 taxa that we felt were too challenging to score correctly. While annotating training data, the 365 annotator could choose “present” or “absent” options to inform the models, but critically they 366 also had an “unknown” option which they were encouraged to select whenever they could not 367 make a confident determination. Ultimately, the proportion of “unknown” annotations, for 368 individual genera and also families, was what we used to determine which taxa were simply too 369 challenging to score consistently for flower presence. There is likely no perfect cutoff, but we 370 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 16 decided that taxa where more than 25% of records were too difficult to label were likely to be 371 problematic for modeling. This decision represents a tradeoff between data quantity and quality. 372 While it is always disappointing to exclude data, we prioritized maintaining stakeholder trust by 373 favoring reliability over volume. We maintain that when attempting to scale challenging tasks 374 such as phenological annotation of herbarium specimens, it is critical to acknowledge both 375 human and technological limits and adjust methods accordingly. 376 Even with substantial taxonomic filtering, model performance was middling on both held-377 out test data and expert validated data. Given the challenges with fitting a single model, we 378 opted for an ensemble approach, which gave us more power to interrogate multiple models and 379 arrive at a consensus. Ensembling is particularly useful because it reduces idiosyncratic errors 380 from any single model. We required a two-thirds majority to score presence from the 3 separate 381 models. This strategy required more up-front effort, but ultimately produced a more rigorous 382 labeling workflow. As in other approaches we have published (Dinnage et al., 2025), our interest 383 is in phenophase presence, and the ensembled models are fairly good at avoiding false 384 positives, but we note that the model still produced a substantial number of false absences. 385 Since we explicitly tuned the models to minimize false positives at the cost of increasing false 386 negatives, it means we are likely throwing out many images that have flowers. 387 The filtering process didn’t end after we completed machine labeling. Out of 11 million 388 records labeled as flowers present, we found that more than 70% lacked a reported 389 latitude/longitude and/or a date. This is consistent with the overall proportion of specimen 390 records on GBIF that have missing coordinates or spatial issues. When we queried GBIF on 391 December 16, 2025, for example, there are 52,254,706 records with associated specimen 392 images for Magnoliopsida and Liliopsida. However, only about 32% of them (16,786,460) have 393 latitude and longitude coordinates and without geospatial issues. This data incompleteness 394 percent is disappointingly high and effectively limits the utility of these data for research, as date 395 and location information is generally required for these data to be research ready (sensu Soltis 396 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 17 et al., 2018). The root cause of this issue appears to be due to the publication of “stub records” 397 to GBIF. These stub records may have taxon names but often lack dates and/or georeferences 398 and they generally are from very large collections. For example, the Museum National d'Histoire 399 Naturelle de Paris published 5,031,777 angiosperm records with images and this provider was 400 one where records were often found to be lacking date and geospatial data. Further, we should 401 note that many of the images we annotated are no longer available on GBIF. Out of the 100 we 402 checked, more than half no longer had resolvable images online; most of those also came from 403 publishers who often supply a large corpus of images. These significant issues bring home the 404 need for efforts to expedite full label digitization to avoid these issues of missing metadata since 405 many labels do contain dates and can be georeferenced. Such efforts are needed to support 406 the best use of primary and secondary data (sensu Soberón and Peterson, 2004), such as 407 phenology annotations generated from images. Creating more persistence for images tied to 408 specimens will also better ensure reproducibility and stability when relating specimen images, 409 core specimen metadata and secondary data derived from those images. 410 Despite setbacks, our ensemble modeling flower annotation pipeline unlocks vast new 411 global phenology data at unprecedented temporal scales. Although in situ phenology data 412 collection networks (Crimmins et al., 2022) and other initiatives to unlock phenology information 413 from sources like iNaturalist (Dinnage et al., 2025) are providing large quantities of 414 contemporary phenology data, data predating these initiatives are less common. They are often 415 in the form of tabular reports that were monitored at one site over many years, although some 416 larger citizen science projects covered multiple sites and years such as the New York 417 Phenology Project (Fuccillo Battle et al., 2022) and Eastern USA Phenology Project (Guralnick 418 et al., 2024) in the middle and late 1800’s. The almost 3 million labeled records that we present 419 here represent a large dataset of historical (pre-2000) phenology now accessible for the 420 research community. However, we note that significant gaps remain in phenology data 421 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 18 coverage across space and time, and that continued digitization and efforts to add missing 422 metadata to existing digitized records will be essential for continuing to close these gaps. 423 All records labeled as flowers present from this effort are available through Phenobase, 424 an open-access datastore of plant phenology information that integrates in situ data from 425 phenology monitoring programs, machine-labeled data from iNaturalist images, and now 426 machine-labeled data from herbarium specimens. Our pipeline produces phenology labels and 427 fields consistent with the Phenobase data structure and Plant Phenology Ontology (Stucky et 428 al., 2018), and uses a controlled vocabulary for other fields, ensuring that data users can search 429 for phenology records at any level of granularity and get consistent results. Future work may 430 include efforts to label herbarium specimens for fruits, however our early explorations found this 431 task to be even more challenging than labeling for flowers, likely requiring greater human effort, 432 and yet more rigorous filtering steps. In sum, despite the challenges, providing information on 433 historical flowering dates across the globe and integrating with other data sources enables new 434 analyses and insights spanning regions and taxa, enabling yet broader and more synthetic 435 questions to be asked across temporal, taxonomic and spatial scales. 436 437

Acknowledgements

438 The authors would like to thank the many individuals who have contributed to collecting, 439 processing, digitizing, and maintaining herbarium specimens, as this effort would not be 440 possible without their dedication. Phenobase is a team effort and the authors acknowledge team 441 members Carrie Seltzer, Ramona Walls, and Taran Lichtenberger. UFIT Research Computing 442 provided computational resources and support. This work was funded by the National Science 443 Foundation (DBI2223512 to Robert Guralnick and DBI2223508 to Daijiang Li). 444 445 Author Contributions 446 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 19 Robert Guralnick, Daijiang Li and Raphael LaFrance initially conceived of this effort. Erin Grady 447 operationalized most aspects of gathering training data, in consultation with Raphael LaFrance 448 and Robert Guralnick. Raphael LaFrance developed the modeling framework for this effort with 449 help from Russell Dinnage. Raphael LaFrance produced final models and descriptive statistics 450 and Erin Grady led expert model validation. Ellen Denny and John Deck helped ensure that 451 data in the right format could be ingested into Phenobase. Erin Grady and Robert Guralnick 452 archived data. Erin Grady and Robert Guralnick wrote the manuscript with help from Raphael 453 LaFrance, Russell Dinnage and Daijiang Li. All authors contributed to drafts and gave final 454 approval for publication. 455 456 Data Availability Statement 457 The ensemble data models and a corresponding JSON file with model metadata data are 458 housed on Zenodo (https://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0). 461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089). 463 464 Supporting Information 465 Additional Supporting Information may be found online in the Supporting Information section at 466 the end of the article. 467 Appendix S1. List of difficult-to-annotate genera and families removed from training and 468 downstream data. 469 Appendix S2. Table S1. Validation results for held-out test data. 470 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 20

References

471 Brenskelle, L., Guralnick, R.P., Denslow, M., and Stucky, B.J. 2020. Maximizing human effort 472 for analyzing scientific images: A case study using digitized herbarium sheets. 2020. 473 Applications in Plant Sciences 8(6): e11370. https://doi.org/10.1002/aps3.11370 474 475 Crimmins, T., Denny, E., Posthumus, E., Rosemartin, A., Croll, R., Montano, M., and Panci, H. 476 2022. Science and Management Advancements Made Possible by the USA National 477 Phenology Network's Nature's Notebook Platform. BioScience 72(9): 908–920. 478 https://doi.org/10.1093/biosci/biac061 479 480 Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. 2018. AutoAugment: Learning 481 augmentation policies from data. In arXiv [cs.CV]. arXiv. 482 https://doi.org/10.48550/arXiv.1805.09501 483 484 Davis, C.C., Willis, C.G., Connolly, B., Kelly, C. and Ellison, A.M. 2015. Herbarium records are 485 reliable sources of phenological change driven by climate and provide novel insights into 486 species' phenological cueing mechanisms. American Journal of Botany 102: 1599-1609. 487 https://doi.org/10.3732/ajb.1500237 488 489 Deacy, W.W., Armstrong, J.B., Leacock, W.B., Robbins, C.T., Gustine, D.D., Ward, E.J., 490 Erlenbach, J.A. and Stanford, J.A. 2017. Phenological synchronization disrupts trophic 491 interactions between Kodiak brown bears and salmon. Proceedings of the National 492 Academy of Sciences of the United States of America 114 (39): 10432-10437. 493 https://doi.org/10.1073/pnas.1705248114 494 495 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 21 Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. 2009. ImageNet: A large-scale 496 hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern 497 Recognition: 248–255. 498 499 Dinnage, R., Grady, E., Neal, N., Deck, J., Denny, E., Walls, R., Seltzer, C. et al. 2025. 500 PhenoVision: A framework for automating and delivering research-ready plant 501 phenology data from field images. Methods in Ecology and Evolution 16(8): 1763-1780. 502 503 Fuccillo Battle, K., Duhon, A., Vispo, C. R., Crimmins, T. M., Rosenstiel, T. N., Armstrong-504 Davies, L. L., and de Rivera, C. E. 2022. Citizen science across two centuries reveals 505 phenological change among plant species and functional groups in the Northeastern US. 506 The Journal of Ecology 110(8): 1757–1774. https://doi.org/10.1111/1365-2745.13926 507 508 Goëau, H., A. Mora-Fallas, J. Champ, N. L. R. Love, S. J. Mazer, E. Mata-Montero, A. Joly, and 509 P. Bonnet. 2020. A new fine-grained method for automated visual analysis of herbarium 510 specimens: A case study for phenological data extraction. Applications in Plant Sciences 511 8(6): e11368. https://doi.org/10.1002/aps3.11368 512 513 Guralnick, R., Crimmins, T., Grady, E., and Campbell, L. Phenological response to climatic 514 change depends on spring warming velocity. 2024. Communications Earth & 515 Environment 5, 634. https://doi.org/10.1038/s43247-024-01807-8 516 517 Hussein, B.R., Malik, O.A., Ong, W.H., and Slik, J.W.F. 2022. Applications of computer vision 518 and machine learning techniques for digitized herbarium specimens: A systematic 519 literature review. Ecological Informatics 69, 101641. 520 https://doi.org/10.1016/j.ecoinf.2022.101641 521 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 22 522 Kudo, G., Nishikawa, Y., Kasagi, T. and Kosuge, S. 2004. Does seed production of spring 523 ephemerals decrease when spring comes early? Ecological Research 19, 255–259. 524 https://doi.org/10.1111/j.1440-1703.2003.00630.x 525 526 Lorieul, T., Pearson, K. D., Ellwood, E. R., Goëau, H., Molino, J.-F., Sweeney, P. W., Yost, J. 527 M., et al. 2019. Toward a large-scale and deep phenological stage annotation of 528 herbarium specimens: Case studies from temperate, tropical, and equatorial floras. 529 Applications in Plant Sciences 7(3): e1233. https://doi.org/10.1002/aps3.1233 530 531 Manangan, A., Brown, C., Saha, S., Bell, J., Hess, J., Uejio, C., Fineman, S., and Schramm, P. 532 2021. Long-term pollen trends and associations between pollen phenology and seasonal 533 climate in Atlanta, Georgia (1992-2018). Annals of allergy, asthma & immunology : 534 official publication of the American College of Allergy, Asthma, & Immunology, 127(4), 535 471–480.e4. https://doi.org/10.1016/j.anai.2021.07.012 536 537 Menzel, A., Sparks, T.H., Estrella, N., Koch, E., Aasa, A., Ahas, R., Alm-Kubler et al. 2006. 538 European phenological response to climate change matches the warming pattern. 539 Global Change Biology, 12: 1969-1976. https://doi.org/10.1111/j.1365-540 2486.2006.01193.x 541 542 Nagai, S., Saitoh, T.M. and Yoshitake, S. 2019. Cultural ecosystem services provided by 543 flowering of cherry trees under climate change: a case study of the relationship between 544 the periods of flowering and festivals. International Journal of Biometeorology 63, 1051–545 1058. https://doi.org/10.1007/s00484-019-01719-9 546 547 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 23 Pearson, K.D., Nelson, G., Aronson, M.F.J., Bonnet, P., Brenskelle, L., Davis, C.C., Denny, 548 E.G. et al. 2020. Machine Learning Using Digitized Herbarium Specimens to Advance 549 Phenological Research, BioScience 70(7), 610–620. 550 https://doi.org/10.1093/biosci/biaa044 551 552 Richardson, A.D., Keenan, T.F., Migliavacca, M., Ryu, Y., Sonnentag, O., Toomey, M. 2013. 553 Climate change, phenology, and phenological control of vegetation feedbacks to the 554 climate system. Agricultural and Forest Meteorology, 169, 156–173. 555 https://doi.org/10.1016/j.agrformet.2012.09.012 556 557 Soberón, J. and Peterson, A.T. 2004. Biodiversity informatics: managing and applying primary 558 biodiversity data. Philosophical Transactions of the Royal Society of London, Series B 559 Biological Sciences 359(1444):689-698. https://doi.org/10.1098/rstb.2003.1439. 560 561 Soltis, P.S. 2017. Digitization of herbaria enables novel research. American Journal of Botany 562 104: 1281-1284. https://doi.org/10.3732/ajb.1700281 563 564 Soltis, P. S., G. Nelson, G., and S.A.James. 2018. Green digitization: Online botanical 565 collections data answering real‐world questions. Applications in Plant Sciences 6(2), 566 e1028. https://doi.org/10.1002/aps3.1028 567 568 Stucky, B. J., Guralnick, R., Deck, J., Denny, E. G., Bolmgren, K., and Walls, R. 2018. The plant 569 phenology ontology: A new informatics resource for large-scale integration of plant 570 phenology data. Frontiers in Plant Science 9: 517. 571 https://doi.org/10.3389/fpls.2018.00517. 572 573 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 24 Visser, M.E. and Gienapp, P. 2019. Evolutionary and demographic consequences of 574 phenological mismatches. Nature Ecology & Evolution 3, 879–885. 575 https://doi.org/10.1038/s41559-019-0880-8 576 577 Williamson, D.R., Prestø, T., Westergaard, K.B., Trascau, B.M., Vange, V., Hassel, K., Koch, W. 578 and Speed, J.D.M. 2025. Long-term trends in global flowering phenology. New 579 Phytologist. https://doi.org/10.1111/nph.70139 580 581 Willis, C.G., Ellwood, E.R., Primack, R.B., Davis, C.C., Pearson, K.D., Gallinat, A.S., Yost, J.M. 582 et al. 2017. Old Plants, New Tricks: Phenological Research Using Herbarium 583 Specimens. Trends in Ecology & Evolution 32(7):531-546. 10.1016/j.tree.2017.03.015 584 585 Yadav, S., Korat, J. R., Yadav, S., Mondal, K., Kumar, A., Homeshvari, and Kumar, S. 2023. 586 Impacts of Climate Change on Fruit Crops: A Comprehensive Review of Physiological, 587 Phenological, and Pest-Related Responses. International Journal of Environment and 588 Climate Change 13 (11):363–371. https://doi.org/10.9734/ijecc/2023/v13i113179. 589 590 Younis, S., Schmidt, M., Weiland, C., Dressler, S., Seeger, B., Hickler, T. 2020. Detection and 591 annotation of plant organs from digitised herbarium scans using deep learning. 592 Biodiversity Data Journal 8: e57090. https://doi.org/10.3897/BDJ.8.e57090 593 594 Zhou, N., Siegel, Z.D., Zarecor, S., Lee, N., Campbell, D.A., Andorf, C.M., Nettleton, D. et al. 595 2018. Crowdsourcing image analysis for plant phenomics to generate ground truth data 596 for machine learning. PLoS Computational Biology 14(7):e1006337. 597 https://doi.org/10.1371/journal.pcbi.1006337 598 599 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 25 Figures 600 601 Figure 1. Examples of false positives, situations where the ensemble model consensus scored 602 the image as “flowers present”, but there are no flowers present in the image. A. and B. are 603 examples where fruits were confused for flowers. This was the most common cause of false 604 positives, accounting for over half of cases. C. is an example where leaves were confused as 605 flowers. Around 30% of false positives were situations where vegetative (non-reproductive) 606 structures were confused as flowers. 607 608 609 610 611 612 613 614 615 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 26 Figure 2. Examples of false negatives, situations where the ensemble model consensus scored 616 the image as “flowers absent”, but there are flowers present in the image. A. is an example of a 617 situation where there are obvious flowers on the sheet and the machine simply missed them. B. 618 is an example of a situation where the floral structure is not connected to the plant. This often 619 caused the machine to miss the flowers in the image. C. is an example where flowers are 620 present, but just really challenging to see. 621 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 27 Figure 3. Percent of all records pushed through the ensemble pipeline labeled as “flowers 622 present”, “flowers absent”, and “equivocal”, highlighting that the majority of records labeled as 623 “flowers present” were unusable because they were missing critical date and/or location 624 information. Only records in the “Flowers present” category that had all critical fields were 625 retained, representing 12.9% of all records. 626 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 28 627 Figure 4. Pie charts showing the top 20 families by number of records in the training dataset 628 (includes training, testing, and validation sets) and in the labeled, present data. 629 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 29 630 631 Figure 5. Number of “flowers present” labeled records per decade, beginning at 1800. The 632 dashed line indicates 2007, the year the USA National Phenology Network was founded and the 633 year before iNaturalist began. 634 635 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 30 Tables 636 Table 1. Description of the three models included in the ensemble modeling approach. 637 Threshold values are cut-offs applied to each model’s continuous output probabilities. The 638 difference between model 2 and model 3 is the loss function used. Mean square error loss 639 functions are more akin to confidence scores and are typically Gaussian, while cross-entropy is 640 better suited for classification tasks and tends to skew continuous probabilities more strongly to 641 0 or 1. 642 Model type Image size requirements Threshold low Threshold high Loss function EffNet CNN 528 x 528 pixels 0.05 0.95 Mean squared error Visual transformer network (ViT) 384 x 384 0.5 0.7 Mean squared error Visual transformer network (ViT) 384 x 384 0.5 0.7 Cross entropy 643 644 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint Grady et al. - Petal to the metal Grady et al. 31 Table 2. Expert validation of machine-labels for flowers present or absent. Bolded values 645 represent situations where the human expert and model consensus were in agreement. Grey 646 cell represents the accuracy relevant to the final, downstream data. 647 648 Expert Annotation Model consensus: Present Model consensus: Absent Present 90.5% (543) 47.7% (143) Absent 9.9% (57) 52.3% (157) 649 650 651 .CC-BY-NC 4.0 International licenseavailable under a (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-12T06:43:03.944938+00:00
License: CC-BY-NC-4.0