Abstract
14
Herbarium specimens represent critical historical records of plant phenology, yet automating 15
annotation of reproductive structures remains challenging given the diversity of floral 16
morphologies, specimen age and quality, and image quality. Here, we present a machine 17
learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium 18
specimens and deliver these data to the phenology research community. After testing multiple 19
strategies for generating training data, we found in-house expert-curated annotations were 20
essential for producing reliable results. Expert validation found relatively strong accuracy for 21
detecting present floral structures, but still had moderately high false negative rates. Applying 22
the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million 23
records labeled with flowers present. However, only 2.9 million of these contained complete 24
metadata necessary for downstream phenology research, highlighting the need for full label 25
digitization efforts. Still, this dataset represents a large compilation of historical herbarium-26
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 2
derived phenology records available as a resource for the phenology community. We end by 27
demonstrating how integrating these machine-labeled records into Phenobase, a publicly-28
available phenology database, expands taxonomic and temporal coverage for large-scale 29
phenological analyses, and discuss remaining challenges and next steps. 30
31
Key Words: computer vision; data integration; herbarium specimens; machine learning; plant 32
phenology; plant flowering 33
34
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 3
Introduction
35
Natural life cycle events, such as plant flowering, fruiting, and leaf-out, occur annually on 36
seasonally-defined schedules shaped by environmental cues. The timing and sequence of 37
these events, which vary depending on species and environment, underpin many biological and 38
human processes, including species interactions (Kudo et al., 2004), carbon cycling 39
(Richardson et al., 2013), resource acquisition (Deacy et al., 2017), agriculture (Yadav et al., 40
2023), tourism (Nagai et al., 2019), and allergy management (Manangan et al., 2023). Because 41
these seasonal changes are critical and fundamentally linked with many socio-ecological 42
processes, a whole branch of inquiry, phenology, is devoted to understanding their drivers and 43
consequences. 44
A fundamental question in phenology is how different species and ecosystems have 45
responded and will continue to respond to accelerating human-driven environmental change. 46
The rate and direction of these phenological shifts are of particular interest as they are often 47
considered one of the first and best indicators of broader ecological disruption (Menzel et al., 48
2006). For example, such shifts can interfere with key interactions that require phenological 49
synchrony, such as pollination by interacting insects, and can have cascading consequences for 50
species’ population health, ecosystems, and people (Visser et al., 2019). Thus, access to data 51
capturing the timing of key life stages across time, space, and taxa is increasingly of interest 52
and importance across many scientific and applied disciplines. 53
Concerted phenology monitoring has been a key means to generate data needed to 54
document phenological change. However, many regions across the globe don’t have phenology 55
monitoring programs, and in areas that do, the temporal extent is often limited. For example, the 56
USA National Phenology Network began coalescing observational monitoring in 2007 57
(Crimmins et al., 2022), but national-scale phenology data before this date are limited and often 58
of short duration (with some exceptions, see for example, Hough’s efforts from 1851-1859; 59
Guralnick et al., 2025). Beyond direct phenology monitoring, there are new resources, such as 60
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 4
iNaturalist, producing digital vouchers where phenological information can be secondarily 61
derived (Dinnage et al., 2025). Such contemporary contributory science platforms, like 62
iNaturalist, harness community members across the globe to generate large quantities of 63
observations linked to images and now include means to annotate phenological states. While 64
these resources are helping to further fill spatial and taxonomic data gaps and providing new 65
opportunities for phenology research at scale, they are even more limited in temporal extent 66
compared to monitoring data. 67
A critical component for defining and predicting phenological change is understanding 68
phenology and its trends over decadal to century-level timescales. Herbaria, libraries that house 69
collections of pressed plant vouchers from the field, contain specimens dating back sometimes 70
hundreds of years and have the potential to offer valuable historical phenology data, 71
complementing contemporary monitoring initiatives (Davis et al., 2015; Willis et al., 2017). 72
Recent efforts to image and digitize herbarium specimens have allowed for broader use of these 73
plant collections (Soltis 2017), with information on over 108 million angiosperm specimens (~40 74
million with associated images) available on the Global Biodiversity Information Facility (GBIF) 75
as of October 2025. Ideally, each of these herbarium specimens contains valuable information 76
about the location, collection date, species identity, and occasionally the phenological phase of 77
that plant individual. For the vast majority of plants lacking phenology annotations, it requires 78
significant effort to classify which, if any, phenological phases are visibly present on specimens. 79
Automating phenology annotation for digitized herbarium specimens is a clear next step, 80
one that would unlock vast amounts of historical data and enable phenological change research 81
across diverse taxa and environments. Recent machine learning efforts to classify phenological 82
phases in photos of live plants have been successful (Dinnage et al., 2025); however, 83
herbarium specimens present a unique set of challenges. Dried and pressed plant specimens 84
often lose color and structure over time, reproductive structures can be lost during processing, 85
some taxa have structures that are undetectable from digitized images alone or without 86
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 5
dissection, and specimens contain non-plant material such as labels and herbarium stamps that 87
may interfere with machine detection (Pearson et al., 2020). Additionally, accuracy of machine 88
phenology annotation is highly dependent on the annotator who creates training data used to 89
calibrate models, and curating a reliable and taxonomically-diverse training dataset requires 90
extensive effort and careful consideration. 91
Several works have applied machine learning to herbarium specimens, with greater 92
success in species identification than phenological recognition tasks (Hussein et al., 2022). 93
Phenological efforts have often focused on small, curated species and image sets (Goëau et al., 94
2020; Younis et al., 2020) and many have reported challenges annotating specific phenological 95
structures such as flowers (Lorieul et al., 2019; Younis et al., 2020). More recent efforts have 96
successfully addressed specific phenology research questions using models trained by 97
compositing many datasets from multiple different sources (Williamson et al., 2025). Our aim 98
here, however, was not to produce a standalone research product or answer a particular 99
question, but to develop a trustable, usable, and integratable data resource for the broader 100
research community. Doing so required careful curation of our own training data, which allowed 101
us to leverage information gathered during training to reduce noise in the final data product, as 102
we discuss belowt. 103
Here, we present an automated phenology annotation pipeline that accurately detects 104
flowers present on digitized herbarium specimens. With the goal of maximizing geographic and 105
taxonomic coverage across the angiosperm tree of life, this effort unlocks vast quantities of 106
previously limited historical phenology information. We outline the challenges, successes, 107
complexities and strategic decisions involved in developing the pipeline, emphasizing the need 108
for high-quality training data generated by botanically-knowledgeable individuals and rigorous 109
filtering on multiple dimensions both pre- and post-modeling. We discuss the trade off between 110
strict data filtering for improved model accuracy and data retention to maximize the volume of 111
downstream historical records annotated with phenology information. We highlight a surprising 112
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 6
and disappointing discovery that the majority of specimens lack metadata sufficient for 113
phenological annotation, which we discovered after all the rest of the effort. Finally, we 114
demonstrate the vast scale of previously-unavailable historical phenology data that are now 115
available to stakeholders, and discuss the integration of these data into Phenobase 116
(https://phenobase.netlify.app/), an online database of millions of research-ready, community-117
sourced phenology data from field images and from structured in situ phenology monitoring 118
programs. 119
120
Methods
121
Gathering specimen image and metadata 122
We opted to gather data for this work from GBIF, which provides access to billions of species 123
occurrence records and, in many cases, associated images. We searched GBIF in mid-2023 for 124
all “Magnoliopsida” and “Liliopsida”, using filters to select the basis of record as “preserved 125
specimen” and media type as “images”. This produced 33,691,939 occurrence records with 126
35,314,854 corresponding multimedia records. We did not filter out records missing coordinates 127
or location information so that when metadata of a record is updated in the future, its associated 128
phenological information will have already been extracted. We downloaded this dataset and 129
associated images in October 2024 and, using a nearest pixel filter, programmatically resized 130
each image to 1024 pixels along the shorter edge to reduce storage and satisfy GPU memory 131
limitations. Because services are often down and media links can change, the number of 132
images downloaded is a subset of the total number of records in the original GBIF dataset. In 133
sum, after contending with link issues, removing small or low quality images, and images 134
associated with taxa that we decided to filter out (described below), we collected 22,447,563 135
images for downstream labeling and model training. 136
137
Labeling training data 138
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 7
A challenging aspect of developing a reliable automated pipeline for phenophase annotation on 139
herbarium specimens across the angiosperm tree of life was generating enough high-quality 140
training data. Our goal was to produce a trusted historical phenology data resource for broad 141
use, and we therefore deliberately chose to manually generate our own training data rather than 142
utilizing data from multiple external sources. As mentioned above, manually scoring pressed 143
plant specimens for phenophases is time consuming and difficult, and a key challenge is 144
balancing the need for botanical expertise with the capacity of expert annotators. Other 145
initiatives have employed volunteers, undergraduate students, and paid non-experts with 146
relative success (Brenskelle et al., 2020; Zhou et al., 2018), however these focused on a few 147
key taxa with easily-distinguishable structures. Limiting annotations to a few taxa at the genus 148
or family level allows for more in-depth, targeted training and familiarity with target taxa which 149
likely lead to better results. We were unsure whether employing non-experts would be 150
successful when scoring thousands of diverse taxa. 151
We first tested whether a paid undergraduate student could produce accurate 152
phenophase annotations across all flowering plants. We employed a student to score a random 153
selection of specimens for presence or absence of open flowers. If the student was uncertain or 154
unable to make an informed decision about the digitized image, they had the option to mark 155
“unknown”. We then also had a botanically-knowledgeable individual validate these annotations 156
to assess their accuracy. As reported in Results, we found the undergraduate’s open flower 157
annotations were not at a level usable for training a model. Based on these results and other 158
work that suggests that one of the most important variables in annotation accuracy is simply the 159
annotator themselves (Brenskelle et al., 2020), we made the conscious choice to employ one in-160
house botanically-knowledgeable individual to generate all training data. This decision greatly 161
reduced our capacity to generate high quantities of training data, but ensured high-quality and 162
consistency, which was our priority. 163
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 8
In addition to the first set of specimens annotated to validate the undergraduate 164
annotations, we annotated a second additional set of specimens for open flowers to increase 165
training data. While compiling the second set, we randomly selected specimens in families 166
underrepresented in the first set in an effort to promote taxonomic evenness. Additional records 167
were added to the training data iteratively through repeated bouts of image validation 168
throughout model training. 169
170
Filtering training data 171
To ensure accurate and reliable results, we opted to filter out taxa pre-model-training that were 172
challenging to manually annotate. First, we removed genera that had more than 5 records in the 173
training dataset and an “unknown” annotation rate greater than or equal to 25%. After removing 174
these genera, we removed families using the same criteria. Many of these removed taxa were 175
those with small reproductive structures, such as Chenopodiaceae, or those with flowers and 176
fruits that are difficult to distinguish from each other in images alone, like Poaceae. These taxa 177
were removed permanently from our pipeline, both from training and downstream data. We 178
made this choice strategically to balance data quality and quantity. Throughout this process, we 179
found it critical to acknowledge and respond to limitations, both in human expertise, machine 180
capabilities, and overall capacity. Final lists of removed genera and families are available in 181
Supporting Information. 182
183
Model training and post-training thresholding 184
Model Ensembling Approach 185
One challenge with training models on all herbarium data is the vast and complex floral diversity 186
across the whole of the group, challenging models to properly abstract features that can apply 187
broadly and that are not easily confusable with other objects. We realized single models often 188
performed modestly well, but that ensembled models might perform better than any single one 189
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 9
given this variability. We tried several different neural networks, architectures, and model 190
parameters and ultimately determined that an ensemble modeling approach, combining three 191
models, was most effective based on downstream validation statistics. Note that all three 192
models were fine tuned versions of models that were previously pretrained on ImageNet data 193
(Deng et al., 2009). 194
Despite differences across the three models (Effnet, ViTMSE and VITCE, summarized in 195
Table 1), all training followed the same general steps. First, expert-labeled records were split 196
into 60% training, 20% testing, and 20% validation sets. After low-quality images were filtered 197
out, the images were resized to match the size requirements of the specific model (Table 1). 198
Images in the training dataset were augmented using random horizontal and vertical flips and 199
“auto augment” (Cubuk et al., 2018), which uses multiple techniques including random rotation, 200
color jitter, sharpness adjustment, and image shear. Next, all models were color adjusted to use 201
the ImageNet color mean and standard deviation. During training, each model was saved as to 202
five checkpoints The final best model was chosen as the checkpoint with the best F1 score on 203
the holdout validation dataset. The F1 score is the harmonic mean of precision (how many 204
positive model predictions are actually positive?) and recall (how many actual positives did the 205
model detect?). We chose the F1 score because it is particularly useful for assessing predictive 206
performance for the positive class (“flowers present”), which is the focus of this effort. This 207
overall approach is standard for model selection and ensures comparable scoring across model 208
approaches (e.g. EffNet CNN versus ViT) and that models are tuned, selected, and ensembled 209
based on their best validation performance. 210
We ran models using a single A100 GPU on the University of Florida’s HiPerGator. Each 211
model was trained for at least 200 epochs, and we saved five checkpoints during training. At the 212
end of every epoch, we evaluated model performance on a held-out validation set that was 213
never used for backpropagation, which we used to assure model generalization and select the 214
best checkpoint. We used several metrics for evaluating the model, including F1 score, 215
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 10
accuracy, validation set loss, precision, and recall. After training, each checkpoint from all three 216
models was post-processed to determine the probability thresholds that maximized accuracy. 217
We divided the range [0.0–1.0] into three regions defined by two thresholds. Images with a 218
score above the high threshold were considered a positive indication of the trait, images with 219
scores below the low threshold were considered a negative indication of the trait. Images with 220
scores between the high and low thresholds were classified as “equivocal”, meaning the model 221
could not decide if the trait was present or absent. We adjusted those two thresholds until the 222
images in the test (holdout) dataset accuracy was at the maximum. We also guarded against 223
equivocal records being more than 30% of the test dataset. We used a batch size of 32, the 224
initial learning rate was 5e-5 with a weight decay parameter of 0.01, and an AdamW optimizer. 225
For each image, each model generated a final outcome of positive, negative, or 226
equivocal. To ensemble our three models, we used a two-thirds “majority rule” approach, 227
focusing in particular on predictions of “flowers present” (e.g., where at least two models of the 228
three predicted present with high certainty). Since herbarium specimens often contain only a 229
portion of a plant individual and absence on the whole plant cannot be inferred, we focused on 230
the accuracy of “present” labels and only retained images machine-labeled as “present” in 231
downstream data. All other cases were scored “flowers not present” or “equivocal”. Note that the 232
ensemble itself can be undecided (equivocal), for instance with [positive, negative, equivocal] or 233
two or more equivocal votes from the three model outputs. 234
235
Model validation 236
We validated ensemble results both using held-out test data and through additional 237
expert validation. For the expert validation, we machine-labeled a random subset of images and 238
validated 600 specimens machine-labeled as “present” for flowers and 300 specimens machine-239
labeled as “absent” for flowers. Consistent with the Plant Phenology Ontology (Stucky et al., 240
2018), “flowers” here refers to any floral structure, including both flower buds and open flowers. 241
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 11
Although the models were originally trained on “open flowers”, we opted to conduct validation at 242
the broader “flower” level, as these tasks are difficult and we felt “flower” better reflected the 243
features detected by the ensemble models. To further understand instances where the model 244
consensus was incorrect, we anecdotally looked at false positives and false negatives to 245
understand general patterns in the mistakes (see Results). 246
247
Machine labeling and integration with other sources 248
After model validation, the remaining unlabeled downloaded images were passed through the 249
ensemble model pipeline. We retained only images labeled as “present” by at least two out of 250
three models and data fields were standardized for integration in Phenobase, which has a 251
defined set of required and optional fields. As we discuss below, we dropped records that we 252
labelled but were missing these key fields that are required for phenology research. Key fields 253
here include: scientific name, event date, and decimal latitude and longitude. The remaining 254
data were made publicly available through the Phenobase web portal 255
(https://phenobase.netlify.app/). 256
257
Evaluating specimens missing critical fields 258
A key challenge and major limitation during this process was that a large number of specimens 259
did not have date or location information included in their metadata. As these two pieces of 260
information are critical for meaningful phenology data, this resulted in the loss of substantial 261
data. To further investigate this issue, we randomly selected 100 specimens labeled as “flowers 262
present” that were missing both date and location information. We manually annotated these 263
specimens for whether or not they included specific (day, month and year) date information and 264
georeferencable locality information on the specimen label. 265
266
Results
267
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 12
Labeling training data 268
We worked with an undergraduate student to annotate a total of 3,525 specimens for open 269
flowers, providing them with resources for the task and asking them to use online resources 270
when needed. When validating these annotations, the undergraduate and one of our team 271
members (lead author EG), who served as a botanical expert, were in agreement 69% of the 272
time. Given this result, we chose to not include the undergraduate annotations in our final 273
training dataset. EG then generated a total of 9,004 annotations from individual specimens for 274
open flowers. After removing 49 genera and an additional seven families (Appendix S1; see 275
Supporting Information with this article), 1660 genera and 261 families remained. Our final 276
training dataset contained 7,205 specimen records, 4319 annotated as open flowers present 277
and 2886 annotated as open flowers absent. Once expert-labeled records were split, 4324 278
images were included in the training set, 1442 in the testing set, and 1439 in the validation set. 279
280
Model fitting and validation 281
We utilized three models (Table 1) and ensembled their results to predict the presence of 282
flowers. Our training model accuracy scores across all three models varied between 88.3% and 283
92.5%. In general, the vision transformers models (ViT) performed better (92.2% and 92.5% 284
accuracy) than the EffNet model (88.3% success). Overall, from the full ensemble, validation on 285
the held-out test data found 96.3% accuracy for detecting the presence of flowers and 86.8% 286
accuracy for detecting the absence of flowers (Appendix S2). However, expert validation on 287
records not in the initial training, testing, or validation processes found the ensemble models to 288
be 90.5% accurate for detecting the presence of flowers and 52.3% accurate for the absence of 289
flowers (Table 2). The low accuracy for detecting absences was expected and likely due both to 290
expert-validated data being more diverse than training data and that we were less concerned 291
with training for absences in our overall process. Further investigation of the false positives from 292
the expert validation found that the majority of false positives were fruits mistaken for flowers, 293
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 13
although some sheets had no reproductive structures present and one sheet had a fruit 294
illustration (Figure 1). Investigation into false negatives from the expert validation found a variety 295
of reasons for the mistakes, many explainable, including situations where flowers weren’t 296
attached to the main plant, flowers that may have been confused as fruits, and flowers that were 297
very challenging to see (Figure 2). 298
299
Machine labeling and integration with other sources 300
After downloading and filtering out problematic images (images that could not be downloaded, 301
images that cannot be read, or images that were either less than 10 KB or greater than 32 MB) 302
and images of hard-to-annotate genera and families (Appendix S1), we were left with 303
22,447,563 images of digitized herbarium specimens. These images were fed through the 304
ensemble modeling pipeline and annotated for presence or absence of flowers. Of these 305
records, 11,131,820 were labeled as flowers present (meaning at least ⅔ models scored as 306
present), but only 2,913,353 of these contained valid latitude, longitude, and date information 307
(Figure 3). 5,748,788 images were scored as flowers absent (meaning at least ⅔ models scored 308
as absent), and 5,566,955 were equivocal (meaning there was no consensus among the 309
models). Only records scored as flowers present and with the needed metadata were retained 310
and integrated into Phenobase (https://phenobase.netlify.app/). 311
312
Data coverage and evaluating specimens missing critical information 313
The 2,913,35 records annotated as flowers present represented 383 plant families and 8,883 314
genera (Figure 4). They represent 320 years of collecting, with the earliest record dated in 1501 315
(Figure 5). We evaluated a small subset of the 8,218,467 records that were missing locality and 316
date information to determine exactly what fields were missing and if those issues could be 317
resolved with more complete digitization of labels. Out of the 100 random specimens that were 318
missing date and latitude/longitude information, 65 had a date (specific day and year) on the 319
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 14
specimen label while 35 did not; 85 had a locality that is georeferenceable and localizable to a 320
unit lower than country, while 15 did not. A total of 57 out of 100 had both a georeferenceable 321
specific locality and date information on the label. Only 44 of the 100 records had images 322
retrievable from the original URL, while for 56 records the image URL failed. 323
324
Discussion
325
While machine-automated annotation is a clear next step in addressing bottlenecks and 326
delivering large quantities of valuable data to the research community, some annotation tasks 327
are harder than others. Further, tasks that are relatively straightforward in one context, such as 328
annotating flowers in photographs of live plants, can be much more challenging in other 329
contexts, such as attempting to do the same for dried, pressed, and aged plants on herbarium 330
sheets. Other groups have produced herbarium data using machine labeling approaches to 331
answer their own research questions about machine learning, certain taxa, or regions, and 332
many showcase reasonably high precision and recall (Goëau et al., 2020; Lorieul et al., 2019; 333
Williamson et al., 2025; Younis et al., 2020). Here, however, our focus was on creating a 334
reliable resource at the broadest possible scale and for the phenology research community as a 335
whole; we were not focused on answering specific research questions. Our first priority was that 336
stakeholders could trust, understand, and easily use these data. A key challenge for us was to 337
develop a set of solutions that we felt met a reasonable bar for trustworthiness, which lead to 338
rigorous data filtering, and to clearly communicate where these solutions may fall short. Here we 339
are not wholly concerned with the size of the data product or any other metric of success. 340
Rather, creating trustworthy data in this context meant carefully vetting the quality of training 341
data, evaluating validation statistics, and being thoughtful about the data structure and 342
documentation made available to stakeholders so they can best use these data and understand 343
their value and limitations. 344
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 15
We also caution that the most appealing solutions for annotating flowers may be difficult 345
at the taxonomic, spatial and temporal scale we attempted here. Our initial hope when 346
generating training data was that we could produce flower annotations utilizing a team of 347
people, likely undergraduate students, with limited botanical training. While this approach might 348
work when focusing on a small subset of plants, such as a genus or family, when scaling to all 349
angiosperms we found it untenable. A vast, complicated diversity of inflorescence morphologies 350
interacts with herbarium sheet taphonomic processes to make it especially challenging to detect 351
flower presence, even for an expert botanist. We argue that the ideal solution for creating high 352
quality annotations is delegating annotation tasks to botanists working on taxonomic groups 353
where they have expertise, but this is also impractical given capacity and logistical constraints. 354
Our solution here was using one expert with ample practice scoring sheets, which has both 355
benefits and drawbacks. One major benefit is that there is likely more consistency and higher 356
quality with this method than a crowdsourced solution. This is especially important in dealing 357
with the many cases where flower presence was uncertain, which was critical for removing 358
certain families and genera from downstream annotation labeling. However, a major drawback 359
in using only one annotator to generate training data is that this method is much less scalable, 360
generating significantly fewer annotated records to use for training. Finally, more focused 361
expertise in certain taxonomic groups may have produced less “unknowns” and higher quality 362
training data for those groups that could lead to inclusion in downstream modeling. 363
A key step in our process was making thoughtful, informed decisions about removing 364
taxa that we felt were too challenging to score correctly. While annotating training data, the 365
annotator could choose “present” or “absent” options to inform the models, but critically they 366
also had an “unknown” option which they were encouraged to select whenever they could not 367
make a confident determination. Ultimately, the proportion of “unknown” annotations, for 368
individual genera and also families, was what we used to determine which taxa were simply too 369
challenging to score consistently for flower presence. There is likely no perfect cutoff, but we 370
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 16
decided that taxa where more than 25% of records were too difficult to label were likely to be 371
problematic for modeling. This decision represents a tradeoff between data quantity and quality. 372
While it is always disappointing to exclude data, we prioritized maintaining stakeholder trust by 373
favoring reliability over volume. We maintain that when attempting to scale challenging tasks 374
such as phenological annotation of herbarium specimens, it is critical to acknowledge both 375
human and technological limits and adjust methods accordingly. 376
Even with substantial taxonomic filtering, model performance was middling on both held-377
out test data and expert validated data. Given the challenges with fitting a single model, we 378
opted for an ensemble approach, which gave us more power to interrogate multiple models and 379
arrive at a consensus. Ensembling is particularly useful because it reduces idiosyncratic errors 380
from any single model. We required a two-thirds majority to score presence from the 3 separate 381
models. This strategy required more up-front effort, but ultimately produced a more rigorous 382
labeling workflow. As in other approaches we have published (Dinnage et al., 2025), our interest 383
is in phenophase presence, and the ensembled models are fairly good at avoiding false 384
positives, but we note that the model still produced a substantial number of false absences. 385
Since we explicitly tuned the models to minimize false positives at the cost of increasing false 386
negatives, it means we are likely throwing out many images that have flowers. 387
The filtering process didn’t end after we completed machine labeling. Out of 11 million 388
records labeled as flowers present, we found that more than 70% lacked a reported 389
latitude/longitude and/or a date. This is consistent with the overall proportion of specimen 390
records on GBIF that have missing coordinates or spatial issues. When we queried GBIF on 391
December 16, 2025, for example, there are 52,254,706 records with associated specimen 392
images for Magnoliopsida and Liliopsida. However, only about 32% of them (16,786,460) have 393
latitude and longitude coordinates and without geospatial issues. This data incompleteness 394
percent is disappointingly high and effectively limits the utility of these data for research, as date 395
and location information is generally required for these data to be research ready (sensu Soltis 396
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 17
et al., 2018). The root cause of this issue appears to be due to the publication of “stub records” 397
to GBIF. These stub records may have taxon names but often lack dates and/or georeferences 398
and they generally are from very large collections. For example, the Museum National d'Histoire 399
Naturelle de Paris published 5,031,777 angiosperm records with images and this provider was 400
one where records were often found to be lacking date and geospatial data. Further, we should 401
note that many of the images we annotated are no longer available on GBIF. Out of the 100 we 402
checked, more than half no longer had resolvable images online; most of those also came from 403
publishers who often supply a large corpus of images. These significant issues bring home the 404
need for efforts to expedite full label digitization to avoid these issues of missing metadata since 405
many labels do contain dates and can be georeferenced. Such efforts are needed to support 406
the best use of primary and secondary data (sensu Soberón and Peterson, 2004), such as 407
phenology annotations generated from images. Creating more persistence for images tied to 408
specimens will also better ensure reproducibility and stability when relating specimen images, 409
core specimen metadata and secondary data derived from those images. 410
Despite setbacks, our ensemble modeling flower annotation pipeline unlocks vast new 411
global phenology data at unprecedented temporal scales. Although in situ phenology data 412
collection networks (Crimmins et al., 2022) and other initiatives to unlock phenology information 413
from sources like iNaturalist (Dinnage et al., 2025) are providing large quantities of 414
contemporary phenology data, data predating these initiatives are less common. They are often 415
in the form of tabular reports that were monitored at one site over many years, although some 416
larger citizen science projects covered multiple sites and years such as the New York 417
Phenology Project (Fuccillo Battle et al., 2022) and Eastern USA Phenology Project (Guralnick 418
et al., 2024) in the middle and late 1800’s. The almost 3 million labeled records that we present 419
here represent a large dataset of historical (pre-2000) phenology now accessible for the 420
research community. However, we note that significant gaps remain in phenology data 421
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 18
coverage across space and time, and that continued digitization and efforts to add missing 422
metadata to existing digitized records will be essential for continuing to close these gaps. 423
All records labeled as flowers present from this effort are available through Phenobase, 424
an open-access datastore of plant phenology information that integrates in situ data from 425
phenology monitoring programs, machine-labeled data from iNaturalist images, and now 426
machine-labeled data from herbarium specimens. Our pipeline produces phenology labels and 427
fields consistent with the Phenobase data structure and Plant Phenology Ontology (Stucky et 428
al., 2018), and uses a controlled vocabulary for other fields, ensuring that data users can search 429
for phenology records at any level of granularity and get consistent results. Future work may 430
include efforts to label herbarium specimens for fruits, however our early explorations found this 431
task to be even more challenging than labeling for flowers, likely requiring greater human effort, 432
and yet more rigorous filtering steps. In sum, despite the challenges, providing information on 433
historical flowering dates across the globe and integrating with other data sources enables new 434
analyses and insights spanning regions and taxa, enabling yet broader and more synthetic 435
questions to be asked across temporal, taxonomic and spatial scales. 436
437
Acknowledgements
438
The authors would like to thank the many individuals who have contributed to collecting, 439
processing, digitizing, and maintaining herbarium specimens, as this effort would not be 440
possible without their dedication. Phenobase is a team effort and the authors acknowledge team 441
members Carrie Seltzer, Ramona Walls, and Taran Lichtenberger. UFIT Research Computing 442
provided computational resources and support. This work was funded by the National Science 443
Foundation (DBI2223512 to Robert Guralnick and DBI2223508 to Daijiang Li). 444
445
Author Contributions 446
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 19
Robert Guralnick, Daijiang Li and Raphael LaFrance initially conceived of this effort. Erin Grady 447
operationalized most aspects of gathering training data, in consultation with Raphael LaFrance 448
and Robert Guralnick. Raphael LaFrance developed the modeling framework for this effort with 449
help from Russell Dinnage. Raphael LaFrance produced final models and descriptive statistics 450
and Erin Grady led expert model validation. Ellen Denny and John Deck helped ensure that 451
data in the right format could be ingested into Phenobase. Erin Grady and Robert Guralnick 452
archived data. Erin Grady and Robert Guralnick wrote the manuscript with help from Raphael 453
LaFrance, Russell Dinnage and Daijiang Li. All authors contributed to drafts and gave final 454
approval for publication. 455
456
Data Availability Statement 457
The ensemble data models and a corresponding JSON file with model metadata data are 458
housed on Zenodo (https://doi.org/10.5281/zenodo.17079402). Images used in training, 459
validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460
this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0). 461
Training data and final ensemble output can be found on Zenodo 462
(https://doi.org/10.5281/zenodo.17675089). 463
464
Supporting Information 465
Additional Supporting Information may be found online in the Supporting Information section at 466
the end of the article. 467
Appendix S1. List of difficult-to-annotate genera and families removed from training and 468
downstream data. 469
Appendix S2. Table S1. Validation results for held-out test data. 470
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 20
References
471
Brenskelle, L., Guralnick, R.P., Denslow, M., and Stucky, B.J. 2020. Maximizing human effort 472
for analyzing scientific images: A case study using digitized herbarium sheets. 2020. 473
Applications in Plant Sciences 8(6): e11370. https://doi.org/10.1002/aps3.11370 474
475
Crimmins, T., Denny, E., Posthumus, E., Rosemartin, A., Croll, R., Montano, M., and Panci, H. 476
2022. Science and Management Advancements Made Possible by the USA National 477
Phenology Network's Nature's Notebook Platform. BioScience 72(9): 908–920. 478
https://doi.org/10.1093/biosci/biac061 479
480
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. 2018. AutoAugment: Learning 481
augmentation policies from data. In arXiv [cs.CV]. arXiv. 482
https://doi.org/10.48550/arXiv.1805.09501 483
484
Davis, C.C., Willis, C.G., Connolly, B., Kelly, C. and Ellison, A.M. 2015. Herbarium records are 485
reliable sources of phenological change driven by climate and provide novel insights into 486
species' phenological cueing mechanisms. American Journal of Botany 102: 1599-1609. 487
https://doi.org/10.3732/ajb.1500237 488
489
Deacy, W.W., Armstrong, J.B., Leacock, W.B., Robbins, C.T., Gustine, D.D., Ward, E.J., 490
Erlenbach, J.A. and Stanford, J.A. 2017. Phenological synchronization disrupts trophic 491
interactions between Kodiak brown bears and salmon. Proceedings of the National 492
Academy of Sciences of the United States of America 114 (39): 10432-10437. 493
https://doi.org/10.1073/pnas.1705248114 494
495
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 21
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. 2009. ImageNet: A large-scale 496
hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern 497
Recognition: 248–255. 498
499
Dinnage, R., Grady, E., Neal, N., Deck, J., Denny, E., Walls, R., Seltzer, C. et al. 2025. 500
PhenoVision: A framework for automating and delivering research-ready plant 501
phenology data from field images. Methods in Ecology and Evolution 16(8): 1763-1780. 502
503
Fuccillo Battle, K., Duhon, A., Vispo, C. R., Crimmins, T. M., Rosenstiel, T. N., Armstrong-504
Davies, L. L., and de Rivera, C. E. 2022. Citizen science across two centuries reveals 505
phenological change among plant species and functional groups in the Northeastern US. 506
The Journal of Ecology 110(8): 1757–1774. https://doi.org/10.1111/1365-2745.13926 507
508
Goëau, H., A. Mora-Fallas, J. Champ, N. L. R. Love, S. J. Mazer, E. Mata-Montero, A. Joly, and 509
P. Bonnet. 2020. A new fine-grained method for automated visual analysis of herbarium 510
specimens: A case study for phenological data extraction. Applications in Plant Sciences 511
8(6): e11368. https://doi.org/10.1002/aps3.11368 512
513
Guralnick, R., Crimmins, T., Grady, E., and Campbell, L. Phenological response to climatic 514
change depends on spring warming velocity. 2024. Communications Earth & 515
Environment 5, 634. https://doi.org/10.1038/s43247-024-01807-8 516
517
Hussein, B.R., Malik, O.A., Ong, W.H., and Slik, J.W.F. 2022. Applications of computer vision 518
and machine learning techniques for digitized herbarium specimens: A systematic 519
literature review. Ecological Informatics 69, 101641. 520
https://doi.org/10.1016/j.ecoinf.2022.101641 521
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 22
522
Kudo, G., Nishikawa, Y., Kasagi, T. and Kosuge, S. 2004. Does seed production of spring 523
ephemerals decrease when spring comes early? Ecological Research 19, 255–259. 524
https://doi.org/10.1111/j.1440-1703.2003.00630.x 525
526
Lorieul, T., Pearson, K. D., Ellwood, E. R., Goëau, H., Molino, J.-F., Sweeney, P. W., Yost, J. 527
M., et al. 2019. Toward a large-scale and deep phenological stage annotation of 528
herbarium specimens: Case studies from temperate, tropical, and equatorial floras. 529
Applications in Plant Sciences 7(3): e1233. https://doi.org/10.1002/aps3.1233 530
531
Manangan, A., Brown, C., Saha, S., Bell, J., Hess, J., Uejio, C., Fineman, S., and Schramm, P. 532
2021. Long-term pollen trends and associations between pollen phenology and seasonal 533
climate in Atlanta, Georgia (1992-2018). Annals of allergy, asthma & immunology : 534
official publication of the American College of Allergy, Asthma, & Immunology, 127(4), 535
471–480.e4. https://doi.org/10.1016/j.anai.2021.07.012 536
537
Menzel, A., Sparks, T.H., Estrella, N., Koch, E., Aasa, A., Ahas, R., Alm-Kubler et al. 2006. 538
European phenological response to climate change matches the warming pattern. 539
Global Change Biology, 12: 1969-1976. https://doi.org/10.1111/j.1365-540
2486.2006.01193.x 541
542
Nagai, S., Saitoh, T.M. and Yoshitake, S. 2019. Cultural ecosystem services provided by 543
flowering of cherry trees under climate change: a case study of the relationship between 544
the periods of flowering and festivals. International Journal of Biometeorology 63, 1051–545
1058. https://doi.org/10.1007/s00484-019-01719-9 546
547
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 23
Pearson, K.D., Nelson, G., Aronson, M.F.J., Bonnet, P., Brenskelle, L., Davis, C.C., Denny, 548
E.G. et al. 2020. Machine Learning Using Digitized Herbarium Specimens to Advance 549
Phenological Research, BioScience 70(7), 610–620. 550
https://doi.org/10.1093/biosci/biaa044 551
552
Richardson, A.D., Keenan, T.F., Migliavacca, M., Ryu, Y., Sonnentag, O., Toomey, M. 2013. 553
Climate change, phenology, and phenological control of vegetation feedbacks to the 554
climate system. Agricultural and Forest Meteorology, 169, 156–173. 555
https://doi.org/10.1016/j.agrformet.2012.09.012 556
557
Soberón, J. and Peterson, A.T. 2004. Biodiversity informatics: managing and applying primary 558
biodiversity data. Philosophical Transactions of the Royal Society of London, Series B 559
Biological Sciences 359(1444):689-698. https://doi.org/10.1098/rstb.2003.1439. 560
561
Soltis, P.S. 2017. Digitization of herbaria enables novel research. American Journal of Botany 562
104: 1281-1284. https://doi.org/10.3732/ajb.1700281 563
564
Soltis, P. S., G. Nelson, G., and S.A.James. 2018. Green digitization: Online botanical 565
collections data answering real‐world questions. Applications in Plant Sciences 6(2), 566
e1028. https://doi.org/10.1002/aps3.1028 567
568
Stucky, B. J., Guralnick, R., Deck, J., Denny, E. G., Bolmgren, K., and Walls, R. 2018. The plant 569
phenology ontology: A new informatics resource for large-scale integration of plant 570
phenology data. Frontiers in Plant Science 9: 517. 571
https://doi.org/10.3389/fpls.2018.00517. 572
573
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 24
Visser, M.E. and Gienapp, P. 2019. Evolutionary and demographic consequences of 574
phenological mismatches. Nature Ecology & Evolution 3, 879–885. 575
https://doi.org/10.1038/s41559-019-0880-8 576
577
Williamson, D.R., Prestø, T., Westergaard, K.B., Trascau, B.M., Vange, V., Hassel, K., Koch, W. 578
and Speed, J.D.M. 2025. Long-term trends in global flowering phenology. New 579
Phytologist. https://doi.org/10.1111/nph.70139 580
581
Willis, C.G., Ellwood, E.R., Primack, R.B., Davis, C.C., Pearson, K.D., Gallinat, A.S., Yost, J.M. 582
et al. 2017. Old Plants, New Tricks: Phenological Research Using Herbarium 583
Specimens. Trends in Ecology & Evolution 32(7):531-546. 10.1016/j.tree.2017.03.015 584
585
Yadav, S., Korat, J. R., Yadav, S., Mondal, K., Kumar, A., Homeshvari, and Kumar, S. 2023. 586
Impacts of Climate Change on Fruit Crops: A Comprehensive Review of Physiological, 587
Phenological, and Pest-Related Responses. International Journal of Environment and 588
Climate Change 13 (11):363–371. https://doi.org/10.9734/ijecc/2023/v13i113179. 589
590
Younis, S., Schmidt, M., Weiland, C., Dressler, S., Seeger, B., Hickler, T. 2020. Detection and 591
annotation of plant organs from digitised herbarium scans using deep learning. 592
Biodiversity Data Journal 8: e57090. https://doi.org/10.3897/BDJ.8.e57090 593
594
Zhou, N., Siegel, Z.D., Zarecor, S., Lee, N., Campbell, D.A., Andorf, C.M., Nettleton, D. et al. 595
2018. Crowdsourcing image analysis for plant phenomics to generate ground truth data 596
for machine learning. PLoS Computational Biology 14(7):e1006337. 597
https://doi.org/10.1371/journal.pcbi.1006337 598
599
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 25
Figures 600
601
Figure 1. Examples of false positives, situations where the ensemble model consensus scored 602
the image as “flowers present”, but there are no flowers present in the image. A. and B. are 603
examples where fruits were confused for flowers. This was the most common cause of false 604
positives, accounting for over half of cases. C. is an example where leaves were confused as 605
flowers. Around 30% of false positives were situations where vegetative (non-reproductive) 606
structures were confused as flowers. 607
608
609
610
611
612
613
614
615
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 26
Figure 2. Examples of false negatives, situations where the ensemble model consensus scored 616
the image as “flowers absent”, but there are flowers present in the image. A. is an example of a 617
situation where there are obvious flowers on the sheet and the machine simply missed them. B. 618
is an example of a situation where the floral structure is not connected to the plant. This often 619
caused the machine to miss the flowers in the image. C. is an example where flowers are 620
present, but just really challenging to see. 621
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 27
Figure 3. Percent of all records pushed through the ensemble pipeline labeled as “flowers 622
present”, “flowers absent”, and “equivocal”, highlighting that the majority of records labeled as 623
“flowers present” were unusable because they were missing critical date and/or location 624
information. Only records in the “Flowers present” category that had all critical fields were 625
retained, representing 12.9% of all records. 626
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 28
627
Figure 4. Pie charts showing the top 20 families by number of records in the training dataset 628
(includes training, testing, and validation sets) and in the labeled, present data. 629
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 29
630
631
Figure 5. Number of “flowers present” labeled records per decade, beginning at 1800. The 632
dashed line indicates 2007, the year the USA National Phenology Network was founded and the 633
year before iNaturalist began. 634
635
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 30
Tables 636
Table 1. Description of the three models included in the ensemble modeling approach. 637
Threshold values are cut-offs applied to each model’s continuous output probabilities. The 638
difference between model 2 and model 3 is the loss function used. Mean square error loss 639
functions are more akin to confidence scores and are typically Gaussian, while cross-entropy is 640
better suited for classification tasks and tends to skew continuous probabilities more strongly to 641
0 or 1. 642
Model type Image size
requirements
Threshold
low
Threshold
high
Loss
function
EffNet CNN 528 x 528
pixels
0.05 0.95 Mean
squared error
Visual transformer network
(ViT)
384 x 384 0.5 0.7 Mean
squared error
Visual transformer network
(ViT)
384 x 384 0.5 0.7 Cross
entropy
643
644
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Grady et al. - Petal to the metal Grady et al. 31
Table 2. Expert validation of machine-labels for flowers present or absent. Bolded values 645
represent situations where the human expert and model consensus were in agreement. Grey 646
cell represents the accuracy relevant to the final, downstream data. 647
648
Expert Annotation Model consensus: Present Model consensus: Absent
Present 90.5% (543) 47.7% (143)
Absent 9.9% (57) 52.3% (157)
649
650
651
.CC-BY-NC 4.0 International licenseavailable under a
(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made
The copyright holder for this preprintthis version posted December 22, 2025. ; https://doi.org/10.64898/2025.12.19.695552doi: bioRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.