⚙
AI-generated deep summary
by claude@2026-07, 2026-07-06
· read from full text
ⓘ
This paper describes how the DiSSCo research infrastructure uses “Machine Learning as a Service” to extract quantitative specimen data, focusing on automating interpretation of scientific collection material rather than biomedical discovery. At a high level, it presents an infrastructure/technical approach for turning specimen-related information into structured quantitative outputs using ML services. The main finding is that an ML-as-a-service workflow can operationalize specimen data extraction within a distributed system for scientific collections, with the explicit caveat that the work is an infrastructure-focused research idea/preprint rather than a biomedical study with clinical outcomes. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.
Abstract
The Distributed System for Scientific Collections (DiSSCo) is a research infrastructure to integrate European natural science collections (NSCs) digitally. The aim is to facilitate and enhance the access, management and analysis of collection assets in one unified digital collection. The Machine Annotation Services (MAS) are essential components of DiSSCo’s Digital Specimen Architecture (DSArch). These services automate the annotation of digital objects to enable labeling and categorization of NSC's digital assets. To further advance this, a Machine Learning as a Service (MLaaS) approach was developed which provides researchers with the access to pre-trained machine learning models for complex tasks such as instance segmentation and morphological analysis of datasets. MLaaS enhances the DiSSCo’s scalability and flexibility and allows the integration of machine learning tools in close alignment with the FAIR (Findable, Accessible, Interoperable, Reusable) principles. This study employs DiSSCO's MLaaS framework for the quantitative analysis of herbarium specimens. Machine learning models such as Mask R-CNN and YOLO11 are comparatively applied to detect and generate the pixel-level masks of plant organs in herbarium sheets. Subsequently, these models are used to reconstruct the scale in the herbarium sheet and to calculate the surface area of identified plant organs. Based on our finding that YOLO11 performs better than the Mask R-CNN for our use case, we deployed a YOLO11-based service as MAS in DSArch to open up natural science collections on scale for research fields such as plant phenology and climate change science.
Full text
1,505 characters
· extracted from
oa-doi-fallback
· click to expand
Preprint
ARPHA Preprints
https://doi.org/10.3897/arphapreprints.e160486 (29 May 2025)
https://doi.org/10.3897/arphapreprints.e160486 (29 May 2025)
Published in: Research Ideas and Outcomes https://doi.org/10.3897/rio.11.e160367
Other versions:
- Preprint InfoPreprint Info
- CiteCite
- MetricsMetrics
- CommentComment
- RelatedRelated
- CitedCited
ARPHA Preprints
doi:
10.3897/arphapreprints.e160486
First posted
29 May 2025
Authors
Rajapreethi Rajendran
- Corresponding author
Senckenberg – Leibniz Institution for Biodiversity and Earth System Research, Frankfurt am Main, Germany
Senckenberg – Leibniz Institution for Biodiversity and Earth System Research, Frankfurt am Main, Germany
Senckenberg – Leibniz Institution for Biodiversity and Earth System Research, Frankfurt am Main, Germany
Naturalis Biodiversity Center, Leiden, Netherlands
Naturalis Biodiversity Center, Leiden, Netherlands
Distributed System of Scientific Collections - DiSSCo, Leiden, Netherlands
Naturalis Biodiversity Center, Leiden, Netherlands
Distributed System of Scientific Collections - DiSSCo, Leiden, Netherlands
Naturalis Biodiversity Center, Leiden, Netherlands
DiSSCo, Leiden, Netherlands
Conflict of interest
The authors have declared that no competing interests exist.
This is an open access preprint distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.