Extraction of Quantitative Specimen Data using Machine Learning as a Service in the DiSSCo Research Infrastructure

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-06 · read from full text

This paper describes how the DiSSCo research infrastructure uses “Machine Learning as a Service” to extract quantitative specimen data, focusing on automating interpretation of scientific collection material rather than biomedical discovery. At a high level, it presents an infrastructure/technical approach for turning specimen-related information into structured quantitative outputs using ML services. The main finding is that an ML-as-a-service workflow can operationalize specimen data extraction within a distributed system for scientific collections, with the explicit caveat that the work is an infrastructure-focused research idea/preprint rather than a biomedical study with clinical outcomes. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

The Distributed System for Scientific Collections (DiSSCo) is a research infrastructure to integrate European natural science collections (NSCs) digitally. The aim is to facilitate and enhance the access, management and analysis of collection assets in one unified digital collection. The Machine Annotation Services (MAS) are essential components of DiSSCo’s Digital Specimen Architecture (DSArch). These services automate the annotation of digital objects to enable labeling and categorization of NSC's digital assets. To further advance this, a Machine Learning as a Service (MLaaS) approach was developed which provides researchers with the access to pre-trained machine learning models for complex tasks such as instance segmentation and morphological analysis of datasets. MLaaS enhances the DiSSCo’s scalability and flexibility and allows the integration of machine learning tools in close alignment with the FAIR (Findable, Accessible, Interoperable, Reusable) principles. This study employs DiSSCO's MLaaS framework for the quantitative analysis of herbarium specimens. Machine learning models such as Mask R-CNN and YOLO11 are comparatively applied to detect and generate the pixel-level masks of plant organs in herbarium sheets. Subsequently, these models are used to reconstruct the scale in the herbarium sheet and to calculate the surface area of identified plant organs. Based on our finding that YOLO11 performs better than the Mask R-CNN for our use case, we deployed a YOLO11-based service as MAS in DSArch to open up natural science collections on scale for research fields such as plant phenology and climate change science.
Full text 1,505 characters · extracted from oa-doi-fallback · click to expand
Preprint ARPHA Preprints https://doi.org/10.3897/arphapreprints.e160486 (29 May 2025) https://doi.org/10.3897/arphapreprints.e160486 (29 May 2025) Published in: Research Ideas and Outcomes https://doi.org/10.3897/rio.11.e160367 Other versions: - Preprint InfoPreprint Info - CiteCite - MetricsMetrics - CommentComment - RelatedRelated - CitedCited ARPHA Preprints doi: 10.3897/arphapreprints.e160486 First posted 29 May 2025 Authors Rajapreethi Rajendran - Corresponding author Senckenberg – Leibniz Institution for Biodiversity and Earth System Research, Frankfurt am Main, Germany Senckenberg – Leibniz Institution for Biodiversity and Earth System Research, Frankfurt am Main, Germany Senckenberg – Leibniz Institution for Biodiversity and Earth System Research, Frankfurt am Main, Germany Naturalis Biodiversity Center, Leiden, Netherlands Naturalis Biodiversity Center, Leiden, Netherlands Distributed System of Scientific Collections - DiSSCo, Leiden, Netherlands Naturalis Biodiversity Center, Leiden, Netherlands Distributed System of Scientific Collections - DiSSCo, Leiden, Netherlands Naturalis Biodiversity Center, Leiden, Netherlands DiSSCo, Leiden, Netherlands Conflict of interest The authors have declared that no competing interests exist. This is an open access preprint distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00