From literature to biodiversity data: mining arthropod organismal and ecological traits with machine learning

preprint OA: closed
Full text JSON View at publisher

Abstract

The fields of taxonomy and biodiversity research have witnessed an exponential growth in published literature. This vast corpus of articles holds information on the diverse biological traits of organisms and their ecologies. However, access to and extraction of relevant data from this extensive resource remain challenging. Advances in text and data mining (TDM) and Natural Language Processing (NLP) techniques offer new opportunities for liberating such information from the literature. Testing and using such approaches to annotate articles in machine actionable formats is therefore necessary to enable the exploitation of existing knowledge in new biology, ecology, and evolution research. Here we explore the potential of these methods to annotate and extract organismal and ecological trait data for the most diverse animal group on Earth, the arthropods. The article processing workflow uses manually curated trait dictionaries with trained NLP models to perform labelling of entities and relationships of thousands of articles. A subset of manually annotated documents facilitated the formal evaluation of the performance of the workflow in terms of entity recognition and normalisation, and relationship extraction, highlighting several important technical challenges. The results are made available to the scientific community through an interactive web tool and queryable resource, the ArTraDB Arthropod Trait Database. These methodological explorations provide a framework that could be extended beyond the arthropods, where TDM and NLP approaches applied to the taxonomy and biodiversity literature will greatly facilitate data synthesis studies and literature reviews, the identification of knowledge gaps and biases, as well as the data-informed investigation of ecological and evolutionary trends and patterns.
Full text 1,893 characters · extracted from oa-doi-fallback · click to expand
Preprint ARPHA Preprints https://doi.org/10.3897/arphapreprints.e153174 (18 Mar 2025) https://doi.org/10.3897/arphapreprints.e153174 (18 Mar 2025) Published in: Biodiversity Data Journal https://doi.org/10.3897/BDJ.13.e153070 Other versions: - Preprint InfoPreprint Info - CiteCite - MetricsMetrics - CommentComment - RelatedRelated - CitedCited ARPHA Preprints doi: 10.3897/arphapreprints.e153174 First posted 18 Mar 2025 Authors Dalle Molle Institute for Artificial Intelligence Research (IDSIA USI-SUPSI), Lugano, Switzerland SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland Department of Ecology and Evolution, University of Lausanne, Lausanne, Switzerland SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland Dalle Molle Institute for Artificial Intelligence Research (IDSIA USI-SUPSI), Lugano, Switzerland SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland Plazi, Bern, Switzerland Digital Society Initiative, University of Zurich, Zurich, Switzerland Dalle Molle Institute for Artificial Intelligence Research (IDSIA USI-SUPSI), Lugano, Switzerland SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland Robert M Waterhouse - Corresponding author Department of Ecology and Evolution, University of Lausanne, Lausanne, Switzerland SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland Conflict of interest The authors have declared that no competing interests exist. Disclaimer: This article is (co-)authored by any of the Editors-in-Chief, Managing Editors or their deputies in this journal. Supporting agencies SNF - Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung This is an open access preprint distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00