RECODE - Relational Ecological COrpus for Data Extraction

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Ecology, conservation biology and related disciplines are inherently data-based, with the success of many research projects and initiatives (e.g., protected areas, monitoring plans etc . ) being directly dependent on the availability of location and trait information on species and populations. Unfortunately, this data is often either non-existent or available only as unstructured text within publications, especially for megadiverse taxa such as many invertebrate orders. With the emergence of large language models, there have been many attempts to automatically parse such data in machine-readable formats with variable success, either using prompt engineering or training models fit-for-purpose through named entity recognition and relation extraction. Model training has proven more efficient for complex data relations, but it needs labelled corpora, i.e. curated training data containing examples of this information for models to statistically learn from. This is a time-consuming process and, to our knowledge, no standard datasets exist upon which to train new and increasingly better models being released at an increasingly fast pace. Here we describe RECODE, a manually annotated corpus of ecological and taxonomic literature, aimed at training and fine-tuning models for automated extraction of occurrence and trait data from unstructured text. All documents presented at this stage have been annotated and validated by experts familiar with the traits of the test taxa (spiders and insects).
Full text 2,212 characters · extracted from oa-doi-fallback · click to expand
Preprint ARPHA Preprints https://doi.org/10.3897/arphapreprints.e177631 (11 Nov 2025) https://doi.org/10.3897/arphapreprints.e177631 (11 Nov 2025) Published in: Biodiversity Data Journal https://doi.org/10.3897/BDJ.14.e177365 Other versions: - Preprint InfoPreprint Info - CiteCite - MetricsMetrics - CommentComment - RelatedRelated - CitedCited ARPHA Preprints doi: 10.3897/arphapreprints.e177631 First posted 11 Nov 2025 Authors Vasco Veiga Branco - Corresponding author Finnish Museum of Natural History LUOMUS, University of Helsinki, Helsinki, Finland Centre for Ecology, Evolution and Environmental Changes (cE3c) & CHANGE - Global Change and Sustainability Institute, Faculdade de Ciências, Universidade de Lisboa, Lisbon, Portugal LASIGE and Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, Lisbon, Portugal University of Helsinki, Helsinki, Finland Finnish Museum of Natural History LUOMUS, University of Helsinki, Helsinki, Finland LASIGE and Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, Lisbon, Portugal Department of Botany and Zoology, Faculty of Science, Masaryk University, Brno, Czech Republic Finnish Museum of Natural History LUOMUS, University of Helsinki, Helsinki, Finland Centre for Ecology, Evolution and Environmental Changes (cE3c) & CHANGE - Global Change and Sustainability Institute, Faculdade de Ciências, Universidade de Lisboa, Lisbon, Portugal Centre for Ecology, Evolution and Environmental Changes (cE3c) & CHANGE - Global Change and Sustainability Institute, Faculdade de Ciências, Universidade de Lisboa, Lisbon, Portugal Finnish Museum of Natural History LUOMUS, University of Helsinki, Helsinki, Finland Centre for Ecology, Evolution and Environmental Changes (cE3c) & CHANGE - Global Change and Sustainability Institute, Faculdade de Ciências, Universidade de Lisboa, Lisbon, Portugal Conflict of interest The authors have declared that no competing interests exist. This is an open access preprint distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0