⚙
AI-generated deep summary
by claude@2026-07, 2026-07-05
· read from full text
ⓘ
This paper proposes using large language models (LLMs) to extract species interaction observations from billions of unstructured text comments submitted by citizen scientists worldwide, leveraging their multilingual, spatiotemporally rich coverage. The authors present a workflow demonstrated through two case studies in which LLMs identify biotic interactions from comments, aiming to expand the breadth and resolution of interaction data for ecological analysis. A key limitation is the reliance on text comment data and on the models’ ability to correctly infer interactions from unstructured descriptions rather than from structured field observations. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.
Abstract
Species interactions can determine species’ population sizes, geographic ranges, evolutionary trajectories, and responses to environmental change. Yet, despite their importance to many fundamental and applied questions, information on species interactions is often lacking due to constraints in data collection. Billions of text comments that have been submitted by millions of citizen scientists around the world have the potential to fill these gaps. Comments can be used to identify biotic interactions using advanced large language models (LLMs), providing a novel source of interaction data that is unusually high in spatiotemporal coverage, breadth, and resolution. This novel approach opens new avenues to evaluate species interactions on a broader scale, and to characterize and conserve biodiversity under pressing global change. Highlights - Although species interactions are central to biodiversity dynamics, progress in resolving their fundamental properties and forecasting their shifts under global change has been hindered by persistent data limitations - Citizen science platforms contain billions of observer text comments that often contain valuable information about species interactions, but the unstructured format of the information and the size of the datasets make these comments difficult to use - Large language models (LLMs) provide an unparalleled opportunity to collect and analyze species interactions from such comments - Using two case studies, we present a workflow that leverages LLMs to automatically collect species interaction observations from citizen science comments in multiple languages around the world - Such a novel source of data greatly expands the data coverage and resolution of species interactions across space and time and can help to answer both long-standing ecological questions and new, pressing questions about ecological responses to global change
Full text
1,992 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Species interactions can determine species’ population sizes, geographic ranges, evolutionary trajectories, and responses to environmental change. Yet, despite their importance to many fundamental and applied questions, information on species interactions is often lacking due to constraints in data collection. Billions of text comments that have been submitted by millions of citizen scientists around the world have the potential to fill these gaps. Comments can be used to identify biotic interactions using advanced large language models (LLMs), providing a novel source of interaction data that is unusually high in spatiotemporal coverage, breadth, and resolution. This novel approach opens new avenues to evaluate species interactions on a broader scale, and to characterize and conserve biodiversity under pressing global change.
Highlights
- Although species interactions are central to biodiversity dynamics, progress in resolving their fundamental properties and forecasting their shifts under global change has been hindered by persistent data limitations
- Citizen science platforms contain billions of observer text comments that often contain valuable information about species interactions, but the unstructured format of the information and the size of the datasets make these comments difficult to use
- Large language models (LLMs) provide an unparalleled opportunity to collect and analyze species interactions from such comments
- Using two case studies, we present a workflow that leverages LLMs to automatically collect species interaction observations from citizen science comments in multiple languages around the world
- Such a novel source of data greatly expands the data coverage and resolution of species interactions across space and time and can help to answer both long-standing ecological questions and new, pressing questions about ecological responses to global change
Competing Interest Statement
The authors have declared no competing interest.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.