Adaptive sampling for ecological monitoring using biased data: A stratum-based approach

preprint OA: closed CC-BY-4.0

Abstract

Indicators of biodiversity change across large extents of geographic, temporal and taxonomic space are frequent products of various types of ecological monitoring and other data collection efforts. Unfortunately, many such indicators are based on data that are highly unlikely to be representative of the intended statistical populations: they are biased with respect to their estimands. Where there is full control over sampling processes, individual units within a population have known response propensities, but these are unknown in the absence of any statistical design. This could be due to the voluntary nature of surveys or because of data aggregation. In these cases some degree of sampling bias is inevitable and we must do something to ameliorate it. One such option is poststratification to adjust for uneven surveying of strata assumed to be important for unbiased estimation. We propose that a similar strategy can be used for the prioritisation of future data collection: that is, an adaptive sampling process focused on actively increasing representativeness defined in terms of response propensities. This is easily achieved by monitoring the proportional allocation of sampled units in strata relative to that expected under simple random sampling. The allocation of new units is thus that which reduces the departure from randomness (or, equivalently, that equalising response propensities across population units), allowing an estimator to approach that level of error expected under random sampling. We describe the theory supporting this straightforward strategy, and demonstrate its application using the National Plant Monitoring Scheme, a UK-focused, structured citizen science monitoring programme with uneven uptake.
Full text 2,646 characters · extracted from oa-doi-fallback · 2 sections · click to expand

Abstract

Indicators of biodiversity change across large extents of geographic, temporal and taxonomic space are frequent products of various types of ecological monitoring and other data collection efforts. Unfortunately, many such indicators are based on data that are highly unlikely to be representative of the intended statistical populations. Where there is full control over sampling processes, individual spatial units within a geographical population have known inclusion probabilities, but these are unknown in the absence of any statistical design. This could be due to the voluntary nature of surveys and/or because of dataset aggregation. In these cases some degree of sampling bias is inevitable and, depending on error tolerance relative to some real-world goal, we may need to ameliorate it. One option is poststratification to adjust for uneven surveying of strata assumed to be important for unbiased estimation. We propose that a similar strategy can be used for the prioritisation of future data collection: that is, an adaptive sampling process focused on increasing representativeness defined in terms of inclusion probabilities. This is easily achieved by monitoring the proportional allocation of sampled units in strata relative to that expected under simple random sampling. The allocation of new units is thus that which reduces the departure from randomness (or, equivalently, that equalising unit inclusion probabilities), allowing an estimator to approach that level of error expected under random sampling. We describe the theory supporting this, and demonstrate its application using sample locations from the UK National Plant Monitoring Scheme, a citizen science monitoring programme with uneven uptake, and data on the true distribution of the plant Calluna vulgaris. This in silico example demonstrates how the successful application of the method depends on the extent to which proposed strata capture correlations between inclusion probabilities and the response of interest. DOI https://doi.org/10.32942/X2MG82 Subjects Applied Statistics, Ecology and Evolutionary Biology, Life Sciences, Other Ecology and Evolutionary Biology

Keywords

survey error, survey quality, poststratification, weighting, response propensity, R-indicators, time-trends Dates Published: 2024-09-10 03:29 Last Updated: 2025-06-20 19:42 Older Versions License CC BY Attribution 4.0 International Additional Metadata Conflict of interest statement: None Data and Code Availability Statement: https://doi.org/10.5281/zenodo.13736327 Language: English

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0