USGS Science⌕ Search

USGS · 70028775

Two-stage sequential sampling: A neighborhood-free adaptive sampling procedure

Abstract

Designing an efficient sampling scheme for a rare and clustered population is a challenging area of research. Adaptive cluster sampling, which has been shown to be viable for such a population, is based on sampling a neighborhood of units around a unit that meets a specified condition. However, the edge units produced by sampling neighborhoods have proven to limit the efficiency and applicability of adaptive cluster sampling. We propose a sampling design that is adaptive in the sense that the final sample depends on observed values, but it avoids the use of neighborhoods and the sampling of edge units. Unbiased estimators of population total and its variance are derived using Murthy's estimator. The modified two-stage sampling design is easy to implement and can be applied to a wider range of populations than adaptive cluster sampling. We evaluate the proposed sampling design by simulating sampling of two real biological populations and an artificial population for which the variable of interest took the value either 0 or 1 (e.g., indicating presence and absence of a rare event). We show that the proposed sampling design is more efficient than conventional sampling in nearly all cases. The approach used to derive estimators (Murthy's estimator) opens the door for unbiased estimators to be found for similar sequential sampling designs. ?? 2005 American Statistical Association and the International Biometric Society.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M. Salehi, D. R. Smith. 2005. Two-stage sequential sampling: A neighborhood-free adaptive sampling procedure. https://doi.org/10.1198/108571105x28183

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Dynamic population models with temporal preferential sampling to infer phenology

To study population dynamics, ecologists and wildlife biologists typically use relative abundance data, which may be subject to temporal preferential sampling. Temporal preferential sampling occurs when the times at which observations are made and the latent process of interest are conditionally dependent. To account for preferential sampling, we specify a Bayesian hierarchical abundance model that considers the dependence between observation times and the ecological process of interest. The proposed model improves relative abundance estimates during periods of infrequent observation and accounts for temporal preferential sampling in discrete time. Additionally, our model facilitates posterior inference for population growth rates and mechanistic phenometrics. We apply our model to analyze both simulated data and mosquito count data collected by the National Ecological Observatory Network. In the second case study, we characterize the population growth rate and relative abundance of several mosquito species in the Aedes genus. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme↗

Techniques to improve ecological interpretability of black box machine learning models

Statistical modeling of ecological data is often faced with a large number of variables as well as possible nonlinear relationships and higher-order interaction effects. Gradient boosted trees (GBT) have been successful in addressing these issues and have shown a good predictive performance in modeling nonlinear relationships, in particular in classification settings with a categorical response variable. They also tend to be robust against outliers. However, their black-box nature makes it difficult to interpret these models. We introduce several recently developed statistical tools to the environmental research community in order to advance interpretation of these black-box models. To analyze the properties of the tools, we applied gradient boosted trees to investigate biological health of streams within the contiguous USA, as measured by a benthic macroinvertebrate biotic index. Based on these data and a simulation study, we demonstrate the advantages and limitations of partial dependence plots (PDP), individual conditional expectation (ICE) curves and accumulated local effects (ALE) in their ability to identify covariate–response relationships. Additionally, interaction effects were quantified according to interaction strength (IAS) and Friedman’s H 2 "> H 2 statistic. Interpretable machine learning techniques are useful tools to open the black-box of gradient boosted trees in the environmental sciences. This finding is supported by our case study on the effect of impervious surface on the benthic condition, which agrees with previous results in the literature. Overall, the most important variables were ecoregion, bed stability, watershed area, riparian vegetation and catchment slope. These variables were also present in most identified interaction effects. In conclusion, graphical tools (PDP, ICE, ALE) enable visualization and easier interpretation of GBT but should be supported by analytical statistical measures. Future methodological research is needed to investigate the properties of interaction tests. Supplementary materials accompanying this paper appear on-line.

Journal of Agricultural, Biological, and Environme↗

Extreme value-based methods for modeling elk yearly movements

Species range shifts and the spread of diseases are both likely to be driven by extreme movements, but are difficult to statistically model due to their rarity. We propose a statistical approach for characterizing movement kernels that incorporate landscape covariates as well as the potential for heavy-tailed distributions. We used a spliced distribution for distance travelled paired with a resource selection function to model movements biased toward preferred habitats. As an example, we used data from 704 annual elk movements around the Greater Yellowstone Ecosystem from 2001 to 2015. Yearly elk movements were both heavy-tailed and biased away from high elevations during the winter months. We then used a simulation to illustrate how these habitat effects may alter the rate of disease spread using our estimated movement kernel relative to a more traditional approach that does not include landscape covariates. Supplementary materials accompanying this paper appear online.

Journal of Agricultural, Biological, and Environme↗