USGS ScienceSearch

SEARCH · USGS Science

Results for “Environmental and Ecological Statistics”

Search indexed USGS publications on groundwater, aquifers, geologic maps, mineral resources and earthquakes. Explore source records by subject and place.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bringing Bayesian models to life

Bringing Bayesian Models to Life empowers the reader to extend, enhance, and implement statistical models for ecological and environmental data analysis. We open the black box and show the reader how to connect modern statistical models to computer algorithms. These algorithms allow the user to fit models that answer their scientific questions without needing to rely on automated Bayesian software. We show how to handcraft statistical models that are useful in ecological and environmental science including: linear and generalized linear models, spatial and time series models, occupancy and capture-recapture models, animal movement models, spatio-temporal models, and integrated population-models.

Book

Statistical facilitation in environmental science: Integrating results from complementary statistical analyses can improve ecological interpretations

Professionals working in biological conservation seek to understand, manage, and restore populations of native organisms using many techniques. A common approach for this discipline is using long-term data collections to inform decision making. However, several quantitative issues complicate statistical analysis of monitoring datasets and can reduce the utility of results for conservation decision making. Integrating results from multiple analyses applied to the same dataset (i.e., approaching the same biological problem using different techniques) is one way to address concerns related to field data that violate statistical assumptions. This process allows data analysts, researchers, and managers to assemble insights based on the weight of evidence. Here we tested whether three different statistical techniques [(1) multiple logistic regression on original data, (2) multiple logistic regression on standardized data (i.e., mean of 0 and standard deviation of 1), and (3) random forest analysis] identified a similar hierarchy for selecting natural and anthropogenic habitat regressors. Our examination of how environmental variables affected Plains Minnow ( Hybognathus placitus ), a state-threatened fish, is relevant to other taxa and locations. We gained useful information from redundancies (i.e., agreements across analyses). New directions also emerged by addressing ambiguities (i.e., disagreements among results across analyses). When multiple analyses were integrated into one ecological story, a clearer interpretation emerged. Viewing different statistical tests as facilitators that provide mutual advantages can advance the understanding and application of statistical analyses applied to non-experimental field datasets.

Kansas

A method for assigning species into groups based on generalized Mahalanobis distance between habitat model coefficients

Habitat association models are commonly developed for individual animal species using generalized linear modeling methods such as logistic regression. We considered the issue of grouping species based on their habitat use so that management decisions can be based on sets of species rather than individual species. This research was motivated by a study of western landbirds in northern Idaho forests. The method we examined was to separately fit models to each species and to use a generalized Mahalanobis distance between coefficient vectors to create a distance matrix among species. Clustering methods were used to group species from the distance matrix, and multidimensional scaling methods were used to visualize the relations among species groups. Methods were also discussed for evaluating the sensitivity of the conclusions because of outliers or influential data points. We illustrate these methods with data from the landbird study conducted in northern Idaho. Simulation results are presented to compare the success of this method to alternative methods using Euclidean distance between coefficient vectors and to methods that do not use habitat association models. These simulations demonstrate that our Mahalanobis-distance- based method was nearly always better than Euclidean-distance-based methods or methods not based on habitat association models. The methods used to develop candidate species groups are easily explained to other scientists and resource managers since they mainly rely on classical multivariate statistical methods. ?? 2008 Springer Science+Business Media, LLC.

Environmental and Ecological Statistics

Using structural equation modeling to investigate relationships among ecological variables

Structural equation modeling is an advanced multivariate statistical process with which a researcher can construct theoretical concepts, test their measurement reliability, hypothesize and test a theory about their relationships, take into account measurement errors, and consider both direct and indirect effects of variables on one another. Latent variables are theoretical concepts that unite phenomena under a single term, e.g., ecosystem health, environmental condition, and pollution (Bollen, 1989). Latent variables are not measured directly but can be expressed in terms of one or more directly measurable variables called indicators. For some researchers, defining, constructing, and examining the validity of latent variables may be the end task of itself. For others, testing hypothesized relationships of latent variables may be of interest. We analyzed the correlation matrix of eleven environmental variables from the U.S. Environmental Protection Agency's (USEPA) Environmental Monitoring and Assessment Program for Estuaries (EMAP-E) using methods of structural equation modeling. We hypothesized and tested a conceptual model to characterize the interdependencies between four latent variables-sediment contamination, natural variability, biodiversity, and growth potential. In particular, we were interested in measuring the direct, indirect, and total effects of sediment contamination and natural variability on biodiversity and growth potential. The model fit the data well and accounted for 81% of the variability in biodiversity and 69% of the variability in growth potential. It revealed a positive total effect of natural variability on growth potential that otherwise would have been judged negative had we not considered indirect effects. That is, natural variability had a negative direct effect on growth potential of magnitude -0.3251 and a positive indirect effect mediated through biodiversity of magnitude 0.4509, yielding a net positive total effect of 0.1258. Natural variability had a positive direct effect on biodiversity of magnitude 0.5347 and a negative indirect effect mediated through growth potential of magnitude -0.1105 yielding a positive total effects of magnitude 0.4242. Sediment contamination had a negative direct effect on biodiversity of magnitude -0.1956 and a negative indirect effect on growth potential via biodiversity of magnitude -0.067. Biodiversity had a positive effect on growth potential of magnitude 0.8432, and growth potential had a positive effect on biodiversity of magnitude 0.3398. The correlation between biodiversity and growth potential was estimated at 0.7658 and that between sediment contamination and natural variability at -0.3769.

Environmental and Ecological Statistics

Efficient statistical mapping of avian count data

We develop a spatial modeling framework for count data that is efficient to implement in high-dimensional prediction problems. We consider spectral parameterizations for the spatially varying mean of a Poisson model. The spectral parameterization of the spatial process is very computationally efficient, enabling effective estimation and prediction in large problems using Markov chain Monte Carlo techniques. We apply this model to creating avian relative abundance maps from North American Breeding Bird Survey (BBS) data. Variation in the ability of observers to count birds is modeled as spatially independent noise, resulting in over-dispersion relative to the Poisson assumption. This approach represents an improvement over existing approaches used for spatial modeling of BBS data which are either inefficient for continental scale modeling and prediction or fail to accommodate important distributional features of count data thus leading to inaccurate accounting of prediction uncertainty.

Environmental and Ecological Statistics

Survey methods for assessing land cover map accuracy

The increasing availability of digital photographic materials has fueled efforts by agencies and organizations to generate land cover maps for states, regions, and the United States as a whole. Regardless of the information sources and classification methods used, land cover maps are subject to numerous sources of error. In order to understand the quality of the information contained in these maps, it is desirable to generate statistically valid estimates of accuracy rates describing misclassification errors. We explored a full sample survey framework for creating accuracy assessment study designs that balance statistical and operational considerations in relation to study objectives for a regional assessment of GAP land cover maps. We focused not only on appropriate sample designs and estimation approaches, but on aspects of the data collection process, such as gaining cooperation of land owners and using pixel clusters as an observation unit. The approach was tested in a pilot study to assess the accuracy of Iowa GAP land cover maps. A stratified two-stage cluster sampling design addressed sample size requirements for land covers and the need for geographic spread while minimizing operational effort. Recruitment methods used for private land owners yielded high response rates, minimizing a source of nonresponse error. Collecting data for a 9-pixel cluster centered on the sampled pixel was simple to implement, and provided better information on rarer vegetation classes as well as substantial gains in precision relative to observing data at a single-pixel.

Environmental and Ecological Statistics

A statistical evaluation of non-ergodic variogram estimators

Geostatistics is a set of statistical techniques that is increasingly used to characterize spatial dependence in spatially referenced ecological data. A common feature of geostatistics is predicting values at unsampled locations from nearby samples using the kriging algorithm. Modeling spatial dependence in sampled data is necessary before kriging and is usually accomplished with the variogram and its traditional estimator. Other types of estimators, known as non-ergodic estimators, have been used in ecological applications. Non-ergodic estimators were originally suggested as a method of choice when sampled data are preferentially located and exhibit a skewed frequency distribution. Preferentially located samples can occur, for example, when areas with high values are sampled more intensely than other areas. In earlier studies the visual appearance of variograms from traditional and non-ergodic estimators were compared. Here we evaluate the estimators' relative performance in prediction. We also show algebraically that a non-ergodic version of the variogram is equivalent to the traditional variogram estimator. Simulations, designed to investigate the effects of data skewness and preferential sampling on variogram estimation and kriging, showed the traditional variogram estimator outperforms the non-ergodic estimators under these conditions. We also analyzed data on carabid beetle abundance, which exhibited large-scale spatial variability (trend) and a skewed frequency distribution. Detrending data followed by robust estimation of the residual variogram is demonstrated to be a successful alternative to the non-ergodic approach.

Environmental and Ecological Statistics

Nonstationary demographic state-space models using unreplicated counts for species undergoing environmental stressors

A fundamental task in ecological statistics is to estimate abundance and growth rate distributions from wildlife monitoring data to inform conservation management. Modeling time series of wildlife populations presents a number of challenges from both statistical and ecological perspectives, including discreteness; lack of replication; nonstationarity; and observation, demographic, and other phenomenological processes. Nonstationary dynamics are often exhibited by populations undergoing environmental stressors. Models must account for these characteristics to produce reliable estimates of abundance and trends, yet estimation can be challenging with unreplicated data. We propose nonstationary demographic state-space models using unreplicated counts for populations undergoing environmental stressors. A reduced growth rate model matches the complexity of the unreplicated count data, and a fecundity bound on growth rate distributions allows the separation of processes affecting growth rates like environmental stressors from those affecting abundance external to growth rates like migration. NDSSMs allow for the embedding of nonstationary model components, and we explore the use of changepoints, volatility clustering, and migration processes. We apply the proposed nonstationary models in case studies of herons affected by predator/competitor reestablishment and three bat species affected by a fungal pathogen causing white-nose syndrome. Nonstationary models outperform stationary models and generalized linear mixed effects models according to model scoring and visual inspection of predictions, and provide estimates more consistent with published values. Incorporating migration improves model fit universally, even with approximate one-way immigration, most likely because populations are extirpated, recolonized, and increase multiple-fold over the upper bound set by species fecundity. In addition, estimates of the timing and severity of the environmental stressor differed for models with migration. Including nonstationary and demographic components in a fecundity-bounded growth rate model improves inference and benefits interpretability of hyperparameters. In turn, this adjusts uncertainties in predictions of abundance and growth rates over time, providing the ingredients needed for informed conservation analysis and for directing future monitoring of at-risk species.

Journal of Agricultural, Biological and Environmen

Representing general theoretical concepts in structural equation models: The role of composite variables

Structural equation modeling (SEM) holds the promise of providing natural scientists the capacity to evaluate complex multivariate hypotheses about ecological systems. Building on its predecessors, path analysis and factor analysis, SEM allows for the incorporation of both observed and unobserved (latent) variables into theoretically-based probabilistic models. In this paper we discuss the interface between theory and data in SEM and the use of an additional variable type, the composite. In simple terms, composite variables specify the influences of collections of other variables and can be helpful in modeling heterogeneous concepts of the sort commonly of interest to ecologists. While long recognized as a potentially important element of SEM, composite variables have received very limited use, in part because of a lack of theoretical consideration, but also because of difficulties that arise in parameter estimation when using conventional solution procedures. In this paper we present a framework for discussing composites and demonstrate how the use of partially-reduced-form models can help to overcome some of the parameter estimation and evaluation problems associated with models containing composites. Diagnostic procedures for evaluating the most appropriate and effective use of composites are illustrated with an example from the ecological literature. It is argued that an ability to incorporate composite variables into structural equation models may be particularly valuable in the study of natural systems, where concepts are frequently multifaceted and the influence of suites of variables are often of interest. ?? Springer Science+Business Media, LLC 2007.

Environmental and Ecological Statistics

Inference for finite-sample trajectories in dynamic multi-state site-occupancy models using hidden Markov model smoothing

Ecologists and wildlife biologists increasingly use latent variable models to study patterns of species occurrence when detection is imperfect. These models have recently been generalized to accommodate both a more expansive description of state than simple presence or absence, and Markovian dynamics in the latent state over successive sampling seasons. In this paper, we write these multi-season, multi-state models as hidden Markov models to find both maximum likelihood estimates of model parameters and finite-sample estimators of the trajectory of the latent state over time. These estimators are especially useful for characterizing population trends in species of conservation concern. We also develop parametric bootstrap procedures that allow formal inference about latent trend. We examine model behavior through simulation, and we apply the model to data from the North American Amphibian Monitoring Program.

Environmental and Ecological Statistics

Bayesian analysis of Jolly-Seber type models

We propose the use of finite mixtures of continuous distributions in modelling the process by which new individuals, that arrive in groups, become part of a wildlife population. We demonstrate this approach using a data set of migrating semipalmated sandpipers ( Calidris pussila ) for which we extend existing stopover models to allow for individuals to have different behaviour in terms of their stopover duration at the site. We demonstrate the use of reversible jump MCMC methods to derive posterior distributions for the model parameters and the models, simultaneously. The algorithm moves between models with different numbers of arrival groups as well as between models with different numbers of behavioural groups. The approach is shown to provide new ecological insights about the stopover behaviour of semipalmated sandpipers but is generally applicable to any population in which animals arrive in groups and potentially exhibit heterogeneity in terms of one or more other processes.

Environmental and Ecological Statistics

Comparing methods to estimate the proportion of turbine-induced bird and bat mortality in the search area under a road and pad search protocol

Estimating bird and bat mortality at wind facilities typically involves searching for carcasses on the ground near turbines. Some fraction of carcasses inevitably lie outside the search plots, and accurate mortality estimation requires accounting for those carcasses using models to extrapolate from searched to unsearched areas. Such models should account for variation in carcass density with distance, and ideally also for variation with direction (anisotropy). We compare five methods of accounting for carcasses that land outside the searched area (ratio, weighted distribution, non-parametric, and two generalized linear models ( glm )) by simulating spatial arrival patterns and the detection process to mimic observations which result from surveying only, or primarily, roads and pads (R&P) and applying the five methods. Simulations vary R&P configurations, spatial carcass distributions (isotropic and anisotropic), and per turbine fatality rates. Our results suggest that the ratio method is less accurate with higher variation relative to the other four methods which all perform similarly under isotropy. All methods were biased under anisotropy; however, including direction covariates in the glm method substantially reduced bias. In addition to comparing methods of accounting for unsearched areas, we suggest a semiparametric bootstrap to produce confidence-based bounds for the proportion of carcasses that land in the searched area.

Environmental and Ecological Statistics

Model-based surveillance system design under practical constraints with application to white-nose syndrome

Infectious diseases are powerful ecological forces structuring ecosystems, causing devastating economic impacts and disrupting society. Successful prevention and control of pathogens requires knowledge of the current scope and severity of disease, as well as the ability to forecast future disease dynamics. Assessment of the current situation as well as prediction of the future conditions, rely on spatially referenced information regarding the presence or absence of a pathogen, and the prevalence of the pathogen in the population. In particular, knowledge about the location of the disease front is foundational for deploying disease countermeasures to prevent further disease spread and focusing control efforts to reduce disease intensity in affected areas. In this paper, we develop a model-based approach to designing sampling strategies for wildlife disease surveillance at the disease front. Specifically, we use a mechanistic spatio-temporal model based on an underlying partial differential equation to track the disease dynamics and predict the disease prevalence in the future. We also devise an optimal surveillance system design at the disease front that takes into account the practical constraints of sampling. We evaluate the effectiveness of our proposed design via a simulation study and demonstrate the application of the proposed approach by designing a surveillance strategy for the pathogen that causes white-nose syndrome.

Environmental and Ecological Statistics

A hierarchical model for eDNA fate and transport dynamics accommodating low concentration samples

Environmental DNA (eDNA) sampling is an increasingly important tool for answering ecological questions and informing aquatic species management; however, several factors currently limit the reliability of ecological inference from eDNA sampling. Two particular challenges are (1) determining species source location(s) and (2) accurately and precisely measuring low concentration eDNA samples in the presence of multiple sources of ecological and measurement variability. The recently introduced eDNA Integrating Transport and Hydrology (eDITH) model provides a framework for relating eDNA measurements to source locations in riverine networks, but little empirical work has been done to test and refine model assumptions or accommodate low concentration samples, that can be systematically undermeasured. To better understand eDNA fate and transport dynamics and our ability to reliably quantify low concentration samples, we developed a hierarchical model and used it to evaluate a fate and transport experiment. Our model addresses several low concentration challenges by modeling the number of copies in each PCR replicate as a latent variable with a count distribution and conditioning detection and quantification on replicate copy number. We provide evidence that the eDNA removal rate declined through time, estimating that over 80% of eDNA was removed over the first 10 m, traversed in 41 s. After this initial period of rapid decay, eDNA decayed slowly with consistent detection through our farthest site 1 km from the release location, traversed in 67.8 min. Our model further allowed us to detect extra-Poisson variation in the allocation of copies to replicates. We extended our hierarchical model to accommodate a continuous effect of inhibitors and used our model to provide evidence for the inhibitor hypothesis and explore the potential implications. While our model is not a panacea for all challenges faced when quantifying low-concentration eDNA samples, it provides a framework for a more complete accounting of uncertainty.

Environmental and Ecological Statistics

A closure test for time-specific capture-recapture data

The assumption of demographic closure in the analysis of capture-recapture data under closed-population models is of fundamental importance. Yet, little progress has been made in the development of omnibus tests of the closure assumption. We present a closure test for time-specific data that, in principle, tests the null hypothesis of closed-population model M(t) against the open-population Jolly-Seber model as a specific alternative. This test is chi-square, and can be decomposed into informative components that can be interpreted to determine the nature of closure violations. The test is most sensitive to permanent emigration and least sensitive to temporary emigration, and is of intermediate sensitivity to permanent or temporary immigration. This test is a versatile tool for testing the assumption of demographic closure in the analysis of capture-recapture data.

Environmental and Ecological Statistics

Uncertainty, learning, and the optimal management of wildlife

Wildlife management is limited by uncontrolled and often unrecognized environmental variation, by limited capabilities to observe and control animal populations, and by a lack of understanding about the biological processes driving population dynamics. In this paper I describe a comprehensive framework for management that includes multiple models and likelihood values to account for structural uncertainty, along with stochastic factors to account for environmental variation, random sampling, and partial controllability. Adaptive optimization is developed in terms of the optimal control of incompletely understood populations, with the expected value of perfect information measuring the potential for improving control through learning. The framework for optimal adaptive control is generalized by including partial observability and non-adaptive, sample-based updating of model likelihoods. Passive adaptive management is derived as a special case of constrained adaptive optimization, representing a potentially efficient suboptimal alternative that nonetheless accounts for structural uncertainty.

Environmental and Ecological Statistics

Echelon approach to areas of concern in synoptic regional monitoring

Echelons provide an objective approach to prospecting for areas of potential concern in synoptic regional monitoring of a surface variable. Echelons can be regarded informally as stacked hill forms. The strategy is to identify regions of the surface which are elevated relative to surroundings ( R elative ELEVATIONS or RELEVATIONS ). These are areas which would continue to expand as islands with receding (virtual) floodwaters. Levels where islands would merge are critical elevations which delimit echelons in the vertical dimension. Families of echelons consist of surface sectors constituting separate islands for deeper waters that merge as water level declines. Pits which would hold water are disregarded in such a progression, but a complementary analysis of pits is obtained using the surface as a virtual mould to cast a counter-surface (bathymetric analysis). An echelon tree is a family tree of echelons with peaks as terminals and the lowest level as root. An echelon tree thus provides a dendrogram representation of surface topology which enables graph theoretic analysis and comparison of surface structures. Echelon top view maps show echelon cover sectors on the base plane. An echelon table summarizes characteristics of echelons as instances or cases of hill form surface structure. Determination of echelons requires only ordinal strength for the surface variable, and is thus appropriate for environmental indices as well as measurements. Since echelons are inherent in a surface rather than perceptual, they provide a basis for computer-intelligent understanding of surfaces. Echelons are given for broad-scale mammalian species richness in Pennsylvania.

Environmental and Ecological Statistics

Modelling heterogeneity in the recoveries of marked animal populations with covariates of individual animals, groups of animals or recovery time

A general framework is developed for modelling rates of survival and recovery of marked animal populations in terms of auxiliary information collected at the time of marking. The framework may be used to estimate differences in survival or recovery among individual animals, groups of animals, and recovery times. Analyses of the recoveries of tagged fish and banded bird populations are used to illustrate the specification and selection of various models.

Environmental and Ecological Statistics